ConMIL: interactive and contrastive text-guided multiple instance learning for whole slide image classification.
basic_science · Level V
Where this comes from
- Record sourced from PubMed, PMID 42745549.
- Also identified by DOI 10.1093/bioinformatics/btag679.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Whole-slide image (WSI) classification in computational pathology typically relies on Multiple Instance Learning (MIL) for weakly supervised analysis. Recent pathology vision-language models have inspired text-guided approaches, but these methods typically use text for representation alignment or region localization, rather than directly incorporating semantic signals into MIL attention weighting. Furthermore, these approaches often rely on static prompts and provide limited insight into the learned nonlinear transformations performed by the classifier. We propose ConMIL, an interactive contrastive text-guided MIL framework for WSI classification. At its core, ConMIL introduces a contrastive semantic-guided attention mechanism that uses paired positive and negative pathology-specific text embeddings to directly modulate MIL attention weighting. This mechanism is complemented by human-in-the-loop prompt refinement to improve semantic specificity and a Kolmogorov-Arnold Network (KAN) classifier that enables visualization and quantitative inspection of learned nonlinear transformations. Experiments on CAMELYON16, TCGA-BRCA, and BRACS demonstrate that ConMIL consistently outperforms representative MIL baselines while producing pathology-consistent attention heatmaps and inspectable nonlinear transformations. The source code for ConMIL is available at https://github.com/anxuanhan/ConMIL. Supplementary data are available at Bioinformatics online.