ConMIL: interactive and contrastive text-guided multiple instance learning for whole slide image classification.

Han, Anxuan; Jolley, Alexandra; Butler, Lisa M; Chen, Weitong · Bioinformatics · 2026

basic_science · Level V

Where this comes from

Abstract

Whole-slide image (WSI) classification in computational pathology typically relies on Multiple Instance Learning (MIL) for weakly supervised analysis. Recent pathology vision-language models have inspired text-guided approaches, but these methods typically use text for representation alignment or region localization, rather than directly incorporating semantic signals into MIL attention weighting. Furthermore, these approaches often rely on static prompts and provide limited insight into the learned nonlinear transformations performed by the classifier. We propose ConMIL, an interactive contrastive text-guided MIL framework for WSI classification. At its core, ConMIL introduces a contrastive semantic-guided attention mechanism that uses paired positive and negative pathology-specific text embeddings to directly modulate MIL attention weighting. This mechanism is complemented by human-in-the-loop prompt refinement to improve semantic specificity and a Kolmogorov-Arnold Network (KAN) classifier that enables visualization and quantitative inspection of learned nonlinear transformations. Experiments on CAMELYON16, TCGA-BRCA, and BRACS demonstrate that ConMIL consistently outperforms representative MIL baselines while producing pathology-consistent attention heatmaps and inspectable nonlinear transformations. The source code for ConMIL is available at https://github.com/anxuanhan/ConMIL. Supplementary data are available at Bioinformatics online.