M2OTCA: Multiple-magnification optimal transport-based cross-attention learning for whole slide image classification.
basic_science · Level V
Where this comes from
- Record sourced from PubMed, PMID 42001680.
- Also identified by DOI 10.1016/j.media.2026.104082.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Accurate and automated whole slide image (WSI) classification is important in early diagnosis of cancer, which is typically achieved using multiple instance learning (MIL). Existing MIL approaches for WSI classification still face practical challenges, particularly related to large numbers of patches cropped from a WSI with only slide-level labels. In these cases, training often suffers from overfitting due to the weak supervision provided by the slide-level labels. Therefore, it is crucial to exploit more information from limited slide-level data for WSI analysis. However, existing approaches concentrate on only single-magnification feature mining, which fails to capture feature consistency among different magnifications. To address this shortcoming, we here propose a multiple-magnification optimal transport-based cross-attention (M2OTCA) MIL framework for WSI classification, in which optimal transport (OT) is applied to match feature distributions of a WSI across different magnifications for selecting magnification-specific informative patches. To alleviate the high computational complexity, we effectively summarize the magnification-specific basic structural distribution at each magnification by condensing constituting instances as region structural prototypes, and modeling the original OT learning as cross-magnification region structural prototypes OT learning. Moreover, to generate a faithful explanation for cross-magnification OT learning, we derive region structural prototype gradients at the high magnification prototypes (20×) as the indicators to enhance low magnification prototypes (10× and 5×) for efficient feature co-expressions. Then, prototypes at each magnification are aggregated for magnification-specific predictions which are concatenated for final prediction. Extensive experiments on four representative WSI datasets show the superiority of M2OTCA compared to other state-of-the-art methods.