VIGIL: Vision-Language Guided Multiple Instance Learning Framework for Ulcerative Colitis Histological Healing Prediction.
Where this comes from
- Record sourced from PubMed, PMID 42594012.
- Also identified by DOI 10.1109/TBME.2026.3723583.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Ulcerative colitis (UC) is characterized by chronic inflammation with alternating remission-relapse cycles, necessitating precise histological healing (HH) evaluation to improve clinical outcomes. To address the limitations of annotation-intensive deep learning methods and suboptimal multi-instance learning (MIL) in HH prediction, we propose VIGIL, the first vision-language-guided MIL framework that integrates white-light endoscopy (WLE) and endocytoscopy (EC). VIGIL introduces a dual-branch MIL module, KS-MIL, based on top-K typical frame selection and similarity-weighted metric learning to effectively model relationships among frame features. By incorporating diagnostic report text and a specially designed multi-level alignment and supervision strategy between image-text pairs, VIGIL establishes joint vision-language guidance during training to capture richer disease-related semantic information. Furthermore, VIGIL employs a multi-modal masked relation fusion (MMRF) strategy to uncover latent diagnostic correlations between WLE and EC representations. Comprehensive experiments on a real-world clinical dataset demonstrate VIGIL's superior performance, achieving 92.69% accuracy and 94.79% AUC, outperforming existing state-of-the-art methods. The proposed VIGIL framework establishes an effective vision-language guided MIL paradigm for UC HH prediction, reducing annotation burdens while improving prediction reliability. The research outcomes provide new insights for non-invasive UC diagnosis and hold theoretical significance and clinical value for advancing intelligent healthcare development.