Towards more efficient and better multi-view and multi-modal retinopathy assisted diagnosis.
basic_science · Level V
Where this comes from
- Record sourced from PubMed, PMID 41655395.
- Also identified by DOI 10.1016/j.artmed.2026.103376.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Fundus images are widely used in early retinopathy examination to prevent visual impairment caused by retinopathy. The retinopathy examination process based on fundus images can be mainly summarized in three steps: (1) ophthalmologists obtain comprehensive fundus information by jointly analyzing multi-view fundus images; (2) ophthalmologists obtain complementary lesion information by contrastingly analyzing multi-modal fundus images; (3) ophthalmologists diagnose retinopathy categories and write specialized fundus reports. To simulate the clinical fundus image examination process, we introduce an efficient multi-view and multi-modal fundus image joint ancillary diagnosis framework that can simultaneously accept fundus images of different views and modalities for pathology classification and symptom report generation tasks. In our framework, we propose jointly employing self-attention in intra-view local and inter-view sparse global windows to extract comprehensive fundus information among different views. We propose a multi-modal fusion transformer via shunted multi-scale cross-attention to model lesions of various scales by splitting attention granularity at query and queried modalities to fuse complementary lesion information among different modalities. The experimental results of retinopathy classification and report generation tasks indicate that our proposed method is superior to other benchmarking methods, achieving a classification accuracy of 83.96% and a report generation CIDEr of 0.934.
Medical subject headings
- Retinal Diseases
- Fundus Oculi
- Diagnosis, Computer-Assisted
- Image Interpretation, Computer-Assisted