A clinician aligned vision language framework for stepwise interpretation in fundus fluorescein angiography.

Su, Zichang; Liu, Xiaocong; Guan, Bingtao; Shao, An; Liu, Xindi; Yan, Yan; Luo, Ziyao; Li, Zhikang et al. · NPJ Digit Med · 2026

prospective_cohort · Level II

Where this comes from

Abstract

Fluorescein fundus angiography (FFA) is essential for diagnosing retinal vascular diseases, yet its interpretation is expertise-intensive. Here, we present Clin-FFA-VLM, a multimodal vision-language framework that mirrors retina specialists' cognitive workflow by decomposing FFA interpretation into three stages: lesion-aware visual perception, clinical report generation, and diagnostic decision support. Trained and tested on a multi-center dataset of 13,178 FFA images with expert label, 21,717 FFA images with 1790 clinical reports and diagnosis across 7 retinal diseases, Clin-FFA-VLM achieves an F1 of 0.834 for lesion detection, an entity-level F1 of 0.73 for report generation, and a diagnostic F1 of 0.77 by jointly reasoning over images and self-generated reports. External validation across two independent hospitals confirmed its generalizability (F1 of 0.78 and 0.70). In a prospective reader study with 200 FFA cases, Clin-FFA-VLM significantly improved diagnostic accuracy for medical students and residents (p < 0.05), bridging the gap between automated systems and clinical practice.