Using deep learning to identify inherited retinal diseases based on wide-field retinal imaging data.

Joskowicz, Leo; Buchbinder, Tim; Chodorov, Eldan; Hoogi, Assaf; Matos, Katherine; Rivera, Antonio; Sharon, Dror; Banin, Eyal et al. · PLoS One · 2026

retrospective_cohort · Level III

Where this comes from

Abstract

To evaluate a novel image-based deep learning method for the automated identification of inherited retinal diseases (IRDs) and to explore the feasibility of predicting selected causative gene groups using a multimodal analysis of wide-field fundus autofluorescence (FAF) and pseudocolor fundus (pCF) images. The method was evaluated using a retrospective dataset of patient studies containing FAF and pCF images, as well as genetic tests for IRD. Patients with confirmed IRD for which both wide-field FAF and pCF images and genetic tests for IRD performed at Hadassah University Medical Center were included. The dataset consisted of 409 patients (330 patients with IRD with the 25 most commonly affected genes in our population and patients without IRD, and 79 patients without IRD). Nine EfficientNet-V2-m convolutional neural networks were trained for the following three classification tasks: a binary IRD vs. non-IRD classification, and classification into two groups of five causative genes (Groups 1 and 2). For each task, three models were trained on the FAF images only, the pCF images only, and both the FAF and pCF images. The performance of the models was then evaluated and compared using 5-fold cross-validation. Accuracy, precision, F1 scores, AUC, and confusion matrices. The multimodal classification models that were trained on both the FAF and pCF images yielded the best results. The binary classification model had a mean (±SD) accuracy of 0.95 ± 0.01, a mean precision of 0.92 ± 0.01, and a mean F1 score of 0.90 ± 0.02. The Group 1 classification model had a mean accuracy of 0.92 ± 0.03, a mean precision of 0.93 ± 0.03, and a mean F1 score of 0.89 ± 0.03. Finally, the Group 2 classification model had a mean accuracy of 0.85 ± 0.03, a mean precision of 0.87 ± 0.04, and a mean F1 score of 0.83 ± 0.04. Our results indicate that determining whether a patient has IRD can be performed with high accuracy within this retrospective cohort based on FAF and pCF images using image-based deep learning classifiers. This image-based approach may assist clinicians during the patient's initial visit by providing decision support prior to genetic testing. It may also help prioritize patients for genetic workup, particularly in settings in which genetic testing is not readily available. Further prospective and external validation is required before clinical implementation.

Medical subject headings