Learning like a radiologist: a medical vision-language model for radiological image analysis via curriculum learning.
basic_science · Level V
Where this comes from
- Record sourced from PubMed, PMID 42249121.
- Also identified by DOI 10.1038/s41746-026-02713-3.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Medical vision-language models (MVLMs) offer promise in clinical practice but face limitations in generalizability, data quality, and clinically meaningful evaluation. We propose RadiSim-CL, an MVLM trained via curriculum learning by simulating the three-phase pathway of a radiologist: foundational knowledge understanding, anatomical knowledge, and advanced diagnostic reasoning. To support this, we curate RadiSim, a 12-million image-text pair dataset aligned to these phases. We evaluate the model using a five-stage coarse-to-fine validation framework: (1) modality recognition, (2) anatomical recognition, (3) anatomical localization, (4) abnormality and disease diagnosis, and (5) disease differentiation and grading. This framework spans 24 zero-shot subtasks across MR, CT, and DR imaging. RadiSim-CL achieves comparable performance to state-of-the-art baselines in both foundational and anatomical tasks, and demonstrates superior capabilities in complex reasoning (e.g., an AUC of 0.953 for brain tumor diagnosis and an accuracy of 0.764 for meningioma grading). Ablation studies further confirm the curriculum's effectiveness. RadiSim-CL thus offers a scalable, clinically aligned solution to enhance diagnostic precision.