Radiology Board-Style Examinations and Large Language Models: A Scoping Review of Model Performance.
systematic_review · Level I
Where this comes from
- Record sourced from PubMed, PMID 41616963.
- Also identified by DOI 10.1016/j.jacr.2026.01.017.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Large language models (LLMs) are increasingly being evaluated for their ability to answer official radiology board-style examination questions. Understanding their accuracy, limitations, and potential applications in education is essential for assessing their utility in the field. A scoping review was conducted in October 2025 across PubMed, Scopus, and Web of Science, following Preferred Reporting Items for Systematic Reviews and Meta-Analyses guidelines. Studies were included if they evaluated LLMs on official radiology board-style examination questions. After screening 205 unique records, 29 studies met the inclusion criteria. Data were extracted on study characteristics, including LLM type and version, input modality, language, examination type, answer format, comparison with humans, and reported outcomes. The reviewed studies evaluated multiple LLMs, predominantly Chat Generative Pre-trained Transformer (GPT)-based models (GPT-3.5, GPT-4, GPT-4 Turbo, GPT-4o), as well as Claude, Gemini, Llama 3, and Mixtral. Text-only evaluations generally yielded higher accuracy (≈65%-90%) compared with multimodal tasks (45%-89%). GPT-4 and its variants consistently outperformed earlier versions, occasionally exceeding average human performance. Open-source models such as Llama 3 70B and Mixtral achieved comparable results to proprietary models, offering advantages in local deployment and privacy. Few studies directly compared LLM performance with human radiologists. LLMs demonstrate promising performance in answering text-based radiology board-style examination questions, particularly GPT-4-based models. Nevertheless, significant limitations persist in multimodal tasks and complex reasoning scenarios.
Medical subject headings
- Large Language Models
- Radiology
- Specialty Boards
- Educational Measurement
- Clinical Competence