Artificial Intelligence in Radiology: Performance of ChatGPT-4v and GPT-4o on Diagnostic Radiology in-Training (DXIT) Examination Questions.

Martini, Reema S; Sang, Alan; Saunders, Pedro; Bala, Wasif; Li, Hanzhou; Moon, John T; Balthazar, Patricia · J Am Coll Radiol · 2026

other · Level V

Where this comes from

Abstract

The purpose of this study is to examine the performance of Chat Generative Pre-trained Transformer (GPT)-4vision (GPT-4v) and GPT-4omni (GPT-4o) on the ACR's Diagnostic Radiology in-Training (DXIT) examination, comparing performance on image-based and text-only questions. In all, 1,136 publicly available DXIT examination questions were input into GPT-4v and GPT-4o with a prompt asking the large language model to provide its answer, rationale, and confidence level (0-100). Accuracy of each model across different categories was then analyzed, with χ<sup>2</sup> tests to compare proportions, t tests to compare means, and receiver operating characteristic curves to evaluate confidence levels. GPT-4o and GPT-4v achieved accuracies of 73.5% and 69.3%, respectively (P < .0001) while scoring 55.6% and 50.3% on image-based questions (P < .0001). Receiver operating characteristic curves of confidence levels and correctness produced areas under the curve of 0.64 and 0.66 for GPT-4o and GPT-4v, respectively. GPT-4o outperformed GPT-4v on nearly every metric, with both models outperforming the national average performance of postgraduate year 3 radiology residents (61.9%) on the 2022 DXIT examination. However, performance on image-based questions remains significantly worse than text-only questions, and both models score below radiology trainees from the same cohort. Both models exhibit limited ability to predict correctness using an intrinsic confidence level. Use of ChatGPT for test preparation and image interpretation must therefore be approached with caution.

Medical subject headings