Determinants of Reader-LLM Interaction in Thoracic Radiology: Impact of Model Confidence and Reader Expertise.

Song, Jiyoung; Jeong, Won Gi; Han, Dae Hee; Ko, Hongseok; Yoon, Soon Ho; Kim, Hyungjin; Naruetook, Phakhanith; Leong, Wai Ling et al. · Radiology · 2026

retrospective_cohort · Level III

Where this comes from

Abstract

Background Large language models (LLMs) provide diagnostic suggestions and rationales, but determinants of successful reader-LLM interaction remain unclear. Purpose To determine how LLM attributes and reader expertise independently and jointly influence the selective integration of diagnostic advice in human-LLM collaboration. Materials and Methods In this retrospective study, 10 readers evaluated 100 chest imaging cases (radiography, CT, MRI, or PET) from the Korean Society of Thoracic Radiology Weekly Case platform (January 2018 to December 2020), unaided (session 1) and randomized to LLM-assisted setups (session 2) at the reader-case level: high accuracy (76% [379 of 500 reader-case pairs]) using OpenAI's GPT-5 or low accuracy (27% [133 of 500]) using OpenAI's GPT-4o (August 2025), providing ranked diagnostic options with rationales against a reference standard established by the case author. The primary outcome was adequate interaction (accepting correct or rejecting incorrect suggestions). Data were analyzed using multivariable generalized estimating equations, adjusted for reader expertise, diagnostic correctness and reader confidence (session 1), model confidence (score assigned to the correct diagnosis), and reference panel-assessed rationale quality. Results A total of 100 patients were included (mean age, 50.0 years ± 16.3 [SD]; 59 male). After multivariable adjustment, model confidence (odds ratio [OR], 3.82 [95% CI: 1.58, 9.25]; <i>P</i> = .003) and reader expertise (OR, 2.06 [95% CI: 1.38, 3.07]; <i>P</i> < .001) were independently associated with adequate interaction, with a weaker effect of confidence among experts (OR, 0.79 [95% CI: 0.67, 0.94]; <i>P</i> = .008). Higher rationale quality reduced rejection of correct suggestions (OR, 0.79 [95% CI: 0.67, 0.93]; <i>P</i> = .005) but increased acceptance of incorrect suggestions (OR, 1.71 [95% CI: 1.47, 1.99]; <i>P</i> < .001). Higher reader expertise (OR, 0.54 [95% CI: 0.41, 0.70]; <i>P</i> < .001) and reader confidence (OR, 0.80 [95% CI: 0.67, 0.94]; <i>P</i> = .007) were protective, reducing acceptance of incorrect suggestions. Conclusion Successful reader-LLM collaboration is associated with model confidence and reader expertise, with rationale quality facilitating correct advice uptake but increasing overreliance on incorrect suggestions. © RSNA, 2026 <i>Supplemental material is available for this article.</i>

Medical subject headings