Human-in-the-Loop Large Language Model-Augmented Diagnostic Reasoning in Thoracic Imaging: Impact of Radiologic Expertise.

Song, Jiyoung; Ko, Hongseok; Han, Dae Hee; Hwang, Eui Jin; Cho, Hye Soo; Lee, Ji Young; Jeong, Won Gi; Yoon, Soon Ho et al. · AJR Am J Roentgenol · 2026

retrospective_cohort · Level III

Where this comes from

Abstract

<b>BACKGROUND.</b> Technical and regulatory constraints limit application of large language models (LLMs) for augmenting diagnostic reasoning in radiology. Reader-mediated text-based workflows may provide a practical alternative. <b>OBJECTIVE.</b> The purpose of this study was to evaluate the impact on diagnostic performance of LLM assistance using reader-generated free-text image descriptions and to assess the effect of reader expertise on this LLM-augmented diagnostic workflow. <b>METHODS.</b> This retrospective study included 93 cases (encompassing radiographic, CT, MRI, and PET/CT images) from the Korean Society of Thoracic Radiology quiz platform from January 2014 to December 2017. Five differential diagnoses (the correct diagnosis and four distracters) were assembled for each case. Ten readers (five thoracic radiologists and five radiology residents) independently interpreted cases. In session 1, readers selected the most likely diagnosis and provided a free-text description of key findings. An LLM (Gemini 3.0 Pro) received as input the free-text description-without case images-and generated output that ranked the five differential diagnoses along with explanatory rationales for the top-three options. In session 2, readers were provided the LLM output from their own free-text description and reselected a most likely diagnosis. LLM performance using images-without free-text descriptions-was also assessed. LLM accuracy was determined using top-ranked diagnoses. Reader groups were compared using generalized estimating equations. <b>RESULTS.</b> LLM accuracy was 52.7% when images were inputted and 63.9% when reader-generated descriptions were inputted. LLM accuracy was greater using descriptions generated by thoracic radiologists than by residents (67.3% vs. 60.4%; <i>p</i> < .001). From session 1 to session 2, accuracy increased from 56.3% to 65.6% for thoracic radiologists and from 42.4% to 58.5% for residents, respectively. Improvement in accuracy between sessions was greater for residents than thoracic radiologists (16.1 vs 9.2 percentage points; <i>p</i> = .02). Residents, compared with thoracic radiologists, had a greater rate of accepting LLM-favored diagnoses (73.2% vs 48.9%; <i>p</i> < .001), including a greater rate of switching to an incorrect diagnosis after a misleading LLM output (60.6% vs 32.4%; <i>p</i> = .009). <b>CONCLUSION.</b> The text-based LLM-assisted workflow yielded improved reader accuracy, although this was heavily influenced by reader expertise. <b>CLINICAL IMPACT.</b> The utility of human-in-the-loop workflows arises from dynamic reader-LLM interactions shaped by the expertise of the operator formulating model inputs and critically evaluating model outputs.