Evaluating large language models for lay summaries of radiology reports using tailored prompting strategies and mixed-method assessment.

Tasneem, Nanziba; van der Pol, Christian B; Zahoor, Ambreen; Juggath, Nitin; McGowan, Kyle; Lokker, Cynthia; Saha, Ashirbani · PLOS Digit Health · 2026

other

Where this comes from

Abstract

Radiology reports are often filled with medical jargon that limits patient understanding. Lay summaries can improve understanding but are time-consuming for healthcare providers to create. The objective of this study is to explore the use of tailored prompts for five Large Language Models (LLMs) in generating lay summaries from radiology reports. Using 100 reports from the publicly available "BioNLP 2023 report summarization" dataset, lay summaries were generated by each LLM, under select prompting styles [Few-Shot (GPT-4), Generated Knowledge (GPT-4o mini, Gemini 1.5 - Pro, Gemini 1.5 - Flash), and Zero-Shot (Llama 3.1)] informed by a pilot work. The summaries were evaluated using a mixed-method framework: subjective assessment (Likert statements) by blinded experts (n = 2 radiology fellows) and Large Reasoning Models (LRMs) [(Gemini 2.5 - Pro (LRM 1); GPT-oss-120b (LRM 2)], and readability metrics (Flesch-Kincaid Grade Level and Flesch Reading Ease). Using percentage agreement of Likert statements, the LLM-prompt combinations' performances were ranked, and Friedman and post-hoc Nemenyi tests were conducted. Gemini 1.5 - Flash and - Pro (generated knowledge) were rated highest by human experts and LRMs for generating actionable lay summaries that require minimal supervision [P < 4.97 × 10-2 (Rater 1); P < 9.03 × 10-21 (Rater 2), P < 6.90 × 10-15 (LRM 1), P < 2.760 × 10-5 (LRM 2). GPT-4 (few-shot) achieved the highest human-rated accuracy (98%), while Gemini 1.5 - Flash (LRM 1-rated: 95%) and Gemini 1.5 - Pro (LRM 2-rated: 91%) ranked first in LRM-rated accuracy. Gemini 1.5 - Pro produced the most accessible summaries (Flesch-Kincaid Grade Level: 7.55 ± 1.38, Flesch Reading Ease: 67.84 ± 7.78). Strong agreement was observed between experts and LRMs [0.96% (LRM 1) and 3.4% (LRM 2) complete disagreement]. Overall, this study highlights Gemini-models with generated knowledge prompts and the potential of LRM evaluators in assessing LLM-generated lay summaries.