Multi-center benchmarking of large language models for clinical decision support in lung cancer screening.
cross_sectional · Level IV
Where this comes from
- Record sourced from PubMed, PMID 41274285.
- Also identified by DOI 10.1016/j.xcrm.2025.102465 and PMC identifier 12765833.
- Licence recorded as CC BY-NC-ND.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Large language models (LLMs) are increasingly explored for clinical applications, but their ability to generate management recommendations for lung cancer screening remains uncertain. In this cross-sectional, multi-center study, 148 anonymized low-dose computed tomography (CT) reports from three healthcare institutions are used to assess the readability, accuracy, and consistency of four widely adopted models (GPT-3.5, GPT-4, Claude 3 Sonnet, and Claude 3 Opus). Among them, Claude 3 Opus produces the most readable recommendations, while GPT-4 achieves the highest clinical accuracy. Importantly, performance dose not differ significantly across institutions, underscoring the robustness of these models to variations in reporting templates and their utility in diverse healthcare settings. In an exploratory analysis, two state-of-the-art models, proprietary GPT-4o and its open-source counterpart DeepSeek-R1, show comparable performance to GPT-4, outperforming GPT-3.5. These findings highlight the potential role of LLMs to enhance clinical decision support in lung cancer screening across diverse healthcare settings.
Medical subject headings
- Lung Neoplasms
- Early Detection of Cancer
- Benchmarking
- Decision Support Systems, Clinical
- Language