Validating Radiology Artificial Intelligence Model Performance on Photon-Counting CT Images Using Large Language Models for Ground Truth Extraction.

Ng, Yee Seng; Kanani, Mohammed M; King, William E; Miller, Zachary D; Brown, Lauryn; Wishart, Joey; Lindgren-Ruby, Alex; Medverd, Jonathan R et al. · J Am Coll Radiol · 2026

retrospective_cohort · Level III

Where this comes from

Abstract

The aim of this study was to evaluate the feasibility of using large language models (LLMs) to automate ground truth label extraction from radiology reports, enabling scalable assessment and monitoring of radiologic artificial intelligence (AI) tools. The framework was tested by validating AI model performance on a newly installed photon-counting CT (PCCT) scanner. Four FDA-cleared deep learning-based computer-aided detection and triage tools targeting pulmonary embolism, intracranial hemorrhage, cervical spinal fractures, and vertebral compression fractures were retrospectively analyzed. Radiology reports from examinations acquired using the new PCCT scanner and conventional scanners were processed using an LLM (Llama 3.3) to extract binary ground truth labels. AI outputs were compared with these labels to estimate performance metrics. Discrepant cases were adjudicated by three human annotators, with interrater reliability measured using Fleiss's κ test. Performance metrics were recalculated after partial human correction of LLM errors. LLM-extracted labels enabled rapid performance assessment across all four diagnostic tasks. There were no statistically significant differences in performance between the PCCT and non-PCCT cohorts. In discrepant cases, the agreement between LLM labels and final human annotations (κ = 0.731) was comparable with interreader agreement (κ = 0.720), supporting the reliability of LLM labeling. LLMs can be used to automate ground truth label extraction from radiology reports, offering a scalable and efficient alternative to manual annotation. This method supports rapid local validation of AI tools, even in response to input drift from new imaging hardware.

Medical subject headings