Reasoning or reciting? A temporal contamination audit of large language models in clinical medicine.
other · Level V
Where this comes from
- Record sourced from PubMed, PMID 42281378.
- Also identified by DOI 10.1093/jamia/ocag069.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Evaluate whether large language models reason or simply regurgitate training data in clinical diagnosis. We audited 2000 clinical case reports from PubMed Central: 1000 from 2021 to 2022 (within training data) and 1000 from 2025 (after training cutoffs). Five frontier LLMs generated diagnoses evaluated by an independent AI judge validated against physician consensus (n = 10 000 evaluations). Diagnostic accuracy was virtually identical across temporal cohorts (66.8% contaminated vs 66.9% clean), directly contradicting the memorization hypothesis. Lexical similarity was uniformly low (mean ROUGE-L 0.057), and semantic similarity measured by BERTScore showed no memorization signal (F1 0.8182 contaminated vs 0.8195 clean, Δ = +0.0013), confirming that models generate novel reasoning rather than regurgitating training data. This large-scale audit, using both lexical and semantic similarity metrics, provides compelling evidence that LLMs engage in genuine clinical reasoning rather than regurgitating memorized training data. Models demonstrated equivalent accuracy on cases they could not have seen during training, suggesting they have internalized generalizable medical knowledge rather than memorizing specific cases.