Reasoning or reciting? A temporal contamination audit of large language models in clinical medicine.
Journal:
Journal of the American Medical Informatics Association : JAMIA
Published Date:
Jun 11, 2026
Abstract
OBJECTIVE: Evaluate whether large language models reason or simply regurgitate training data in clinical diagnosis. MATERIALS AND METHODS: We audited 2000 clinical case reports from PubMed Central: 1000 from 2021 to 2022 (within training data) and 1000 from 2025 (after training cutoffs). Five frontier LLMs generated diagnoses evaluated by an independent AI judge validated against physician consensus (nā=ā10ā000 evaluations). RESULTS: Diagnostic accuracy was virtually identical across temporal cohorts (66.8% contaminated vs 66.9% clean), directly contradicting the memorization hypothesis. Lexical similarity was uniformly low (mean ROUGE-L 0.057), and semantic similarity measured by BERTScore showed no memorization signal (F1 0.8182 contaminated vs 0.8195 clean, Ī = +0.0013), confirming that models generate novel reasoning rather than regurgitating training data. DISCUSSION: This large-scale audit, using both lexical and semantic similarity metrics, provides compelling evidence that LLMs engage in genuine clinical reasoning rather than regurgitating memorized training data. CONCLUSION: Models demonstrated equivalent accuracy on cases they could not have seen during training, suggesting they have internalized generalizable medical knowledge rather than memorizing specific cases.
Authors
Keywords
No keywords available for this article.