Confidence-Accuracy Alignment in Cardiology Knowledge: Comparing Medical-Specific and General-Purpose Large Language Models Using ACCSAP.
Journal:
The American journal of cardiology
Published Date:
May 7, 2026
Abstract
Large language models (LLMs) are increasingly integrated into healthcare, yet their clinical reliability depends not only on accuracy but also on confidence calibration. General-purpose models have demonstrated strong performance on medical knowledge tasks, while medical-specific models are designed to offer domain alignment. Whether specialization improves clinically meaningful reliability remains unclear. Cardiology, with its complex case-based reasoning, provides a high-stakes test domain. To compare general-purpose and medical-specific LLMs on a standardized cardiology knowledge benchmark, with emphasis on diagnostic accuracy, confidence calibration, uncertainty, and fidelity. A total of 365 text-based questions from the Adult Clinical Cardiology Self-Assessment Program (ACCSAP) were evaluated after exclusion of image-dependent items. ChatGPT-4o and Gemini 2.5 Pro represented general-purpose models, while MedGemma 27B served as a medically fine-tuned comparator. Models received a standardized structured prompt eliciting stepwise reasoning, final answer selection, and self-reported confidence, uncertainty, and fidelity. Statistical comparisons included chi-square testing, nonparametric analyses, correlation coefficients, and Brier scores for calibration. Accuracy differed significantly by model (chi-square = 58.26, p < 0.001): Gemini achieved 87% (95% CI 84% to 91%), ChatGPT 85% (81% to 89%), and MedGemma 67% (62% to 71%). All models reported high confidence, but calibration was modest. Mean confidence differed only slightly between correct and incorrect responses (absolute differences <3%). Brier scores indicated imperfect calibration (Gemini 0.115, ChatGPT 0.137, MedGemma 0.262). ChatGPT demonstrated the strongest confidence-accuracy correlation (r = 0.80, p = 0.005), while Gemini and MedGemma showed weak or nonsignificant alignment. MedGemma exhibited higher uncertainty and lower fidelity across categories. Performance varied by subspecialty, with generalist models outperforming in integrative domains. General-purpose LLMs outperformed a medical-specific model on text-based cardiology assessment, suggesting that large-scale general training may confer advantages in complex clinical reasoning. However, all models showed clinically limited confidence calibration, indicating that self-reported certainty is an unreliable indicator of correctness. Until uncertainty estimation improves, LLM use in cardiology should remain supportive and clinician-supervised.
Authors
Keywords
No keywords available for this article.