Semantic Clinical Knowledge Augmentation Improves Medical Question Answering Across Commercial and Open Large Language Models.

Journal: IEEE journal of biomedical and health informatics
Published Date:
(1)

Abstract

OBJECTIVE: To evaluate whether semantic clinical knowledge augmentation through Semantic Clinical Artificial Intelligence (SCAI) improves medical question answering across commercial and open large language models, and to examine how a reasoning distilled model performs on the same benchmark. MATERIALS AND METHODS: We evaluated Google Gemini, Microsoft Copilot, Meta Llama 3 70B, and DeepSeek-R1-Distill-Llama-70B on official text only United States Medical Licensing Exam (USMLE) sample questions from Step 1 (n = 87), Step 2 CK (n = 103), and Step 3 (n = 123), with and without SCAI. SCAI comprised HD-NLP parsing, graph and knowledge graph embeddings, a trained SCAI LLM semantic knowledge reasoner, and a semantic knowledge base; the current implementation contained 33,023,902 triples across 246 relation labels. Responses were scored against the answer key with secondary human verification. All models produced a response to every text only item, so incorrect responses were counted as confabulations. RESULTS: Across 313 items, accuracy increased from 90.7% to 97.1% for Gemini, 82.4% to 92.0% for Llama 3 70B, and 55.3% to 81.8% for DeepSeek-R1-Distill; Copilot changed from 93.6% to 94.2%. Within model gains were statistically significant for Gemini, Llama 3 70B, and DeepSeek-R1-Distill on all three steps, but not for Copilot. The best single model result was Llama 3 70B plus SCAI on Step 3 (121/123, 98.4%). Gemini plus SCAI was the numerically strongest commercial system. CONCLUSIONS: SCAI was associated with materially better accuracy and lower confabulation (False Positive Results) across several heterogeneous LLMs, supporting a semantically grounded pipeline in which a dedicated SCAI LLM generates clinical context for a host LLM while suggesting that strong general reasoning performance does not automatically translate into healthcare readiness.

Authors

Keywords

No keywords available for this article.