Comparative analysis of diagnostic prediction and clinical reasoning of large language models in complex endodontic case scenarios.
Journal:
Journal of dentistry
Published Date:
Dec 21, 2025
Abstract
OBJECTIVE: To compare the diagnostic prediction accuracy and clinical reasoning of four state-of-the-art large language models (LLMs) - GPT-4o, Claude 3.7 Sonnet, DeepSeek R1, and Gemini 2.0 Flash - in complex endodontic case analysis using structured prompts. METHODS: Nine complex endodontic case reports from peer-reviewed journals were converted into standardized narratives with concealed diagnoses. Each LLM generated three prioritized (Bayesian-ranked) diagnostic possibilities with corresponding justifications using structured prompts. Two endodontic specialists independently abstracted the gold-standard diagnoses into a standardized reference format and evaluated the LLM outputs using a rubric assessing clinical plausibility, radiographic justification, and terminological accuracy. Statistical analysis included Kruskal-Wallis testing, Bonferroni-adjusted pairwise comparisons, and the Diagnostic Agreement Index. RESULTS: Gemini 2.0 Flash and Claude 3.7 Sonnet generally achieved higher interpretive reasoning quality than DeepSeek R1 across all diagnostic ranks. However, post-hoc Bonferroni-adjusted comparisons revealed that their differences relative to GPT-4o were not statistically significant (adjusted p = 1.000), whereas DeepSeek R1 differed significantly from both Claude and Gemini (adjusted p < 0.05) and approached significance compared with GPT-4o (adjusted p = 0.059). CONCLUSIONS: This study demonstrates significant variation in the interpretive reasoning performance of contemporary LLMs. Claude 3.7 Sonnet and Gemini 2.0 Flash outperformed DeepSeek R1, although their performance did not differ significantly from GPT-4o. CLINICAL SIGNIFICANCE: This study demonstrates that model selection and structured prompting critically influence the reliability of AI-assisted endodontic diagnoses, with Gemini 2.0 and Claude 3.7 showing superior reasoning quality, highlighting the need for specialized training and human oversight for safe clinical integration.
Authors
Keywords
No keywords available for this article.