Evidence-based evaluation of large language models in advanced clinical decision-making for removable prosthodontics.

Journal: The Journal of prosthetic dentistry
Published Date:

Abstract

STATEMENT OF PROBLEM: Artificial intelligence is widely used to answer questions about prosthodontic treatments, but responses regarding removable prosthodontics remain limited with gaps. PURPOSE: Large language models (LLMs) have been increasingly used by clinicians and trainees to access clinical information. However, their ability to make accurate, evidence-based decisions in removable prosthodontics remains unclear. This study aimed to evaluate and compare the performance of 5 LLMs in responding to advanced clinical questions related to removable complete and partial dentures. MATERIAL AND METHODS: A curated set of 10 advanced, evidence-based clinical questions covering key domains in removable prosthodontics was developed based on current consensus statements, systematic reviews, and professional guidelines. Each question was posed independently to 5 LLMs using single-turn prompts and default settings. Two board-certified prosthodontists independently evaluated each response using a predefined 10-point rubric assessing scientific accuracy, comprehensiveness, clarity, and clinical relevance, with guideline-based reference answers serving as the standard. Intra-rater reliability was evaluated by repeat scoring after a 4-week interval. Statistical analyses included reliability testing, a linear mixed-effects model, and nonparametric comparisons among models (α=.05). RESULTS: DeepSeek V3 recorded the highest mean scores, averaged across both evaluators and both scoring sessions (8.1/10), followed by ChatGPT-4o (7.2/10), Google Gemini Advanced (6.7/10), Microsoft Copilot (6.3/10), and ChatGPT-4.0 (6.0/10). Statistically significant differences were observed among the models (P<.001). The Cronbach α and intraclass correlation coefficients (ICC) indicated high internal consistency and reliability. CONCLUSIONS: Contemporary LLMs can provide clinically relevant responses to advanced questions in removable prosthodontics, extending into higher-level diagnostic and treatment-planning domains, with DeepSeek V3 demonstrating the highest performance. However, their output should be interpreted with caution and used only as adjunctive decision-support tools under expert clinician oversight.

Authors

Keywords

No keywords available for this article.