Accuracy and Completeness of Four Artificial Intelligence Chatbots in Assisting Diagnosis and Treatment of Post-traumatic Endodontic Sequelae.

Journal: Journal of endodontics
Published Date:

Abstract

INTRODUCTION: This study aimed to evaluate the accuracy and completeness of the answers generated by four freely available AI chatbots regarding the management of endodontic sequelae in traumatized permanent teeth and the regenerative endodontic procedure, analyzing their performance by question type and over time. METHODS: Four AI chatbots (ChatGPT 3.5, Copilot, Perplexity, and DeepSeek) were evaluated with 14 structured questions developed by endodontists, applied at baseline (day 1) and repeated after seven days (day 8). Nine questions involved clinical cases of traumatized immature and mature permanent teeth (diagnosis, treatment, and treatment with prior diagnosis), and five were conceptual about regenerative endodontic procedures (REP). Two blinded endodontists independently evaluated response accuracy (1-6 points) and completeness (1-3 points) based on agreement with the expert benchmark, with results reported as median ± interquartile range. Reproducibility was analyzed with weighted Kappa, while Kruskal-Wallis with Dunn's post hoc and Wilcoxon signed-rank tests were used for comparisons. RESULTS: On baseline, DeepSeek exceeded in diagnosis while Perplexity consistently outperformed others in treatment-related domains. Regarding response accuracy, Perplexity showed higher score values compared to ChatGPT and Microsoft Copilot, and similar scores to DeepSeek. In terms of completeness, no significant differences were observed among the Ais chatbots (p > 0.05). Over time, only DeepSeek demonstrated a significant increase in completeness scores on day 8 (mean ± standard deviation: 1.56 ± 0.8) compared to baseline (mean ± standard deviation: 1.34 ± 0.55), with p < 0.05. Conceptual REP questions achieved high accuracy across chatbots. CONCLUSIONS: DeepSeek showed superior performance in diagnosis, while Perplexity outperformed others in treatment domains. All chatbots performed well on conceptual REP questions. Although chatbots demonstrated moderate-to-high performance in selected domains, none consistently matched the expert benchmark across all evaluated scenarios. Therefore, chatbot-generated information should be interpreted with caution and verified against current evidence-based guidelines before clinical application.

Authors

Keywords

No keywords available for this article.