Evaluating Large Language Models in Endodontic Irrigants: A Structured Assessment of Accuracy, Readability, Hallucination Subtypes and Clinical Risk.
Journal:
Journal of endodontics
Published Date:
Aug 28, 2026
Abstract
BACKGROUND: Large language models (LLMs) are increasingly used by clinicians and learners for endodontic information, yet their reliability for irrigation-related knowledge remains unclear. This study evaluated the accuracy, readability, hallucination profile, clinical risk, and temporal stability of 4 general-purpose LLMs on endodontic irrigation questions, including false-premise prompts. METHODS: A literature-based reference set was developed for sodium hypochlorite, calcium hypochlorite, ethylenediaminetetraacetic acid, and chlorhexidine. ChatGPT-5.2, Claude Sonnet 4.5, Gemini 3 Pro, and DeepSeek 3.2 were each asked 100 questions comprising 80 factual items and 20 contradiction-seeking items. Responses were independently scored by 2 blinded endodontists for accuracy, hallucination subtype, and clinical risk. Five readability indices were calculated, and model stability was reassessed 10 days later. RESULTS: Inter-rater agreement was almost perfect (weighted κ = 0.92 for accuracy; κ = 0.97 for hallucination). Claude Sonnet 4.5 and Gemini 3 Pro showed the highest accuracy (1.75 and 1.74) and the lowest hallucination rates (7% and 9%). DeepSeek 3.2 showed the lowest accuracy (1.08), the highest hallucination rate (29%), and critical-risk outputs in 16% of responses. Gemini showed the highest test-retest stability (weighted κ = 0.95). Hallucination strongly correlated with clinical risk (ρ = 0.93; P < 0.001). Readability analysis showed a 2-tier pattern: Gemini and DeepSeek produced more accessible text, whereas ChatGPT and Claude generated denser outputs. CONCLUSIONS: LLM performance in endodontic irrigation is model- and irrigant-dependent. Hallucination profiling is clinically relevant, and high initial accuracy does not guarantee temporal stability. LLM outputs should be used only as clinician-verified adjuncts.
Authors
Keywords
No keywords available for this article.