Cross specialty evaluation of large language model responses on intentional replantation.

Journal: Journal of stomatology, oral and maxillofacial surgery
Published Date:

Abstract

BACKGROUND: The expanding use of LLMs (Large Language Models) for rapid, practical information access has accelerated their integration into medicine and dentistry. Yet inaccuracies and fabricated citations raise concerns about reliability in clinical contexts. This study compares ChatGPT 4.0 and OpenEvidence on questions regarding intentional replantation and examines how endodontists & oral and maxillofacial surgeons evaluate the perceived accuracy and completeness of their responses. METHODS: Twenty clinically oriented questions addressing both the endodontic and surgical dimensions of intentional replantation were generated using the "alsoasked.com" tool and subsequently reviewed by specialists in endodontics and oral and maxillofacial surgery. Each question was submitted individually to both models. ChatGPT 4.0 was provided with a specialized prompt to emulate expert-level use, whereas OpenEvidence was queried without additional guidance. The resulting 40 responses were evaluated for FKGL (Flesch-Kincaid Grade Level) readability and textual similarity. Perceived accuracy and completeness were analyzed using linear mixed-effects models, with LLM type and evaluator specialty as fixed effects and rater identity as a random effect. RESULTS: Both LLMs produced responses with advanced readability, and no significant difference in FKGL scores was observed (p = 0.092). Textual similarity remained low (0-9%), reflecting a high degree of originality across outputs. The LMM analyses demonstrated that LLM type had a significant main effect on both perceived accuracy and completeness (p < 0.001), with OpenEvidence receiving higher scores than ChatGPT 4.0 for both outcomes. The evaluator's clinical specialty did not show a significant independent main effect. However, a significant interaction between LLM type and clinical specialty was observed for perceived accuracy, whereas this interaction did not reach statistical significance for completeness. CONCLUSIONS: OpenEvidence's higher perceived accuracy and completeness suggest domain-specific, structured LLMs may better generate high-quality clinical information. This study also provides a rare cross-specialty comparison of evaluations by endodontists and oral-maxillofacial surgeons.

Authors

Keywords

No keywords available for this article.