Comparative Assessment of Performance of Domain-Specific and Primed Versus Non-Primed Artificial Intelligence Chatbots for Clinical Decision-Making in Injuries to the Primary Dentition: An Exploratory Pilot Study.
Journal:
Dental traumatology : official publication of International Association for Dental Traumatology
Published Date:
Jul 27, 2026
Abstract
BACKGROUND/AIMS: To compare the performance of a domain-specific dental trauma chatbot with general-purpose artificial intelligence chatbots under primed and non-primed conditions for clinical decision-making in injuries to the primary dentition. METHODS: Fifteen standardized clinical case scenarios were developed for evaluating three chatbot systems: a domain-specific model (Dental Trauma Evo) and two general-purpose models (ChatGPT 5.2 and Perplexity Pro). General-purpose chatbots were evaluated under primed and non-primed conditions, where priming involved providing a summarized guideline document prior to scenario input. Chatbots were required to generate responses addressing diagnosis, immediate management, follow-up intervals, and radiographic recommendations. Two Pediatric dentists independently evaluated responses using a binary scoring system based on International Association of Dental Traumatology (IADT) guidelines. Statistical comparisons were performed using McNemar's test and agreement analysis with kappa statistics. RESULTS: All chatbots demonstrated complete diagnostic accuracy across scenarios. The domain-specific chatbot achieved 100% accuracy in immediate management decisions, while minor inaccuracies were observed among general-purpose models. The most substantial differences were observed in follow-up interval and radiographic recommendations. Priming significantly improved the performance of general-purpose chatbots, with ChatGPT and Perplexity demonstrating marked gains in follow-up interval accuracy and radiographic recommendations. In these domains, primed general-purpose models demonstrated greater agreement with IADT guideline recommendations than the domain-specific chatbot. CONCLUSIONS: Guideline-based priming was associated with improved performance of general-purpose AI chatbots in clinical decision-making for injuries to the primary dentition, particularly for follow-up scheduling and radiographic recommendations. While domain-specific models provide reliable management guidance, contextual exposure to clinical guidelines enables general large language models to demonstrate a high level of agreement with established recommendations in certain aspects of trauma care in primary dentition. These findings highlight the potential role of prompt-guided AI systems as supportive tools in dental traumatology decision-making.
Authors
Keywords
No keywords available for this article.