Evaluation of Artificial Intelligence Chatbots in Providing Brucellosis-Related Health Information: A Multidimensional Quality Assessment.
Journal:
Zoonoses and public health
Published Date:
Apr 10, 2026
Abstract
INTRODUCTION: AI-based chatbots are increasingly used in accessing health information. However, there are significant differences in the accuracy, transparency of sources, readability and reliability of the information provided by these systems. In infectious diseases with heterogeneous clinical courses, requiring long-term follow-up and where patient information is critical, such as brucellosis, the quality of digital information sources is of particular importance. This study aims to compare the multidimensional performance of different AI-based chatbots in the delivery of health information related to brucellosis. METHODS: Eight chatbots (ChatGPT-4o, Gemini 2.5 Pro, Claude 3.5 Sonnet, Microsoft Copilot, Perplexity AI, Grok-1.5, Mistral Le Chat and DeepSeek) were evaluated using standardized clinical questions. Clinical accuracy, source transparency, readability/patient-friendliness, ethical safety and perceived trust level were analyzed using the QUEST, DISCERN, JAMA criteria, PEMAT-P, Ateşman Readability Index and Trust/Confidence scales. Scores were normalized to create a heat map. Subgroup analyses were also conducted. RESULTS: Significant performance differences were found among the chatbots. DeepSeek achieved the highest scores in clinical accuracy, structured information presentation and source transparency. Claude stood out with higher perceived trust in addition to accuracy and transparency. It was observed that readability and perceived trust do not always coincide, and response length alone is not an indicator of quality. CONCLUSION: This study demonstrates that AI-based chatbots should not be evaluated in a clinical context using a single 'best practice' approach. For public health-critical diseases such as brucellosis, selective and controlled chatbot integration tailored to the intended use may offer a safer and more effective approach.
Authors
Keywords
No keywords available for this article.