Which bot is best? Accuracy and readability of generative artificial intelligence assistant responses on adrenal nodules.

Journal: American journal of surgery
Published Date:

Abstract

INTRODUCTION: Patients increasingly use generative artificial intelligence assistants (chatbots) for medical information. We examined surgeon-perceived accuracy, reliability, and readability of chatbot responses to adrenal nodule queries. METHODS: Six commonly asked adrenal nodule questions were input into five chatbots. Blinded answers were reviewed by 10 endocrine surgeons for correctness and reliability (6-point Likert scale) and content structure (3-point Likert scale). One-way ANOVAs and Tukey adjusted post hoc analyses tested differences. Inter-rater reliability across surgeon ratings was assessed by two-way intraclass correlation coefficient. Reading grade levels were assessed using Lexile Text Analyzer (MetaMetrics, Inc). RESULTS: Correctness, reliability, and jargon use differed significantly (p ≤ 0.05). Perplexity scored highest for correctness, reliability, and thoroughness but used the most jargon. Gemini scored lowest on all domains, but used the least jargon. Mean Lexile reading level was 8th-10th grade. CONCLUSION: Chatbot responses about adrenal nodules vary significantly. More accurate, reliable, and thorough answers contained excess jargon, but all responses exceeded recommended patient education levels.

Authors

Keywords

No keywords available for this article.