Accuracy, Completeness, and Clarity of an AI-Based Chatbot for the EAU Neuro-Urology Guidelines.

Journal: European urology focus
Published Date:

Abstract

This study aimed to externally validate the performance of the European Association of Urology (EAU) Guidelines Bot in neuro-urology by assessing the accuracy, completeness, and clarity of chatbot-generated answers to guideline-based questions and to compare its performance with that of a general-purpose large language model (ChatGPT 5.5). A cross-sectional validation study was conducted using 47 questions derived from the EAU Neuro-Urology Guidelines. Each question was linked to a specific recommendation and classified by recommendation strength (strong vs weak). Questions were independently submitted to both the EAU Guidelines Bot and ChatGPT 5.5 without additional prompting. Two expert urologists independently evaluated each response for accuracy, completeness, and clarity using a five-point Likert scale; discrepancies were resolved by a third reviewer. Overall, 45 questions (95.7%) were linked to strong recommendations and two (4.3%) to weak recommendations. The EAU Guidelines Bot and ChatGPT 5.5 achieved identical mean accuracy scores (4.96 ± 0.20), with all responses rated as highly accurate (Likert 4-5). ChatGPT 5.5 indicated significantly higher completeness scores than did the EAU Guidelines Bot (4.74 ± 0.44 vs 4.57 ± 0.54; p = 0.011), whereas clarity scores were not significantly different (4.83 ± 0.38 vs 4.77 ± 0.43; p = 0.083). High-quality completeness was observed in 46/47 EAU Guidelines Bot responses (97.9%) and 47/47 ChatGPT responses (100%). Score discrepancies between systems were identified in ten of 47 questions (21.3%) and were limited to completeness and clarity domains. Performance remained uniformly high across recommendation grades, with no meaningful differences observed. The EAU Guidelines Bot showed excellent accuracy, completeness, and clarity when applied to neuro-urology guideline-based questions. Its performance was comparable to that of ChatGPT 5.5, with both systems providing highly accurate guideline-concordant responses. Although ChatGPT 5.5 generated more comprehensive answers, the EAU Guidelines Bot maintained closer adherence to the original guideline recommendations. Although not a substitute for clinical judgment, the tool appears to be a reliable adjunct for rapid access to evidence-based neuro-urological guidance.

Authors

Keywords

No keywords available for this article.