Performance of ChatGPT, Claude, and AMBOSS on the European Board of Urology In-Service Assessment and Alignment With the European Association of Urology 2025 Guidelines: Comparative Study.
Journal:
JMIR formative research
Published Date:
Sep 2, 2026
Abstract
BACKGROUND: Recent advances in AI, particularly large language models, have generated growing interest in their application to medical education and examination preparation. However, the accuracy, reasoning quality, and adherence to clinical guidelines of these tools in postgraduate urology assessments remain unclear. OBJECTIVE: This study aimed to evaluate the performance of 3 AI tools, ChatGPT (GPT-4.0), Claude (version 4.5), and AMBOSS, on European Board of Urology (EBU)-style multiple-choice questions, with a particular focus on accuracy, insight, concordance, and adherence to European Association of Urology (EAU) guidelines. METHODS: A total of 200 single-best-answer questions from the EBU In-Service Assessment workbook (2021-2022) were input into each AI model. Models were prompted to select an answer and provide an explanation. Two urologists with post-Fellowship of the Royal College of Surgeons (FRCS) training independently assessed the outputs. Accuracy was defined as correct answer selection. Concordance was defined as the logical alignment between the answer and its explanation. Insight was evaluated across 3 domains-nonobvious deduction, discriminative reasoning, and clinical validity-and was graded as low, moderate, or high. RESULTS: ChatGPT demonstrated the highest accuracy (171/200, 85.5%), compared to Claude and AMBOSS (both 159/200, 79.5%; P=.14). Concordance was also significantly higher for ChatGPT (190/200, 95%) than for Claude (176/200, 88%) and AMBOSS (152/200, 76%; P<.001). Nonobvious deduction was predominantly low to moderate across all models, reflecting the recall-based nature of many questions. ChatGPT and Claude showed stronger discriminative reasoning, while AMBOSS demonstrated limited exclusion of alternative options. Clinical validity was high overall, with ChatGPT showing the greatest consistency with EAU guidelines. There was substantial agreement between the 2 reviewers (weighted κ coefficient >0.61). CONCLUSIONS: AI tools can achieve high accuracy on EBU-style assessments; however, differences in reasoning quality and guideline adherence are evident. ChatGPT demonstrated superior performance across all evaluated domains, supporting its role as a potential adjunct in postgraduate urology education.
Authors
Keywords
No keywords available for this article.