Comparison of large language models in the management of pediatric ureteropelvic junction obstruction: a comparative analysis of 125 clinical scenarios.

Journal: Pediatric surgery international
Published Date:

Abstract

PURPOSE: To evaluate and compare the clinical accuracy, reliability, comprehensiveness and readability of three prominent Large Language Models (LLMs) (ChatGPT, Gemini, and Copilot) in the management of pediatric ureteropelvic junction obstruction. METHODS: A total of 125 unique clinical scenarios representing various stages and complexities of pediatric ureteropelvic junction obstruction(UPJO) were developed. Responses from ChatGPT (GPT-5.3 version, OpenAI), Gemini(Google), and Copilot GPT 5 (Microsoft) were independently evaluated by two expert pediatric urologists across four domains: Clinical Accuracy, Completeness, Reliability, and Readability (scored 1-5). Initial inter-rater discrepancies were resolved through a consensus-building process. Statistical analysis included Kruskal-Wallis tests for performance comparison and weighted Cohen's kappa for inter-rater agreement. RESULTS: ChatGPT demonstrated significantly higher scores in clinical accuracy (4.13 +/- 1.31) and reliability (4.08 +/- 1.30) compared to its counterparts, showing the closest alignment with international guidelines. Conversely, Copilot achieved the highest readability score (4.13 +/- 0.65) but exhibited a "readability-accuracy paradox," where professional formatting masked frequent clinical inaccuracies (3.49 +/- 1.67). Gemini provided comprehensive content but was hindered by structural deficits and the lowest readability score (2.88 ± 0.83). The expert consensus process successfully improved inter-rater agreement (κ = 0.532) to (κ = 0.586). CONCLUSION: While LLMs show promise as decision-support tools in pediatric surgery, their performance is inconsistent. ChatGPT is currently the most robust model for guideline-based management of UPJO. However, the "deceptive confidence" of models like Copilot poses a risk of misinformation. Future integration should explore multimodal capabilities, including the analysis of imaging and ongoing validation against standardized reporting frameworks.

Authors

Keywords

No keywords available for this article.