Performance of 5 Large Language Models in Perioperative Consultation for Pediatric Hypospadias: Cross-Sectional Comparative Study.
Journal:
Journal of medical Internet research
Published Date:
Jul 29, 2026
Abstract
BACKGROUND: Hypospadias is a common congenital malformation requiring surgery. Caregivers face substantial perioperative information needs, and large language models (LLMs) offer a potential health education channel, but their performance in pediatric urology and the relation between citation accuracy and clinical content safety lack systematic evaluation. OBJECTIVE: This study aimed to evaluate 5 LLMs (ChatGPT-4o, Gemini-2.5-Pro, OpenEvidence, Zhipu Qingyan, and DeepSeek) for pediatric hypospadias perioperative consultation, and to characterize the dimensions clinicians and caregivers prioritize. METHODS: A noninterventional cross-sectional study was conducted at a tertiary hospital in April 2025. From a 40-item bank, 10 high-priority questions were selected by an independent caregiver screening cohort (N=34, cohort A) and classified into 3 risk levels. Twenty-three pediatric urology experts (6 dimensions) and 36 primary caregivers (cohort B; 4 dimensions) evaluated responses by double-blind forced-ranking (reverse-scored, 5=best). Friedman tests with Kendall W assessed overall differences; paired Wilcoxon tests with Bonferroni correction (adjusted α=.005) and rank-biserial r with Hodges-Lehmann 95% CIs were used post hoc. Reference authenticity was independently verified by 2 reviewers (XH and WH) using a 5-category scheme (V/PV/F/G/NR [V: Verifiable, PV: Partially Verifiable, F: Fabricated, G: Guideline-Based, Nonspecific, and NR: No References]), with consensus after canonical-source reverification (Cohen κ=0.702 preadjudication). An 8-reviewer clinical safety audit (7 senior specialists plus 1 European Association of Urology [EAU]-anchored intermediate-title clinician) applied a 4-level severity scheme (None/Mild/Moderate/Severe). RESULTS: Models differed significantly (caregiver: χ²4=77.5, W=0.538, P<.001; expert: χ²4=62.2, W=0.676, P<.001). Gemini-2.5-Pro ranked first (expert median 5.0, IQR 3.0-5.0; caregiver 4.0, IQR 3.0-5.0). DeepSeek ranked second (4.0 both), with superior Empathy versus ChatGPT-4o (r=-0.343; P<.001). OpenEvidence scored lowest (2.0 both), despite high citation accuracy. Expert-caregiver agreement was strong (Spearman ρ=0.89; P=.04). Citation accuracy diverged sharply: OpenEvidence was fully verifiable (V=100%, F=0%), whereas DeepSeek and Zhipu Qingyan showed the highest fabrication (F of 33% and 24%, respectively); Gemini-2.5-Pro fabricated none but used nonspecific guideline citations (G=85%). The safety audit yielded 78 flags, including 9 Severe-level flags across 5 question-model combinations; OpenEvidence carried the largest Severe burden (5 of 9) and the highest severity-weighted score, whereas Gemini-2.5-Pro had the lowest. Bibliographic accuracy and clinical safety were dissociable, and the ranking held under poststratification weighting. CONCLUSIONS: High citation accuracy does not guarantee clinical safety. In the first dual-perspective evaluation, no model was uniformly best. Gemini-2.5-Pro was most comprehensive but relied on nonspecific guidelines. DeepSeek scored highest on caregiver-rated Empathy, yet it had the highest fabrication rate. OpenEvidence produced the most verifiable citations but carried the heaviest Severe-flag burden. These dimension-level priorities, the dissociation between citation quality and safety, and the portable evaluation framework can inform future pediatric medical-artificial intelligence (AI) development. For perioperative use, AI should follow a tiered human-machine collaboration model with mandatory clinician oversight in high-risk scenarios.
Authors
Keywords
No keywords available for this article.