Comparison of the performance of ChatGPT Plus, Gemini Pro, and experienced nurses in pediatric emergency triage using physician reference assessment: A prospective observational study.
Journal:
Journal of pediatric nursing
Published Date:
Jun 10, 2026
Abstract
AIM: We aimed to evaluate the performance of ChatGPT Plus and Gemini Pro, two large language model-based artificial intelligence (AI) tools, to predict actual triage levels of patients in the pediatric emergency department and compare them with the decisions of experienced nurses. DESIGN AND METHODS: This single-center, prospective observational study was conducted between September 15 and October 15, 2025. Nurses received standardized refresher training on the Emergency Severity Index (ESI) before data collection. A triage nurse assessed each patient while a pediatric emergency medicine physician simultaneously determined the gold-standard triage level; AI triage levels were recorded by a separate researcher. Patients were classified as emergent (ESI-1 and 2), urgent (ESI-3), or non-urgent (ESI-4 and 5). Agreement was evaluated using the Quadratic Weighted Kappa (κ_w) coefficient. RESULTS: A total of 378 patients were included. While near-perfect agreement was observed between the physician and nurses, agreement for ChatGPT Plus and Gemini Pro was moderate (κ_w 0.892, 0.586, and 0.544, respectively). The nurses had a higher triage accuracy than AI tools (92% compared to 65% for ChatGPT Plus and 63% for Gemini Pro, p < 0.001). AI tools showed a tendency toward under-triage of emergent patients and over-triage of both urgent and non-urgent patients (p = 0.038 and p < 0.001). CONCLUSION: These findings demonstrate that general AI tools are not yet sufficiently accurate to classify pediatric cases and cannot replace experienced nurses. IMPLICATIONS TO PRACTICE: Future hybrid AI models incorporating real-time clinical observations and age-specific physiological parameters may enhance applicability of these technologies in pediatric emergency triage.
Authors
Keywords
No keywords available for this article.