Critical Care-Specific vs General-Purpose Large Language Models in Emergency Intensive Care Unit Diagnosis: Single-Center Retrospective Paired Comparative Study.

Journal: Journal of medical Internet research
Published Date:

Abstract

BACKGROUND: The emergency intensive care unit (EICU) manages the most critically ill patients, where rapid and accurate diagnosis is essential yet challenging. Diagnostic error rates in this setting are more than twice as high as in general wards, with serious consequences for patient outcomes. Large language models (LLMs) have attracted growing interest as decision-support tools; however, direct comparative evidence between critical care-specialized and general-purpose LLMs across the admission-to-discharge diagnostic workflow remains limited. OBJECTIVE: This study aimed to compare the top-1 diagnostic accuracy of a critical care-specific LLM (Qiyuan 3.0.1) with 3 general-purpose models (GPT‑5.1, DeepSeek V3.1, and Qwen3‑32B) for EICU diseases, providing evidence for intelligent tool selection. METHODS: This single‑center retrospective paired study enrolled 184 consecutive EICU patients (April 2025-March 2026). Two standardized datasets were constructed: an initial dataset (first 24 hours of admission) and a final dataset (complete clinical course). All 4 models received identical zero‑shot prompts and generated diagnoses independently under masked conditions. The gold standard was the consensus diagnosis by 3 senior intensivists (>10 years' EICU experience; Fleiss κ=0.82). The primary end point was final‑stage top-1 accuracy; secondary end points were initial‑stage top-1 accuracy and the number of correctly matched diagnoses among the first 3 outputs at the final stage. Overall comparisons used Cochran Q test, followed by paired McNemar tests with Bonferroni correction; intergroup differences for top-3 counts were assessed by Friedman rank sum test. RESULTS: Final-stage top-1 accuracy varied significantly across models (Cochran Q=20.32; P<.001): Qiyuan 3.0.1 reached 64.1% (118/184), followed by GPT-5.1 (109/184, 59.2%), DeepSeek V3.1 (105/184, 57.1%), and Qwen3-32B (95/184, 51.6%). Corrected pairwise comparisons (α=.0083) confirmed Qiyuan 3.0.1, GPT-5.1, and DeepSeek all outperformed Qwen3-32B significantly (all adjusted P<.008), while no statistical gaps were detected between Qiyuan 3.0.1 and the 2 top-performing general models. Though overall initial-stage accuracy differed significantly (Cochran Q=13.87; P<.001), no pairwise comparisons yielded significant results after correction. All models shared a median of 2 (IQR 1.0-2.0) correct top-3 diagnoses with no intergroup disparity (Friedman χ23=3.34; P=.34). Notably, all models' top-1 accuracy stayed below 70%, and stratified analysis revealed heterogeneous performance across core EICU diseases including sepsis, severe pneumonia, and gastrointestinal bleeding. CONCLUSIONS: Under the specific data conditions of this study, the critical care-specialized Qiyuan 3.0.1 performed comparably to leading general‑purpose LLMs (GPT‑5.1 and DeepSeek V3.1), supporting its potential for further specialty‑oriented exploration. Nevertheless, absolute accuracy below 70% precludes its direct deployment as an independent diagnostic standard. Bridging the gap from preliminary evaluation to clinical translation requires multicenter external validation, prospective human‑machine collaboration trials, and deeper model optimization-including sustained fine‑tuning on critical care corpora, transparent reasoning pathway design, and systematic safety boundary assessment.

Authors

Keywords

No keywords available for this article.