Real-world diagnostic and management decisions in rhinology by a large language model: a retrospective study.

Journal: European archives of oto-rhino-laryngology : official journal of the European Federation of Oto-Rhino-Laryngological Societies (EUFOS) : affiliated with the German Society for Oto-Rhino-Laryngology - Head and Neck Surgery
Published Date:

Abstract

BACKGROUND: Large language models (LLMs) have shown promising performance in medical knowledge assessment; however, their capacity to assist real-world clinical decision-making in rhinology, particularly when integrating clinical, endoscopic, and radiologic data remains insufficiently investigated. OBJECTIVE: To evaluate the diagnostic classification and management recommendations of a large language model using real-world rhinologic cases that incorporate comprehensive clinical and written CT reports findings. METHODS: This retrospective, single-center diagnostic accuracy study included 301 adults with sinonasal disease evaluated at a tertiary care center. Structured clinical cases encompassing symptoms, endoscopic findings, and written CT reports findings were entered into ChatGPT-4 and compared with physician diagnoses and management plans serving as the reference standard. Diagnostic performance was assessed using sensitivity, specificity, accuracy, area under the curve, and Cohen's kappa coefficient. RESULTS: Diagnostic performance varied by disease phenotype, with a marked discrepancy between CRSwNP and CRSsNP. The model demonstrated high sensitivity and accuracy for chronic rhinosinusitis with nasal polyps (sensitivity, 95.3%; accuracy, 89.0%), whereas performance for chronic rhinosinusitis without nasal polyps showed lower sensitivity (58.3%) but high specificity (97.7%). This clinically important discrepancy indicates that the model performs more reliably in phenotypes with distinct endoscopic and radiologic features than in those requiring more nuanced clinical interpretation. In contrast, management recommendations demonstrated consistently high accuracy for both medical and surgical decisions (overall accuracy, 92.7%), with substantial agreement between AI-generated and physician recommendations. CONCLUSION: When evaluated using real-world rhinologic cases integrating clinical, endoscopic, and written CT reports findings, a large language model demonstrated strong diagnostic agreement for CRSwNP and consistently reliable management recommendations. However, further validation is required before routine clinical implementation.

Authors

Keywords

No keywords available for this article.