ChatGPT Versus DeepSeek in Resectability Categorization of Chronic Rhinosinusitis: A Comparison Based on EPOS 2020.
Journal:
World journal of otorhinolaryngology - head and neck surgery
Published Date:
Oct 8, 2026
Abstract
OBJECTIVE: To evaluate and compare the performance of ChatGPT-5.5 (2026) and DeepSeek-V3 (2026) in classifying CRS resectability according to the EPOS 2020 guidelines, employing various prompting strategies. METHODS: In this retrospective study, 150 chronic rhinosinusitis cases were reviewed, and three rhinology experts classified resectability according to the EPOS 2020 criteria. GPT-5.5 and DeepSeek-V3 were evaluated using default-knowledge, in-context, and chain-of-thought (CoT) prompting based on accuracy, sensitivity, specificity, PPV, NPV, and F1 score. Prespecified subgroup analyses were conducted according to nasal polyp status and surgical eligibility to assess the consistency of prompting-strategy effects. ENT resident performance was compared with and without AI assistance. Statistical analyses included chi-square or Fisher's exact, Student's t or Mann-Whitney U, and ANOVA tests. RESULTS: DeepSeek-V3 achieved higher accuracy than GPT-5.5 under default-knowledge (81.3% vs. 72.7%, p < 0.05) and in-context prompting (86.7% vs. 78.0%, p < 0.01). Notably, CoT prompting significantly improved both models, yielding comparable peak accuracies (95.3% vs. 96.0%, p = 1.00). Subgroup analyses revealed consistent effects across GPT-5.5 subgroups, whereas DeepSeek-V3 exhibited heterogeneous response patterns. Furthermore, CoT-assisted ENT residents demonstrated significantly greater accuracy than unassisted peers (GPT-5.5: 97.3% vs. 72.8%, p < 0.001; DeepSeek-V3: 98.0% vs. 72.8%, p < 0.001), alongside a 43.0% reduction in completion time (95% CI: 41.5%-44.4%; p < 0.001). CONCLUSIONS: ChatGPT-5.5 and DeepSeek-V3, when using the CoT prompting strategy, achieved similarly high accuracy in resectability categorization. While AI-generated reports improved surgeons' accuracy and efficiency, these models should be further validated as support tools before widespread clinical implementation.
Authors
Keywords
No keywords available for this article.