Performance of nine large language models on real-world polysomnography interpretation for obstructive sleep apnea: a multi-dimensional comparative analysis.
Journal:
European archives of oto-rhino-laryngology : official journal of the European Federation of Oto-Rhino-Laryngological Societies (EUFOS) : affiliated with the German Society for Oto-Rhino-Laryngology - Head and Neck Surgery
Published Date:
Aug 7, 2026
Abstract
PURPOSE: The diagnostic and therapeutic capabilities of large language models in interpreting real-world polysomnography reports for obstructive sleep apnea remain insufficiently characterized, with prior studies limited by small samples, single-model designs, and unidimensional outcome measures. This study compared the clinical performance of nine contemporary large language models in diagnosing obstructive sleep apnea and generating treatment recommendations from polysomnography reports, benchmarked against expert consensus. METHODS: Two hundred twelve polysomnography records from a university sleep laboratory were retrospectively classified as simple (nā=ā153) or complex (nā=ā59) based on AASM criteria. De-identified reports, originally in Turkish, were submitted to nine large language models using an English prompt within a one-week frozen evaluation window. Model outputs were independently scored by two sleep medicine specialists using a four-dimensional rubric encompassing diagnostic accuracy, recommendation quality, safety, and parameter coverage. Generalized estimating equations accounted for within-case clustering across 1,908 model-case assessments. RESULTS: Fully correct diagnostic rates ranged from 75.0% to 86.3%, with Claude 4.1 Opus, ChatGPT-5, and Gemini 2.5 Pro forming a statistically indistinguishable top tier. All models demonstrated significant performance degradation on complex cases (11.8-25.7 percentage point decline). Safety rates exceeded 88% across all models. Moderate obstructive sleep apnea was the most challenging diagnostic category. Hypoxemia disproportionately impaired diagnostic accuracy in lower-ranked models. CONCLUSION: Contemporary large language models demonstrate promising yet imperfect capacity for real-world polysomnography report interpretation in obstructive sleep apnea, with performance varying by model and case complexity. These findings support selected models as potential adjunctive decision-support tools under specialist oversight.
Authors
Keywords
No keywords available for this article.