Comparative evaluation of large language models and clinicians in real-world glaucoma clinical reasoning.
Journal:
Graefe's archive for clinical and experimental ophthalmology = Albrecht von Graefes Archiv fur klinische und experimentelle Ophthalmologie
Published Date:
Aug 6, 2026
Abstract
PURPOSE: Clinical decision-making in glaucoma is complex and requires integration of heterogeneous information, including patient history, examination findings, and risk stratification. While artificial intelligence (AI) has shown strong performance in image-based ophthalmic tasks, its capability in specialty-specific clinical reasoning remains insufficiently explored. METHODS: Performance was evaluated by glaucoma specialists using a predefined rubric across three clinically oriented domains: medical accuracy (40%), key-point recall (30%), and logical completeness (30%). The weighted composite score was used as a descriptive summary of case-based reasoning quality. RESULTS: AI models showed structured clinical reasoning performance in this case-based dataset, with weighted mean scores overlapping with those of attending ophthalmologists and exceeding those of some lower-performing trainees. These findings should be interpreted as exploratory performance patterns rather than evidence of equivalence. Inter-individual variability was substantial among human clinicians, particularly residents. AI systems often included safety-critical diagnostic and management elements, while the best-performing human clinician achieved the highest individual score overall. CONCLUSION: In this limited 34-case evaluation, large language model-based AI systems produced structured glaucoma-related reasoning with performance that overlapped with attending ophthalmologists but did not establish clinical equivalence. These systems require specialist oversight and further validation before clinical use, but may have potential as supervised decision-support and educational tools.
Authors
Keywords
No keywords available for this article.