Multimodal foundation model assistance for differentiating primary open-angle glaucoma in highly myopic eyes.

Journal: Asia-Pacific journal of ophthalmology (Philadelphia, Pa.)
Published Date:

Abstract

PURPOSE: To develop and evaluate a multimodal foundation-model-assisted system for differentiating primary open-angle glaucoma (POAG) from non-glaucoma in highly myopic eyes and assessing whether artificial intelligence (AI) assistance changes ophthalmologist diagnostic performance. DESIGN: Retrospective diagnostic model-development study with internal validation and paired sequential multi-reader evaluation. PARTICIPANTS: Model development used 603 eye-level cases (502 POAG, 101 non-glaucoma); a fixed 150-case set from 149 patients was used for internal test reporting and the paired reader evaluation. METHODS: A RETFound-based multimodal classifier used color fundus photographs, optical coherence tomography (OCT) images, and 24 structured OCT/retinal nerve fiber layer (RNFL) variables. Six ophthalmologists reviewed each case without and then with AI assistance after a washout interval. MAIN OUTCOME MEASURES: Model area under the receiver operating characteristic curve (AUC); reader sensitivity, specificity, accuracy, F1 score, confidence, reading time, diagnostic switching, and inter-reader agreement. RESULTS: The final multimodal model achieved an AUC of 0.911 (95% confidence interval [CI], 0.862-0.950) on the fixed 150-case set. In the sequential paired evaluation, unaided versus AI-assisted mean sensitivity was 68.0% versus 75.3%, accuracy was 74.2% versus 79.6%, F1 score was 0.790 versus 0.844, and mean specificity was 91.3% in both phases; Fleiss kappa was 0.583 versus 0.709. Reader-level diagnostic differences were not significant after multiplicity adjustment and were exploratory. CONCLUSIONS: In this internal retrospective study, the AI-assisted phase showed numerically higher mean sensitivity, accuracy, confidence, and inter-reader agreement than the unaided phase, with unchanged mean specificity. Reader-level diagnostic differences were not statistically significant after multiplicity adjustment and should be considered exploratory. Prospective multicenter validation is required before clinical deployment.

Authors

Keywords

No keywords available for this article.