A Comparison of Machine Learning and Human Graders for Glaucoma Diagnosis from Fundus Images for Population Screening.

Journal: Ophthalmology
Published Date:

Abstract

PURPOSE: To compare the accuracy of vertical cup-disc ratios (VCDR), ascertained by machine learning (ML) versus human graders, from fundus images for glaucoma detection. This study utilizes population-based data, with a disease prevalence and case-mix that is closer to a real-world setting than conventional case-control studies, with the aim of developing improved glaucoma screening tests. DESIGN: Cross-sectional analysis of a population-based study. PARTICIPANTS: 6,304 participants of the EPIC-Norfolk Eye Study with color fundus images gradable by humans and ML in both eyes. METHODS: VCDR was independently estimated from two-dimensional fundus images of EPIC-Norfolk Eye Study participants by trained human graders (H-VCDR) and an externally trained, open access, ML model (ML-VCDR). A neural network trained on 81,830 ophthalmologist-labeled images was used to generate pseudo-labels for over 100,000 UK Biobank images, on which ML-VCDR was subsequently trained. Glaucoma status was ascertained by tertiary center specialist examination. Predictive performance of VCDR for glaucoma status was examined using logistic regression. ML-VCDR estimates were additionally compared to a popular open-source ML model (AutoMorph) and scanning laser ophthalmoscopy (Heidelberg Retinal Tomography (HRT)). MAIN OUTCOME MEASURES: Area Under the Receiver Operated Characteristic Curve (AUROC), explained variance (McFadden's pseudo-R2). RESULTS: Of 6,304 participants (mean age 68 years; 57% women), 696 had glaucoma or suspect status in at least one eye. For left eyes, H-VCDR and ML-VCDR explained 17% (95% CI 14.7 - 20.4) and 31% (95% CI 27.9 - 33.9) of glaucoma status variance and had an area under the ROC curve (AUROC) of 79% (95% CI 76.8 - 81.2) and 88% (95% CI 86.3 - 89.1), respectively. Right eye H-VCDR and ML-VCDR explained 20% (95% CI 16.9 - 22.6) and 35% (95% CI 32.4 - 37.9) of the variance, and had an AUROC of 81% (95% CI 78.6 - 82.5) and 90% (95% CI 88.7 - 91.0), respectively. ML-VCDR also performed better than AutoMorph and HRT at predicting glaucoma status from VCDR estimations. CONCLUSIONS: In this population-based setting, ML far outperformed trained human graders at predicting specialist-ascertained glaucoma status from fundus images. This provides promise for ML-supported strategies for glaucoma population screening.

Authors

Keywords

No keywords available for this article.