Diabetic retinopathy screening using artificial intelligence: a comprehensive performance comparison on datasets from diverse populations and imaging modalities.
Journal:
Physiological measurement
Published Date:
Sep 4, 2026
Abstract
Objective
Diabetic retinopathy (DR) is the leading cause of preventable blindness in adults and poses significant challenges in low- and middle-income regions due to limited access to skilled clinicians and diagnostic facilities. Automated screening solutions using artificial intelligence (AI) have emerged as an efficient alternative, achieving high diagnostic accuracy. However, these solutions are often developed using data from specific populations obtained using relatively expensive high-end devices. This study addresses the potential scope limitations by evaluating the AI-based screening tool retina.help across eight datasets representing diverse populations, imaging modalities, and geographic regions.
Approach
The datasets include both public and private sources, with images captured using tabletop and handheld fundus cameras. Key performance metrics for detecting binary referable DR - sensitivity, specificity, and area under the receiver operating characteristic curve (AUROC) - were calculated on an image-by-image basis.
Main results
Retina.help demonstrated high accuracy on tabletop images, achieving AUROC values of 0.97 on the BRSET and DeepDRiD datasets. Handheld device performance was more variable, with AUROC ranging from 0.88 (Filipino dataset) to 0.99 (Finnish dataset). Sensitivity declined with increased retinal pigmentation, as evidenced by lower values for datasets from Tanzania (62.2%) and Brazil (76.7%) compared to Finland (89.9%). Images from handheld devices often yielded lower sensitivity due to challenges related to low-contrast images. Nonetheless, retina.help generalized well across diverse datasets, showcasing its robustness.
Significance
The study highlights the impact of imaging equipment, demographics, and image quality on diagnostic performance. These findings underscore the need for benchmarking AI-based DR screening tools using standardized datasets that encompass diverse populations and imaging conditions. Such evaluations can guide the development of equitable, reliable and robust screening solutions.
.
Authors
Keywords
No keywords available for this article.