Assessing the generalisability of foundation models to ultra-wide field retinal imaging for diabetic retinopathy screening in Denmark and Greenland.

Journal: International journal of medical informatics
Published Date:

Abstract

BACKGROUND: Foundation models have shown promising performance in ophthalmology image analysis, but their ability to generalize to unseen imaging types and populations remains unknown. We evaluated the generalizability of ophthalmology foundation models to ultra-wide field (UWF) retinal images for diabetic retinopathy (DR) screening in a Danish and a Greenlandic population. METHODS: Three ophthalmology foundation models (RETFound DINOv2, VisionFM, and EyeCLIP) were fine-tuned and evaluated using 6,374 UWF retinal images from 1,760 participants in Denmark and 6,558 images from 1,146 participants in Greenland. Binary DR classification (normal vs. any retinopathy) was performed under four experimental settings: fine-tuning on the Danish dataset, fine-tuning on the Greenlandic dataset, external validation of Danish-fine-tuned models on the Greenlandic dataset, and sequential fine-tuning from the Danish to the Greenlandic dataset. Model discrimination and calibration were assessed. RESULTS: DR prevalence differs between the Danish and Greenlandic datasets, with 45% and 14% of all images having DR, respectively. When fine-tuned and evaluated within the same population, discrimination was similar in Denmark and Greenland, with RETFound DINOv2 achieving the highest AUROC (0.76 [95% CI: 0.73, 0.78] and 0.76 [0.73, 0.80], respectively). External validation on the Greenlandic dataset showed worse performance across models (AUROC 0.59-0.62). Sequential fine-tuning improved discrimination (AUROC 0.70-0.78). However, calibration remained poor across all settings, with calibration intercepts ranging from -1.69 to 0.37 and slopes from 0.25 to 0.78. CONCLUSION: Foundation models showed limited generalizability when applied to unseen imaging contexts and populations, with disparities in model performance observed across Danish and Greenlandic populations. Local fine-tuning improved discrimination, but did not resolve calibration issues, underscoring the importance of careful calibration evaluation to ensure clinical relevance.

Authors

Keywords

No keywords available for this article.