Impact of reference standard quality on the diagnostic accuracy of AI tools for MCI: A systematic review and meta-analysis.

Journal: Archives of gerontology and geriatrics
Published Date:

Abstract

BACKGROUND: Artificial intelligence (AI)-based tools have shown promise for identifying mild cognitive impairment (MCI), but substantial heterogeneity limits interpretation of pooled diagnostic performance. One potentially important but underexplored source of heterogeneity is variation in reference standard quality. OBJECTIVE: To evaluate the diagnostic accuracy of AI-based tools for MCI and explore whether reference standard quality may contribute to heterogeneity across studies. METHODS: Following PRISMA-DTA guidelines (PROSPERO: CRD420251106411), we searched eight databases for diagnostic accuracy studies of AI-based tools for MCI. Study quality was assessed using QUADAS-2. Pooled estimates were generated using a bivariate random-effects model, with subgroup analyses used to explore potential sources of heterogeneity. RESULTS: Eight studies (11 datasets, N = 8411) were included. The pooled AUC was 0.92 (95% CI: 0.90-0.94), although substantial heterogeneity was observed (Specificity I2 = 85.67%). In subgroup analyses, studies using robust clinical reference standards (High-Quality, k = 7) showed higher diagnostic performance (DOR = 44.71; 95% CI: 21.53-92.85) than studies using proxy measures such as brief cognitive screening tools (Lower-Quality, k = 4; DOR = 25.29; 95% CI: 8.41-76.03). However, this subgroup difference should be interpreted cautiously given the small number of included studies. CONCLUSION: Variation in reference standard quality was associated with differences in observed AI diagnostic performance for MCI and may represent an important source of heterogeneity. AI-based tools showed more favorable performance when validated against robust clinical criteria, although the limited number of studies precludes definitive conclusions.

Authors

Keywords

No keywords available for this article.