Bias, fairness, and equity in artificial intelligence systems used in dental imaging: A systematic review.

Journal: International journal of medical informatics
Published Date:

Abstract

BACKGROUND: Artificial intelligence (AI) is increasingly used in dental imaging for automated interpretation of dental images and as a clinical decision-support system. Although reported diagnostic accuracies are high, limited attention has been paid to bias, fairness, and equity within such AI-enabled dental systems. In this review, bias refers to systematic errors arising from non-representative or imbalanced training or test datasets; fairness refers to consistent and equitable AI model performance across diverse population groups; and equity refers to the provision of comparable diagnostic value across groups that differ in age, sex, ethnicity, dentition stage, or socioeconomic background. AIM AND OBJECTIVES: The aim of this systematic review was to assess original research on the use of AI in dental imaging, particularly with regard to diagnostic accuracy, methodological quality, and reporting on bias, fairness, and equity. The specific objectives were to: (1) assess diagnostic accuracy; (2) examine demographic reporting and subgroup analyses; (3) determine how bias, fairness, and equity are addressed in model development and validation; and (4) identify methodological priorities for more equitable dental imaging AI research. METHODS: A comprehensive literature search was conducted in accordance with PRISMA 2020 across PubMed/MEDLINE, Scopus, Web of Science, and Google Scholar for studies published from January 2010 to March 2025. The review protocol has been registered in PROSPERO (CRD420261336733; registered 10 March 2026). Original research articles applying AI to dental imaging were included. Risk of bias was assessed using an adapted QUADAS-2 tool, and equity-related reporting was evaluated according to demographic description, dataset representativeness, and subgroup performance analysis. Narrative synthesis was undertaken because of heterogeneity in datasets, models, and outcome measures. RESULTS: Ten original studies met the inclusion criteria. All included studies used deep learning models applied to 2D dental imaging modalities such as panoramic radiographs, periapical radiographs, bitewing radiographs, and intraoral photographs; no eligible CBCT-based studies were identified. AI models demonstrated high diagnostic performance for tooth detection and tooth numbering, and encouraging results for caries detection, gingival assessment, plaque detection, and impacted tooth detection. Most studies demonstrated low methodological risk of bias. However, none of the included studies performed demographic subgroup analysis, only two reported limited demographic summaries, and most lacked sufficient reporting to support any meaningful equity assessment. Specifically, 0/10 studies conducted subgroup analysis, 2/10 provided partial demographic reporting, and 8/10 provided no meaningful demographic reporting. CONCLUSION: AI systems used in dental imaging show strong technical capability but lack adequate evaluation of bias, fairness, and equity. Future research should use representative, multi-centre datasets, report demographic characteristics transparently, and incorporate subgroup performance analysis to support fair and equitable clinical deployment. A fair validation procedure should include an independent test set from demographically diverse groups with subgroup-specific performance reporting, and a representative dataset should reflect the characteristics of the intended real-world clinical population.

Authors

Keywords

No keywords available for this article.