Diagnostic Performance of Machine Learning Models for Predicting Bloodstream Infections in Large Cohorts: A Systematic Review and Meta-analysis.
Journal:
Infectious diseases and therapy
Published Date:
Aug 28, 2026
Abstract
INTRODUCTION: Machine learning (ML) models have been increasingly applied to support the diagnosis of bloodstream infections (BSI), particularly through the analysis of large routinely collected healthcare datasets. However, it remains uncertain whether large sample sizes alone are sufficient to achieve high diagnostic accuracy. We performed a systematic review and meta-analysis to evaluate the diagnostic performance of ML-based models for BSI in large cohorts. METHODS: MEDLINE, Embase, and Web of Science were systematically searched from inception to December 31, 2025 (Prospero registration CRD420251080948). We included observational studies evaluating ML models for BSI diagnosis with ≥ 1000 patients or BSI episodes and reporting sufficient data to reconstruct diagnostic accuracy measures. Pooled sensitivity and specificity were estimated using a bivariate random-effects Reitsma model. Secondary analyses included pooled area under the summary receiver operating characteristic curve (SROC-AUC), predictive values, subgroup analyses, and meta-regression. RESULTS: Twenty-four retrospective studies were included. Most studies evaluated adult hospitalized patients and used routinely collected structured data, including laboratory, vital sign, and administrative variables. Tree-based ensemble algorithms were the most commonly used architectures. The pooled sensitivity and specificity were 76.6% (95% CI 66.1-84.7%) and 84.5% (95% CI 75.4-90.7%), respectively, with an SROC-AUC of 0.87. The pooled prevalence of BSI was 9.5%, corresponding to pooled positive and negative predictive values of 34.2% and 97.2%, respectively. Significant inter-study heterogeneity was observed (I2 = 74%). CONCLUSIONS: ML-based models demonstrated good overall diagnostic performance for predicting BSI, but accuracy remains insufficient to solidly support clinical decision-making at the present time. These findings may suggest that increasing dataset size alone could not overcome limitations related to the limited clinical granularity of routinely extracted structured data.
Authors
Keywords
No keywords available for this article.