Can ensemble methods improve predictive performance of existing models estimating chronic kidney disease among patients with diabetes?
Journal:
International journal of medical informatics
Published Date:
Mar 21, 2026
Abstract
BACKGROUND: Clinical prediction models often suffer from poor model transportability and/or subgroup performance resulting from using a single data source. We aimed to determine whether ensemble methods can combine multiple existing models to improve predictive performance when compared to component models. METHODS: As a case study, we used electronic medical records from the Canadian Primary Care Sentinel Surveillance Network (CPCSSN) to test ensemble methods for models estimating the risk of developing chronic kidney disease (CKD) among people with diabetes in a cohort of 37,604 individuals. We considered 13 models identified from prior systematic reviews and combined their unique risk estimates using many strategies (e.g., averaging or mixture-of-experts). We assessed discrimination, precision, recall, calibration, net reclassification index, and integrated discrimination improvement. RESULTS: Ensemble methods performed well, but no better than the best performing component model. Among ensemble methods, the averaging or selection process with the best performance weighted the predictions from all component models by their development cohort size (AUROC: 0.827 [95% CI: 0.821 to 0.833]). However, this did not exceed the best performing component model (AUROC: 0.826 [95% CI: 0.820 to 0.832]). Similarly, based on the NRI>0, estimated risks based on the ensemble methods were often worse than the best performing component model. CONCLUSIONS: This study suggests ensemble methods may not improve predictive performance, though further research should confirm these findings. SUMMARY TABLE: Many clinical prediction models exist that predict the same outcome, but commonly suffer poor performance when applied in new settings. Ensemble methods provide a method of combining multiple models developed across diverse settings to potentially improve predictive performance. When applied in primary care electronic medical records, we found that ensemble models based on existing clinical prediction models could match, but did not surpass the performance of the best performing component model. Ensemble methods may not be necessary to combine existing models; rather, the best performing component model can be used.
Authors
Keywords
No keywords available for this article.