Interpretable machine learning for cardiovascular disease risk prediction in cancer survivors: development and internal validation.
Journal:
Cardio-oncology (London, England)
Published Date:
Jul 20, 2026
Abstract
BACKGROUND: Cardiovascular disease (CVD) is a major concern among cancer survivors. However, the intersection of cancer and CVD has only recently gained broader attention, and substantial evidence gaps remain. This study aimed to identify risk factors associated with incident CVD in cancer survivors and to develop a machine learning model for CVD risk prediction. METHODS: In this retrospective study, we included 2,500 patients receiving systemic antitumor therapy at Jilin Cancer Hospital; 188 incident CVD events were observed. Variables spanning demographic, clinical, tumor- and treatment-related, and laboratory domains were collected. The dataset was randomly split into a training set (75%) and an internal validation set (25%). Feature selection was performed using LASSO regression. Six prediction models were developed, including five machine learning algorithms (GBM, CoxBoost, XGBoost, SVM, and RSF) and a traditional Cox proportional hazards model. Model performance was evaluated using time-dependent receiver operating characteristic curves (AUC), calibration analyses, and decision curve analysis. Model interpretability was assessed using SHapley Additive exPlanations (SHAP). RESULTS: In univariate analyses, 48 variables were associated with CVD risk (P < 0.20). LASSO regression identified 19 predictors for model development. Key predictors included elevated systolic blood pressure, specific cancer types, anthracycline use, and a history of hypertension. The XGBoost model demonstrated the best predictive performance, with an average AUC of 0.666, sensitivity of 69.8%, specificity of 60.7%, and overall accuracy of 61.4% in the internal validation set. The model showed good calibration and yielded a positive net benefit across a range of clinical thresholds. SHAP analysis indicated that cancer type, lack of anti-HER2 therapy, elevated systolic blood pressure, advanced T stage, advanced TNM stage, and higher uric acid levels were the most influential predictors. CONCLUSION: This study developed and internally validated interpretable machine learning models to predict CVD risk among cancer survivors. The models demonstrated good discrimination and calibration and outperformed traditional methods. By enabling individualized risk quantification and providing transparent interpretation of key predictors, this approach offers a practical tool to support personalized surveillance and prevention strategies in cardio-oncology.
Authors
Keywords
No keywords available for this article.