Machine learning-based analysis of risk factors and construction of a predictive model for hyperuricemia in Chinese health examination population.
Journal:
Scientific reports
Published Date:
Jul 21, 2026
Abstract
Hyperuricemia (HUA) imposes a growing public health burden, calling for better risk stratification tools. In this cross-sectional study of 4906 Chinese adults undergoing routine health checks (overall HUA prevalence: 26.0%), we built machine learning-based predictive models using a stratified 80/20 data split. To avoid variable selection bias, we applied LASSO regression with tenfold cross-validation, which identified 12 core predictors from routine clinical and demographic data. Among four algorithms tested, the Gradient Boosting (GB) model showed the best discrimination (AUC = 0.770) and good calibration (slope = 0.982, intercept = - 0.007, Brier score = 0.157). SHAP analysis revealed serum creatinine (Scr), HDL cholesterol (HDL-C), and body mass index (BMI) as the top predictors. Notably, SHAP interaction plots uncovered a nonlinear rise in risk above a Scr threshold and a compounded risk when high Scr coincided with low HDL-C. In summary, the well-calibrated GB model offers a reliable, data-driven tool for HUA risk screening, with insights into marker interactions to guide targeted prevention and early intervention.
Authors
Keywords
No keywords available for this article.