Unsupervised clustering approaches for phenotype classification and cardiometabolic risk assessment: A cross-sectional study of NHANES 1999-2006.
Journal:
Annals of epidemiology
Published Date:
Aug 12, 2026
Abstract
PURPOSE: Clear guidelines on using dual-energy X-ray absorptiometry (DXA) are lacking. DXA-derived phenotypes based on whether the person was above or below the median fat- and muscle-mass compared to a reference population were previously constructed. Whether this cutoff can be improved with unsupervised clustering techniques was the study objective. METHODS: Data were from the National Health and Nutrition Examination Survey (1999-2006 cycles, n=5,566; split into 70/30% training and test datasets), a representative U.S. SAMPLE: Phenotypes based on partitioning deciles of fat- and muscle-mass adjusted for age and sex by k-means, and hierarchical clustering were identified. Model fit was assessed using the silhouette and elbow method. Performance of logistic regression models to identify unfavorable cardiometabolic risks was assessed with the area under the receiver operating characteristic curves (ROC-AUC), stratified by sex and incorporating weighting and the complex sampling design. RESULTS: Optimal models were 2-means k-clusters, 4-means k-clusters, and 5 hierarchical clustering phenotypes. ROC-AUCs from 2-means k-clusters (0.52 to 0.63) were the lowest. Performance of the hierarchical clustering and the 4-means k-cluster phenotypes was higher, but not statistically significantly different from the median-split. CONCLUSIONS: While unsupervised clustering methods improved performance, ROC-AUCs were moderate. Future work investigating other health outcomes is needed.
Authors
Keywords
No keywords available for this article.