Identification and importance analysis of osteoporosis risk factors in perimenopausal women based on the Transformer-LSTM hybrid model and SMOTE.

Journal: Climacteric : the journal of the International Menopause Society
Published Date:

Abstract

OBJECTIVE: This study aimed to identify the core risk factors for osteoporosis in perimenopausal women and propose an analytical framework based on the Transformer-LSTM hybrid model and SMOTE algorithm. METHOD: Cohort data on perimenopausal osteoporosis from the Department of Gynecology, Peking University Third Hospital were preprocessed, including cleaning invalid samples, converting feature types and imputing missing values. The SMOTE algorithm was then used to generate synthetic samples to balance the class distribution. Subsequently, a hybrid model integrating Transformer and long short-term memory (LSTM) was constructed: Transformer was employed to capture long-range correlations between features, while LSTM extracted sequence dependencies, thereby improving classification performance. Random errors were controlled through 20 independent experiments. Feature importance was calculated using the permutation method, and multi-dimensional visualization analyses were conducted using lollipop charts, horizontal bar charts, cumulative contribution curves and radar charts. RESULTS: The results showed that total type 1 collagen N-terminal propeptide (TP1NP) is the most critical physiological indicator affecting abnormal bone mineral density. Factors such as 'educational level', 'body weight (kg)' and 'whether menopause occurred before the age of 45' also exerted significant influences. Notably, the top 25 core features could explain 80% of the associated risk. The proposed Transformer-LSTM hybrid model achieved an overall classification accuracy of 89.52 ± 1.48%, a macro F1-score of 0.867 ± 0.018 and a macro area under the curve (AUC) of 0.923 ± 0.012 - representing improvements of 3.76% in accuracy, 2.5% in F1-score and 3.6% in AUC compared with the single Transformer model, and of 7.38% in accuracy, 6.4% in F1-score and 7.8% in AUC compared with the single LSTM model. For the minority osteoporosis class, the model's recall rate reached 0.879 ± 0.025, which was 11.4% higher than that of the random forest model and 19.0% higher than that of the logistic regression model. CONCLUSION: This study provides a scientific basis and quantitative reference for the early identification, prevention, clinical assessment and health management of osteoporosis. The significant performance improvements of the proposed model (3.76-19.0% across key metrics) validate its superiority in capturing complex feature correlations and identifying minority class samples, addressing the core challenges of imbalanced clinical data and limited feature mining capability of traditional models. The identification of core risk factors and their quantitative importance not only enhances the interpretability of machine learning-based osteoporosis risk assessment but also enables targeted screening and personalized intervention for perimenopausal women, which is of great significance for improving the early diagnosis rate of osteoporosis and reducing the social and economic burden of related fractures.

Authors

Keywords

No keywords available for this article.