A Systematic Evaluation of Cohort Selection Criteria and Their Impact on Machine Learning Model Performance and Demographic Disparities in COVID-19 Outcomes: Cohort Study.

Journal: JMIR formative research
Published Date:

Abstract

BACKGROUND: Cohort selection criteria play a critical role in shaping machine learning (ML) model performance and the equity of clinical outcome predictions across demographic groups. In practice, cohort definitions are often influenced by variable and sometimes inconsistent data processing decisions, which may introduce bias and limit the generalizability of ML models. During the COVID-19 pandemic, rapid cohort construction further increased concerns about transparency and fairness in ML-based analyses. OBJECTIVE: This study aimed to systematically examine how cohort selection and data processing decisions influence ML performance and demographic equity in predicting COVID-19-related in-hospital mortality. METHODS: Using data from the National COVID Cohort Collaborative (N3C), we evaluated 2 sets of cohorts. Set 1 consisted of 16 cohorts derived from 4 primary data processing decisions, including COVID-19 case identification, inpatient inclusion, diagnosis date selection, and admission timestamp availability. Set 2 expanded this design to 64 cohorts by additionally applying provider ID and location identifier filtering. Model performance was assessed using the area under the receiver operating characteristic curve (AUC) across multiple training-testing cohort combinations. Three ML models-logistic regression, random forest, and gradient boosting-were evaluated using 3 analytical approaches: maximum AUC classification, direct AUC regression, and AUC gap analysis. Performance was further examined across demographic subgroups defined by gender, race, and ethnicity. RESULTS: This study analyzed data from the N3C, including patients with a first positive COVID-19 diagnosis between August 1, 2020, and December 31, 2021. Data preprocessing, cohort construction, and model development were completed prior to analysis. Cohort selection decisions had a substantial impact on ML model performance. Admission time inclusion or exclusion emerged as the most influential factor in Set 1 and consistently affected model accuracy across analytical approaches. In Set 2, this decision remained important, while additional criteria, particularly provider ID filtering, also significantly influenced results. The importance of specific decisions varied across models and evaluation strategies. Analyses across demographic subgroups showed that data processing decisions affected predictive performance differently by gender, race, and ethnicity. CONCLUSIONS: Seemingly minor cohort selection and data processing decisions can meaningfully affect both predictive accuracy and demographic equity in ML-based COVID-19 outcome prediction. These findings highlight the risk of bias introduced by differences in cohort definitions and underscore the need for transparent, standardized, and equity-aware cohort selection practices to support fair and reproducible ML research in health care.

Authors

Keywords

No keywords available for this article.