Hypothesis Generation via Interpretable Machine Learning: A Case Study on Risk Factors for Postradiation Therapy Lung Cancer Recurrence.

Journal: Advances in radiation oncology
Published Date:

Abstract

PURPOSE: Interpretability is highly desirable for oncologic outcome prediction, as it increases the level of transparency and trustworthiness of the model. This model characteristic is particularly relevant in the setting of modest sample size. Existing work has focused on Shapley Additive Explanations to provide post-hoc explanations on black-box models. These models are not intrinsically interpretable. In this study, we investigated the applicability of an intrinsically interpretable glass-box model, Explainable Boosting Machine (EBM), for hypothesis generation from an early-stage lung cancer data set. METHODS AND MATERIALS: We applied EBM to a stripped data set on postradiation therapy lung cancer recurrence, aiming to extract as much information as possible by using pristine EBM configurations. We compared the key features ranked by EBM with those identified through univariate statistical analysis. Additionally, we benchmarked its performance against logistic regression and random forest models, while also evaluating the hypotheses generated by EBM. This study was approved by an institutional review board at the University of Pennsylvania. RESULTS: EBM identified primary tumor size and body mass index as the most prognostic features, aligning with the results of the univariate analysis. Its interpretability provides safeguards against misinterpretation; the model revealed potential age-related bias in this single-arm data set and possible confounding interactions between race and body mass index. EBM yielded competitive performances and more interpretable insights compared with logistic regression and random forest but was not immune from generalizability challenges arising from limited data. CONCLUSIONS: The modest performance prevents EBM from being used as a clinical decision support tool, when applied to limited data. However, its interpretable, glass-box nature makes it useful for hypothesis generation.

Authors

Keywords

No keywords available for this article.