When high accuracy misleads: Stability limits of supervised feature importance in QSAR biodegradation.

Journal: Chemosphere
Published Date:

Abstract

Supervised machine learning excels at target prediction but can mischaracterize structure-biodegradability associations when feature importance is treated as ground truth. Using the QSAR Biodegradation dataset (1055 chemicals; 41 descriptors), we compare targeted supervised models (random forest, XGBoost, logistic regression), unsupervised methods (feature agglomeration, highly variable gene selection), and non-targeted supervised approaches (Spearman correlation). We evaluate cross-validated accuracy and ranking stability via a top-10 selection protocol and a leave-top-1-out perturbation. XGBoost attains the highest accuracy (0.8569) yet exhibits ranking instability; random forests are similarly unstable. In contrast, unsupervised and non-targeted supervised methods achieve strong accuracy (≈0.819-0.849) with perfect stability. Results caution against equating high predictive accuracy with reliable feature importance and support stability-aware, label-agnostic selection for interpretable materials science.

Authors

Keywords

No keywords available for this article.