Symbolic and domain-generalized machine learning for interpretable solubility modeling in supercritical CO₂.
Journal:
Scientific reports
Published Date:
Jul 19, 2026
Abstract
Accurate prediction of drug solubility in supercritical CO₂ remains challenging due to the limited generalizability of compound-specific correlations and the black-box nature of most machine learning models. This study proposes a domain-aware symbolic regression framework that discovers closed-form analytical expressions for ln(y), enabling interpretable and transferable solubility modeling across chemically diverse pharmaceutical compounds. A leave-one-drug-out (LODO) validation strategy is employed to rigorously assess extrapolation performance on unseen drug domains across a curated dataset of 196 experimental data points spanning 9 antihypertensive compounds. Under this evaluation, the domain-adversarial neural network (DANN) component of the proposed framework achieved RMSE = 0.33 ± 0.08, MAE = 0.24 ± 0.06, and R² = 0.89 ± 0.05, outperforming standard machine learning baselines including Random Forest, XGBoost, and Multi-Layer Perceptron in cross-domain generalization. The symbolic regression component additionally recovered compact closed-form analytical expressions that accurately represent the full experimental dataset (in-sample expression fit: R² = 0.962, RMSE = 0.031, MAE = 0.024) while remaining physically interpretable and analytically tractable. The combined framework demonstrates that domain-aware learning and symbolic discovery can jointly address the dual challenge of predictive robustness under domain shift and model interpretability in supercritical pharmaceutical solubility modeling.
Authors
Keywords
No keywords available for this article.