Physics-based machine learning for enhanced drug formulation development.

Journal: Journal of controlled release : official journal of the Controlled Release Society
Published Date:

Abstract

Formulation design is constrained by scarce and heterogeneous experimental data, which limits the accuracy and generalizability of conventional AI models. Here, we introduce a physics-based machine learning (PBML) approach that integrates physics-based modeling with data-driven learning to improve drug formulation development. Our approach predicts key formulation properties across two different systems, including physical stability of amorphous solid dispersions (ASDs) and molecular hygroscopicity. For ASDs, molecular dynamics (MD)-derived descriptors that explicitly encode non-covalent interactions (drug-polymer interaction energies, hydrogen-bond networks) and mobility (diffusion coefficients) markedly outperform empirical experimental parameters on the same dataset, improving generalization to unseen APIs under grouped cross-validation (75.2% vs. 66.1%). For hygroscopicity, hygroscopic/nonhygroscopic labels were first assigned based on MD simulation results. A classification model was then trained using these MD-derived labels and validated to have strong performance on external experimental datasets (accuracy = 0.967; F1-score = 0.957; AUC-ROC = 0.931). Across both systems, the synergy between MD-derived descriptors and Tabular Prior-data Fitted Network (TabPFN) performs well on limited formulation datasets. In addition, SHapley Additive exPlanations (SHAP) analyses align feature importance with known experimental mechanisms. For instance, stronger drug-polymer attraction and lower API mobility stabilize ASDs, while higher surface polarity and electrostatic potential (ESP) variance drive hygroscopicity, thereby improving interpretability at the representation level. Overall, our PBML framework provides a data-efficient and mechanism-grounded approach that can enhance decision-making in formulation design. This approach has the potential to extend the utility of limited datasets in drug formulation design and to reduce experimental burden.

Authors

Keywords

No keywords available for this article.