A structured oral formulation database for machine learning: Uncovering data-informed design strategies to facilitate effective formulation development.

Journal: International journal of pharmaceutics
Published Date:

Abstract

Oral dosage form development still relies heavily on empirical trial-and-error, while the high prevalence of poorly soluble drug candidates increases the need for structured data support. To address this limitation, we constructed the Computational Pharmaceutics Intelligent Manufacturing Database (CPIMD), an ML-oriented database that integrates physicochemical properties of active pharmaceutical ingredients, qualitative excipient compositions, release categories, and in vitro dissolution data from 683 marketed oral dosage forms approved by the PMDA of Japan. A standardized workflow for data cleaning, feature encoding and dissolution profile digitalization was used to transform raw information into a machine learning ready dataset. Using CPIMD, we applied unsupervised clustering to characterize four major formulation pattern clusters defined by drug properties, release categories, and associated excipient combinations. Together, these clusters summarize a formulation pattern matrix linking API physicochemical properties, release objectives, and associated functional excipients within the current dataset. A proof-of-concept random forest model further showed that binary excipient features could predict release type with 97.1% test accuracy, supporting the utility of CPIMD for downstream ML applications. CPIMD provides a structured data foundation for predictive modeling, preliminary excipient screening, and data-informed oral formulation development.

Authors

Keywords

No keywords available for this article.