Artificial intelligence-driven cross-endpoint stacked ensemble learning for separate but meta-feature coupled prediction of six human pharmacokinetic endpoints with RDKit molecular representation and applicability domain control on real TDC benchmark data.
Journal:
Naunyn-Schmiedeberg's archives of pharmacology
Published Date:
Aug 7, 2026
Abstract
Machine learning prediction of human pharmacokinetic (PK) properties usually trains one endpoint at a time, discarding the correlations linking absorption, distribution, metabolism and excretion (ADME) parameters. We developed an artificial-intelligence cross-endpoint stacked ensemble predicting six human PK endpoints within a single modeling framework, coupled through cross-target meta-features, with permutation interpretability and applicability-domain control. The six benchmark sets are independently curated and not compound aligned, so prediction is separate per endpoint, not joint across matched molecules. Six Therapeutics Data Commons (TDC) datasets are as follows: human intestinal absorption (HIA, n = 578), Caco-2 permeability (n = 906), volume of distribution (VDss, n = 1111), plasma protein binding (PPBR, n = 1614), hepatocyte clearance (CL, n = 1020) and half-life (t½, n = 665). Compounds were represented by 34 RDKit descriptors. Per endpoint, four base learners produced leakage-free out-of-fold predictions; pooled across all six endpoints, these formed a 24-dimensional cross-target meta-feature matrix feeding a Ridge or logistic meta-learner. Baselines were a same-endpoint stack and single-task random forest on Bemis-Murcko scaffold splits, with bootstrap intervals and paired tests. The ensemble reached HIA AUROC 0.935 (MCC 0.64), Caco-2 MAE 0.386 (R2 0.60), VDss ρ 0.71, PPBR MAE 9.3%, CL ρ 0.34 and t½ ρ 0.49. It improved on single-task learning for Caco-2 (p = 0.002) and half-life (ρ 0.49 versus 0.36, p = 0.020), tied elsewhere and degraded nothing. The half-life gain exceeded the same-endpoint stack (ρ 0.35), isolating a cross-target effect. Cross-target stacking matches or exceeds single-task learning across six PK endpoints, with the clearest benefit for half-life.
Authors
Keywords
No keywords available for this article.