Biological rationales from language models enable leakage-resistant forecasts of target-indication success

Journal: bioRxiv
Published Date:

Abstract

Identifying which target-indication (T-I) hypotheses can translate into clinical success remains a central challenge in drug discovery. We present PRIORITI (Prospective Rationale-Informed Outcome Reasoning for Integrated Target-indication Intelligence), a leakage-resistant framework that separates upstream biological evidence synthesis from outcome prediction. Given only a target gene and disease indication, a domain-instructed large language model (LLM) synthesizes drug-agnostic human genetics across prespecified channels. These evidence were embedded and combined with static entity context derived from gene summaries and disease definitions. The resulting representation is used to fit a supervised model trained on 6,784 temporally curated historical T-I outcomes, which is one of the largest cohorts assembled for this task. On a held-out historical test set of 1,696 T-I pairs, PRIORITI achieved a ROC-AUC of 0.883 and an expected calibration error (ECE) of 0.023, outperforming both a static target and indication context baseline and a human-curated genetic-evidence reference. On a strictly post-cutoff out-of-time benchmark of 765 T-I pairs, the largest such evaluation cohort reported to date, performance remained robust, with ROC-AUC 0.817, PR-AUC 0.357 and ECE 0.044, exceeding static target and indication context baseline and direct zero-shot LLM forecasting. A label-blinded LLM-based explainer further converted locked predictions into auditable prioritization rationales. Together, PRIORITI provides a calibrated, biology-first probability together with decision-oriented rationales for T-I prioritization before downstream translational investment.

Authors

  • Zhang
  • W.; Xu
  • J.; Zhang
  • Z.; Liu
  • M.; Sun
  • R.; Al-lazikani
  • B.; Shen
  • X.; Kopetz
  • S.; Wu
  • L.; Zhao
  • B.; Wu
  • C.

Categories