A closed-loop reinforcement learning framework for rapid compound directed optimization

Journal: bioRxiv
Published Date:

Abstract

Generative artificial intelligence (AI) holds transformative potential for drug discovery, yet existing architectures typically operate in open loops without experimental feedback. Here we introduce rapid compound directed optimization (RCDO), a closed-loop reinforcement learning framework that accelerates the optimization process by bridging dry-lab computation with wet-lab feedback. RCDO couples a three-dimensional structure-guided generative model with a multi-level reward system updated after each design cycle using experimental measurements from all synthesized compounds, including inactive or developability-failed compounds. By continuously aligning the generative model with accumulated wet-lab measurements, RCDO substantially compresses optimization timelines. We evaluated RCDO through retrospective benchmarking against historical optimization trajectories and prospective wet-lab campaigns targeting ROR1, NLRP3, and NSD3. Across prospective evaluations, RCDO rapidly resolved key optimization bottlenecks within two to three design cycles: improving the oral exposure of an ROR1 inhibitor by 40-fold while maintaining antitumor efficacy, reducing CYP2C19 inhibition of an NLRP3 antagonist by 20-fold while preserving inflammasome activity, and boosting the binding affinity of an NSD3 hit by 18-fold. By directly coupling wet-lab feedback to generative learning, RCDO establishes an efficient platform for compound directed optimization, transforming AI-driven drug discovery from static generation into continuous experimental adaptation.

Authors

  • Wang
  • H.; Lu
  • D.; Lyu
  • W.; Xiu
  • S.; Shi
  • C.; Zhou
  • X.; Xi
  • B.; Feng
  • W.; Xiao
  • Y.; Chen
  • Y.; Zhang
  • H.; Li
  • Q.; Huang
  • B.; Liu
  • Z.

Categories