Integrating Developmental Neurotoxicity Prediction with Adverse Outcome Pathway Reasoning

Journal: bioRxiv
Published Date:

Abstract

Experimental assessment of developmental neurotoxicity (DNT) is costly, time-consuming, and difficult to scale across large chemical inventories. New approach methodologies (NAMs), including computational approaches based on machine learning, deep learning, and large language models (LLMs), offer new ways to screen and prioritize chemicals. However, these emerging computational approaches have largely focused on predicting toxicity endpoints, with less attention to the mechanisms underlying those predictions. For DNT assessment, understanding this mechanistic basis is equally important. In this study, we integrated LLM inference with the adverse outcome pathway (AOP) framework to connect DNT predictions with plausible mechanisms. We constructed an AOP-informed benchmark linking ChEMBL target-activity measurements for 5,861 compounds to 14 AOPs representing five neurodevelopmental outcomes and evaluated ten LLMs for the prediction of molecular initiating events (MIEs), key events (KEs), key event relationships (KERs), and adverse outcomes (AOs), as well as the generation of mechanistic explanations. Explanation quality was evaluated using both a naive LLM judge and a rubric-based LLM judge. Across the zero-shot models, the highest observed macro-F1 scores were 0.818 for MIEs and 0.463 for AOs, but only 0.056 for KEs and 0.034 for KERs, identifying downstream pathway reconstruction as the principal challenge. Few-shot prompting substantially improved performance, with Gemini 3.8 Flash under four-shot prompting achieving macro-F1 scores of 0.956 for AOs, 0.813 for MIEs, 0.492 for KEs, and 0.457 for KERs. Both LLM-judge configurations indicated improved mechanistic explanation quality with few-shot examples, although strong outcome prediction did not consistently ensure recovery of the recorded pathway. Our results demonstrate the potential of AOP-informed LLMs to support DNT assessment by linking predicted neurodevelopmental outcomes to plausible biological mechanisms and enabling predictions to be evaluated at both the outcome and pathway levels.

Authors

  • Fan
  • Y.; Xiao
  • S.; Gao
  • F.