Leveraging large language models for metabolic engineering design.
Journal:
Trends in biotechnology
Published Date:
Apr 23, 2026
Abstract
Establishing efficient cell factories involves a continuous process of trial and error due to metabolic complexity. This complexity makes predicting effective engineering targets a challenging task. Therefore, successful previous designs are vital for future cell factory development. In this study, we developed a method using large language models to extract metabolic engineering strategies from research articles. We created a database containing over 29 006 metabolic engineering entries, 1210 products, and 751 organisms. Using this database, we trained a deep learning model to predict engineering targets for cell factories. Our model outperformed traditional algorithms, demonstrated strong generalization to unseen products and multigene combinations, and was experimentally validated with geraniol overproduction in yeast, leading to the identification of several novel targets. Our study provides a valuable dataset, a chatbot, and an engineering target prediction model for the metabolic engineering field and exemplifies an efficient method for leveraging existing knowledge for future predictions.
Authors
Keywords
No keywords available for this article.