Leveraging large language models for metabolic engineering design.

Journal: Trends in biotechnology
Published Date:

Abstract

Establishing efficient cell factories involves a continuous process of trial and error due to metabolic complexity. This complexity makes predicting effective engineering targets a challenging task. Therefore, successful previous designs are vital for future cell factory development. In this study, we developed a method using large language models to extract metabolic engineering strategies from research articles. We created a database containing over 29 006 metabolic engineering entries, 1210 products, and 751 organisms. Using this database, we trained a deep learning model to predict engineering targets for cell factories. Our model outperformed traditional algorithms, demonstrated strong generalization to unseen products and multigene combinations, and was experimentally validated with geraniol overproduction in yeast, leading to the identification of several novel targets. Our study provides a valuable dataset, a chatbot, and an engineering target prediction model for the metabolic engineering field and exemplifies an efficient method for leveraging existing knowledge for future predictions.

Authors

Keywords

No keywords available for this article.