Machine learning-driven optimization of specific, compact, and efficient base editors via single-round diversification.
Journal:
Nucleic acids research
Published Date:
Jun 8, 2026
Abstract
Base editing shows great potential in research and clinical applications. Current iterations of the deaminases used to create precise single-nucleotide changes via base editing exhibit undesirable effects, including off-targeting, off-base editing, and bystander editing. Current deaminases are derived from either larger eukaryotic deaminases, which exhibit high levels of Cas-independent DNA targeting, or from evolved variants of the smaller Escherichia coli TadA protein (ecTadA), which exhibits off-base editing. To overcome the limitations inherent to using a single protein sequence for engineering, we diversified newly identified TadA orthologs by DNA shuffling to yield millions of training sequences for measuring base editor efficiency. We trained generative models on the performance data from the pools of variants and drew on information-theoretic insights to efficiently explore the sequence space to generate diverse and high-performing deaminases. From a single round of diversification, we created a small set of novel and specific cytosine and adenosine deaminases that were markedly distinct in sequence from published base editor deaminases. We found that our model-created deaminases generally outperform those we identified through typical directed evolution. The novel compact deaminases identified here show high on-base activity, comparable to the leading published base editors, and with demonstrably lower off-base activity.
Authors
Keywords
No keywords available for this article.