Controlled Protein Design via Statistical Energy Functions: A Rossmann Fold Case Study.

Journal: Journal of chemical information and modeling
Published Date:

Abstract

The field of protein design has advanced significantly through the integration of computational techniques, including both statistically controlled models and deep learning approaches. While deep learning has gained widespread attention, statistical models remain indispensable for elucidating the physical properties of protein structures. This study investigates the continued relevance of statistical models in protein design by designing a Rossmann fold protein comprising ten motifs distributed across three layers─a representative domain found in numerous protein structures. Two statistical energy functions were employed: SCUBA (Side chain Unknown Backbone Arrangement) for scaffold design and ABACUS2 (A Backbone-based Amino Acid Usage Survey) for sequence design. By applying these functions with tailored restraints to incorporate desired structural features, we generated 300 low-energy sequence candidates. Subsequent filtering revealed that 69% of these sequences exhibited high-confidence fold similarity based on AlphaFold2 predictions. Experimental validation was performed on nine selected sequences, guided by protein structure predictions. A single crystal structure, resolved at 1.8 Å, demonstrated a main-chain deviation of 2.602 Å from the designed model, which adhered to the imposed constraints. These findings underscore the controllability achievable with statistical models in protein design, offering valuable insights for future applications.

Authors

Keywords

No keywords available for this article.