Examine neural network models for protein sequencing based features prediction.
Journal:
Journal of computer-aided molecular design
Published Date:
Aug 26, 2026
Abstract
Protein functions are essential for understanding life at the molecular level. High-throughput sequencing generates vast amounts of raw protein sequences, but only about 1 percent have been carefully annotated with their functions. The experimental research needed to annotate these functions is costly and time-consuming, and lags behind the rapid growth in the number of sequences. This situation has led to the development of computational methods to predict protein functions. One proposed solution is an autoencoder framework that accurately predicts protein functions using both network and protein sequence data. Specifically, the framework encodes protein sequence information related to domains, families, and patterns-into a lengthy, sparse binary vector. The autoencoder has been rigorously tested both experimentally and statistically using two datasets. The results demonstrate that it outperforms other neural network models, including convolutional neural networks, recurrent neural networks, long short-term memory networks, and bidirectional long short-term memory networks. In comparison, the autoencoder achieved accuracy, precision, recall, and F1 scores of 0.94, 0.92, 0.92, and 0.92, respectively, on the protein meta- data and protein structure sequence datasets. The remainder of this study focuses on classifying feature types in protein sequence datasets.
Authors
Keywords
No keywords available for this article.