Peptide language pragmatic analysis and two-stage hierarchical learning framework for therapeutic peptide prediction.

Journal: BMC biology
Published Date:

Abstract

BACKGROUND: Therapeutic peptides exert pivotal effects in diverse biological processes, and have attracted significant interest in the field of biomedicine in recent years. However, most existing methods often fail to adequately capture the intricate interactions among amino acid residues and the contextual dependencies within peptide sequences, which hampers the extraction of deep semantic representations and ultimately restricts predictive performance. Moreover, the task of multi-functional therapeutic peptide prediction is inherently constrained by the challenge of imbalanced multi-label classification resulting from long-tailed distribution patterns. RESULTS: In this study, we propose a two-stage hierarchical deep learning framework, named TPpred-PepPA, for the prediction of multi-functional therapeutic peptides based on pragmatic analysis. Specifically, ProtT5 is employed to extract deep semantic representations that capture residue-level contextual information. In the first stage, a transformer-based network is utilized to perform shared representation learning, wherein the encoder model captures the intricate inter-residue interaction to characterize the contextual semantics of peptide sequences. In the second stage, the framework is fine-tuned by incorporating task-specific classifiers and optimizing the classification decision with Asymmetric Loss. Then the dynamic thresholding strategy is utilized to address the long-tail distribution problem, enabling more accurate prediction performance of multi-functional therapeutic peptide. Moreover, we adopted the SHAP analysis and motif identification to interpret feature contributions and identify key functional peptide fragments, respectively. Our experimental results indicate that TPpred-PepPA significantly outperforms all current baseline methods in identifying multi-functional therapeutic peptides and exhibits robust performance in recognizing rare functional categories. CONCLUSION: We developed TPpred-PepPA, a two-stage hierarchical deep learning framework based on the ProtT5 pre-trained large language model. Compared with existing methods, TPpred-PepPA achieves state-of-the-art predictive performance and provides valuable interpretability for the discovery of multi-functional therapeutic peptides. Finally, a web server has been established and is accessible at http://bliulab.net/TPpred-PepPA .

Authors

Keywords

No keywords available for this article.