AntiCapt: Fine-Tuned Nucleotide Language Models for Predicting and Designing Anticancer Aptamers

Journal: bioRxiv
Published Date:

Abstract

Over the past decade, aptamers have emerged as promising therapeutics, with cancer therapeutics as a major research area. Existing computational methods, however, either predict aptamer-target pairs or design aptamers against a specific target. In this study, we present AntiCapt, a computational method for identifying single-stranded DNA (ssDNA) anticancer aptamers (ACAs) using sequence descriptors and fine-tuned nucleotide language models. The main dataset comprises 1,021 experimentally validated ACAs and an equal number of non-aptamer sequences. Comparative analysis revealed distinct patterns in the nucleotide, dinucleotide, and trinucleotide compositions associated with ACAs. We developed machine learning (ML) models using composition, autocorrelation, binary profiles and structural features. Among these, the best-performing composition-based model achieved an AUC of 0.93 on an independent dataset, while chemical descriptor-and structural feature-based models achieved a maximum AUC of 0.89. ML models using pretrained and fine-tuned NLM embeddings were also developed, with fine-tuned HyenaDNA embeddings achieving the highest performance, with an independent AUC of 0.94. Furthermore, we developed a model for discriminating anticancer aptamers from general aptamers. The best-performing models were integrated into AntiCapt, a web server and a standalone tool for predicting, designing, and genome-scale scanning anticancer aptamers (https://webs.iiitd.edu.in/raghava/anticapt/).

Authors

  • Bajiya
  • N.; Mehta
  • N. K.; Raghava
  • G. P. S.

Categories