GRNPred: A Multimodal Graph Transformer with Masked Gene Expression Pretraining for Gene Regulatory Network Inference
Journal:
bioRxiv
Published Date:
Apr 29, 2026
Abstract
Gene regulatory network (GRN) inference is a key problem in systems biology that aims to identify transcription factor (TF)-target gene interactions from high-dimensional gene expression data, but it remains challenging due to limited labeled data, class imbalance, and complex nonlinear regulatory relationships. To address this, we propose GRNPred, a multimodal graph transformer framework that integrates gene expression, functional annotations, semantic gene descriptions, regulatory motif priors, and co-expression network topology. GRNPred uses a two-stage training strategy: first, a self-supervised pretraining phase where a graph transformer learns transcriptional context through masked gene-expression reconstruction on TF-centered subgraphs, and second, a supervised fine-tuning phase for TF-target edge prediction using known regulatory annotations. By leveraging transformer-based attention, the model captures long-range and context-dependent interactions that traditional methods struggle to model. Extensive evaluation across seven benchmark datasets and three regulatory network constructions shows that GRNPred outperforms state-of-the-art approaches, achieving up to 0.94 AUROC and 0.93 AUPRC while maintaining strong robustness across diverse biological settings.