Probing Genomic Foundation Models with Splice-Variant Perturbations

Journal: bioRxiv
Published Date:

Abstract

As benchmarks saturate, traditional validation methods fail to capture the probabilistic representations of genomic foundation models. This study probes self-supervised generalists (NTv3 650M, Evo2 40B, Genos 10B) alongside task-specific specialists (SpliceAI, AlphaGenome, Borzoi) using a mechanistically focused paradigm centered on expert-curated single-nucleotide substitutions within the splicing regions of the exon-rich OPA1 gene. Extending this analysis across the broader corpus of OPA1 mutations, variants of uncertain significance (VUS), and an independent 65-gene dataset of spliceogenic mutations demonstrates that generalists can match task-specific networks. Training data scale, genetic diversity, and evolutionary context are key for capturing RNA splicing syntax, whereas sparse routing may constrain performance. Mechanistically, foundation models exhibit a compute-asymmetry: the Evo2 40B model frontloads compute into massive parameter scale to resolve single-nucleotide variants within a narrow 425-bp context window, whereas the bidirectional NTv3 650M model and the Genos 10B mixture-of-experts (MoE) architecture backload compute to inference, requiring spatial logit aggregation and expansive context windows of up to 94 kb. Furthermore, classification accuracy across these foundation models benefitted from locus-specific thresholding rather than mutation-class calibration, exposing representation boundaries of the models.

Authors

  • Alavi
  • M. V.

Categories