ARSENAL: Learning Transferable Regulatory DNA Representations with Targeted Short-Context Language Models
Journal:
bioRxiv
Published Date:
Jul 12, 2026
Abstract
DNA language models (DNALMs) aim to learn representations of genomic sequence for variant interpretation, regulatory prediction, and sequence design. Most DNALMs are trained on whole genomes and long contexts, but regulatory DNA poses a distinct challenge: functional elements are sparse, context dependent, and encoded by short transcription factor motif syntax embedded in extensive background sequence. We introduce ARSENAL, a short-context masked DNA language model pretrained on ENCODE candidate cis-regulatory elements. ARSENAL recovers diverse transcription factor motifs de novo and improves zero-shot regulatory variant effect prediction relative to other DNALM foundation models. ARSENAL embeddings also improve supervised regulatory sequence models at predicting chromatin accessibility and regulatory variant scoring. Finally, ARSENAL serves as an efficient generative prior, enabling multi-objective regulatory sequence design with supervised oracles. ARSENAL shows that targeted self-supervised pretraining on regulatory regions can learn biologically meaningful and transferable regulatory representations without genome-scale training, long contexts or task-specific labels.