Genomic foundation model embeddings encode higher-order viral genome architecture beyond sequence composition: a benchmark of Evo 2
Journal:
bioRxiv
Published Date:
Jul 16, 2026
Abstract
Genomic foundation models such as Evo 2 are increasingly applied to microbial genomics, yet how well their representations capture viral genome organisation, and how reliably they generate viral sequence, remain poorly characterised. We present a reproducible benchmark of Evo 2 on viral genomes. Using a pre-registered RefSeq viral corpus (19,429 genomes, organised by Baltimore class and host domain), we evaluated three axes: linear probes decoding Baltimore class, host domain and viral family from mean-pooled embeddings; ridge-regression probes recovering genomic features, including higher-order architectural properties such as gene density, coding fraction and gene overlap; and generative completion of fragmented genomes, scored on a leakage-safe set of eukaryote-infecting viruses (excluded from Evo 2's training corpus by design) against a bacteriophage comparator. All probes used cross-validation with sequence-identity clustered folds, benchmarked against both a GC-and-length control and a 6-mer composition representation. From its optimal intermediate layer, the 20B embedding classified Baltimore class at 0.96 accuracy and host domain at 0.99, exceeding both baselines; for viral family, however, 6-mer composition (0.89) matched the embedding (0.91. Most informatively, the embedding decoded coding fraction, gene density and gene overlap (R = 0.61, 0.77 and 0.64) far beyond 6-mer composition (0.10, 0.38 and 0.27), evidencing genuine encoding of genome architecture rather than nucleotide composition (p < 0.001). Performance scaled with model size. In generation, perplexity was lower for bacteriophages (1.18 bits/nt) than for held-out eukaryotic viruses (1.80). Evo 2 encodes functional viral genome architecture beyond composition, while taxonomic and generative behaviour partly reflect composition and training exposure.