Evo 2 as a classification machine: evidence from in-context learning and mechanistic interpretability

Journal: bioRxiv
Published Date:

Abstract

In-context learning (iCL) is an emergent capability of Large Language Models (LLMs), allowing them to perform new tasks at inference time using prompt-injected examples. While extensively studied in LLMs, the boundaries of its capabilities and underlying mechanisms remain poorly characterised in genomic Language Models (gLMs). Here, we map the operating regime of Evo 2, a nucleotide-level foundation gLM, across five binary classification tasks spanning biological and artificial sequences. We observe robust iCL on shorter natural sequences (F1=0.902 for miRNA, 0.785 for Toxins), degrading with sequence length, and collapsing at kilobase scale. Strikingly, we find no benefit in model scaling, as the 7B model systematically outperforms the 40B variant. We further find that perplexity, a widely used gLM performance proxy, poorly predicts accuracy. Mechanistic interpretability reconciles these observations: the logit-lens profiling suggests a prediction-generalisation trade-off, while Jacobian Scope indicates models might track prompts' structure rather than the signal-carrying content.

Authors

  • Bertolini Agnoletto
  • L.; Curion
  • F.; Petrillo
  • M.; Leoni
  • G.; Ronco
  • M.; Ruiz Serra
  • V.; Consoli
  • S.; Ceresa
  • M.

Categories