VINTER: a generative vision-language model for annotating and interrogating spatial tissues

Journal: bioRxiv
Published Date:

Abstract

Spatial transcriptomics links gene expression to tissue structure, yet current models learn embeddings whose meaning is assigned afterwards or read histology without measured expression. Tissue decisions are therefore rarely made from morphology, genes and neighbourhood together, or examined through them. Here we introduce VINTER, a generative vision-language model aligned on 1.45 million histology-caption pairs that reads histology, measured gene expression and spatial neighbourhood as one sequence. It returns the probability of every tissue class, and because genes enter as text, its molecular evidence can be edited with the image held fixed. Reading all three sources together, VINTER annotated tissue more accurately than spatial-domain methods and a spatial foundation model, and transferred across cohorts without target labels. Editing only the neighbourhood resolved the follicular B-cell and plasma-cell programmes behind its reading of tertiary lymphoid structure (TLS) maturation in lung and kidney cancer. Its class probabilities exposed what single labels conceal in triple-negative breast cancer: immune heterogeneity within stroma and cancer hidden beneath benign labels. Without retraining, the same model transferred to Visium HD and single-cell Xenium data and resolved follicular architecture within colorectal cancer TLS. VINTER thus turns spatial annotation into an interrogable reading of tissue across platforms and resolutions.

Authors

  • Wu
  • J.; Xu
  • W.; Zhuang
  • Y.; Zhang
  • Y.; Loza Lopez
  • M. d. J.; yang
  • y.; Nakai
  • K.

Categories