From modality-specific to compositional foundation models for cell biology.
Journal:
Cell systems
Published Date:
Feb 18, 2026
Abstract
Deriving principles governing cell biology from single-cell measurements across modalities, called multimodal modeling, can advance our understanding of cellular states in health and disease. Realizing the full potential of multimodal models requires learning generalizable representations across data types, diseases, and biological contexts. This perspective examines the potential of compositional AI as a modular design approach for constructing multimodal foundation models that unify biological modalities-such as chromatin accessibility, protein abundance, spatial transcriptomics, microscopy imaging, and textual annotations-into cohesive representations of cellular behavior. We present key deep learning modeling approaches, along with transformer-based attention strategies to implement them, while addressing challenges posed by limited data availability and structural differences between modality representations. We also discuss how to connect and align partially overlapping multimodal measurements to build a comprehensive representation space. By synthesizing these technical advancements, we chart a path toward agentic virtual cell models, offering insights into opportunities, limitations, and future directions for leveraging multimodal AI to decode the complexity of cellular systems.
Authors
Keywords
No keywords available for this article.