Accessible and reproducible deployment reveals the practical boundaries of single-cell foundation models

Journal: bioRxiv
Published Date:

Abstract

Single-cell foundation models (scFMs) have been widely promoted as a unifying paradigm for transcriptomic analysis, yet whether large-scale pretraining translates into reproducible biological advantages remains unclear. Their adoption is further hindered by heterogeneous implementations, preprocessing requirements, and computational environments. Here we develop a unified, automated, and reproducible framework for standardized deployment and controlled evaluation of scFMs across datasets, computational environments, training regimes, and downstream analyses, substantially lowering the technical barriers to their use. Leveraging this framework, we systematically investigate thirteen scFMs alongside established methods across nearly one hundred datasets spanning diverse biological contexts. Our analyses reveal clear practical boundaries to scFM utility. First, increased model scale, architectural complexity, pretraining corpus size, or input encoding does not consistently translate into superior downstream performance. Instead, measurable properties of embedding geometry provide a model-agnostic, representation-level explanation for differences in zero-shot performance across diverse model families. Second, the benefits of pretrained representations depend strongly on the biological and supervision regime: scFMs provide their clearest advantages under extremely limited supervision, particularly for rare-cell annotation and open-set detection of source-absent cell states, whereas established methods remain competitive or preferable in most other settings. Task-matched analyses further show that scFM representations transfer inconsistently to spatial-domain recovery, while their gene embeddings capture broad functional relatedness without reliably recovering context-specific regulatory relationships. Together, these results establish that scFM utility is neither universal nor determined simply by model scale alone, but varies with learned representation geometry, biological context, and supervision. By combining reproducible deployment with large-scale empirical and mechanistic investigation, our framework provides a principled foundation for determining when foundation-model pretraining offers genuine practical value and when simpler approaches remain sufficient.

Authors

  • Hou
  • S.; Yang
  • P.; Ma
  • W.; Xiang
  • J.; Wang
  • J. X.; Wan
  • H.; Ma
  • Y.; Zhou
  • X.

Categories