▸ Chatomics Field GuideWhat They Don't Teach You →

Glossary · AI and Claude Code

Foundation model (biology)

scGPT and Geneformer promise one pretrained model for every single-cell task, but your count matrix decides whether that promise holds for your data.

By Ming "Tommy" Tang, Director of Bioinformatics in Big Pharma · Reviewed October 2026 · 2 min read

Also: single-cell foundation model, scGPT, Geneformer

Definition

A foundation model in biology is a large neural network, usually a transformer, pretrained with self-supervised learning on a very large unlabeled dataset and then reused for many downstream tasks through zero-shot prediction, fine-tuning, or in silico perturbation. In single-cell work, genes and their expression values are treated as tokens, like words in a sentence, and the model learns relationships between genes within a cell. scGPT (pretrained on over 33 million human cells with masked gene prediction) and Geneformer (rank-value encoding, 30 million cells in V1 and 104 million in V2) are the two names you will meet first.

You will meet this term in a paper that claims a pretrained model beats your usual Seurat or Scanpy workflow at cell type annotation, batch integration, or perturbation prediction. Or a collaborator will ask whether you can "just run scGPT" on the new dataset.

The term decides how much you trust a result you did not build. A foundation model is not a method with a known failure profile like a negative binomial DE test. It is a large pretrained artifact, and what it learned depends on the data it saw and how that data was processed.

Why it matters

The pretraining corpus is public single-cell data, and public single-cell data is messy. Datasets come from different pipelines (Cell Ranger versions, STARsolo, Alevin, kallisto/bustools) that produce incompatible count matrices. Depth differences across batches can look like biology. A model trained on that mixture can encode those artifacts as if they were signal.

The concrete case: you fine-tune on your own tumor dataset and the embedding separates samples cleanly. Before you call that a discovery, check whether the separation tracks per-cell depth or batch. Even 100 million cells is a sliver of the roughly 30 trillion cells in one human body, so rare cell types and regional tumor heterogeneity may sit outside what the model has seen. Plan to validate on your own data, not on the paper's benchmark.

Where people get it wrong

People treat "foundation model" as a synonym for "better embedding" or for any LLM applied to biology. It is neither. The label describes a training recipe: large-scale self-supervised pretraining, then reuse. It says nothing about whether the model beats a simple baseline on your task, and the sources behind this page do not give a systematic benchmark comparison. Run a baseline you already trust (your normal normalization, PCA, and integration) and compare on labels you know. Also do not assume the model has seen your preprocessing: Geneformer uses rank-value encoding and scGPT uses its own gene, expression, and condition tokens, so feeding either a matrix from a different pipeline without reading the documentation is a quiet way to get plausible nonsense.

A concrete example

You have a tumor scRNA-seq dataset with two batches, and the second batch was sequenced much deeper. You embed the cells with a pretrained foundation model and see two clean clusters. Before trusting it, check whether the clusters track batch and library size rather than cell state, and compare against your standard pipeline's embedding.

python
import scanpy as sc

# adata.obsm["X_fm"] holds the foundation model embedding you computed
# (see the model's own docs; Geneformer: geneformer.readthedocs.io)
sc.pp.calculate_qc_metrics(adata, inplace=True)
sc.pp.neighbors(adata, use_rep="X_fm")
sc.tl.umap(adata)

# If batch and depth panels look like the clusters, the embedding is tracking artifacts
sc.pl.umap(adata, color=["batch", "total_counts", "n_genes_by_counts", "cell_type_manual"])

Related terms

Questions people ask

What is a foundation model in bioinformatics?

It is a large AI model pretrained on a huge unlabeled dataset with self-supervised learning, then adapted to many tasks by transfer learning or fine-tuning. In single-cell biology, the model treats genes and expression values as tokens and learns how genes relate within a cell.

What is the difference between scGPT and Geneformer?

scGPT is decoder-inspired, pretrained on over 33 million human cells with masked gene prediction, and uses gene, expression value, and condition tokens. Geneformer is encoder-only and uses rank-value encoding, where genes are ranked by expression relative to the cell's transcriptome. V1 used 30 million cells and V2 used 104 million.

Can I use a single-cell foundation model without fine-tuning?

Yes, that is zero-shot use: the pretrained model makes predictions directly on your data. Treat the output as a hypothesis and check it against a baseline pipeline and known labels, because zero-shot quality depends on how close your data and preprocessing are to what the model saw.

What tasks do single-cell foundation models support?

Cell type annotation, batch correction and integration, gene function prediction, gene network inference, perturbation response prediction, and imputation of sparse RNA-seq data. Other models include scBERT, scFoundation, GeneCompass, and Nicheformer, which adds spatial transcriptomics.

Related reading on the blog

Sources

  1. scGPT: toward building a foundation model for single-cell multi-omics using generative AI — scGPT training scale, masked gene prediction, and token types.
  2. Single-cell foundation models: bringing artificial intelligence into cell biology — Genes as tokens, downstream tasks, and other models such as scFoundation and GeneCompass.
  3. Foundation models in bioinformatics — General definition and architecture types.
  4. Geneformer GitHub repository — Rank-value encoding, V1 and V2 training sizes, zero-shot, fine-tuning, and in silico perturbation.