▸ Chatomics Field GuideWhat They Don't Teach You →

Glossary · Single-Cell and Spatial

Cell cycle scoring

A phase call for every cell that tells you whether your clustering split is biology or just who happened to be dividing.

By Ming "Tommy" Tang, Director of Bioinformatics in Big Pharma · Reviewed September 2026 · 3 min read

Also: CellCycleScoring, S score G2M score

Definition

Cell cycle scoring assigns each cell an S-phase score and a G2/M-phase score from the average expression of canonical cell cycle marker genes, then classifies the cell into G1, S, or G2M phase from those two scores. In Seurat's CellCycleScoring, cells with high S markers (PCNA, MCM6, and others) are called S phase, cells with high G2/M markers (TOP2A, MKI67, and others) are called G2M, and cells with neither are called G1, or quiescent. Both Seurat and Scanpy use the same 97-gene Tirosh et al. 2015 marker list; Scanpy's version additionally scores each cell against a randomly matched background gene set to control for expression-level bias.

You meet this the moment PCA or UMAP splits your cells along a gradient that lines up suspiciously well with proliferation markers instead of cell type. Cell cycle scoring is how you check that suspicion computationally instead of squinting at MKI67 and TOP2A on a feature plot: it gives every cell an S-phase score, a G2/M-phase score, and a discrete phase call you can plot, filter, or regress on.

The decision it feeds is whether and how to correct for cell cycle before clustering. Get it wrong and you either bury real biology by regressing out signal that overlaps with your actual cell-type or treatment effect, or you leave a cluster of proliferating cells from several lineages masquerading as one weird "doublet-like" cluster.

Why it matters

Case: a PBMC or tumor dataset produces one cluster co-expressing markers from two lineages, say T cell and myeloid genes together. Score cell cycle before calling it a doublet cluster. If that cluster's average S/G2M score is elevated relative to the rest, cells are separating by "am I dividing," not by lineage identity, and it is not a doublet artifact to filter out, it is every lineage's proliferating subset pooled together by shared cycling genes.

The correction you pick then matters. Regressing S.Score and G2M.Score together (vars.to.regress = c('S.Score','G2M.Score')) erases all proliferation-linked variation, including the real distinction between cycling and quiescent cells, which is often exactly the axis a stem-cell or progenitor study needs to keep. Regressing CC.Difference (S.Score minus G2M.Score) instead keeps the cycling-vs-quiescent split intact and only removes which phase a proliferating cell happens to be in. Pick the wrong one of these two and you either destroy a biological axis or leave phase noise sitting in your PCs.

Where people get it wrong

The mistake: treating the phase label (G1, S, or G2M) as the signal and the numeric score as an afterthought, when it is the reverse. The classifier only compares S score against G2M score against neither, with no real cutoff for "is this cell cycling at all." A cell with S.Score = 0.02 and G2M.Score = 0.01 gets called S phase the same as one with S.Score = 0.9, so counting "percent of cluster in S phase" without checking the raw score distribution will overstate how much genuine proliferation difference exists between groups.

A second mix-up: cell cycle scoring is a special case of module-score scoring with a fixed, curated marker list, not any gene set you hand it. It is not interchangeable with running an arbitrary gene set through the same machinery and expecting comparably clean phase separation; the marker list quality is doing most of the work, and forcing every cell into G1/S/G2M regardless of whether it is actually cycling is a known weakness of this family of methods (Seurat, Revelio, CS1CC).

A concrete example

Run CellCycleScoring on log-normalized data before scaling, check whether phase actually separates your clusters on a DimPlot, then choose CC.Difference regression instead of full S/G2M regression if you need to preserve the distinction between cycling and quiescent cells.

r
library(Seurat)
s.genes <- cc.genes$s.genes
g2m.genes <- cc.genes$g2m.genes
pbmc <- CellCycleScoring(pbmc, s.features = s.genes, g2m.features = g2m.genes, set.ident = FALSE)

# look before you regress: does Phase line up with a cluster split?
DimPlot(pbmc, group.by = "Phase")

# keep the cycling-vs-quiescent axis, remove only phase-specific variation
pbmc$CC.Difference <- pbmc$S.Score - pbmc$G2M.Score
pbmc <- ScaleData(pbmc, vars.to.regress = "CC.Difference", features = rownames(pbmc))

Related terms

Questions people ask

Should I regress out cell cycle before or after clustering?

Before, on log-normalized data ahead of scaling and PCA. Bioconductor guidance is to regress on the logcounts matrix, then repeat feature selection, integration, and clustering, because PCs computed before regression still carry the cell cycle signal.

Does cell cycle scoring work the same in Scanpy and Seurat?

Both use the same 97-gene Tirosh et al. 2015 list split into S-phase and G2M-phase genes. Scanpy's score_genes_cell_cycle scores each set against a randomly matched control gene set, while Seurat's CellCycleScoring assigns a phase by comparing each cell's S score against its G2M score.

Should I regress cell cycle genes out or just remove them from the analysis?

Bioconductor guidance favors excluding cell cycle genes, for example dropping them from the highly variable gene list, over linear regression, because removal is more predictable and less likely to introduce artifacts in unrelated genes.

Will regressing out cell cycle hurt real biology in proliferating populations?

Yes, if you regress S.Score and G2M.Score together, since that removes all cycling-versus-quiescent variation along with phase differences. Use CC.Difference regression instead when the cycling-vs-not distinction matters, such as in stem or progenitor cell studies.

Why does my UMAP still separate by cell cycle after I regressed it out?

Regression likely happened after PCA was already computed on the unregressed matrix. Regress on the log-normalized data first, then recompute PCA and UMAP downstream of that correction.

Related pages

Sources

  1. Cell-Cycle Scoring and Regression • Seurat — marker gene lists, phase assignment logic, CC.Difference vs full regression
  2. Cell-Cycle Scoring and Regression - Scanpy - Read the Docs — Scanpy's score_genes_cell_cycle algorithm and gene list
  3. Cell cycle regression for scRNA-seq data - Bioconductor Support — gene removal vs regression, regress before PCA guidance
  4. From G1 to M: a comparative study of methods for identifying cell cycle phases — marker-based methods forcing cells into predetermined phases