Glossary · Single-Cell and Spatial
Regressing out variables
Regressing out percent.mt, nCount or cell cycle looks like free cleanup, but it can delete the biology you came for.
By Ming "Tommy" Tang, Director of Bioinformatics in Big Pharma · Reviewed October 2026 · 2 min read
Also: vars.to.regress, regress out cell cycle
Definition
Regressing out a variable means fitting a linear model of each gene's expression on a nuisance covariate, such as percent.mt, nCount_RNA or a cell cycle score, and keeping the residuals as the new expression values. Everything the covariate explains is subtracted, and downstream PCA and clustering see only what is left. The method cannot tell technical signal from biological signal that happens to correlate with the covariate. Think of it as sanding off a stain: it also takes some of the wood if the stain is in the grain.
You meet this term the first time you copy a Seurat or Scanpy tutorial. There is a vars.to.regress argument or a regress_out call sitting in the normalization step, usually with percent.mt or cell cycle scores in it. It is easy to run and easy to leave in without asking what it does to your data.
It decides what your clusters are made of. Regression changes the expression values that feed PCA, so whatever you remove never reaches clustering. Remove a nuisance and clusters get cleaner. Remove something that is partly real biology and you erase a population you were looking for.
Why it matters
The choice changes which clusters exist. If detection rate tracks batch or prep date, regressing it can shrink differences between clusters and show that an apparently novel state was an artifact. If you regress nCount on a tissue with plasma cells or hepatocytes, which genuinely carry more RNA, you erase real biology. If you regress cell cycle in a tumor where cycling differs by condition, you can hide a proliferative subclone or introduce spurious signal. Decide from plots, not habit.
Where people get it wrong
People treat vars.to.regress as a standard cleanup list and paste in percent.mt, nCount_RNA and the cell cycle scores every time. With SCTransform, depth is already modeled, so adding nCount_RNA is potentially redundant. People also confuse regression with batch correction: regression edits expression values per gene, while Harmony edits PCA embeddings and leaves the values alone. Scanpy's own docs warn that regress_out can overcorrect.
A concrete example
Before regressing anything, color the UMAP by nUMI, nFeature and percent.mt and check whether a cluster boundary follows depth more than known markers. If it does and the cell type has no reason to carry more RNA, regress. If the cluster is a type like plasma cells, leave it. Then run the normalization with only the covariate you justified.
# Seurat: regress only percent.mt, after checking it splits clusters for technical reasons
pbmc <- SCTransform(pbmc, vars.to.regress = "percent.mt", verbose = FALSE)
# Several covariates, e.g. before Harmony
# SCTransform(obj, vars.to.regress = c("percent.mt", "nCount_RNA"))Related terms
Questions people ask
- Should I regress out percent.mt?
Only after you check whether it tracks a technical problem or a cell type. High percent.mt marks stress and apoptosis, so regressing it removes that signature. But the sources give no cutoff for when to do it. Filter clearly damaged cells first, then color your UMAP by percent.mt and regress only if clusters split along it without a biological reason.
- Should I regress out cell cycle in Seurat?
Not by default. The OSCA book says cell cycle adjustment is not necessary for routine applications, because cycle is usually minor next to cell type identity. In tumors or developing tissue, cycling is part of the biology, and regressing it can hide real proliferative subclones.
- What does vars.to.regress do in SCTransform?
It tells SCTransform to include the listed metadata columns as covariates in its regression during normalization and variance stabilization. The default is NULL, so nothing extra is removed unless you ask. SCTransform already corrects for sequencing depth, so adding nCount_RNA is potentially redundant.
- What is the difference between scanpy regress_out and Seurat vars.to.regress?
Both fit a linear model per gene and keep the residuals. Scanpy's
scanpy.pp.regress_outis a separate step on the AnnData object, and its docs note it can overcorrect in some circumstances. In Seurat the regression is an argument inside the normalization call.- Does regressing out replace batch correction?
No. Regression removes the effect of covariates from expression values. Harmony works on PCA embeddings and leaves the original values unchanged, and it is compatible with upstream regression. They solve different problems, so decide on each separately.
Related pages
Related reading on the blog
Sources
- Using sctransform in Seurat — vars.to.regress parameter in SCTransform, default NULL, and the percent.mt example.
- scanpy.pp.regress_out, Scanpy stable documentation — Linear regression implementation and the documented tendency to overcorrect.
- Chapter 17: Cell cycle assignment - Orchestrating Single-Cell Analysis — Cell cycle adjustment is not necessary for routine applications; regressBatches approach.
- Overclustering in scRNA-seq: How to Check and Fix It — Plasma cells and hepatocytes carry more RNA; cell cycle can be real biology.
- Normalization and variance stabilization of single-cell RNA-seq data using regularized negative binomial regression — Regression of technical metrics and use of residuals downstream.