▸ Chatomics Field GuideWhat They Don't Teach You →

Sanity check · Single-Cell RNA-seq

How to Detect Batch Effects in Single-Cell RNA-seq

Before you pick Harmony, RPCA or scVI, find out whether batch is separable from your condition at all, because no method fixes a confounded design.

By Ming "Tommy" Tang, Director of Bioinformatics in Big Pharma · Reviewed October 2026 · 6 min read

You have Cell Ranger output for several samples, you merged them, ran PCA and UMAP, and the plot shows one island per sample. Or the opposite: everything looks beautifully mixed after Harmony, and you cannot tell whether that is a good sign or whether the correction just erased the treatment effect you were funded to find.

What is at stake is false biology in both directions. Skip the check and you report clusters that are really "the Tuesday prep". Correct blindly and you can merge a real disease-specific state into its healthy neighbor. If batch is confounded with condition, nothing you run will separate them.

This page gives you an order of operations for the next hour: cheap plots first, then a quantitative score, then the one question that decides everything, which is whether any batch contains more than one condition.

What it looks like when it's happening

  • On the uncorrected UMAP, each sample or sequencing run forms its own island, even for cell types you know are shared, such as T cells.
  • PC1 or PC2 separates processing date, lane or operator rather than condition or cell type.
  • Clusters are almost pure for one sample: a table of cluster by sample shows most clusters with more than 90% of cells from a single batch.
  • A 'new' cluster appears in only one batch and its top markers are stress genes, mitochondrial genes or ribosomal genes rather than lineage markers.
  • A batch with much higher median UMIs per cell sits apart from the others, and the axis separating it correlates with nCount_RNA or nFeature_RNA.
  • After integration, the condition difference disappears along with the batch difference, and cells from healthy and disease samples overlap completely in every cluster.
  • Condition A samples all come from one date and condition B from another, so the batch and condition columns in your metadata are identical.

Why it happens

A batch effect is any systematic technical difference between groups of samples: flowcell, library prep, lab, operator, date or reagent lot. In droplet scRNA-seq it enters at capture, reverse transcription, amplification and sequencing, and it shifts many genes a little in the same direction for every cell in that batch. PCA is built to find the largest coordinated variance, so a consistent small shift across thousands of genes can easily outrank the variance that separates cell types.

Sequencing depth makes it worse. Per-cell depth is a major technical effect: cells with fewer UMIs have systematically more zeros, and dropout is not random across cells. If one batch was sequenced deeper, the main axis of variation can be depth rather than cell state, and the batch forms its own cluster. Depth-imbalanced batches also make correction methods over- or under-correct, depending on the depth regime.

In nearly every scRNA-seq design, batch is the same thing as sample, so the real question is not whether a batch effect exists but how much integration to apply. Integration makes cells cluster by cell type and smooths donor-specific differences. No integration keeps that variation but costs you alignment. The right choice depends on whether you want shared cell types or shifted states.

The hard stop is confounding. If condition A was processed in batch 1 and condition B in batch 2, then the condition difference equals batch plus biology and the two cannot be separated by any correction method. The honest answer is a redesign, ideally with each batch containing both conditions. That only works if batch was recorded at all: flowcell, lane, prep date, kit lot and lab need to be in your metadata, because an unannotated batch cannot be corrected.

The checks

Run them in order. Each one tells you what healthy looks like and what the problem looks like.

0/8 checked · saved in this browser

  1. Before any plotting, build a table of batch (date, lane, prep, operator) by biological condition. You are looking for whether any batch contains more than one condition. Do this first because it decides whether correction is even meaningful.

    r
    # meta: seurat_obj@meta.data with columns sample, batch, condition
    table(seurat_obj$batch, seurat_obj$condition)
    table(seurat_obj$sample, seurat_obj$batch)
    Healthy
    Each batch contains cells from at least two conditions, ideally balanced, and each condition appears in more than one batch.
    Red flag
    The table is diagonal: every batch holds exactly one condition. That is full confounding, and no integration method can rescue the comparison.
  2. List the metadata columns you have for flowcell, lane, prep date, kit lot and lab, and compare them with what the wet lab can tell you. Ask the person who ran the experiment; do not infer batch from sample names alone.

    Healthy
    Every sample has an explicit batch annotation for each processing step, and you can say which steps are shared between samples.
    Red flag
    Only a sample ID exists. You cannot tell whether two samples shared a prep day, so you cannot tell batch from donor.
  3. Run PCA and UMAP on the merged, normalized data with no correction, then color the embedding by batch, by sample and by condition. This is the standard visual detection: check whether clustering follows batch origin rather than biology. Bioconductor's batchelor vignette does the same with NoCorrectParam() then runPCA() and runTSNE() or runUMAP(), then plotTSNE(combined, colour_by="batch").

    r
    library(Seurat)
    obj <- merge(obj_list[[1]], obj_list[-1])
    obj <- NormalizeData(obj) |> FindVariableFeatures() |> ScaleData() |> RunPCA()
    obj <- RunUMAP(obj, dims = 1:30)
    DimPlot(obj, group.by = "batch")
    DimPlot(obj, group.by = "condition")
    DimPlot(obj, group.by = "sample")
    Healthy
    Shared cell types such as T cells or monocytes sit together across batches, with some batch-driven offset inside each island but no island that is one batch only.
    Red flag
    Every batch forms its own island, including cell types that must be shared. Or PC1 separates date or lane instead of lineage.
  4. For each of the first PCs, compute how much variance in the PC coordinates is explained by batch (a simple linear model R-squared is enough; this is the idea behind principal component regression). Do the same with log depth. If a depth variable explains the same PC as batch, the 'batch effect' may partly be a depth effect.

    r
    emb <- Embeddings(obj, "pca")[, 1:20]
    r2 <- function(y, x) summary(lm(y ~ x))$r.squared
    data.frame(
      PC = colnames(emb),
      batch = apply(emb, 2, r2, x = obj$batch),
      condition = apply(emb, 2, r2, x = obj$condition),
      depth = apply(emb, 2, r2, x = log10(obj$nCount_RNA))
    )
    Healthy
    Batch R-squared is small on the leading PCs, and the PCs with high condition or cell-type signal are not the ones with high batch signal.
    Red flag
    PC1 has a high batch R-squared and the same PC also tracks log depth. You are looking at a technical axis, and correcting it will behave differently depending on depth regime.
  5. Cluster the uncorrected data, then compute for each cluster the fraction of cells from each batch. Compare with the expected composition of the experiment. Composition can differ for real reasons, so pair this with the batch-by-condition table from check 1.

    r
    obj <- FindNeighbors(obj, dims = 1:30) |> FindClusters(resolution = 0.5)
    prop.table(table(obj$seurat_clusters, obj$batch), margin = 1)
    Healthy
    Major clusters contain cells from most batches. Batch-specific clusters, if any, are small and have a plausible biological explanation.
    Red flag
    Most clusters are dominated by one batch, or the same cell type appears as several clusters, one per batch.
  6. Process each sample separately (normalize, PCA, cluster, annotate by markers) and check that the same cell types are recovered in each. Because batch is nearly always the same thing as sample, this per-sample view tells you whether the cell types are truly shared before any integration touches them. Then compare with the integrated clusters.

    Healthy
    Per-sample clustering finds the same major lineages with the same marker genes in every sample.
    Red flag
    A cluster exists in one sample alone and its markers are stress, mitochondrial or ribosomal genes. Or integrated clusters merge populations that per-sample analysis kept clearly separate.
  7. Compute kBET on the PCA embedding, using batch as the label. It runs a chi-squared test on whether each fixed-size neighborhood has the same batch composition as the full dataset, so a low rejection rate indicates well-mixed batches. Run it before and after correction. NMI between k-means clusters and batch labels is an alternative, where values closer to 1 indicate better mixing as the paper frames it. I could not confirm an official rejection-rate cutoff, so judge the before-versus-after change and always read it next to the UMAP.

    Healthy
    Rejection rate drops substantially after correction, and the corrected embedding mixes batches within cell types while cell-type clusters stay distinct.
    Red flag
    Rejection stays high after correction, or it drops to near zero because the method also merged distinct cell types and condition-specific states, which kBET alone will not reveal.
  8. After integration, rerun the cross-tab of cluster by condition, and check that known condition-specific markers still separate cells within a cell type. Compare to the uncorrected result. Harmony adjusts only the PCA embedding and leaves expression values untouched, so you can always return to raw counts for DE within the integrated cluster labels.

    Healthy
    Known biology, such as an interferon response in treated cells, is still visible within a cell type, and condition-specific clusters you can defend from markers still exist.
    Red flag
    Condition-specific populations vanish and every cluster is a balanced mix of conditions. Strong integration, for example a very low Harmony lambda, may have removed the effect you wanted.

What to do about it

Redesign or add a bridging batch

When: Batch is perfectly confounded with condition: each batch holds a single condition.

Stop analysing and talk to the wet lab. Rerun or add samples so that each batch contains cells from both conditions, ideally by multiplexing conditions into one run. Record flowcell, lane, prep date, kit lot and lab for every sample this time.

Caveat: Costs time and money, and sometimes the material is gone. If so, report the comparison as exploratory and say plainly that condition and batch cannot be separated.

Harmony on the PCA embedding

When: Cell types are shared across batches, the batch effect is moderate, and your aim is a common map and consistent cluster labels.

Run RunHarmony on the PCA embedding with your batch variable, then build neighbors, UMAP and clusters on the harmony reduction. Defaults are theta=2, lambda=1, sigma=0.1; smaller lambda corrects more aggressively. The reference example uses lambda = .1 for stronger correction. Start at defaults and only tighten if per-sample checks show batches still separated.

obj <- RunHarmony(obj, group.by.vars = "batch")
obj <- RunUMAP(obj, reduction = "harmony", dims = 1:30)
obj <- FindNeighbors(obj, reduction = "harmony", dims = 1:30) |> FindClusters(resolution = 0.5)

Caveat: Harmony corrects the embedding only, not the expression matrix. Over-correction can blend real condition-specific states, so rerun check 8 after every change.

Seurat RPCA integration

When: Batches share few cell types, or a substantial fraction of cells in one dataset have no match in the other, and you want a conservative method.

Run FindIntegrationAnchors with reduction = "rpca" on the list of per-batch objects, then IntegrateData. RPCA is more conservative and faster than CCA. If rare or divergent cell types fail to align, raise k.anchor from its default of 5; the Seurat vignette suggests 20.

immune.anchors <- FindIntegrationAnchors(object.list = ifnb.list, anchor.features = features, reduction = "rpca")
immune.combined <- IntegrateData(anchorset = immune.anchors)

Caveat: Raising k.anchor increases alignment strength, which can pull together cells that should stay separate. Seurat v5 uses an IntegrateLayers workflow instead, so follow the vignette for your installed version.

scVI with batch as a covariate

When: You have many batches or large datasets and want a model-based latent space that handles count data directly.

Train scVI on raw counts with batch as the covariate and use the latent representation for neighbors, UMAP and clustering. See scvi-tools.org for the current API; the parameters were not available for this page, so no settings are quoted here.

Caveat: The latent space is a corrected embedding, not corrected expression. Check it with the same per-sample and condition checks as any other method.

Skip integration and model batch in the test

When: Your question is about donor-specific or condition-shifted states, or you are doing differential expression.

Keep cells unintegrated for state discovery. For DE, aggregate to pseudobulk per sample and put batch in the design matrix, for example design <- model.matrix(~ replicate + group), rather than testing on a pre-corrected matrix.

Caveat: Unintegrated clusters split by donor, so annotation is harder. Pre-corrected matrices used for tests give overly optimistic p-values, so do not feed integrated values into DESeq2 or limma.

When not to "fix" it

Do not correct when the separation is the biology you came to find. If you are looking for activated or shifted states, such as activated B cells in disease, integration may fold them into the healthy cluster. Likewise, if a batch-specific cluster has real lineage markers and shows up in per-sample analysis, it is a cell population, not an artifact. In those cases keep the data unintegrated, show per-sample views, and use batch as a covariate in testing. Also do not correct when batch is fully confounded with condition: the output will look clean and mean nothing.

Five things experienced analysts do here

  1. Write the batch-by-condition table into your analysis notebook before touching any integration function, and look at it again whenever someone says 'just integrate it'.
  2. Always show the uncorrected UMAP colored by sample next to the corrected one. A reviewer who can see the before can judge the after.
  3. Cluster each sample alone first. If a cell type is not found in a sample by itself, integration will not honestly create it.
  4. Treat low-depth and high-depth batches as a separate problem: check log depth against the leading PCs before blaming 'batch' and choosing a correction strength.
  5. Start Harmony at its defaults and tighten lambda only when a specific check fails. Every step of extra strength is a step toward erasing condition effects.

Questions people ask

How do I check for batch effects in scRNA-seq?

Run PCA and UMAP on merged, uncorrected data and color by batch, sample and condition. Then cross-tabulate cluster by batch and tabulate batch against condition. If cells group by processing date or lane instead of cell type, you have a batch effect worth handling.

Can I use ComBat for scRNA-seq batch effects?

For single-cell work the usual choice is an embedding or model-based method such as Harmony, Seurat RPCA or scVI, which fit sparse count data and give you a shared map. This page does not cover ComBat. Whatever you use, put batch in the design for DE rather than testing on a pre-corrected matrix.

When should I not correct a batch effect in scRNA-seq?

Skip integration when you are looking for new or shifted states, or when donor-specific variation is part of the question. Also be careful when batch is confounded with condition, because correction then cannot separate the two. The fix there is a redesign, not a better algorithm.

Does Harmony change my expression values?

No. Harmony adjusts only the PCA coordinates. Your raw counts stay as they are, so after getting integrated cluster labels you can run differential expression on the counts within each cell type, with batch as a covariate.

What kBET value means batches are mixed well?

A low rejection rate means well-mixed neighborhoods, but no official cutoff appears in the sources used here. Compare the rate before and after correction and read it alongside the UMAP and your condition-specific markers.

Related pages

Related reading on the blog

Sources

  1. Batch Effect: To Correct or Not for Bulk RNA-seq Data — Put batch in the design matrix instead of pre-correcting.
  2. Correcting batch effects in single-cell RNA-seq data (Bioconductor batchelor vignette) — Uncorrected PCA/t-SNE colored by batch and principal component regression.
  3. Seurat Integration with RPCA — RPCA commands and k.anchor guidance.
  4. A test metric for assessing single-cell RNA-seq batch correction (kBET) — kBET and NMI definitions.
  5. RunHarmony Parameter Reference — Harmony default parameters and the lambda = .1 example.

Part of the Batch effects series.