▸ Chatomics Field GuideWhat They Don't Teach You →

Glossary · RNA-seq

Batch correction (ComBat, limma removeBatchEffect)

Batch correction can rescue a PCA plot and ruin a p-value in the same script, so the real decision is where in the analysis you apply it.

By Ming "Tommy" Tang, Director of Bioinformatics in Big Pharma · Reviewed October 2026 · 3 min read

Also: ComBat, removeBatchEffect, RUV

Definition

Batch correction is any procedure that estimates the systematic technical differences between groups of samples (flowcell, library prep, lab, operator, date, reagent lot) and removes or adjusts for them. limma::removeBatchEffect() fits a linear model of expression on the biological design plus batch and subtracts the batch component from a log-expression matrix. ComBat (sva) uses an empirical Bayes framework to adjust for batch differences in both mean and variance, and ComBat_seq does the same on raw counts with a negative binomial model, so the output stays integer. Think of it as sanding off a shared technical tilt so the biological signal is easier to see, which only works if batch and biology were not the same thing to begin with.

You will meet batch correction the first time a PCA plot separates your samples by sequencing date instead of by treatment. Someone suggests running ComBat or limma::removeBatchEffect(), the clusters tighten, and the plot looks publishable. What that fix does to your statistics depends on where you apply it.

The decision is not "correct or not". It is whether you are making a picture or running a test. Plots can use a corrected matrix. Differential expression should model batch directly. The two tools in the title sit on opposite sides of that line, and mixing them up is how analyses get overly optimistic p-values.

Why it matters

Leek et al. showed that batch can explain more variation than the biological factor of interest in many high-throughput studies, and can even reverse apparent group differences. If you ignore it, you get fake differential expression and classifiers that "predict" batch. If you pre-correct the matrix and then test, limma computes the wrong residual degrees of freedom and standard errors, so p-values come out too optimistic. A concrete case: control vs KO with three replicates, where PC1 separates replicate 1, 2 and 3. If those replicates share technical handling, the right move is ~ replicate + group in the design. Running removeBatchEffect() first and testing on the output makes the result look stronger than the data justify. And if all controls were sequenced in one batch and all treated samples in another, no method can separate the two.

Where people get it wrong

The most common mistake is treating removeBatchEffect() as a preprocessing step before differential expression. It was built for visualization (PCA, heatmaps), and for testing you should put batch in the design matrix instead. A second mistake is calling any set of replicates a batch. Replicate 1, 2, 3 is only a batch variable if the samples shared a technical event such as the same sequencing date, plate or processing run. A third is feeding log-transformed data to original ComBat and then handing the result to DESeq2 or edgeR, which need integer counts. Original ComBat output can be non-integer or negative. ComBat_seq keeps counts as integers, but the same design-in-the-model advice still applies to testing. Finally, correction cannot fix a fully confounded design.

A concrete example

You have bulk RNA-seq with a control vs KO comparison and three replicates processed on different days. PC1 separates the replicates and PC2 separates the conditions. Test with batch in the design, and use a corrected matrix only for the plot. If you need corrected integer counts for a tool that wants counts, ComBat_seq takes the raw count matrix with batch and optional group. Afterwards, check that known non-DE genes did not become significant and that your positive controls (for example the knocked-out gene) still go the right way.

r
library(limma)
library(sva)

# 1. Differential expression: batch goes in the design, matrix stays uncorrected
design <- model.matrix(~ replicate + group)
fit <- lmFit(v, design)

# 2. Visualization only: protect the biology, remove the batch, then plot PCA
design_bio <- model.matrix(~ group)
expr_corrected <- removeBatchEffect(expr, batch = batch, design = design_bio)

# 3. If you need corrected integer counts, ComBat_seq works on raw counts
adj_counts <- ComBat_seq(counts, batch = batch, group = group)

Related terms

Questions people ask

What is the difference between ComBat and limma removeBatchEffect?

ComBat uses empirical Bayes to adjust for differences in both mean and variance across batches. removeBatchEffect fits a linear model and removes the batch-associated mean shift from a log-expression matrix. Include batch in the limma design for testing, and consider ComBat when batch variances differ a lot.

Should I use ComBat or put batch in the limma design?

For differential expression with known batches, put batch in the design matrix. That is usually preferred over a two-step procedure of ComBat followed by testing. Limma with batch in the design assumes equal batch variances, so large variance differences are the case where ComBat may be more appropriate.

When should I correct for batch effects?

Correct or model batch when batch reflects a shared technical event and your design is not fully confounded with condition. Do not call biological replicates a batch unless they share a technical event. If condition and batch are the same thing, no analysis can fix it, so randomize conditions across lanes, dates and operators in the design.

Can I run differential expression on ComBat-corrected data?

Not on log-transformed ComBat output with count-based tools, because it can be non-integer or negative. ComBat_seq returns integer counts that edgeR and DESeq2 accept. Modeling batch in the design is still the cleaner route for testing.

Related pages

Related reading on the blog

Sources

  1. Batch Effect: To Correct or Not for Bulk RNA-seq Data — removeBatchEffect is for visualization; batch belongs in the design matrix for DE; replicates are not automatically a batch.
  2. ComBat: Adjust for batch effects using an empirical Bayes framework — Parametric vs non-parametric priors and function signature.
  3. ComBat_seq: Adjust for batch effects in RNA-seq count data — ComBat_seq parameters and defaults.
  4. removeBatchEffect: Remove Batch Effect — Arguments, log-expression input, and the design argument.
  5. ComBat-Seq: batch effect adjustment for RNA-Seq count data — Negative binomial model and integer-preserving output.
  6. Batch effects: ComBat or blocking in limma? — Mean and variance adjustment in ComBat vs blocking in limma.