▸ Chatomics Field GuideWhat They Don't Teach You →

Sanity check · Bulk RNA-seq

How to Choose a Normalization Method in Bulk RNA-seq

Median-of-ratios, TMM, CPM and TPM answer different questions, and the wrong one fails quietly when your treatment shifts the whole transcriptome.

By Ming "Tommy" Tang, Director of Bioinformatics in Big Pharma · Reviewed October 2026 · 5 min read

You have a count matrix from STAR or salmon, a treatment and a control, and a DESeq2 or edgeR script that runs without errors. Somewhere in the pipeline a normalization was applied, probably by default, and you have not checked whether it fits your experiment. Or you were handed a TPM table and told to "just run DESeq2".

What is at stake: library-size normalization assumes most genes do not change. When that is false (MYC activation, a failed rRNA depletion, a drug that shuts down transcription), the method quietly shifts every gene's fold change by the same wrong amount. You get a confident volcano plot built on the wrong baseline. The same plot also appears if you feed a tool the wrong kind of numbers.

In the next hour you can decide which normalization belongs to which task, confirm that your input matrix is the right type, test the "most genes unchanged" assumption on your own data, and know when you need spike-ins instead of a smarter algorithm.

What it looks like when it's happening

  • You ran DESeq2 and the MA plot cloud sits visibly above or below zero instead of centered on the horizontal line.
  • Thousands of genes are called significant in one direction only, and the treatment is known to affect global transcription or cell composition.
  • DESeq2 on a TPM or CPM table runs but returns odd dispersion fits, or a colleague handed you a 'normalized' table and the values are not integers.
  • PC1 correlates with total reads or number of detected genes rather than with condition.
  • Low-expression genes light up as 'up-regulated' in the samples that were sequenced deeper.
  • Heatmaps of the same genes look different when drawn from CPM, TPM and DESeq2 normalized counts, and you are not sure which is honest.
  • A qPCR or spike-in readout says global RNA output dropped, but your RNA-seq shows most genes unchanged.

Why it happens

A sequencer reads a fixed amount of material per library. Counts are therefore relative: a gene's count depends on how much of the library belongs to every other gene. If a few very abundant genes go up, everything else gets a smaller share of reads and looks down, even though its true abundance did not change. That is composition bias, and it is the reason raw counts cannot be compared across samples.

Median-of-ratios (DESeq2) and TMM (edgeR) both fix composition bias by estimating a per-sample scaling factor from the genes assumed not to change. Median-of-ratios takes the geometric mean per gene across samples, divides each sample by it, and takes the median ratio per sample. TMM trims the extreme log-ratios and averages the rest. Both are anchored to the same idea: the bulk of the transcriptome is stable. Under typical conditions they behave similarly, and the literature shows they are equivalent under specific conditions. Median-of-ratios is also tolerant of up versus down imbalance, as long as most genes are flat.

The assumption fails when the biology is a global shift. MYC activation amplifies transcription across the genome. A failed rRNA depletion changes what fraction of reads are informative. Strong cell-type composition changes in a tissue move large blocks of genes together. No algorithm can recover the absolute scale from the data alone: it is impossible to normalize for global shifts unless the data includes spike-ins or control genes that should not change.

The second source of trouble is mixing outputs. DESeq2 and edgeR do not test normalized counts. They model raw counts and put the normalization into the GLM as an offset, because dividing counts by a factor would distort the mean-variance relationship the count model needs. CPM and TPM are already divided and already throw away the count scale. They are good for plotting and for filtering, not for the test. TPM adds gene length on top, which matters for comparing genes within a sample and is irrelevant to composition correction.

The checks

Run them in order. Each one tells you what healthy looks like and what the problem looks like.

0/6 checked · saved in this browser

  1. Before any model, check that the matrix going into DESeq2 or edgeR holds integer counts from the quantifier, not TPM, CPM, FPKM or a batch-corrected table. With salmon, import with tximport to get gene-level counts rather than reading the TPM column. Non-integers are the giveaway.

    r
    stopifnot(all(counts_mat == round(counts_mat)))
    range(colSums(counts_mat))  # library sizes in the expected range for your design
    Healthy
    Every value is a non-negative integer (tximport estimated counts are rounded by DESeqDataSetFromTximport), and column sums look like the depth you ordered.
    Red flag
    Decimals, values that sum to about one million per column (TPM or CPM), or negative values (a batch-corrected log matrix). Stop and go back to the raw counts.
  2. Estimate size factors and compare them with raw column sums. They should track depth loosely but not identically. Print them next to the condition labels so you can see whether one group is systematically scaled.

    r
    dds <- DESeqDataSetFromMatrix(countData = data, colData = meta, design = ~ sampletype)
    dds <- estimateSizeFactors(dds)
    data.frame(sf = sizeFactors(dds), libsize = colSums(counts(dds)), meta)
    Healthy
    Size factors roughly proportional to library size, with no split between conditions beyond what depth explains.
    Red flag
    Size factors that separate cleanly by condition while library sizes are similar. Either composition really differs between groups or the assumption is failing, and the next checks tell you which.
  3. Run PCA on a variance-stabilized or log-CPM matrix. Color by condition, and separately correlate PC1 with total reads, number of detected genes and %rRNA. Depth artifacts show up here before they show up as false positives.

    Healthy
    PC1 or PC2 separates condition; the correlation of PC1 with library size, detected genes and %rRNA is weak.
    Red flag
    PC1 tracks total reads or %rRNA, or depth differs between conditions. DESeq2 and edgeR handle moderate depth differences but fail when depth is very low or confounded with condition, and no normalization choice repairs a confounded design.
  4. Draw the MA plot of your main contrast (plotMA on the DESeq2 results, or plotMD in edgeR). The bulk of the genes, especially the mid-to-high expression ones, should form a cloud centered on log fold change 0. Look at where the center sits, not at the number of significant dots.

    Healthy
    A symmetric cloud centered at zero, with significant genes spread above and below.
    Red flag
    The whole cloud sits above or below zero, or is tilted with expression level. That means the scaling factor was estimated from genes that are not stable, or the biology is a global shift.
  5. If you have spike-ins, plot their log fold change between conditions after your normalization: they should center on zero. Without spike-ins, take a set of genes you have independent reason to believe are stable (housekeeping genes are an imperfect proxy) and see whether they shift. A shift in spike-ins or trusted genes while the rest is called unchanged means the normalization absorbed a real global change.

    Healthy
    Spike-ins or trusted control genes center on zero fold change, and the transcriptome-wide result agrees with orthogonal readouts like total RNA yield or qPCR per cell.
    Red flag
    Controls move with the condition, or an orthogonal measurement says global RNA output changed while your table says nothing did.
  6. Run the contrast with both DESeq2 and edgeR (calcNormFactors with method TMM, or the default TMMwsp if you have many zeros). Compare the size factors and the overlap of significant genes. This is a sensitivity check, not a vote.

    r
    dge <- DGEList(counts = counts_mat, group = meta$sampletype)
    dge <- calcNormFactors(dge, method = "TMM")
    dge$samples$norm.factors
    Healthy
    The two methods give similar scaling and largely overlapping hit lists, as expected when both rely on the same assumption.
    Red flag
    Large disagreement in scaling points to unusual composition or many zeros. Disagreement between two methods that share an assumption does not tell you which is right; it tells you the assumption is under strain.

What to do about it

Use the right quantity for the right job

When: Always, as the default for a standard design where most genes are stable.

Give DESeq2 or edgeR raw integer counts and let the model handle normalization through offsets. Use DESeq2 normalized counts (counts(dds, normalized=TRUE)) or log-CPM for heatmaps and PCA. Use CPM or TPM for filtering low-count genes and for quick exploratory comparisons, and use TPM when you compare different genes within one sample. Do not run a test on any of those.

Caveat: Normalized counts are for visualization only. Plotting them is fine, but resist feeding them back into a count model.

Re-estimate the scaling from a control gene set

When: You have housekeeping genes or other genes you can defend as stable, but no spike-ins, and the MA cloud is off center.

Pass the control genes to edgeR so the TMM factors are computed from them: calcNormFactors(dge, method="TMM") accepts a controlGenes argument. In DESeq2, supply the control genes to estimateSizeFactors through controlGenes. Re-draw the MA plot and confirm the controls now sit at zero.

Caveat: Housekeeping genes are not guaranteed stable under a strong perturbation. If your treatment affects them, you have swapped one wrong assumption for another.

Raise the TMM trim for moderate global shifts

When: A sizeable fraction of genes moves in one direction, as in some transcription factor activation experiments, but not the whole genome.

Increase logratioTrim from its default of about 0.15 to 0.3, which lets up to roughly 30% of genes shift in one direction instead of 15%. Then recheck the MA plot.

Caveat: This stretches the assumption; it does not remove it. The packet does not give a validated rule for choosing the value, so treat it as a sensitivity analysis and report it.

Design the experiment with spike-ins

When: The biology is expected to change global RNA output (transcription inhibitors, MYC-driven systems, drug treatments) and absolute change is part of the question.

Add an external RNA spike-in to each sample in proportion to cell number or input, before library prep, and compute scaling factors from the spike-in counts. Spike-in normalized data can detect global shifts that standard methods miss.

Caveat: It has to be planned before the library is made; you cannot add it afterward. The evidence base is mostly single-cell work, and added spike-ins bring their own pipetting noise.

Fix the design before blaming the normalization

When: PC1 follows depth, rRNA fraction or batch, and that variable is entangled with condition.

Include batch in the model design rather than pre-correcting the matrix, and re-sequence or rebalance if depth is confounded with condition. See the batch effect post for when to correct and when not to.

Caveat: If condition and batch are perfectly confounded, no model can separate them.

When not to "fix" it

If your experiment is a standard comparison, such as knockout versus control in the same cell line with similar RNA content, and the MA plot is centered with similar scaling from median-of-ratios and TMM, leave the defaults alone. Adding control genes, widening the trim or switching methods on a healthy dataset only adds degrees of freedom to tune toward a result you like. Likewise, a skew in direction is not automatically a bug. A strong transcription factor can really push many genes up, and median-of-ratios tolerates an up versus down imbalance as long as most genes stay flat. Check against an orthogonal readout before concluding the normalization failed.

Five things experienced analysts do here

  1. Write down, before you open the data, whether you expect a global change in RNA output. If yes, plan spike-ins now, because the data alone cannot tell you afterward.
  2. Treat the MA plot center as the primary normalization diagnostic. The significant genes are the least informative part of it.
  3. Keep two matrices in your project and label them: raw integer counts for testing, and normalized values for plotting. Never let one overwrite the other.
  4. Filter low-count genes with CPM if you like, but run the test with the GLM and its own normalization. edgeR's manual makes exactly this split.
  5. When two methods that share an assumption agree, say so in your methods as a sensitivity check, not as proof. Agreement cannot detect a global shift.

Questions people ask

Is TMM or DESeq2 median-of-ratios better for differential expression?

For most designs they perform similarly and are equivalent under specific conditions, so pick the tool that fits your design and replicate number. Edge cases are shared: both assume most genes do not change. Choose between them on modeling features, not on normalization.

Can I use TPM or CPM as input to DESeq2 or edgeR?

No. DESeq2 needs raw unnormalized counts, because pre-normalized values discard the count scale it uses to estimate dispersion. Use tximport to get gene-level counts from salmon, and keep TPM for plots and within-sample comparisons.

Which normalization should I use for a heatmap?

Use DESeq2 normalized counts or log-transformed CPM, then scale genes across samples. TPM is acceptable for comparing genes within a sample. These values are for display only, not for the statistical test.

Does TMM correct for gene length?

No, and it does not need to. Gene length does not enter the TMM calculation and is irrelevant to composition bias correction. Gene length matters only when comparing different genes within one sample, which is where TPM applies.

What do I do if my treatment changes transcription globally?

Standard normalization cannot fix it, because it is impossible to normalize for global shifts without spike-ins or control genes. Plan external spike-ins, or use a defensible control gene set and report it as an assumption. Raising TMM's logratioTrim allows more one-directional change but is only a partial fix.

Related pages

Related reading on the blog

Sources

  1. hbctraining DGE Workshop: Count Normalization — DESeq2 uses raw counts with offsets; normalized counts for visualization; tolerance to up/down imbalance.
  2. Bioconductor csaw Book: Normalizing for Technical Biases — TMM and composition bias.
  3. Bioconductor Support: DESeq2 vs edgeR Normalization Approaches — Offsets inside the GLM rather than dividing counts.
  4. RNASeq Normalization for Samples with Global Shift in Gene Expression — Global shifts need spike-ins or control genes; logratioTrim adjustment; controlGenes.
  5. In Silico Comparison of TMM, RLE, and MRN Normalization Methods (PMC5025571) — TMM and RLE equivalence under specific conditions.
  6. Bioconductor Support: Starting DE Analysis from TPM Data — Why TPM cannot be used with DESeq2.
  7. Spike-in Normalization for Single-Cell RNA-seq — Spike-ins detect global transcriptional shifts missed by standard methods.

Part of the Normalization choice series.