Sanity check · CUT&RUN and CUT&Tag
How to Choose a Normalization Method in CUT&RUN and CUT&Tag
Your treated sample has a fraction of the control's histone mark, and your bigWigs look identical: the normalization you picked erased the biology.
By Ming "Tommy" Tang, Director of Bioinformatics in Big Pharma · Reviewed October 2026 · 6 min read
You treated cells with an inhibitor, ran CUT&RUN for a histone mark in treated and control, and made bigWigs with the default bamCoverage settings. The tracks look the same. Or you were handed a pipeline that "normalizes", and you never checked what to. Either way, the number on the y-axis of every heatmap and aggregate plot comes from one decision made early, and most people make it without noticing.
The stake is simple. Library-size normalization (CPM, RPKM, RPGC, BPM) forces each sample to the same total signal. If the biology is a global gain or loss of a mark, or a change in cell-type composition, that forcing hides the effect or invents the opposite one. A differential result built on the wrong scaling is wrong in a way no p-value will reveal.
This page lets you decide in an hour: which normalization fits your question, how to confirm your E. coli spike-in is usable, how to see whether the choice changes your conclusion, and which methods must never be fed into the same downstream tool.
What it looks like when it's happening
- Aggregate profiles over TSSs or peaks from treated and control overlap almost perfectly, although the antibody target is expected to drop globally (for example H3K27me3 after EZH2 inhibition).
- Heatmaps from two conditions look equally intense because each bigWig was scaled to the same total read count, and the only differences are in shape, not height.
- Most peaks appear 'up' in one condition and a handful appear 'down', a lopsided volcano that points to a scaling offset rather than binding changes.
- The fraction of reads aligning to E. coli differs by several-fold between replicates of the same condition.
- Nearly zero E. coli reads in a CUT&Tag library, so the spike-in scale factor is undefined or enormous.
- Normalized bigWigs from different pipelines (RPGC from one collaborator, spike-in scaled from another) are placed in the same heatmap and the color scales are not comparable.
- A count matrix of CPM or TPM values was passed to a differential tool that expected raw counts.
Why it happens
Every library-size method rests on one assumption: most of the genome does not change between samples, so total reads are a fair yardstick. CPM, RPKM, RPGC and BPM all divide by depth in some form. DESeq2's median-of-ratios and edgeR's TMM are smarter, but they anchor on the same idea: a stable majority of regions. When that holds, they work well. When it fails, they do exactly the wrong thing.
It fails in CUT&RUN more often than people expect, because the assay is used to test perturbations of chromatin itself. An EZH2 inhibitor can reduce H3K27me3 across the genome. A degron on a broadly bound factor can remove it from most of its sites. Cell-type composition changes do the same through a different route. Scale each library to the same total and the loss disappears, while the few regions that stayed put are pushed up and look gained.
Spike-in normalization breaks this dependency. The E. coli DNA carried by the CUT&RUN fusion protein is external to the perturbed biology, so the ratio of endogenous to spike-in reads reflects how much target was recovered per unit of spike-in, not how the sample compares to the rest of its own library. That is why it is the defensible choice for cross-sample comparison. It only works if the spike-in is present at a usable level (about 1% of reads, with 0.2 to 5% acceptable) and was added consistently, which is a bench decision that cannot be added to already-sequenced libraries.
CUT&Tag is different. pAG-Tn5 is highly purified and depleted of residual E. coli DNA, and exogenous spike-in has not been optimized for this assay. So the built-in spike-in that makes CUT&RUN comparisons honest is not available, and the field has not settled what to use instead. Finally, the methods answer different questions. CPM and TPM are for display and within-sample or same-gene comparisons, median-of-ratios and TMM are for differential testing on raw counts, and SCTransform belongs to single-cell data. Mixing their outputs is a category error, not a style choice.
The checks
Run them in order. Each one tells you what healthy looks like and what the problem looks like.
0/7 checked · saved in this browser
Align each library to the reference genome and, separately, to E. coli K12 MG1655. Filter multi-mapping reads and ENCODE DAC exclusion regions before counting, as the spike-in workflow requires. Then compute the fraction of total reads that map to E. coli for every sample. Use
samtools flagstaton each BAM for the raw numbers.bashfor s in sample1 sample2 sample3; do echo $s samtools flagstat ${s}.hg38.bam | head -n 7 samtools flagstat ${s}.ecoli.bam | head -n 7 done- Healthy
- E. coli reads sit near 1% of total reads in each CUT&RUN library, within the 0.2 to 5% acceptable range, and the fraction is similar across replicates of the same condition.
- Red flag
- Spike-in fraction well under 0.2% (noisy scale factor) or well above 5% (spike-in is eating your depth), or a several-fold difference between replicates of one condition. In CUT&Tag, near-zero E. coli is expected and means spike-in scaling is not an option.
State in one sentence whether you are (a) displaying signal for a figure, (b) comparing a mark or factor between conditions where a global change is possible, or (c) testing differential binding per region. (a) can use depth-based scaling, (b) needs spike-in in CUT&RUN, and (c) needs raw counts fed to a count-based tool. Check the biology: does your perturbation plausibly change the target globally?
- Healthy
- The method you pick maps to one of these questions, and the global-change question is answered from the experimental design, not from the data.
- Red flag
- You are using one normalized matrix for display, differential testing and clustering. Or the perturbation targets the mark itself and you are still using library-size scaling.
Extract the fragment length (TLEN) from properly paired reads and plot a histogram per sample. Histone marks should show nucleosomal peaks near 180 bp and multiples. Transcription factors give mostly nucleosome-sized fragments plus variable amounts of shorter ones. Paired-end data is required for this; single-end cannot do it.
bashsamtools view -f 66 -q 30 sample1.hg38.bam | awk '{t=$9; if(t<0) t=-t; if(t>0 && t<1000) print t}' | sort -n | uniq -c > sample1.fraglen.txt- Healthy
- Fragment-size profiles have the same shape across samples of the same target, with nucleosome-spaced peaks about 180 bp apart for histone marks.
- Red flag
- One sample has a different size profile from its replicates. Scaling by any factor cannot repair a library that was cut or tagmented differently.
For each sample compute the factor implied by total mapped reads and the factor implied by spike-in reads, then put them side by side in a table. If the spike-in-derived factors differ between conditions in a direction that depth-based scaling would erase, your question is the global-change case.
rqc <- data.frame( sample = c("ctrl1","ctrl2","trt1","trt2"), endo = c(endo_ctrl1, endo_ctrl2, endo_trt1, endo_trt2), spike = c(spk_ctrl1, spk_ctrl2, spk_trt1, spk_trt2) ) qc$endo_to_spike <- qc$endo / qc$spike qc$spike_frac <- qc$spike / (qc$endo + qc$spike) qc- Healthy
- Endogenous-to-spike-in ratios agree between replicates and differ between conditions only where biology predicts it.
- Red flag
- Replicates disagree on the ratio (a spike-in dosing problem), or the ratio is flat when you expected a large global change and the antibody enrichment is also unconvincing.
Generate one bigWig per sample with depth normalization and one with the spike-in scale factor, using
bamCoverage. Load both sets in a genome browser at a few positive-control regions and plot an aggregate profile for each. Note that--normalizeUsingoptions are RPKM, CPM, BPM, RPGC and None; for spike-in scaling useNonewith your own--scaleFactor.bash# depth-based, 1x coverage bamCoverage -b sample.bam --normalizeUsing RPGC --effectiveGenomeSize [size] -o sample.rpgc.bw # spike-in scaled: no built-in normalization, your factor from E. coli reads bamCoverage -b sample.bam --normalizeUsing None --scaleFactor [spikein_factor] -o sample.spikein.bw- Healthy
- The two sets agree when no global change is expected. When a global change is expected, the spike-in set shows it and the depth-based set does not.
- Red flag
- The two approaches give opposite directions for the same regions. That is the signature of a global shift hidden by library-size scaling, or of a bad spike-in factor. Resolve it with Checks 1 and 4 before trusting either.
Open the matrix you pass to DESeq2 or edgeR. Values must be integer counts per region, with normalization handled inside the tool by median-of-ratios or TMM. CPM, TPM, RPKM and bigWig-derived values are not valid inputs. If spike-in is in play, apply the factor through the tool's size-factor mechanism rather than feeding scaled values.
- Healthy
- Integer count matrix, one column per sample, no pre-normalization.
- Red flag
- Decimal values, columns that all sum to the same total, or a matrix built from bigWigs. These have been scaled once already and the tool will scale them again.
If you ran the nf-co.re CUT&RUN pipeline, look at
--normalisation_mode. The default isSpikein, with a normalization constant of 10,000 and a 50 bp bin size. Check the run log to confirm the spike-in step found enough E. coli reads. For CUT&Tag data, this default is the wrong fit and should be changed deliberately.bashnextflow run nf-co.re/cutandrun --normalisation_mode Spikein --normalisation_binsize 50- Healthy
- The mode you ran matches the assay and the spike-in counts from Check 1.
- Red flag
- Spike-in mode applied to CUT&Tag libraries, or to CUT&RUN libraries with almost no E. coli reads, giving unstable scale factors across samples.
What to do about it
Scale CUT&RUN signal by E. coli spike-in reads
When: You have CUT&RUN data with a usable E. coli fraction and a question where global change is possible: histone mark loss, global factor redistribution, or perturbations that alter composition.
Align reads to the reference and, separately, to E. coli K12 MG1655. Filter multi-mapping reads and ENCODE DAC exclusion regions, then derive each sample's scale factor from its spike-in read count (nf-core uses a normalization constant of 10,000 and does this by default). Apply it as --scaleFactor in bamCoverage with --normalizeUsing None, or let the pipeline do it. Filter alignments first with samtools view -F 1804 -f 2 -q 30 and duplicate removal.
Caveat: Only valid if the spike-in was added consistently at the bench. It cannot be added retroactively to libraries already sequenced, so a bad dosing means a repeat experiment.
Use depth-based scaling for display only, and say so
When: You need a quick visual of enrichment or a peak-centric heatmap and the biology does not imply a global change, or spike-in is unavailable.
Run bamCoverage with --normalizeUsing RPGC --effectiveGenomeSize [size] for 1x coverage across samples, or CPM or BPM if that suits the figure. Label the figure axis with the scaling used. Do not use the same values for differential testing.
Caveat: These methods assume most signal is constant. For global loss or gain they flatten the very thing you are looking for.
Test differential binding on raw counts with median-of-ratios or TMM
When: You are asking which regions differ between conditions and you can justify that most regions are stable, or you provide spike-in-derived size factors.
Count fragments per consensus region, pass the integer matrix to DESeq2 (median-of-ratios) or edgeR (TMM), and let the tool normalize. Use CPM or TPM only for plots, never as input.
Caveat: Both anchor on a stable majority of regions. In a global-change design with no spike-in, they will mislead.
For CUT&Tag, plan around the missing spike-in
When: Your assay is CUT&Tag, where pAG-Tn5 is depleted of E. coli DNA and exogenous spike-in is not optimized.
Treat comparisons as depth-normalized and state the assumption of a stable majority explicitly. Include replicates and orthogonal controls, such as a region or mark you expect unchanged, to test the assumption. If a global change is central to the study, consider whether CUT&RUN with its built-in spike-in fits the question better.
Caveat: The field has not settled a standard for CUT&Tag. Without a trusted external reference, global changes in absolute signal remain hard to claim.
Keep RNA-seq and single-cell normalizations out of this workflow
When: You are tempted to reuse TPM, TMM, or SCTransform logic on a CUT&RUN matrix because the tool exists and the question sounds similar.
Match method to data type. CPM, TPM and RPKM/FPKM are RNA-seq quantities, median-of-ratios and TMM belong to count-based DE, and SCTransform is for single-cell counts. For genomic coverage use the bamCoverage modes. Never mix outputs from different methods in one heatmap or one downstream model.
Caveat: A single-cell CUT&Tag analysis may legitimately use a single-cell method, but then it is a single-cell question, with its own checks.
When not to "fix" it
If your perturbation does not plausibly change the target genome-wide, for example a knockout of an unrelated gene probed with a mark that is stable across conditions, depth-based scaling or median-of-ratios is fine and spike-in adds noise from dosing variation. Likewise, if your spike-in fraction is far below 0.2% or varies several-fold between replicates, a spike-in scale factor is worse than none: do not apply it, fix the bench step and rerun. Also leave CUT&Tag alone as a spike-in problem. The missing E. coli signal is a property of the enzyme, not a defect in your library.
Five things experienced analysts do here
- Decide the normalization at the bench. Spike-in cannot be bolted onto sequenced libraries, and the standard CUT&RUN protocol recommends 0.5 to 1 ng of E. coli DNA for 500,000 cells, so plan the dose before the run.
- Make every figure state its scaling in the axis label or legend: 'RPGC', 'spike-in scaled', or 'CPM'. A heatmap without a stated normalization cannot be compared to any other heatmap.
- Always generate both a depth-normalized and a spike-in-normalized track for a few control loci. If they disagree, you have found either a global shift or a broken spike-in, and both are worth knowing before a collaborator sees the figure.
- Keep one raw count matrix on disk and derive every normalized view from it. Normalizing a matrix that was already normalized is the most common way to end up with numbers nobody can reproduce.
- When a result flips between DESeq2 and edgeR, or between depth and spike-in scaling, treat it as information about your stable-majority assumption rather than a reason to pick the version that fits the hypothesis.
Questions people ask
- Can I use TMM or DESeq2 normalization for CUT&RUN?
Yes, for differential binding on raw per-region counts, provided most regions are stable between conditions. Both assume a stable majority and fail when a mark is lost or gained globally. In that case use spike-in-derived scaling instead.
- Is CPM the same as TPM for CUT&RUN?
No. CPM corrects only for sequencing depth, TPM also corrects for feature length, and both are RNA-seq quantities. For genomic coverage use the
bamCoveragemodes, and never use CPM or TPM values as input to differential tools.- Can I use E. coli spike-in normalization for CUT&Tag?
Not reliably. pAG-Tn5 is highly purified and depleted of residual E. coli DNA, and exogenous spike-in has not been optimized for the assay. Use depth-based scaling with an explicit stable-majority assumption and orthogonal controls.
- How much E. coli should be in my CUT&RUN library?
About 1% of total reads, with 0.2 to 5% acceptable. Much lower makes the scale factor noisy, much higher costs endogenous depth. Check the fraction per sample and compare replicates before applying it.
- Which normalization should I use for a CUT&RUN heatmap?
If you compare conditions where a global change is possible, use spike-in scaled bigWigs. Otherwise RPGC or CPM from
bamCoverageis fine for display. Label the figure with the method either way.
Related pages
- Guide · How to Choose a Normalization Method in ATAC-seq
- Guide · How to Choose a Normalization Method in ChIP-seq
- Glossary · ChIP-seq
- Glossary · Count matrix
- Glossary · Heatmap
Sources
- Normalization Choice in ChIP-seq: Which Method to Pick — Spike-in vs library-size assumptions, E. coli fraction targets, CUT&Tag incompatibility, spike-in needing bench planning.
- bamCoverage documentation (deepTools 3.5.6) — The RPKM, CPM, BPM, RPGC and None normalization modes and their formulas.
- nf-co.re CUT&RUN Pipeline Parameters — Default spike-in mode, normalization constant 10,000 and 50 bp bin size.
- Benchmarking Peak Calling Methods for CUT&RUN — Why ChIP-seq defaults behave differently on low-background CUT&RUN data.
- CUT&RUNTools 2.0: Pipeline for CUT&RUN and CUT&Tag Data Analysis — Fragment-size distributions and the requirement for paired-end sequencing.
Part of the Normalization choice series.