Sanity check · ChIP-seq
How to Read a P-Value Histogram in ChIP-seq
Flat with a spike, a hill, a U, or a pile-up near 1: each shape of your differential binding p-values points at a specific broken assumption, and you can see it in one line of R.
By Ming "Tommy" Tang, Director of Bioinformatics in Big Pharma · Reviewed October 2026 · 5 min read
You ran a differential binding analysis, got a table of peaks with p-values and adjusted p-values, and a volcano plot that looks fine. Before you hand over a list of "gained" and "lost" sites, plot the raw p-values as a histogram. It takes one line of R and it tells you whether the test the numbers came from was valid for your data.
If you skip it, the failure is quiet. A hill-shaped histogram means your model inflated variance, so you lose real sites. A U shape means something was filtered or corrected after testing. A bad input normalization produces thousands of fake differential peaks that come from copy-number differences, not from your factor. None of these throw a warning.
This page shows how to read the four shapes, how to separate differential-binding p-values from MACS2 peak p-values (they are different objects), and which check to run next for each shape. You can do all of it in the next hour on the count matrix you already have.
What it looks like when it's happening
- `hist(res$pvalue, breaks=20)` shows a hill: low near 0, a bump in the middle, and a thin spike at 0 or none at all
- The histogram is U-shaped, with a tall bar near 0 and another tall bar near 1
- Almost no p-values below 0.05 even though the IP clearly worked and the conditions should differ
- Thousands of peaks are significant, and they cluster in amplified regions or in open chromatin rather than at your factor's known targets
- A tall bar in the first bin that disappears or splits into spikes once you drop low-count peaks
- Replicate ChIP samples show uneven peak heights in a genome browser while the input tracks look nearly identical
- PCA of the consensus peak counts clusters samples by batch or library prep day instead of by condition
Why it happens
A p-value is uniform between 0 and 1 when the null hypothesis is true. In a differential binding test, most consensus peaks do not change between conditions, so most p-values should be uniform, and the few real changes add a spike near zero. The histogram is therefore a flat floor with a spike on the left. Any other shape means the test's assumptions do not match your data.
A hill, with a bump in the middle and a weak spike near zero, usually means the variance is overestimated. In ChIP-seq this often happens when peak heights are heterogeneous across replicates while the input controls are relatively homogeneous. The model cannot explain the ChIP variability with the input, so it inflates dispersion to compensate, and true differences lose power. Batch effects that overlap with your experimental groups, or unequal variance between groups, produce the same hill.
A U shape, with extra mass near 1 as well as near 0, is a conservative-test signature. The usual causes are filtering features after testing, or applying a correction such as fdrtool on top of values that were already adjusted. A plain pile-up near 1 with no spike at zero also points at overstated dispersion or wrong model assumptions, and the test has no power. Low-count peaks add artificial spikes to the histogram, which is why pre-filtering before testing gives a cleaner picture.
One distinction matters for ChIP-seq. MACS2 p-values come from a dynamic Poisson model of local background, used to decide which regions are peaks. They are biased by design and are not test p-values, so a histogram of them is not a diagnostic. The histogram diagnostic applies to the differential test you ran with DESeq2, edgeR or DiffBind. When input normalization fails, copy-number variation and open-chromatin bias show up as false differential peaks and distort that histogram.
The checks
Run them in order. Each one tells you what healthy looks like and what the problem looks like.
0/8 checked · saved in this browser
Use the
pvaluecolumn of the differential binding result, notpadjand not MACS2 output. Use about 20 breaks so the bin width is 0.05 and a uniform floor is easy to see. Look at the left edge, the middle and the right edge separately.rlibrary(DESeq2) dds <- DESeq(dds) res <- results(dds) hist(res$pvalue, breaks = 20, main = "P-value Histogram")- Healthy
- A flat floor across the whole range with a sharp spike in the first bin. The spike height depends on how many peaks truly change, and a flat floor with no spike can be fine when little changes.
- Red flag
- A hill (bump in the middle), a U (tall bars at both ends), or a slope with a pile-up near 1. Each points at the causes in the sections above.
Check where the numbers came from. MACS2
callpeakwrites a p-value and q-value per peak, using a Poisson model with local background, with a default q-value cutoff of 0.05 or a p-value cutoff of 1e-5 if you used-p. Those values select peaks. The differential test p-values come from the count model in DESeq2, edgeR or DiffBind.bashmacs2 callpeak -t treatment.bam -c control.bam -n sample_name -g hs -q 0.05- Healthy
- Your histogram input is a column of per-peak p-values from the differential model, one row per consensus peak.
- Red flag
- The histogram was made from a MACS2 `-log10(pvalue)` column or from peaks that were already thresholded. That histogram is skewed by construction and says nothing about your test.
Write down the order of operations: count matrix, filter low-count peaks, test, adjust. Low-abundance features create artificial spikes, so filter before testing. Filtering on the p-value or on adjusted p-values after testing produces the U shape.
- Healthy
- Filtering uses a criterion independent of the test result (for example, a minimum count across samples) and happens before the model is fit.
- Red flag
- The pipeline removes peaks after the test using any information from the result, or applies a second correction to adjusted values. A U shape follows.
Bin peaks by
baseMean(or average log CPM in edgeR) into a low and a high group, then plot the p-value histogram for each. This shows whether a spike or a bump lives only in the low-count peaks.rhi <- res$baseMean > median(res$baseMean, na.rm = TRUE) par(mfrow = c(1, 2)) hist(res$pvalue[hi], breaks = 20, main = "high count") hist(res$pvalue[!hi], breaks = 20, main = "low count")- Healthy
- Both halves look broadly similar, flat floor plus a spike near zero. The low-count half can be noisier.
- Red flag
- Odd spikes or a U shape appear only in the low-count half. Raise the pre-filter and rerun the test.
Run PCA on the variance-stabilized or normalized peak counts and color points by condition and by every technical variable you have: library prep date, sequencing lane, antibody lot, operator. A hill histogram with samples grouping by batch is the batch-overlaps-group case.
- Healthy
- Samples separate by condition on one of the first components, and technical variables are scattered.
- Red flag
- Samples cluster by batch, or batch and condition are confounded so you cannot separate them. This is the cause to fix before trusting any p-value.
In a genome browser, load normalized tracks for all ChIP replicates and all inputs at several consensus peaks. Compare how much peak height varies between ChIP replicates against how much the input signal varies. Heterogeneous ChIP with homogeneous input is the setting where the model inflates variance and gives a hill.
- Healthy
- Replicate ChIP peaks have similar heights within a condition, and input tracks do not show large copy-number steps or open-chromatin bumps that differ between samples.
- Red flag
- Replicate peak heights differ a lot while inputs look alike, or input tracks differ between conditions in copy number. Normalization against input is failing.
Open the signal at loci where your factor or mark must bind. For an estrogen receptor ChIP, for example, you expect peaks at genes like TFF1 or GREB1. A clean histogram from a failed ChIP only means the test is valid on data that carries no signal.
- Healthy
- Clear enrichment over input at the known positive-control loci in every IP replicate, and a differential result that moves in the direction the biology predicts at those loci.
- Red flag
- No signal at expected targets. The cause is a failed ChIP, a non-specific antibody, or a sample with no binding, and no statistical fix will help.
After one fix (a higher pre-filter, a batch term in the design, a corrected normalization), refit the model and plot the new histogram next to the old one. Change one thing at a time so you know which fix moved the shape.
- Healthy
- The hill flattens toward a uniform floor and the spike at zero becomes more distinct, or stays the same if the original was already healthy.
- Red flag
- The shape does not change, or it changes in a way that makes a different shape worse. Revert the last change and re-examine the design and normalization.
What to do about it
Pre-filter low-count peaks before testing
When: The histogram has odd spikes or a U shape and the low-count split in the checks shows the problem lives in low-abundance peaks.
Remove consensus peaks with very low counts across samples before you fit the model, using a rule that does not use the test result. Then rerun DESeq2 or edgeR and compare the histogram with the earlier one.
Caveat: The threshold is a judgment call, and the sources do not give a universal value. Set it by looking at the histogram, and do not use p-values to choose it.
Add a surrogate variable for hidden batch
When: PCA shows structure unrelated to condition and you do not have a recorded batch variable, or the recorded one does not explain the hill.
Estimate a surrogate variable with sva and add it to the design before testing, then refit and replot.
library(sva)
mod <- model.matrix(~ condition, data = colData(dds))
mod0 <- model.matrix(~ 1, data = colData(dds))
svseq <- svaseq(assay(dds), mod, mod0, n.sv = 1)
dds$sv1 <- svseq$sv[, 1]
design(dds) <- ~ sv1 + condition
dds <- DESeq(dds)
Caveat: If batch and condition are fully confounded, no surrogate variable can separate them. Also use a known batch term in the design rather than pre-correcting the count matrix when you have one.
Fix the input normalization
When: Peak heights vary between ChIP replicates while inputs are homogeneous, or false differential peaks sit in amplified regions or open chromatin.
Make sure every ChIP sample is paired with its matched input, and that the differential tool (for example DiffBind) is actually using the control. Check that input tracks do not show copy-number differences between conditions, then renormalize and refit.
Caveat: Without a matched input there is nothing to normalize against. In that case say so in the report, because the p-values cannot be fully trusted.
Remove post-test filtering and double correction
When: The histogram is U-shaped and your pipeline filters on results or applies a second correction such as fdrtool.
Reorder the pipeline so filtering happens before the test, and use only the correction built into the tool. Refit from the filtered count matrix and replot.
Caveat: Results will change. Peaks that looked significant after a double correction may no longer pass.
Revisit the experiment when the histogram is flat but the biology is wrong
When: The histogram looks healthy but known targets show no enrichment.
Treat it as a data generation problem. Validate the antibody against known biology, ideally with a knockout or a known positive locus, and repeat the ChIP if enrichment is absent.
Caveat: This costs bench time. A good histogram on a failed IP only proves the statistics ran correctly on noise.
When not to "fix" it
A flat histogram with a small or absent spike can be biology: if your perturbation changes binding at few sites, the spike is small, and forcing it bigger with looser filters or extra surrogate variables does harm. Likewise, if batch and condition are confounded by design, adding a batch term removes the signal you are testing, so report the limitation instead. A slightly lumpy histogram from a handful of replicates is expected noise. Fix the shape only when you can name the mechanism, such as a batch you can see in PCA or a filter you can find in the code, because tuning until the plot looks pretty is another way of fooling yourself.
Five things experienced analysts do here
- Plot the histogram before you look at any volcano plot or peak list, and save it next to the results so a reviewer can see it.
- Keep two columns in your head: MACS2 p-values pick peaks, differential-test p-values judge changes. Never put a MACS2 p-value column through this diagnostic.
- Write the order of operations (filter, test, adjust) in the script as comments. U shapes come from a step done in the wrong order.
- Pair every shape you see with one next check: hill goes to PCA and ChIP-versus-input variability, U goes to the pipeline order, pile-up near 1 goes to dispersion and model assumptions.
- Look at known positive-control loci in a genome browser before you trust any histogram, because a clean shape from a failed IP is still a failed IP.
Questions people ask
- What should a p-value histogram look like in ChIP-seq differential binding?
A flat floor across 0 to 1 with a sharp spike in the first bin. The floor comes from peaks that did not change, and the spike comes from peaks that did. The histogram must come from the differential test, not from MACS2.
- Why is my p-value histogram hill-shaped?
A hill usually means overestimated variance, from batch effects that overlap your groups or from heterogeneous ChIP replicates against homogeneous input. The model inflates dispersion to compensate, so power falls. Check PCA for batch and compare replicate peak heights to the input.
- What does a U-shaped p-value histogram mean?
A conservative test, often caused by filtering after testing or by applying a correction like fdrtool to values that were already adjusted. Reorder the pipeline so filtering comes before the test and use only one correction.
- Can I use the MACS2 p-values for this check?
No. MACS2 p-values come from a dynamic Poisson model of local background and are used to call peaks, so they are biased by design. Use the p-values from your DESeq2, edgeR or DiffBind differential test.
- My histogram is flat. Is my ChIP-seq experiment fine?
Not necessarily. A flat histogram means the test is behaving, not that the immunoprecipitation worked. Confirm enrichment at known target loci and validate the antibody before interpreting the result.
Related pages
- Guide · How to Read a P-Value Histogram in ATAC-seq
- Guide · How to Read a P-Value Histogram in Bulk RNA-seq
- Compare · DESeq2 vs edgeR: Which One Should You Use?
- Guide · How to Avoid Pseudoreplication in Single-Cell RNA-seq
- Guide · How to Call Peaks You Can Trust in ChIP-seq
- Glossary · ChIP-seq
- Glossary · Principal component analysis (PCA)
- Glossary · Count matrix
Related reading on the blog
Sources
- Hill in P-value distribution for heterogenous ChIP-seq samples — Hill-shaped histograms in differential ChIP-seq when input normalization fails and ChIP replicates are heterogeneous
- DESeq2 dispersion and p-value histograms — Flat, hill-shaped and U-shaped histogram interpretation, low-count filtering, and SVA for batch effects
- MACS2 peak calling documentation — MACS2 version 2.2.9.1, default q-value cutoff and callpeak command
- HBC Training: Peak calling with MACS2 — Default p-value threshold of 1e-5 and Benjamini-Hochberg q-values in peak calling
- DiffBind: Differential binding analysis of ChIP-Seq peak data — P-values are uniform under the null; histogram inspection as a QC step for differential binding
Part of the P-value histograms series.