Glossary · Statistics, Artifacts and Pitfalls
P-value histogram
The one plot that tells you whether your differential expression test is measuring biology or measuring its own bugs.
By Ming "Tommy" Tang, Director of Bioinformatics in Big Pharma · Reviewed September 2026 · 3 min read
Definition
A p-value histogram is a plot of the raw p-values from many simultaneous hypothesis tests, typically one per gene in a differential expression analysis, binned into a distribution. Under the global null hypothesis, p-values are uniform on [0,1] by construction, so a well-calibrated test with no true effects produces a flat histogram. A well-powered test with real biological signal instead produces a sharp peak near zero (the true positives) sitting on top of a flat floor (the true negatives). Shapes that deviate from those two patterns, a spike at 1, a U-shape, or a hill in the middle, indicate a problem with the test or the data rather than biology.
You meet the p-value histogram right after results() finishes in DESeq2, or right after any tool that runs thousands of hypothesis tests at once, one per gene, peak, or feature. It's the fastest check of whether your statistical model is behaving before you ever look at a volcano plot or an FDR-adjusted gene list.
Skip it and you inherit whatever is broken upstream: bad normalization, a misspecified design formula, low-count noise, or a test that's silently anti-conservative. None of those show up clearly in a volcano plot or an MA plot. They show up as a shape in this one histogram.
Why it matters
FDR and q-value correction assume the null p-values are uniform. If your histogram is U-shaped or hill-shaped because of low-count discreteness in DESeq2, that assumption is violated, and the adjusted p-values coming out the other end are unreliable even though the pipeline runs without error. This is a silent failure: nothing crashes, results() returns a table, and it's easy to hand that table to someone else believing the FDR column means what it says.
The documented fix for the most common case is boring and effective: filter out genes with fewer than 5 raw counts in at least 10% of samples before calling DESeq2, then re-run. Skipping this step because the pipeline "worked" is how underpowered, low-count genes end up inflating or corrupting your gene list.
Where people get it wrong
People see a peak near zero and stop looking, treating it as confirmation of "lots of DE genes" without checking the rest of the histogram. A genuine signal looks like a peak near zero on top of a flat floor across the rest of [0,1]. A hill-shaped or globally left-skewed histogram with no flat floor is not stronger evidence of more DE genes, it's a sign the test is anti-conservative: it's generating excess small p-values regardless of whether the null is true, usually from low-count discreteness, unmodeled correlation between replicates, or double dipping (clustering cells, then testing for differences between the clusters you just defined). Threshold tests like treat/glmTreat are the other trap: their null hypothesis is an interval, not a point, so a left-skewed histogram with density piled toward larger p-values is expected by design, not evidence of a broken test.
A concrete example
After running DESeq2, pull the raw p-values out of the results table and histogram them before looking at anything else, including the MA plot or the FDR column. A good histogram from real RNA-seq data shows a sharp peak near 0 (the DE genes), a roughly flat floor across the rest of the range (the non-DE genes), and a small peak right at 1 from low-count discreteness. If you instead see a U-shape or a hill in the middle, filter low-count genes and re-run before trusting the FDR values.
library(DESeq2)
dds <- DESeqDataSetFromMatrix(countData = counts, colData = coldata, design = ~ condition)
# fix for U-shaped/hill-shaped histograms: keep genes with >=5 counts in >=10% of samples
keep <- rowSums(counts(dds) >= 5) >= 0.1 * ncol(dds)
dds <- dds[keep, ]
dds <- DESeq(dds)
res <- results(dds)
hist(res$pvalue, breaks = 50, col = "grey", main = "p-value distribution", xlab = "p-value")Related terms
Questions people ask
- What does a flat p-value histogram mean?
A perfectly flat histogram across [0,1] means none of your null hypotheses were rejected, no genes show a detectable effect. That's the expected shape when the null is true everywhere, since p-values are constructed to be uniform under the null. It's not a bug, but if you expected a strong biological effect and got a flat line, check your design formula and sample labels before concluding there's no signal.
- Why is my p-value histogram U-shaped?
U-shapes in DESeq2 output are almost always a low-count artifact: genes with very few reads produce discrete p-value distributions that pile up near both 0 and 1. The documented fix is to filter genes before running DESeq2, keeping only those with raw counts of 5 or more in at least 10% of samples. Re-run the test after filtering and the U-shape typically resolves into a flat floor with a peak near zero.
- What does anti-conservative mean for a p-value histogram?
An anti-conservative test produces more small p-values than it should even when the null hypothesis is true, so the histogram is inflated near zero without a genuine biological reason. Causes include unmodeled correlation between samples, too few replicates, or double dipping (clustering cells and then testing for differences between the clusters you just defined). The downstream damage is that FDR and q-value estimates, which assume a mostly-uniform null, understate the real false discovery rate.
- How is a p-value histogram different from a volcano plot?
A volcano plot shows effect size against significance for every gene, one point per gene, and is a results plot. A p-value histogram collapses all the raw p-values into one distribution and is a diagnostic plot: it tells you whether the test itself is well calibrated before you trust anything the volcano plot shows you.
- What causes a hill-shaped p-value histogram after DESeq2?
A hill shape, where the middle of the distribution bulges instead of staying flat, comes from the discreteness of very low count data feeding into the negative binomial test. Removing low-count rows, the same filtering used to fix U-shapes, usually flattens the hill back to the expected shape.
Related pages
Related reading on the blog
Sources
- p-Value Histograms: Inference and Diagnostics — Diagnostic shapes: flat, left-skewed, U-shaped, hill-shaped, and binomial thresholds for meaningful deviation
- Understanding p value, multiple comparisons, FDR and q value — Why p-values are uniform under the null and how FDR/q-value correction builds on that
- Are published RNA seq data analyses often wrong in calculating p-values and FDR? — The expected good-shape histogram: sharp left peak, uniform floor, peak at 1 from low-count discreteness
- U-shape of p-value histograms is removed by filtering for genes that have >5 counts in >10% of samples — The specific DESeq2 filtering fix for U-shaped histograms
- Hill shape p-value histogram after differential miRNA expression analysis with DESeq2 — Hill-shape cause and fix via low-count filtering