Glossary · Statistics, Artifacts and Pitfalls
P-value
A p-value tells you how surprising your data would be if nothing were going on, not whether your result is true or big enough to matter.
By Ming "Tommy" Tang, Director of Bioinformatics in Big Pharma · Reviewed September 2026 · 3 min read
Definition
A p-value is the probability, calculated under the assumption that the null hypothesis is true, of observing a test statistic as extreme as or more extreme than the one you actually got. It is a statement about the data given the null, not a statement about the null given the data: it does not tell you the probability that the null hypothesis is true, and it does not tell you the probability that your data arose from chance alone. In a typical RNA-seq comparison, the null hypothesis is "this gene's expression does not differ between conditions"; a p-value of 0.001 means a difference this extreme would be rare if that null held, nothing more.
You meet the p-value in the output of every differential expression test: the pvalue column next to log2FoldChange in a DESeq2 or edgeR results table, the association p-value in a GWAS summary statistic file, the t-test p-value comparing two groups of samples. It looks like a verdict. It is not one. It is a single number computed under a specific null hypothesis, and it decides nothing about biology on its own.
Where it actually decides something is in how you rank and filter results before spending real money: which genes go into a pathway enrichment analysis, which variant gets followed up with a validation assay, which "significant" hit gets written into the abstract. Get the interpretation wrong and you either bury a real effect under a strict threshold or chase a statistically real but biologically trivial one.
Why it matters
The consequence of misreading a p-value shows up as soon as you scale past one test. Test 20,000 genes for differential expression and roughly 1,000 of them will cross p < 0.05 by chance alone, even if there is no real difference between your conditions anywhere in the genome. Report the raw p-value column instead of the FDR-adjusted one, and you hand your collaborators a gene list that is mostly noise dressed up as discovery.
The same failure mode shows up in single-cell RNA-seq differential expression between clusters. Because you cluster cells and then test for differences between those same clusters using the same expression data, you are double dipping, and p-values collapse to values like 1e-20 for genes that separate the clusters almost by definition. Thousands of genes look "significant." Almost none of that significance is informative about biology; it is an artifact of testing on the same data you used to define the groups.
Where people get it wrong
The most common mistake is treating the raw p-value and the multiple-testing-adjusted p-value (FDR, padj, q-value) as interchangeable, then reporting whichever column happens to have more hits. They answer different questions: the raw p-value is about one test in isolation, FDR is about the expected proportion of false positives among everything you call significant across thousands of simultaneous tests. A close second is treating p = 0.05 as a bright line, since it came from R.A. Fisher's convention in the 1920s, not from any biological law, so p = 0.0499 and p = 0.0501 get treated as night and day when they are statistically almost identical. A third is chasing the smallest p-value as the "best" hit: p-values shrink automatically as sample size grows, so with 10,000 samples you can get p < 1e-5 for a difference too small to matter, while a gene with log2 fold change of 3 sitting at p = 0.06 gets thrown away.
A concrete example
Pull a DESeq2 results table and look at the raw p-value, the BH-adjusted p-value, and the effect size side by side, rather than filtering on one column alone. In a real run you will typically see genes with tiny raw p-values and near-zero log2 fold change (statistically detectable, biologically irrelevant) sitting right next to genes with p around 0.05-0.08 and a log2 fold change above 2 that are worth a second look before you discard them on a strict p < 0.05 cutoff.
res <- results(dds, contrast = c("condition", "treated", "control"))
res_df <- as.data.frame(res)
# padj is already BH-corrected inside results(); this just shows the math
res_df$padj_manual <- p.adjust(res_df$pvalue, method = "BH")
# rank by effect size, not just raw p-value
res_df[order(-abs(res_df$log2FoldChange)), c("log2FoldChange", "pvalue", "padj")][1:10, ]
# genes that clear FDR but have a trivial effect size
subset(res_df, padj < 0.05 & abs(log2FoldChange) < 0.5)Related terms
Questions people ask
- What is a p-value in simple terms?
It is the probability of seeing data as extreme as yours, or more extreme, if there were actually no effect. It tells you how surprising your result is under a 'nothing is happening' assumption, not whether that assumption is false or how big the effect is.
- Why is p < 0.05 used as the significance threshold?
Because R.A. Fisher proposed it as a convenient convention in the 1920s, and the field kept it. It is not derived from biology and does not mark a real boundary between 'true' and 'false': a result at p = 0.049 and one at p = 0.051 are practically indistinguishable.
- Is a smaller p-value more significant or more important?
Not necessarily. P-values shrink as sample size grows, so with enough samples even a trivially small, biologically meaningless difference can produce a very small p-value. Effect size, such as log2 fold change, tells you whether the difference is worth caring about; the p-value only tells you whether it was statistically detectable.
- What is the difference between a p-value and an FDR or q-value?
A raw p-value describes one test in isolation. When you run thousands of tests at once, as in differential expression across a genome, FDR (often reported as
padj) controls the expected proportion of false positives among everything you call significant, and a q-value gives each individual gene its own FDR-equivalent score. Use FDR-adjusted values, not raw p-values, whenever you are testing more than a handful of hypotheses at once.- Why does p-value misinterpretation matter in genomics specifically?
Genomics routinely tests thousands to millions of hypotheses at once (every gene, every variant, every cluster comparison), so the 5% false-positive rate baked into p < 0.05 translates into hundreds or thousands of false hits by chance alone. Single-cell differential expression between clusters compounds this further through double dipping, where clustering and testing use the same data and inflate significance beyond what multiple-testing correction alone can fix.
Related pages
Related reading on the blog
Sources
- Understanding p value, multiple comparisons, FDR and q value — FDR, Benjamini-Hochberg, q-value definitions and the 5%-false-positives-under-random-chance fact
- P-values in genomics: Apparent precision masks high uncertainty — Replication variability of p-values and the GWAS winner's curse
- Establishing an adjusted p-value threshold to control the family-wide type 1 error in genome wide association studies — Genome-wide significance threshold of p < 5x10^-8 in GWAS