▸ Chatomics Field GuideWhat They Don't Teach You →

Glossary · Statistics, Artifacts and Pitfalls

Q-value

The number that tells you what fraction of your "significant" genes are probably noise, and why it is not just a fancier p-value.

By Ming "Tommy" Tang, Director of Bioinformatics in Big Pharma · Reviewed September 2026 · 3 min read

Also: qvalue

Definition

A q-value is the minimum false discovery rate (FDR) at which a given test would still be called significant. Formally, the q-value of a feature is the FDR of the largest results list that includes that feature, so it is a statement about a set of calls, not about one test in isolation. Computing it requires the full distribution of p-values from all tests, plus an estimate of π₀, the proportion of tests where the null hypothesis is actually true, then thresholding on q(t) = mπ̂₀α̂ / #{T≥t}. In R, Storey's original implementation of this is the Bioconductor qvalue package, which has shipped since Bioconductor 1.6.

You meet the q-value the moment you stop testing one gene and start testing ten thousand at once, which is every RNA-seq, ChIP-seq, or GWAS experiment you'll ever run. A raw p-value answers a question about a single test in isolation. Once you're scanning a whole gene list for hits, that question stops being useful, because even under the null you expect hundreds of p-values below 0.05 by chance alone.

The q-value answers the question you actually care about at that point: if I draw the line here and call everything above it significant, what fraction of that list is garbage? Whether you set your differential expression cutoff at q < 0.05 or q < 0.01 decides how many false leads your wet-lab collaborators chase down after your analysis.

Why it matters

Test 10,000 genes at p < 0.05 with no correction and you expect roughly 500 false positives from chance alone, because p-values are uniformly distributed under the null. Bonferroni fixes that by demanding p < 0.05/10,000, which is so strict it throws out most true hits along with the noise. Q-values sit between those extremes: if you control at q < 0.05 and call 100 genes significant, you expect about 5 of them to be false, not 500 out of 10,000 and not zero.

Get the cutoff wrong in either direction and the cost lands downstream. Set it too loose and your collaborators spend months validating genes that were never real. Set it too tight, or use Bonferroni out of habit, and you bury the handful of true hits that would have justified the experiment. The choice of cutoff is itself a judgment call, not a fixed number.

Where people get it wrong

People use "q-value" and "adjusted p-value" interchangeably, but they are not always the same computation. DESeq2's padj column, by default, is a Benjamini-Hochberg adjusted p-value, which assumes pi0 = 1 (every null is true). Storey's q-value instead estimates pi0 from the data, which is usually less than 1 in a real experiment with true differentially expressed genes, so a proper q-value is typically less conservative than BH padj at the same nominal level. Calling both "the q-value" hides a real methodological difference.

The second trap is treating q-values as portable across subsets. A q-value is only valid for the exact list of tests it was computed from. If you filter, re-rank, or cluster your data first and then compute q-values on a subset, you've broken the FDR guarantee, the same double-dipping problem that inflates p-values in single-cell cluster-then-test workflows.

A concrete example

The qvalue package ships with the Hedenfalk et al. p-values as a worked example. You feed it a vector of p-values from many tests, and it returns a q-value for every one, plus an estimate of pi0 you can sanity-check before trusting the results.

r
if (!require("BiocManager", quietly = TRUE)) install.packages("BiocManager")
BiocManager::install("qvalue")

library(qvalue)
data(hedenfalk)
qobj <- qvalue(hedenfalk$p)

summary(qobj)          # check pi0 estimate and q-value counts by threshold
hist(qobj)              # p-value histogram + pi0 line: should be roughly flat near p=1
sig_genes <- which(qobj$qvalues < 0.05)
length(sig_genes)       # genes with FDR <= 5% if you call this whole list significant

Related terms

Questions people ask

What is a q-value in bioinformatics?

It's the minimum false discovery rate at which a test result would still be called significant, computed across the full set of p-values from an experiment. It tells you the expected fraction of false positives in the list of results you call significant at that threshold, not the chance any single test is a false positive.

Q-value vs p-value: what's the difference?

A p-value is the probability of seeing data this extreme under the null hypothesis for one test in isolation. A q-value is the expected proportion of false positives among all results at least as extreme as this one, computed jointly across every test in the experiment. You need thousands of p-values to compute a single meaningful q-value.

Is padj in DESeq2 the same as a q-value?

Not exactly. DESeq2's default padj is a Benjamini-Hochberg adjusted p-value, which assumes all null hypotheses are true (pi0 = 1). A true Storey q-value, from the qvalue package, estimates pi0 from the data instead, which usually makes it less conservative than BH padj at the same nominal cutoff.

How do I compute q-values in R?

Install the Bioconductor qvalue package with BiocManager::install("qvalue"), then pass a vector of p-values to qvalue(). The returned object holds a qvalues vector, a pi0 estimate, and local FDR values via lfdr(); pi0est() and empPvals() are also available if you need those pieces separately.

What does q < 0.05 mean in practice?

It means that if you call every feature with q < 0.05 significant, about 5% of that resulting list is expected to be a false positive. It does not mean any individual gene has a 5% chance of being wrong; that's a common misreading carried over from p-value intuition.

Related pages

Related reading on the blog

Sources

  1. qvalue Bioconductor package — package functions, installation, and definition of q-value/local FDR
  2. StoreyLab/qvalue GitHub repository — pi0 estimator formula and alternative installation
  3. A statistical method for the conservative adjustment of false discovery rate (q-value) — FDR and q-value formulas, and caveat about underestimating true FDR
  4. Understanding p value, multiple comparisons, FDR and q value — p-value vs q-value distinction and Benjamini-Hochberg procedure