▸ Chatomics Field GuideWhat They Don't Teach You →

Glossary · RNA-seq

Contrast (differential expression)

The model fits coefficients; the contrast is the question you ask of them, and a wrong reference level flips the sign of your biology without any warning.

By Ming "Tommy" Tang, Director of Bioinformatics in Big Pharma · Reviewed October 2026 · 2 min read

Also: results contrast

Definition

A contrast is a linear combination of the coefficients of a fitted model that defines one specific comparison between groups. The contrast vector is multiplied by the coefficient vector β to form the numerator of the test statistic, and the denominator comes from the covariance matrix of the coefficients. The design matrix says what each sample is; the contrast says which comparison gets a log2 fold change and a p-value. Think of the fitted model as a set of measurements and the contrast as the question you put to them.

You fit a model once. Then you ask it questions. Each question is a contrast: treated vs untreated, mutant vs wild type, or the harder one, whether a drug effect differs between two genotypes. You will meet the term in results() in DESeq2 and in makeContrasts() in limma and edgeR.

The contrast decides the sign of every log2 fold change in your table and which samples the p-values are about. A wrong contrast does not throw an error. It returns a clean-looking table that answers a different question than the one you think you asked.

Why it matters

The same fitted model can produce many different result tables, and only the contrast decides which one you get. If your factor levels are in the wrong order, "treated vs control" silently becomes "control vs treated", and every up-regulated gene is reported as down. Your pathway enrichment then tells the opposite story, and nothing in the output warns you.

Contrasts also decide what you can test at all. A difference of differences, such as a drug effect in one genotype compared with the drug effect in another, needs a contrast built across several coefficients. If you only run the default result, you answer a simpler question than the one your collaborator cares about.

Where people get it wrong

People treat the contrast as a label on the output, as if it only names the table. It is part of the statistics: it changes the estimate, the standard error and the p-value. A second confusion is coefficient vs contrast. With an intercept design, a coefficient like condition_treated_vs_untreated already is a comparison against the reference level, which is why results(dds, name=...) matches the equivalent contrast= call. The reference level is the first level of the factor, so if you never relevel, alphabetical order picks your control for you. Check levels() before you trust any sign.

A concrete example

A simple case in DESeq2: compare treated to untreated, by contrast or by coefficient name. Then a difference of differences using a cell means design with no intercept, where the numeric vector tests (D - C) - (B - A). The vector order must match the column order of the design, so print the coefficient names before you run it. The last lines show the limma and edgeR way, which always starts from a no-intercept design.

r
# simple pairwise comparison, treated vs untreated
res <- results(dds, contrast = c("condition", "treated", "untreated"))
# equivalent, by coefficient name
res <- results(dds, name = "condition_treated_vs_untreated")

# difference of differences with a cell means design
design(dds) <- ~ 0 + sex_group
dds <- DESeq(dds)
resultsNames(dds)   # confirm the column order is A, B, C, D
res_dod <- results(dds, contrast = c(1, -1, -1, 1))  # (D - C) - (B - A)

# limma / edgeR
design <- model.matrix(~ 0 + group)
cm <- makeContrasts(TvsU = groupTreated - groupUntreated, levels = design)

Related terms

Questions people ask

Why remove the intercept before defining contrasts?

With ~0 + group you get one coefficient per condition, so any comparison is a plain subtraction of named coefficients. With an intercept, the other coefficients are differences from the reference level, which makes complex contrasts easy to get wrong. The no-intercept form is the usual choice in limma and edgeR, and for difference-of-differences in DESeq2.

What is the difference between name and contrast in DESeq2 results()?

results(dds, name='condition_treated_vs_untreated') picks one fitted coefficient, and it is equivalent to the contrast=c('condition','treated','untreated') form for a simple two-level factor. Use contrast when the comparison is not a single coefficient, such as two non-reference levels or a numeric vector.

How do I know which group is the reference level?

It is the first level of the factor. Run levels() on the column, and use relevel() or factor(..., levels=) to put your control first before fitting. Do this before DESeq(), because the coefficient names depend on it.

What is a contrast in limma?

After lmFit() on a design matrix, you build a contrast matrix with makeContrasts(), naming the design columns to subtract, for example GroupA - GroupB. The levels argument ties the names to your design columns, so a typo in a group name fails loudly.

Related reading on the blog

Sources

  1. Analyzing RNA-seq data with DESeq2 — results() with contrast or name, equivalence of the two forms
  2. RNA-seq analysis is easy as 1-2-3 with limma, Glimma and edgeR — makeContrasts after lmFit in limma
  3. RNA-seq workflow: gene-level exploratory analysis and differential expression — reference level and releveling; design matrix works with contrasts
  4. DESeq2 how to specify contrast to test difference of differences — cell means design and numeric contrast vector
  5. edgeR design matrix and contrast questions — removing the intercept with model.matrix(~0+fac)
  6. DESeq2 design and contrast questions (Bioconductor support) — contrast vector times coefficients forms the test numerator