▸ Chatomics Field GuideWhat They Don't Teach You →

Glossary · Craft and Good Practice

Exploratory data analysis (EDA)

The PCA plot, the boxplot, the raw scatter you make before any p-value exists is what tells you whether the p-value is even worth computing.

By Ming "Tommy" Tang, Director of Bioinformatics in Big Pharma · Reviewed September 2026 · 3 min read

Also: EDA

Definition

Exploratory data analysis (EDA) is the practice of examining a dataset through visualization and descriptive statistics, before any formal hypothesis test or model fit, to understand its structure, distributions, and relationships and to catch problems that would invalidate downstream analysis. In genomics this means things like plotting per-sample library size and mitochondrial percentage, running PCA or MDS on variance-stabilized expression values to check clustering and batch structure, and building a sample distance heatmap to spot swapped or duplicated samples. EDA produces no p-value and proves nothing formally; its output is a judgment call about whether the data is trustworthy enough to model.

You meet EDA the moment you have a count matrix and a design table and before you have run a single statistical test. It is the step where you plot library sizes, check percent mitochondrial reads, run PCA on variance-stabilized expression values, and look at a sample distance heatmap, all to answer one question: does this dataset look the way it should before I let a model tell me a story about it.

Skipping straight to differential expression or a fitted model is the single most common way an early-career analyst produces a result that looks clean and is wrong. EDA is not a formality before the "real" analysis. It is the analysis that decides whether the real analysis is trustworthy.

Why it matters

Get EDA wrong and the failure is silent: the differential expression pipeline runs to completion, DESeq2 or edgeR spit out a gene list with clean-looking p-values, and nothing in the output tells you that one sample was actually a mislabeled replicate or that half your "significant" genes are riding a batch effect that happens to correlate with condition. A single outlier sample can fake a strong correlation (r = 0.9 from one point, as the book's outlier lesson shows) or, in a small-n RNA-seq design, dominate a PC and make an entire treatment group look different when it's really one bad library.

The fix costs almost nothing relative to the damage of skipping it: a PCA plot colored by condition and by batch, a sample distance heatmap, and a look at raw QC metrics before fitting anything. If PC1 splits by sequencing batch instead of by biological condition, you have a confound to design around or explicitly model, not a result to publish. Catching that in five minutes of plotting is cheap; catching it after a collaborator has started validating a "hit" gene at the bench is not.

Where people get it wrong

The specific trap in RNA-seq workflows: DESeq2's variance-stabilizing transformation (vst()) or rlog() exists only to make PCA and clustering plots readable, because raw counts have Poisson-dominated noise at low expression that would swamp a PCA on its own. Analysts who don't know this sometimes carry the vst-transformed matrix forward and feed it into the differential expression test itself, or into a downstream tool that expects counts. The actual DE test has to run on raw counts through DESeq2's negative binomial GLM, not on the transformed values. If you see vst or rlog output going anywhere except a plot, that's the tell.

A second, quieter mistake: treating "I made the PCA plot the paper needs" as equivalent to having done EDA. A PCA plot made once, at the end, to satisfy a reviewer is not EDA. EDA is iterative and happens before decisions get made, not after, and it includes ugly non-graphical checks (raw counts per sample, duplication rate, percent mapped) that never make it into a figure but that would have caught the problem a beautiful PCA plot alone might not.

A concrete example

A 12-sample bulk RNA-seq experiment (2 conditions x 2 batches x 3 replicates) comes in as a DESeqDataSet. Before fitting the DE model, run the standard EDA pass: transform for visualization only, then look at PCA colored by both condition and batch.

r
library(DESeq2)

# vst() is for visualization ONLY -- the DE test below uses raw counts in dds
vsd <- vst(dds, blind = FALSE)

# color by the variable you expect to matter, and by the one you're afraid matters
plotPCA(vsd, intgroup = "condition")
plotPCA(vsd, intgroup = "batch")

# sample distance heatmap: catches swaps and duplicates PCA can miss
sampleDists <- dist(t(assay(vsd)))
pheatmap::pheatmap(as.matrix(sampleDists))

# only now, on the raw counts, run the actual test
dds <- DESeq(dds)
res <- results(dds, contrast = c("condition", "treated", "control"))

Related terms

  • Sanity check
  • Metadata
  • Tidy data
  • Contrast
  • Bioinformatics

Questions people ask

What is exploratory data analysis in bioinformatics?

It is the practice of examining a dataset through visualization and simple statistics, without formal hypothesis testing, to understand its structure, spot anomalies like batch effects or sample swaps, and decide whether the data is fit for the modeling step that comes next. In genomics that typically means PCA or MDS plots, sample distance heatmaps, and per-sample QC metrics like library size and percent mitochondrial reads.

Why do EDA before differential expression?

Differential expression models like DESeq2 and edgeR fit every gene the same way and will happily return significant hits driven by a mislabeled sample, a batch effect, or one outlier library. PCA and sample distance plots surface those problems while they are still cheap to fix, before they get baked into a gene list you hand to a wet-lab collaborator.

Is PCA the same thing as EDA?

No. PCA is one EDA technique among several, alongside MDS, hierarchical clustering on sample distances, and basic plots like boxplots and histograms of per-sample metrics. PCA is the most common first move for expression data because it shows clustering, outliers, and batch structure in one plot, but a full EDA pass also includes looking at raw QC numbers PCA won't show you, like fragment length distributions or duplication rates.

Can EDA replace a robust outlier detection method?

For a first pass, yes, and it should always come first, but visual inspection alone does not scale to datasets with many samples or catch subtle outliers a scatter plot compresses away. That is what tools like robust PCA (PcaGrid, PcaHubert) or DESeq2's Cook's distance are for: they formalize the same question your eyes are asking, and they matter most in small-n RNA-seq designs where one bad replicate can swing an entire contrast.

What is a common mistake when doing EDA on RNA-seq data?

Feeding the variance-stabilized or rlog-transformed values, made for visualization, into the actual differential expression test. DESeq2 and edgeR expect raw counts for their negative binomial models; the transformed matrix exists only so PCA and clustering plots aren't dominated by Poisson noise in low-count genes.

Related reading on the blog

Sources

  1. Exploratory Data Analysis - Secondary Analysis of Electronic Health Records — Foundational definition and core objectives of EDA used in the definition and why_it_matters sections
  2. RNA-seq workflow: gene-level exploratory analysis and differential expression — Standard EDA workflow used for the vst/PCA example and the common_confusion about transform-for-visualization-only
  3. pcaExplorer User Guide — Interactive PCA tool referenced as a way to move past static plots for outlier and PC interpretation
  4. Robust principal component analysis for accurate outlier sample detection in RNA-Seq data — Robust PCA methods (PcaGrid, PcaHubert) referenced in the FAQ on formal outlier detection