▸ Chatomics Field GuideWhat They Don't Teach You →

Glossary · Statistics, Artifacts and Pitfalls

Principal component analysis (PCA)

PCA is the first plot you make and the one most often misread: what the axes are, what center and scale do to them, and why a pretty PC1 vs PC2 can still mislead you.

By Ming "Tommy" Tang, Director of Bioinformatics in Big Pharma · Reviewed October 2026 · 3 min read

Also: PCA, principal components

Definition

PCA is a linear, unsupervised dimensionality reduction method that rewrites correlated variables as uncorrelated principal components. PC1 is the direction of maximum variance in the data, and each later component captures the most remaining variance while staying orthogonal to the earlier ones. Computationally, the principal components are the right singular vectors of the SVD of the centered data matrix (X = U Σ V^T), and the variance explained by a component is its eigenvalue divided by the sum of all eigenvalues. Think of it as rotating your 20,000-gene coordinate system so the first few axes point along the biggest spread between samples.

You will meet PCA in the first ten minutes of almost every omics project: a PC1 vs PC2 scatter of your samples, colored by condition or batch. It is the fastest way to see whether replicates group together, whether one sample is an outlier, and whether batch is a bigger source of variation than biology.

Because it is so quick, it is also easy to run without knowing what it did. Whether you centered, whether you scaled, which way the matrix was oriented, and which function drew the plot all change what you see. This entry covers those choices, so that when a PCA looks strange you know where to look first.

Why it matters

PCA is where you find out what dominates your data before any statistical test runs. In bulk RNA-seq, with roughly 20,000 genes and a handful of samples (p much greater than n), a PC1 vs PC2 plot shows whether samples cluster by condition or by batch. If PC1 separates sequencing batches and your condition only shows up on PC3, your differential expression model needs a batch term, and you would not have known without looking.

Getting it wrong is quiet. A PCA on uncentered data gives you variance around the origin rather than around the data's center, so PC1 loads on high-variance features such as housekeeping genes instead of biology. A PCA on matrices with the wrong orientation runs without error and returns nonsense. In single-cell work, the PCs are the input to neighbor graphs, clustering and UMAP, so a bad choice here propagates into every cluster you later name.

Where people get it wrong

The most common mistake is trusting the defaults without checking them. prcomp() centers by default (center = TRUE) but does not scale (scale. = FALSE), and which one you want depends on whether the data are already normalized. Raw counts benefit from scaling; log-normalized data are often run with scale. = FALSE, so high-mean genes are handled by the normalization instead. Without scaling, high-variance features dominate even when they separate your groups poorly, and they can overshadow low-variance genes with a real signal.

The second mistake is trusting the plot function. ggfortify::autoplot() rescales scores by PCi / sd(PCi) * sqrt(n - 1), so the axes are not the raw PCA values. If your plot looks unfamiliar, pull prcomp()$x and plot it with ggplot2 yourself. Also check orientation: genomic matrices usually have genes in rows, but prcomp() needs samples in rows, so you transpose first. Finally, PCA is linear; a clean-looking PC1 vs PC2 does not mean there is no nonlinear structure, and t-SNE or UMAP are the tools for that question, not a replacement for reading PCA first.

A concrete example

A bulk RNA-seq matrix has genes in rows and samples in columns, already log-normalized. You transpose it so samples are rows, run PCA with centering and no scaling, check variance explained, then plot raw scores directly instead of using autoplot(). Color by batch and by condition in two separate plots and compare which one separates along PC1.

r
# expr: genes x samples, log-normalized
pca <- prcomp(t(expr), center = TRUE, scale. = FALSE)

# variance explained = eigenvalue / sum of eigenvalues
var_expl <- pca$sdev^2 / sum(pca$sdev^2)
round(var_expl[1:5], 3)

# raw scores, not autoplot() rescaled ones
scores <- data.frame(pca$x[, 1:2], coldata)
ggplot(scores, aes(PC1, PC2, color = batch)) + geom_point()

Related terms

Questions people ask

What is PCA used for in RNA-seq?

It is used to look at overall sample structure without labels: whether replicates cluster, whether any sample is an outlier, and whether batch or condition drives the largest variation. It also compresses thousands of genes into a few components used for clustering and visualization, especially in single-cell workflows.

Should I use scale = TRUE in prcomp?

It depends on whether the data are already normalized. Centering is essential in most genomic analyses. Raw counts benefit from scaling, while log-normalized data are often run with scale. = FALSE so each gene is not forced to contribute equal variance. Check how your pipeline did it before choosing.

How many principal components should I keep?

There is no fixed number. Lessons from the book mention 10 to 50 components capturing around 80 to 90 percent of variance as a typical range, but the right number depends on study design. Look at a scree plot, and consider PCAtools, which supports the elbow method and Horn's parallel analysis, or a permutation test.

What is the difference between PCA and SVD?

SVD factorizes any matrix into U, Σ and V^T. PCA on a centered matrix gives the principal components as the right singular vectors, because X^T X = V Σ^2 V^T. prcomp() uses SVD internally, which is more numerically stable than eigen-decomposing the covariance matrix.

Why does my PCA plot look different from autoplot to prcomp output?

ggfortify::autoplot() rescales scores by default using PCi / sd(PCi) * sqrt(n - 1), so the axes are not the raw values. Extract pca$x from prcomp() and plot with ggplot2 to see the unscaled scores.

Related pages

Related reading on the blog

Sources

  1. Understanding prcomp() center and scale Arguments for Single-Cell RNA-seq PCA — Defaults of center and scale., and what happens without centering or scaling.
  2. PCA in action — Matrix orientation and when to scale.
  3. PCAtools: everything Principal Component Analysis — Elbow method and Horn's parallel analysis for choosing components.