Chapter 11
Statistics, Artifacts and Pitfalls
- Effect size
Log2 fold change tells you how big the difference is; the p-value only tells you how surprised to be that it's not zero, and mixing the two up is how real biology gets buried under statistical noise.
- False discovery rate (FDR)
The number that decides whether your "significant genes" list is mostly real or mostly noise dressed up as signal.
- P-value
A p-value tells you how surprising your data would be if nothing were going on, not whether your result is true or big enough to matter.
- P-value histogram
The one plot that tells you whether your differential expression test is measuring biology or measuring its own bugs.
- Partial correlation
Two genes look co-regulated until you control for the transcription factor pulling both of their strings, this is how you find out if the link is real.
- Principal component analysis (PCA)
PCA is the first plot you make and the one most often misread: what the axes are, what center and scale do to them, and why a pretty PC1 vs PC2 can still mislead you.
- Q-value
The number that tells you what fraction of your "significant" genes are probably noise, and why it is not just a fancier p-value.
- Singular value decomposition (SVD)
The matrix factorization that actually runs when you call prcomp() or RunPCA(), and the reason your PCA either scales to a million cells or falls over.
- Z-score
A z-score tells you how far a value sits from the mean in standard deviations, and it only means what you think it means if the mean and SD it's built from are stable.