Assay
Bulk RNA-seq: sanity checks and pitfalls
Reads are aligned (STAR, HISAT2) or pseudo-aligned (salmon, kallisto) to produce a gene-by-sample count matrix, then modeled with a negative binomial GLM (DESeq2, edgeR) or voom-transformed for limma. The recurring failures are design-level, not tool-level: batch confounded with condition, TPM fed into DESeq2, samples swapped at the sequencing core, or a PCA nobody looked at before running the differential test. Depth is usually 20-40M reads per sample and replicates matter far more than depth.
Who this is for: Wet-lab biologists with a handful of conditions and 3-6 replicates each, asking which genes change and whether the change survives correction. Usually the first analysis someone runs after leaving the bench, and the one most often analyzed with defaults that do not match the design.
- How to Detect Batch Effects in Bulk RNA-seq
A PCA plot that separates by prep date instead of treatment is telling you the truth about your experiment, not a bug to plot around.
- How to Log-Transform Counts Without Fooling Yourself in Bulk RNA-seq
The pseudocount you pick for low-count genes can move a fold change more than the biology does, and DESeq2's own statistical model never wants a log at all.
- How to Handle Multiple Testing and FDR in Bulk RNA-seq
Zero genes at padj < 0.05 with a clean PCA is a design or filtering problem to diagnose, not a cutoff to loosen.
- How to Choose a Normalization Method in Bulk RNA-seq
Median-of-ratios, TMM, CPM and TPM answer different questions, and the wrong one fails quietly when your treatment shifts the whole transcriptome.
- How to Handle Outlier Samples in PCA in Bulk RNA-seq
PCA can't tell you whether a stray point is a dead library or your best result; your QC metrics and metadata have to do that job for it.
- How to Read a P-Value Histogram in Bulk RNA-seq
A misshapen histogram is the cheapest bug report your DE test will ever hand you, and almost nobody opens it before trusting padj.
- How to Check Library Strandedness in Bulk RNA-seq
One minute with infer_experiment.py or salmon -l A tells you which strand flag to use, before a wrong guess quietly eats your counts.
- Why You Must Not Use TPM for Differential Expression in Bulk RNA-seq
TPM already threw away the count magnitude DESeq2 and edgeR need to estimate dispersion, and rounding it to dodge the integer check does not bring that information back.