Guides
Sanity checks, by assay
Each guide answers one question about one kind of data: what the failure looks like, why it happens, the checks that expose it, and when leaving it alone is the right call.
Bulk RNA-seq
- How to Detect Batch Effects in Bulk RNA-seq
A PCA plot that separates by prep date instead of treatment is telling you the truth about your experiment, not a bug to plot around.
- How to Log-Transform Counts Without Fooling Yourself in Bulk RNA-seq
The pseudocount you pick for low-count genes can move a fold change more than the biology does, and DESeq2's own statistical model never wants a log at all.
- How to Handle Multiple Testing and FDR in Bulk RNA-seq
Zero genes at padj < 0.05 with a clean PCA is a design or filtering problem to diagnose, not a cutoff to loosen.
- How to Choose a Normalization Method in Bulk RNA-seq
Median-of-ratios, TMM, CPM and TPM answer different questions, and the wrong one fails quietly when your treatment shifts the whole transcriptome.
- How to Handle Outlier Samples in PCA in Bulk RNA-seq
PCA can't tell you whether a stray point is a dead library or your best result; your QC metrics and metadata have to do that job for it.
- How to Read a P-Value Histogram in Bulk RNA-seq
A misshapen histogram is the cheapest bug report your DE test will ever hand you, and almost nobody opens it before trusting padj.
- How to Check Library Strandedness in Bulk RNA-seq
One minute with infer_experiment.py or salmon -l A tells you which strand flag to use, before a wrong guess quietly eats your counts.
- Why You Must Not Use TPM for Differential Expression in Bulk RNA-seq
TPM already threw away the count magnitude DESeq2 and edgeR need to estimate dispersion, and rounding it to dodge the integer check does not bring that information back.
Single-Cell RNA-seq
- How to Detect Ambient RNA Contamination in Single-Cell RNA-seq
When hemoglobin, immunoglobulin or a neighboring cell type's markers show up in every cluster, the soup is talking; here is how to prove it before you annotate anything.
- How to Detect Batch Effects in Single-Cell RNA-seq
Before you pick Harmony, RPCA or scVI, find out whether batch is separable from your condition at all, because no method fixes a confounded design.
- How to Find and Remove Doublets in Single-Cell RNA-seq
A bridge cluster co-expressing two lineage markers is not a transitional cell state until you've ruled out that it's two cells sharing one droplet.
- How to Detect Integration Over-Correction in Single-Cell RNA-seq
Harmony and CCA are built to erase batch, point them at a design where batch and disease are the same variable and they will erase your disease effect too.
- How to Log-Transform Counts Without Fooling Yourself in Single-Cell RNA-seq
A pseudocount of 1 versus 1e-9 can turn the same CD19 fold change from 1.24 into 5.64, know which number your tool used before you trust it.
- How to Sanity-Check Marker Genes and Cell Type Labels in Single-Cell RNA-seq
A FindMarkers table full of RPL, RPS, MT- and FOS genes isn't a cell type. Here's how to catch it before the label ships.
- How to Handle Multiple Testing and FDR in Single-Cell RNA-seq
Bonferroni zeroes out your gene list, Benjamini-Hochberg inflates it, and neither number means anything until you fix pseudoreplication first.
- How to Choose a Normalization Method in Single-Cell RNA-seq
LogNormalize, SCTransform, CPM, TPM and dsb answer different questions; picking the wrong one quietly rewrites your clusters and your DE calls.
- How to Tell If You Overclustered in Single-Cell RNA-seq
Louvain and Leiden will hand you as many clusters as your resolution allows; the only test that matters is whether each one has markers you can name and cells from more than one sample.
- How to Avoid Pseudoreplication in Single-Cell RNA-seq
A p-value of 1e-200 from FindMarkers usually means you tested cells instead of donors, and pseudobulk is the fix a reviewer already knows to ask for.
- How to Choose Cell QC Thresholds in Single-Cell RNA-seq
The percent.mt < 5 you copied from a PBMC vignette is quietly deleting the tumor cells, cardiomyocytes, or plasma cells you were funded to study.
- Why Your UMAP Is Misleading You in Single-Cell RNA-seq
Distances between islands, island size and the direction of a smear on your UMAP are artifacts of the embedding, and here is how to test each claim before it reaches a figure legend.
Single-Nucleus RNA-seq
- How to Detect Batch Effects in Single-Nucleus RNA-seq
Before you run Harmony on your frozen-tissue nuclei, find out whether the batch axis is technical noise or the condition you came to study.
- How to Find and Remove Doublets in Single-Nucleus RNA-seq
A cluster that co-expresses two lineage markers isn't always a transitional cell state, and the fix depends on telling homotypic from heterotypic doublets before you delete anything.
- How to Sanity-Check Marker Genes and Cell Type Labels in Single-Nucleus RNA-seq
FindMarkers will happily hand you a stress signature, an ambient RNA gradient, or a pseudoreplication artifact and let you name a cluster after it.
- How to Tell If You Overclustered in Single-Nucleus RNA-seq
Leiden will happily hand you fifty clusters from thirty cell types if the resolution lets it; the object itself can't tell you which cuts are real.
- How to Avoid Pseudoreplication in Single-Nucleus RNA-seq
5,000 nuclei from two donors is not 5,000 replicates: pick pseudobulk or a mixed model before you trust a single p-value out of FindMarkers.
- How to Choose Cell QC Thresholds in Single-Nucleus RNA-seq
Your snRNA-seq filter probably came from a whole-cell PBMC vignette, and in nuclei a high mito fraction means cytoplasmic contamination, not a dying cell.
Spatial Transcriptomics
- How to Detect Batch Effects in Spatial Transcriptomics
A slide is not a replicate, learn to tell tissue-section artifact from real biology before you trust a single cluster.
- How to Detect Integration Over-Correction in Spatial Transcriptomics
The UMAP where every condition mixes perfectly isn't always good news, sometimes it means Harmony erased the exact difference you set out to measure.
- How to Sanity-Check Marker Genes and Cell Type Labels in Spatial Transcriptomics
FindMarkers and rank_genes_groups will happily hand you ribosomal genes and dissociation stress as your top "cell type marker", here is the checklist that catches it before it reaches a figure legend.
- How to Choose a Normalization Method in Spatial Transcriptomics
Library-size normalization removes the exact tissue-architecture signal you're trying to map, and the scRNA-seq default you already know is often the wrong call here.
- How to Tell If You Overclustered in Spatial Transcriptomics
Leiden and Louvain will happily cut your tissue into as many pieces as the resolution allows; the algorithm has no opinion on whether any of those pieces is a real cell type.
- How to Avoid Pseudoreplication in Spatial Transcriptomics
A wall of p-values near zero on pooled spots or cells almost always means you counted the wrong thing as your sample size.
- How to Choose Cell QC Thresholds in Spatial Transcriptomics
Copying the 10% mito cutoff from a PBMC tutorial into your tumor Visium slide will quietly delete every high-mitochondrial tumor spot you came to study.
- Why Your UMAP Is Misleading You in Spatial Transcriptomics
UMAP islands, cluster sizes and arcs are optical illusions in spatial data; the tissue coordinates and a permutation test are the only evidence that counts.
ATAC-seq
- How to Handle Multiple Testing and FDR in ATAC-seq
Zero significant peaks after BH correction is usually the statistics doing exactly what you asked, not a broken pipeline.
- How to Choose a Normalization Method in ATAC-seq
The method you pick can turn 33 significant regions into 24,450 on the same data, so decide it on purpose and check it with an MA plot.
- How to Read a P-Value Histogram in ATAC-seq
Your differential accessibility table is only as good as the shape of its p-values, and one hist() call tells you whether to trust it before you look at a single peak.
- How to Call Peaks You Can Trust in ATAC-seq
No input control exists for ATAC-seq, so ENCODE blacklists, FRiP, and a five-minute browser check of your top peaks are what stand between you and peaks called in centromeres.
- How to Tell If You Sequenced Deep Enough in ATAC-seq
Before you pay for another lane, find out whether your ATAC-seq libraries are out of depth or out of complexity, and whether a third replicate would help more.
Single-Cell ATAC-seq
- How to Detect Batch Effects in Single-Cell ATAC-seq
Your LSI plot separates by processing date instead of cell type, here's the order of checks that tells you whether that's a fixable artifact or a confound no tool can undo.
- How to Find and Remove Doublets in Single-Cell ATAC-seq
A cluster that looks like a transitional cell state is often two nuclei in one droplet, and scATAC-seq doublets need fragment-based tools, not the scRNA-seq playbook.
- How to Sanity-Check Marker Genes and Cell Type Labels in Single-Cell ATAC-seq
A marker with a huge fold-change can still be sequencing depth or dissociation stress in disguise, here's how to tell before you write the label on the UMAP.
- How to Tell If You Overclustered in Single-Cell ATAC-seq
Louvain and Leiden will keep splitting your point cloud for as long as you keep raising resolution, here's how to find out whether the clusters you kept are real.
- How to Call Peaks You Can Trust in Single-Cell ATAC-seq
Your top-scoring peaks might be centromeric repeats, not regulatory elements, and the pooled peak set your pipeline handed you is quietly erasing your rarest cell type.
- How to Avoid Pseudoreplication in Single-Cell ATAC-seq
Five thousand cells from two donors are not five thousand replicates, and your p-values are lying to you about it.
- How to Choose Cell QC Thresholds in Single-Cell ATAC-seq
The QC axes that matter in scATAC-seq aren't the ones your scRNA-seq muscle memory reaches for, and a copy-pasted vignette cutoff will quietly delete a real cell population.
- Why Your UMAP Is Misleading You in Single-Cell ATAC-seq
The islands, the gaps between them, and the smear you're calling a trajectory are navigation aids, not measurements, so treat them that way before you write the discussion section.
ChIP-seq
- How to Choose a Normalization Method in ChIP-seq
RPKM, CPM, and median-of-ratios all assume nothing changed genome-wide, an assumption that quietly kills real histone-mark shifts before you ever see them.
- How to Read a P-Value Histogram in ChIP-seq
Flat with a spike, a hill, a U, or a pile-up near 1: each shape of your differential binding p-values points at a specific broken assumption, and you can see it in one line of R.
- How to Call Peaks You Can Trust in ChIP-seq
A matched input control and the ENCODE blacklist are the two things standing between your peak list and a stack of copy-number and repeat artifacts you'd otherwise write up as binding sites.
CUT&RUN and CUT&Tag
- How to Choose a Normalization Method in CUT&RUN and CUT&Tag
Your treated sample has a fraction of the control's histone mark, and your bigWigs look identical: the normalization you picked erased the biology.
- How to Call Peaks You Can Trust in CUT&RUN and CUT&Tag
MACS2 and an IgG control were built for ChIP-seq's noise floor; point them at CUT&RUN's much cleaner background and you'll either call nothing or call the centromere.
Variant Calling
- How to Fix Chromosome Naming Mismatches in Variant Calling
Your BED file and your VCF look fine and bedtools says they share nothing: the cause is usually one prefix, and the repair is a text edit, not a re-run.
- How to Detect Contamination in Variant Calling
A low mapping rate isn't "bad sequencing" until you've checked what the unmapped reads actually are.
- How to Catch a Genome Build Mismatch in Variant Calling
A BED file from the wrong build gives you confident, wrong overlaps, and the header check that catches it takes thirty seconds.
- How to Catch a Sample Swap in Variant Calling
X heterozygosity, chrY coverage and pairwise genotype concordance take minutes to compute and catch mislabels that no alignment or variant-quality QC report will ever flag.
- How to Tell If You Sequenced Deep Enough in Variant Calling
Mean coverage looks fine, the VCF looks clean, and you still cannot say whether the missing variants are absent or just unsampled.
- How to Avoid 0-Based vs 1-Based Coordinate Errors in Variant Calling
A VCF position and a BED start are not the same number, and the mismatch silently drops the first base of every target interval.
By topic
- Batch effects
- Sample swaps
- Overclustering
- Doublets
- Ambient RNA
- QC thresholds
- Strandedness
- Genome build mismatches
- Chromosome naming
- 0-based vs 1-based coordinates
- Log transformation of counts
- Pseudoreplication
- P-value histograms
- Multiple testing and FDR
- Normalization choice
- TPM vs counts for differential expression
- Marker gene sanity
- UMAP misreading
- Integration over-correction
- Peak calling controls and blacklists
- Contamination
- Sequencing depth and saturation
- Outlier samples in PCA