Assay
Single-Cell RNA-seq: sanity checks and pitfalls
Droplet-based capture (10x Chromium is the default) yields a sparse cell-by-gene UMI matrix that goes through cell QC, normalization, feature selection, PCA, neighbor graph, clustering and 2-D embedding. Every one of those steps has a knob that changes the biology you report: mitochondrial and count thresholds remove real cell types, resolution invents clusters, and integration can erase the condition effect you were funded to find. Doublets, ambient RNA and sample-confounded batches are present in essentially every dataset and are only a problem when nobody checks.
Who this is for: Immunologists, cancer biologists and core-facility analysts who have Cell Ranger output for 4-20 samples and need clusters they can defend to a reviewer. Most are working in Seurat or Scanpy and choosing thresholds, resolutions and integration methods by copying a vignette.
- How to Detect Ambient RNA Contamination in Single-Cell RNA-seq
When hemoglobin, immunoglobulin or a neighboring cell type's markers show up in every cluster, the soup is talking; here is how to prove it before you annotate anything.
- How to Detect Batch Effects in Single-Cell RNA-seq
Before you pick Harmony, RPCA or scVI, find out whether batch is separable from your condition at all, because no method fixes a confounded design.
- How to Find and Remove Doublets in Single-Cell RNA-seq
A bridge cluster co-expressing two lineage markers is not a transitional cell state until you've ruled out that it's two cells sharing one droplet.
- How to Detect Integration Over-Correction in Single-Cell RNA-seq
Harmony and CCA are built to erase batch, point them at a design where batch and disease are the same variable and they will erase your disease effect too.
- How to Log-Transform Counts Without Fooling Yourself in Single-Cell RNA-seq
A pseudocount of 1 versus 1e-9 can turn the same CD19 fold change from 1.24 into 5.64, know which number your tool used before you trust it.
- How to Sanity-Check Marker Genes and Cell Type Labels in Single-Cell RNA-seq
A FindMarkers table full of RPL, RPS, MT- and FOS genes isn't a cell type. Here's how to catch it before the label ships.
- How to Handle Multiple Testing and FDR in Single-Cell RNA-seq
Bonferroni zeroes out your gene list, Benjamini-Hochberg inflates it, and neither number means anything until you fix pseudoreplication first.
- How to Choose a Normalization Method in Single-Cell RNA-seq
LogNormalize, SCTransform, CPM, TPM and dsb answer different questions; picking the wrong one quietly rewrites your clusters and your DE calls.
- How to Tell If You Overclustered in Single-Cell RNA-seq
Louvain and Leiden will hand you as many clusters as your resolution allows; the only test that matters is whether each one has markers you can name and cells from more than one sample.
- How to Avoid Pseudoreplication in Single-Cell RNA-seq
A p-value of 1e-200 from FindMarkers usually means you tested cells instead of donors, and pseudobulk is the fix a reviewer already knows to ask for.
- How to Choose Cell QC Thresholds in Single-Cell RNA-seq
The percent.mt < 5 you copied from a PBMC vignette is quietly deleting the tumor cells, cardiomyocytes, or plasma cells you were funded to study.
- Why Your UMAP Is Misleading You in Single-Cell RNA-seq
Distances between islands, island size and the direction of a smear on your UMAP are artifacts of the embedding, and here is how to test each claim before it reaches a figure legend.