▸ Chatomics Field GuideWhat They Don't Teach You →

Sanity check · Single-Nucleus RNA-seq

How to Choose Cell QC Thresholds in Single-Nucleus RNA-seq

Your snRNA-seq filter probably came from a whole-cell PBMC vignette, and in nuclei a high mito fraction means cytoplasmic contamination, not a dying cell.

By Ming "Tommy" Tang, Director of Bioinformatics in Big Pharma · Reviewed October 2026 · 5 min read

You got your first snRNA-seq object from frozen tissue, opened the Seurat PBMC tutorial, and pasted nFeature_RNA > 200 & nFeature_RNA < 2500 & percent.mt < 5 into your script. Those numbers were chosen for whole cells. In nuclei, they answer a different question, and they can silently delete a whole cell type.

The stakes are quiet. A fixed cutoff does not throw an error. It removes the cardiomyocytes, the tumor cells, or the low-count plasma cells and neutrophils, and your cell-type proportions shift before you have clustered anything. Nobody downstream can tell the cells were ever there.

This page lets you set thresholds per sample from the distributions in the next hour, and then look at what you removed before you accept the result.

What it looks like when it's happening

  • After filtering with `percent.mt < 5` the object loses an entire cluster that the tissue should contain, such as cardiomyocytes or a tumor population.
  • The violin plot of `percent.mt` shows a pile of nuclei at exactly zero and a long right tail, so a median-based rule gives a cutoff at or near zero.
  • A cluster with high `percent.mt` and high counts of cytoplasmic or glial/neuronal-unrelated markers sits next to a clean cluster of the same cell type in UMAP.
  • Marker genes of abundant cell types show up as low-level expression in every cluster, including cell types that should not express them.
  • `nFeature_RNA` and `nCount_RNA` distributions differ sharply between samples, and one global cutoff keeps almost everything in one sample and almost nothing in another.
  • Clusters separate mainly by sequencing depth or by sample, and the top PCs correlate with `nCount_RNA`.
  • Cell counts after filtering are far below what the loaded nuclei number predicted, with no obvious reason.

Why it happens

Nuclei have no cytoplasm, and mitochondria live in the cytoplasm. In a clean nucleus prep, mitochondrial transcripts should be close to absent. In scRNA-seq, a high mito fraction means the membrane broke and cytoplasmic RNA leaked out, so you filter high values. In snRNA-seq, a high value means cytoplasm came along with the nucleus, so it flags contamination. The direction of concern is similar, but the baseline is different: the recommended cutoff in the sources is 3% for snRNA-seq against 5% for scRNA-seq, and many nuclei sit at zero.

Ambient RNA makes this worse. Nuclei extraction lysed cells and released cytoplasmic RNA into the solution that every droplet captures. Background can reach 3 to 35% of total counts per cell, and it varies across replicates and cells. So a per-droplet metric like mito fraction mixes two things: real cytoplasmic carryover on that nucleus and the ambient soup around it.

Biology adds a second layer. Cardiomyocytes and tumor cells have high mitochondrial content because of energy demand, and plasma cells and neutrophils have low counts because they are low-RNA cells. A fixed cutoff cannot tell these apart from damaged droplets. It also cannot adapt to sample-to-sample differences in tissue quality, nuclei prep and sequencing depth, which is why the cutoff has to be set per sample from the distribution.

Depth is the third piece. If one sample or cell type has systematically higher per-cell depth, depth becomes the main axis of variation and creates batch clusters. Filtering too loosely in one sample and too tightly in another builds that imbalance into your data.

The checks

Run them in order. Each one tells you what healthy looks like and what the problem looks like.

0/8 checked · saved in this browser

  1. Before subsetting, compute nCount_RNA, nFeature_RNA and mito percentage per nucleus and plot them per sample as violins, plus a scatter of counts against features. Read the cutoffs from these plots, not from a vignette. Make sure counting included introns, or your counts will be roughly half of what they should be.

    r
    seu[["percent.mt"]] <- PercentageFeatureSet(seu, pattern = "^MT-")
    VlnPlot(seu, features = c("nCount_RNA", "nFeature_RNA", "percent.mt"),
            group.by = "sample", pt.size = 0, ncol = 3)
    FeatureScatter(seu, "nCount_RNA", "nFeature_RNA", group.by = "sample")
    Healthy
    Each sample shows a clear main population, with `percent.mt` concentrated near zero. Samples from the same tissue have broadly similar shapes, even if their medians differ.
    Red flag
    A sample whose distribution is shifted far from the others, a `percent.mt` bulk well above zero, or counts so low that the main population is cut off by the 200 to 2500 vignette window.
  2. Confirm the pattern matches your annotation: ^MT- for human gene symbols, a different prefix for mouse. Count how many genes matched. A zero-match pattern gives percent.mt of 0 for every nucleus, which looks like perfect data and filters nothing.

    r
    sum(grepl("^MT-", rownames(seu)))
    summary(seu$percent.mt)
    Healthy
    A small positive number of mitochondrial genes matched, and `percent.mt` shows a spread with most nuclei low.
    Red flag
    Zero genes matched, or `percent.mt` is exactly 0 for every nucleus in every sample.
  3. For each sample, flag nuclei as outliers on log counts, log features and mito percentage using median absolute deviations. In OSCA, isOutlier() with batch does this per sample. For mito, use type="higher" for contamination and set min_diff=0.5 so low-coverage samples where many nuclei have zero mito do not get a cutoff of zero. The nmads value below is a starting point you tune against the plots in the next checks.

    r
    library(scuttle)
    sce <- addPerCellQCMetrics(sce, subsets = list(Mt = grep("^MT-", rownames(sce))))
    low_lib   <- isOutlier(sce$sum,      type = "lower",  log = TRUE, batch = sce$sample, nmads = 3)
    low_genes <- isOutlier(sce$detected, type = "lower",  log = TRUE, batch = sce$sample, nmads = 3)
    high_mt   <- isOutlier(sce$subsets_Mt_percent, type = "higher", batch = sce$sample,
                           nmads = 3, min_diff = 0.5)
    discard <- low_lib | low_genes | high_mt
    Healthy
    Each sample gets its own thresholds, and the discard rate is a minority of nuclei per sample, with similar rates across comparable samples.
    Red flag
    One sample loses most of its nuclei while others lose few. This points to a prep problem in that sample, or to a cutoff that is wrong for it, and needs a look before you proceed.
  4. Write down the per-sample cutoffs that step 3 produced and compare them with the packet's reference points: about 3% mito for snRNA-seq, at least 300 detected genes and 500 total UMIs per nucleus, and genes seen in fewer than five nuclei removed. Large disagreement means either your sample is unusual or one of the two is wrong. Decide which by looking at the plots.

    Healthy
    The MAD-derived cutoffs land in the same neighborhood as the rules of thumb for most samples.
    Red flag
    A MAD-derived mito cutoff far above 3% in a sample, which usually means heavy cytoplasmic carryover in that prep, not a normal tissue.
  5. Cluster the object without filtering the flagged nuclei (or keep them in a copy), then plot the discard flag on the UMAP and tabulate discard rate per cluster and per annotated cell type. Removal should be spread thinly. A cluster that is almost entirely discarded is the signal to stop and think.

    r
    seu$discard <- discard[colnames(seu)]
    DimPlot(seu, group.by = "discard")
    prop.table(table(seu$seurat_clusters, seu$discard), margin = 1)
    Healthy
    Discarded nuclei scatter across clusters at roughly the same low rate, or concentrate in a debris-like cluster with low features and no clear markers.
    Red flag
    One well-defined cluster with real marker genes is mostly flagged. That may be cardiomyocytes, tumor cells, plasma cells or neutrophils removed for their biology.
  6. For the nuclei with high mito, look at the markers of the dominant cell types in the tissue and at the unspliced fraction if you have it. Contaminated droplets show other cell types' markers at low levels and a lower unspliced fraction. Intact nuclei with a real high-mito phenotype keep their own cell-type identity and a good unspliced fraction.

    Healthy
    High-mito nuclei express markers of a single identity, and unspliced fraction is similar to that of low-mito nuclei of the same type.
    Red flag
    High-mito nuclei express a mix of unrelated cell type markers and have a lower unspliced fraction, which points to ambient RNA or cytoplasm carryover.
  7. Because true nuclei should not carry mitochondrial transcripts, the OSCA workflow uses that expectation to estimate ambience. Run estimateAmbience() on the raw droplets to get the ambient profile, then use controlAmbience() to see how much of each nucleus's expression is explained by it. This is only possible in snRNA-seq, where the absence of cytoplasmic genes is expected.

    Healthy
    Ambient contribution is small and similar across samples, and ambient marker genes are not driving cluster identity.
    Red flag
    Ambient markers explain a large share of expression in some nuclei, or one sample's ambient profile is very different from the others.
  8. Color the final UMAP by nCount_RNA and by sample, and correlate the top PCs with nCount_RNA. If depth is the main axis of variation, the filter is leaving depth imbalance between samples or cell types.

    Healthy
    Clusters correspond to cell types, and `nCount_RNA` varies within clusters without defining them.
    Red flag
    Clusters or UMAP regions that line up with count gradients or with a single sample, which are batch clusters driven by depth.

What to do about it

Switch to per-sample MAD thresholds with a minimum difference on mito

When: You have several samples with different depth or prep quality and you were using one global cutoff.

Run isOutlier() with batch set to sample on log library size and log detected genes (lower tail) and on mito percentage (higher tail, min_diff=0.5). Store the cutoffs per sample in a table and keep that table with the analysis. Apply the combined flag after you have inspected it.

Caveat: MAD cutoffs assume that most nuclei in a sample are good. In a sample that is mostly contaminated, the median is itself contaminated and the cutoff will be too generous.

Add an absolute floor from the snRNA-seq rules of thumb

When: MAD-derived lower cutoffs fall below what a nucleus can plausibly support, for example in a low-depth sample.

Enforce at least 300 detected genes and 500 total UMIs per nucleus, and remove genes detected in fewer than five nuclei. Take the stricter of the floor and the MAD cutoff for each sample.

Caveat: A floor can remove genuinely low-RNA cell types such as plasma cells or neutrophils. Check the discard rate by cluster before accepting it.

Remove ambient RNA instead of filtering on mito alone

When: High-mito and cross-cell-type marker signal is spread across many nuclei, so a filter would throw away most of a sample.

Run an ambient correction tool such as CellBender, DecontX or SoupX on the raw droplets. In the sources, CellBender gave the most precise background estimates and the best marker gene detection. Then recompute QC metrics on the corrected counts and set thresholds again.

Caveat: Correction changes the counts you analyze. Compare marker genes before and after, and keep the uncorrected object so you can reverse the decision.

Use a multi-metric filter for difficult samples

When: Single thresholds keep failing on challenging samples, for example heavily contaminated frozen tissue.

Try QClus, which combines six metrics: unspliced fraction, mitochondrial percentage, nuclear fraction, tissue-specific nuclear markers, tissue-specific cytoplasm markers and non-cluster marker genes. In the paper it ran on 252 samples without a processing failure, where fixed-threshold methods failed on 20 to 53.

Caveat: It needs tissue-specific marker sets and is less transparent than a plotted cutoff. Still inspect the removed nuclei.

Rescue a biologically high-mito population explicitly

When: Step 5 shows a whole cluster with its own markers being removed, and step 6 shows its identity is clean.

Raise the mito cutoff for that sample or cell type, or filter within annotated cell types instead of across the whole object. Record the decision and the reason next to the cutoff table.

Caveat: Raising the cutoff admits contaminated droplets too. Pair it with ambient correction and verify with the cytoplasmic marker check.

When not to "fix" it

If a high-mito cluster is a real cell type, such as cardiomyocytes or tumor cells with high energy demand, do not tighten the filter to make the plot look cleaner. Likewise, low-count cells such as plasma cells or neutrophils are biology, so do not raise the minimum counts to get a tidier violin. Check cluster identity and markers first. If the removed nuclei share a coherent identity and their own markers, keep them and document the exception.

Five things experienced analysts do here

  1. Write the per-sample cutoff table to a file and commit it with the code, so a reviewer can see exactly what each sample lost.
  2. Always report the number of nuclei before and after filtering per sample. A sample that loses much more than its peers is a prep problem, and you should say so.
  3. Run the filter flag through the UMAP before you drop anything. Discarding is a decision you make after looking, not before.
  4. Treat the vignette numbers as examples of a process. Seurat's `percent.mt < 5` and `nFeature_RNA` window come from PBMC whole cells and are not your starting point.
  5. Keep an unfiltered copy of the object. If a cell type later looks missing or under-represented, you can go back and check whether QC removed it.

Questions people ask

What mitochondrial percentage threshold should I use for snRNA-seq?

The sources give 3% as a recommended value, against 5% for scRNA-seq. Use that as a reference, not a rule. Set the cutoff per sample from the distribution with a MAD-based rule and a minimum difference of 0.5, then look at which clusters lose nuclei.

Why is high mitochondrial content a problem in nuclei but fine in some cells?

Nuclei lack cytoplasm, so mitochondrial transcripts should be near absent. A high fraction means cytoplasm or ambient RNA came along, which is contamination. Some cell types, such as cardiomyocytes and tumor cells, do have high mitochondrial content for biological reasons, so check the identity of the cluster before removing it.

What nFeature_RNA cutoff should I use for snRNA-seq?

The packet's reference minimum is 300 detected genes and 500 total UMIs per nucleus. The vignette window of 200 to 2500 is for whole cells and should not be copied. Read lower and upper limits per sample from the nFeature and nCount plots.

Can I use the same QC thresholds for all samples?

Usually not. Samples differ in tissue quality, nuclei prep and depth, so one global cutoff keeps nearly everything in one sample and removes too much in another. Use per-sample MAD-based outlier calls and compare the discard rates.

Should I remove ambient RNA before or after QC filtering?

Plot the raw metrics first, then decide. If contamination is widespread, run CellBender, DecontX or SoupX on the raw droplets and recompute the QC metrics, because filtering on mito alone would remove too many nuclei.

Related pages

Related reading on the blog

Sources

  1. Single-nuclei RNA-seq processing (Chapter 11, Orchestrating Single-Cell Analysis with Bioconductor) — Mito as contamination marker in nuclei, isOutlier with min_diff, estimateAmbience and controlAmbience.
  2. QClus: a droplet filtering algorithm for enhanced snRNA-seq data quality in challenging samples — Six-metric filtering, 252-sample benchmark, ambient RNA at 3 to 35% and unspliced fraction.
  3. Seurat PBMC3k Tutorial: Quality Control — Vignette filter of nFeature 200 to 2500 and percent.mt < 5 comes from whole cells.
  4. Scanpy calculate_qc_metrics documentation — Computing per-cell QC metrics in Scanpy.
  5. Quality control of sc/snRNA-seq (DOtools Bioconductor vignette) — 3% mito, 300 genes, 500 UMIs, genes in fewer than five nuclei, read thresholds from plots.

Part of the QC thresholds series.