▸ Chatomics Field GuideWhat They Don't Teach You →

Sanity check · Single-Cell RNA-seq

How to Detect Ambient RNA Contamination in Single-Cell RNA-seq

When hemoglobin, immunoglobulin or a neighboring cell type's markers show up in every cluster, the soup is talking; here is how to prove it before you annotate anything.

By Ming "Tommy" Tang, Director of Bioinformatics in Big Pharma · Reviewed October 2026 · 6 min read

You finished clustering and opened a violin plot. Hemoglobin is low but nonzero in T cells, B cells and fibroblasts. Or an immunoglobulin gene lights up in every cluster, or a surfactant gene appears in immune cells from a lung sample. A T cell should not express that, and now you do not trust your annotation.

This is ambient RNA: free-floating transcripts from lysed cells that get captured in every droplet alongside the real cell. Background has been reported at 3 to 35% of total counts per cell, and it varies across replicates and cells. If you miss it, marker genes lose specificity, rare populations appear to express things they do not, and differential expression can report the soup instead of the biology.

In the next hour you can build the ambient profile from your own empty droplets, test whether your odd genes match it, estimate how much of each cell is soup, and decide whether to correct with SoupX, CellBender or DecontX.

What it looks like when it's happening

  • Hemoglobin genes (HBB, HBA1, HBA2) detected at low levels in nearly every non-erythroid cluster on a violin or dot plot.
  • Immunoglobulin genes appear in T cells, myeloid cells and fibroblasts, not just plasma cells and B cells.
  • A tissue-specific gene such as a surfactant transcript shows up in immune cells from the same sample.
  • A cluster is annotated as 'doublets' or 'mixed' only because it expresses markers of two unrelated cell types at low levels.
  • The same handful of highly expressed genes appear as low-level background in every cluster, and the pattern is stronger in low-UMI cells.
  • Marker genes for a rare population are detectable in many other clusters, so differential expression gives long, non-specific marker lists.

Why it happens

Tissue dissociation kills and breaks some cells. Their RNA spills into the suspension, and every droplet, including the empty ones, captures a bit of it. A droplet that holds a real cell therefore sequences that cell's mRNA plus a slice of the soup. The soup is dominated by whatever is most abundant in the lysed cells: hemoglobin from red blood cells, immunoglobulin from plasma cells, surfactant from lung epithelium. That is why the same few genes appear everywhere.

The soup profile is measurable. Droplets with very low UMI counts, typically under 100, contain no cell and approximate the ambient mixture, which reflects roughly uniform sampling of the cells in that sequencing batch. Cell Ranger's cell calling already runs EmptyDrops against an ambient profile pooled from low-count barcodes. The raw_feature_bc_matrix you may have deleted holds exactly the information needed to measure the contamination.

The amount of soup is not constant. It differs between samples and between cells within a sample, and it differs by gene, which is why one global subtraction can be wrong for individual genes or cells. Barcode swapping adds a related leak that looks the same on a plot. Correction tools differ in how they model this, and the tradeoff is real: some over-correct when there is little ambient RNA to remove.

Interpretation depends on the biology. A gene whose signal in a cell tracks the ambient profile, and does not scale with that cell's own identity, is soup. A gene high in one cluster and absent elsewhere is a marker. Everything between those two needs the checks below.

The checks

Run them in order. Each one tells you what healthy looks like and what the problem looks like.

0/8 checked · saved in this browser

  1. Pick the genes that bother you (hemoglobin, immunoglobulin, a tissue-specific marker). Plot them as a violin or dot plot across all clusters in Seurat or Scanpy. Note which clusters express them at a low, flat level and which cluster expresses them strongly. A flat low level everywhere plus a strong peak in one true source cluster is the soup signature.

    Healthy
    A source cluster with high expression and other clusters at near zero, or the gene is simply absent from your data.
    Red flag
    The same gene sits at a low, similar level in almost every cluster, including cell types that cannot biologically make it.
  2. Open the web summary for each sample and look at the barcode rank plot and the estimated number of cells. Confirm that raw_feature_bc_matrix is still on disk next to filtered_feature_bc_matrix. SoupX and FastCAR need the raw matrix and empty droplet annotations, so without it your options narrow to tools that work from the filtered data.

    Healthy
    A clear drop-off ('knee') in the barcode rank plot between cell-containing barcodes and the empty-droplet plateau, and a raw matrix you can load.
    Red flag
    No clean separation between cells and background, or only the filtered matrix was kept.
  3. Load the raw matrix and estimate the soup as the pooled counts of low-UMI barcodes with ambientProfileEmpty() from DropletUtils. Compare the top genes in this profile with your suspect genes. If hemoglobin or immunoglobulin genes rank at the top of the soup, you have your explanation.

    r
    library(DropletUtils)
    # m: raw, unfiltered counts matrix (genes x barcodes)
    profile <- ambientProfileEmpty(m, lower=100)
    head(sort(profile, decreasing=TRUE), 30)
    Healthy
    The top soup genes match what you would expect from the most abundant lysed cells in the tissue, and your suspect genes are among them.
    Red flag
    The soup profile has little overlap with the genes you are worried about, which points toward doublets, barcode swapping or real biology.
  4. Run emptyDrops() on the raw matrix and keep barcodes with FDR at or below 0.01. Compare your cell count with Cell Ranger's. The default lower threshold is 100 UMIs with 10,000 Monte Carlo iterations. Note that barcodes passing the test are non-empty but can still include cell fragments and damaged cells.

    r
    out <- emptyDrops(my.counts)
    is.cell <- out$FDR <= 0.01
    cell.counts <- my.counts[, which(is.cell), drop=FALSE]
    Healthy
    Cell counts close to Cell Ranger's, with the extra or missing barcodes concentrated at the low-UMI edge.
    Red flag
    Large disagreement between the two cell sets, or many called cells whose profile looks like the soup profile.
  5. Use ambientContribMaximum() with the ambient profile to compute the maximum proportion of each cell's counts that could come from the soup. Plot the distribution per sample and per cluster. Low-UMI cells are expected to carry the largest share, so compare against total counts.

    r
    max.contrib <- ambientContribMaximum(m, profile)
    Healthy
    A modest ambient share in most cells, similar across samples processed in the same way.
    Red flag
    One sample with a much higher soup share than the others, or a cluster whose cells are mostly soup. The OSCA chapter suggests dropping genes where 10% or more of counts derive from ambient RNA, or cells with more than 25% of reads identified as ambient contamination.
  6. Choose genes that a cell type must not express, for example hemoglobin genes in a non-erythroid population, and run ambientContribNegative() with that gene set. Cells that are known to lack the gene but show counts for it give a direct estimate of the contamination proportion.

    r
    neg.contrib <- ambientContribNegative(m, profile, negative.genes)
    Healthy
    Estimated contamination that is consistent with the ambient share you saw in the previous check.
    Red flag
    Estimates that vary wildly between cells of the same type, which means a single global contamination fraction will not fit.
  7. Compute the nuclear fraction (proportion of unspliced pre-mRNA) per barcode with DropletQC. Droplets that mostly carry ambient cytoplasmic RNA tend to have a different nuclear fraction from intact cells. Remember that intronic reads are 30 to 50% of reads in 10x 3' data, so judge the nuclear fraction relative to your own sample, not an absolute number.

    Healthy
    Cells form a compact nuclear fraction distribution, with empty or damaged barcodes separated from it.
    Red flag
    Barcodes with high UMI counts but a nuclear fraction that matches empty droplets. These look like cells in the cell calls but behave like soup.
  8. Run one correction tool (SoupX, CellBender or DecontX) and redo the plots from the first check. Check that the true marker remains high in its source cluster and the contaminating genes drop elsewhere. Also compare cluster structure and marker lists before and after, since a decontamination that changes the cluster map is doing more than removing soup.

    Healthy
    Contaminating genes fall in clusters that should not express them, true markers stay high in their own cluster, and cluster structure is recognizably the same.
    Red flag
    The true marker also drops in its own cluster, or the clustering changes a lot. That is over-correction. CellBender and DecontX in full mode are reported to be more prone to this when little ambient RNA is present.

What to do about it

SoupX with the raw matrix

When: You have the raw matrix, and the contamination is clearly gene-driven (hemoglobin, immunoglobulin) with a source cluster you can identify.

Supply raw and filtered matrices plus clusters. SoupX estimates the profile from empty droplets (10 UMIs or fewer), finds cells with zero endogenous expression of the contaminating genes (for example hemoglobin in non-erythroid cells), estimates the contamination fraction from them, and subtracts the background counts. Check the estimated fraction against your earlier ambient contribution plots before accepting it.

Caveat: It needs the raw matrix and a reasonable clustering. In the benchmark packet, SoupX in reduced mode was the more cautious option when ambient RNA is minimal.

CellBender

When: You want the most precise estimate of background noise and the best marker gene detection, and you can afford the extra compute.

Run CellBender on the raw matrix and use the cleaned output in place of the filtered counts. In the comparison in lesson 43, CellBender gave the most precise background estimates and the highest improvement for marker gene detection among CellBender, DecontX and SoupX.

Caveat: In full mode it is more prone to over-correction when ambient RNA is absent. Verify with the before-and-after marker check.

DecontX

When: You only have filtered, clustered data, or you work in Scanpy and want a per-cell contamination estimate.

DecontX is a Bayesian method that estimates and removes contamination at the single-cell level. In Python, run pip install decontx-python for a pure Python implementation with Scanpy integration and near-perfect parity with the R version. A contamination probability of 0.05 is the threshold used for filtering decisions.

Caveat: It does not see the empty droplets, so it infers the soup from the clustered cells. Like CellBender in full mode, it can over-correct when there is little ambient RNA.

Gene-specific correction with scCDC

When: Contamination is concentrated in a few genes and varies a lot between cells, so a uniform fraction does not fit.

scCDC detects which genes are contaminated and corrects them specifically, accounting for heterogeneity across genes and cells instead of assuming uniform contamination. Run it on the genes your checks flagged and compare the result with the global methods.

Caveat: It is newer than SoupX and DecontX, so there are fewer tutorials and less community experience.

Filter instead of subtract

When: Only a minority of genes or cells are affected and you want to avoid changing the count matrix.

Following the OSCA guidance, remove genes where 10% or more of counts come from ambient RNA, or remove cells with more than 25% of reads attributed to ambient contamination. Do this after ambientContribMaximum() so the cutoffs are based on your own distribution.

Caveat: You lose genes that may matter biologically, such as a marker you wanted to test. Use it for exclusion from marker lists more than for wholesale removal.

Model the soup instead of subtracting it

When: The goal is differential expression and you worry that corrected counts distort the statistics.

The OSCA chapter notes that simple per-cell subtraction distorts statistical properties unpredictably and prefers model-based approaches that use offsets or observation weights. Keep the raw counts and include the ambient contribution as an offset or weights in the model.

Caveat: It takes more statistical setup than a one-line tool, and ready-made recipes are fewer.

When not to "fix" it

Do not correct when the signal is biology. Erythroid cells, plasma cells and tissue-resident cells really do express hemoglobin, immunoglobulin and their own markers, and a tissue with real red blood cell infiltration will have hemoglobin in the cluster that makes it. If the gene is high in the cluster that should make it and the ambient profile does not explain the low-level signal elsewhere, leave the counts alone. Also skip aggressive correction when the contamination estimate is small and your conclusions do not depend on the affected genes: over-correction can strip real expression, and full-mode CellBender and DecontX are reported as more prone to it when there is little ambient RNA. If you only have a filtered matrix, a tool that cannot see the empty droplets may do more harm than a clear note in the methods.

Five things experienced analysts do here

  1. Keep raw_feature_bc_matrix for every sample. It is the only place the soup lives, and without it you lose the most direct measurement of contamination.
  2. Name the contaminating gene set before you correct, then check those genes before and after. If you cannot say which genes should change, you cannot tell correction from damage.
  3. Compare the ambient share across samples before merging them. A single sample with a much higher soup share is a handling problem, and pooling it hides that.
  4. Be suspicious of any 'doublet' or 'mixed' cluster whose only evidence is low-level expression of a neighbor's markers. Check it against the soup profile first.
  5. Report which correction you ran and what the estimated contamination was, per sample. A reviewer who sees hemoglobin in every cluster will ask, and the answer should already be in your methods.

Questions people ask

Why do all my cells express hemoglobin genes?

Most likely ambient RNA: free transcripts from lysed red blood cells captured in every droplet. Compare the genes with the profile from ambientProfileEmpty() on the raw matrix. If hemoglobin ranks near the top of the soup, that is your answer.

How much ambient RNA is normal in scRNA-seq?

Background has been reported at 3 to 35% of total counts per cell, varying across replicates and cells. In controlled cross-species experiments SoupX estimated 1 to 2% of observed transcripts, with higher rates in complex biological samples. Compare your samples with each other, because the spread between them is more informative than any single cutoff.

Should I use SoupX, CellBender or DecontX?

If you have the raw matrix, start with SoupX or CellBender. CellBender gave the most precise background estimates and the highest marker gene improvement in the comparison used here, but full-mode CellBender and DecontX over-correct more when little ambient RNA is present. DecontX is the practical choice when you only have filtered, clustered data.

Can I run DecontX in Scanpy?

Yes. Install the pure Python implementation with pip install decontx-python; it integrates with Scanpy and is reported to match the R version closely. SoupX and FastCAR need the raw matrix and empty droplet annotations, while DecontX can work from filtered, clustered data.

Related pages

Related reading on the blog

Sources

  1. Chapter 5: Problems with ambient RNA, Orchestrating Single-Cell Analysis — ambientProfileEmpty, ambientContribMaximum, ambientContribNegative and the 10% and 25% filtering thresholds
  2. emptyDrops: Identify empty droplets, DropletUtils documentation — Default lower threshold of 100 UMIs, FDR 0.01 cell calling
  3. SoupX removes ambient RNA contamination from droplet-based single-cell RNA sequencing data — Three-step SoupX correction and the 1 to 2% cross-species estimate
  4. Decontamination of ambient RNA in single-cell RNA-seq with DecontX — Bayesian single-cell contamination estimation
  5. scCDC: a computational method for gene-specific contamination detection and correction — Gene-specific correction
  6. Removing ambient / contamination RNA from droplet scRNA-seq, omicverse — Which tools need the raw matrix and which work from filtered data
  7. DropletQC: improved identification of empty droplets and damaged cells — Nuclear fraction metric
  8. Cell Ranger Gene Expression Algorithm Overview — Cell Ranger cell calling integrates EmptyDrops
  9. Benchmarking ambient RNA removal across droplet and well-plate platforms — Over-correction risk of CellBender and DecontX in full mode

Part of the Ambient RNA series.