▸ Chatomics Field GuideWhat They Don't Teach You →

Sanity check · Single-Cell RNA-seq

Why Your UMAP Is Misleading You in Single-Cell RNA-seq

Distances between islands, island size and the direction of a smear on your UMAP are artifacts of the embedding, and here is how to test each claim before it reaches a figure legend.

By Ming "Tommy" Tang, Director of Bioinformatics in Big Pharma · Reviewed October 2026 · 6 min read

You ran RunUMAP() or sc.tl.umap(), got a beautiful plot with islands and a long tail, and now you are about to write "cluster A is most similar to cluster B because they sit next to each other" or "these cells are differentiating along the tail". That sentence is the pitfall. A UMAP keeps nearby cells nearby. It does not keep far cells at honest distances.

If you miss this, the claims that fail review are the ones built on geometry: which populations are related, which cluster is bigger, which direction a lineage runs. A reviewer who has seen the same plot drawn with different settings will ask for the test behind the picture, and "it looks that way" is not a test.

This page gives you a set of checks you can run in an hour on an existing Seurat or Scanpy object. Each one replaces a visual impression with a number or a table computed in the high-dimensional data.

What it looks like when it's happening

  • Two clusters sit far apart on the UMAP, and your draft says they are the most distinct populations, but you have not computed any distance or differential expression between them.
  • One cluster looks huge on the plot and a rare population looks tiny, and the figure legend or discussion treats area as abundance.
  • A smear or tail runs from one island to another, and you or a collaborator calls it a differentiation trajectory.
  • Two clusters touch through a thin bridge of cells instead of being separated by a clean gap.
  • Coloring the UMAP by total UMI counts or number of genes detected shows a gradient that lines up with cluster boundaries.
  • A cluster is made almost entirely of cells from one sample, and the plot shows it as its own island.
  • The same data rerun with a different seed, n_neighbors or min_dist gives islands in different places, different arrangements and sometimes a different number of visible groups.
  • The top markers of a cluster are ribosomal or mitochondrial genes, or stress genes such as FOS, JUN and HSPA1A, with no lineage marker.

Why it happens

UMAP takes a cell-by-gene matrix with thousands of dimensions, usually reduced to the top PCs first (in Seurat, RunUMAP() takes PCA dimensions as input, for example dims = 1:10), and squeezes it into two. It does this by building a neighbor graph and finding a 2-D layout that keeps each cell next to its neighbors. That objective rewards local fidelity. It has no term that says the gap between two distant groups should mean anything.

So the global arrangement is an outcome of optimization, not a measurement. Two islands can end up far apart in 2-D without being far apart in expression space, and the island that looks big may just have been given more room. This is the same reason the same hand looks completely different from another angle. PCA is the contrast: it is linear, axes are ranked by variance, and distances between points reflect how much two cells differ overall. UMAP and t-SNE trade that interpretability for the ability to show local structure.

Your data adds its own traps. Cells with similar sequencing depth cluster together by library size alone, and dropout is not random: cells with fewer UMIs have more zeros, which can look like on/off biology when depth differs between batches or cell types. Ambient RNA, reported at 3 to 35% of counts per cell, puts the same transcripts into many droplets and blurs marker specificity. Doublets and stressed cells create bridges and extra islands. The embedding shows all of this faithfully as shape, and shape reads as biology to a human eye.

The tool documentation says the same thing. Scanpy warns that UMAPs should not be over-interpreted, and Seurat cautions that visualization cannot fully represent the complexity of the underlying data. Resolution adds another layer: a high-resolution clustering can look over-clustered on the plot while a low-resolution one lumps distinct identities, and neither impression is a test.

The checks

Run them in order. Each one tells you what healthy looks like and what the problem looks like.

0/9 checked · saved in this browser

  1. Plot the same embedding colored by total UMI counts, number of genes detected and percent mitochondrial reads. Then do the same for sample, batch and condition. Look at whether any gradient aligns with a cluster boundary or the base of a smear. This takes a few minutes and catches the most common false island.

    r
    FeaturePlot(seu, features = c("nCount_RNA", "nFeature_RNA", "percent.mt"))
    DimPlot(seu, group.by = c("sample", "condition"))
    Healthy
    Technical covariates are spread across islands without lining up with boundaries. Each biological island contains cells from several samples.
    Red flag
    A clean gradient in UMI counts or genes detected that follows a cluster boundary or runs along a tail, which points to depth rather than cell state.
  2. Build a cluster-by-sample table of cell counts, convert to row proportions, and look for any cluster dominated by one sample. A cluster that is more than 90% one sample or batch is probably a batch artifact, not a cell type. Do this per cluster, not by eye on the plot.

    r
    tab <- table(seu$seurat_clusters, seu$sample)
    round(prop.table(tab, margin = 1), 2)
    Healthy
    Each cluster draws from most of the samples, with proportions that reflect how many cells each sample contributed.
    Red flag
    A cluster that is almost entirely one sample or one batch, even if it looks like a well-separated island.
  3. Compute cluster centroids in the same PCs you gave to UMAP, and compute centroids in the 2-D UMAP coordinates. Take the pairwise distances between centroids in each space and compute the rank correlation. Plot one against the other. This turns 'these look far apart' into a number you can quote.

    r
    pcs <- Embeddings(seu, "pca")[, 1:10]   # same dims as RunUMAP
    um  <- Embeddings(seu, "umap")
    cl  <- seu$seurat_clusters
    cent <- function(m) apply(m, 2, function(x) tapply(x, cl, mean))
    d_pca  <- as.vector(dist(cent(pcs)))
    d_umap <- as.vector(dist(cent(um)))
    cor(d_pca, d_umap, method = "spearman")
    plot(d_pca, d_umap)
    Healthy
    You see that the ordering of pairs is only loosely preserved. Use the PCA distances, or better a differential expression result, as the quantity you report.
    Red flag
    Pairs that look far apart on the UMAP but are close in PCA space, or the reverse. If the correlation is weak, no sentence in your text should rely on UMAP distance.
  4. Run ElbowPlot() and confirm the PCs you pass to the embedding are the ones carrying signal. The Seurat tutorial uses RunUMAP(pbmc, dims = 1:10) and describes top 10 to 30 PCs as typical, but the right number is dataset-specific. Rerun the embedding with a few different values around your choice.

    r
    ElbowPlot(seu, ndims = 40)
    seu <- RunUMAP(seu, dims = 1:10)
    DimPlot(seu)
    Healthy
    The islands you plan to talk about are present at several reasonable PC counts.
    Red flag
    An island or bridge that appears only at one PC setting. It is a product of that setting, not a stable feature.
  5. Regenerate the embedding with a different random seed, then with a few different values of n_neighbors and min_dist. The packet does not give recommended values, so choose a spread that is clearly smaller and clearly larger than your current setting. Compare which islands and neighborhoods persist, and which only exist at one setting.

    python
    import scanpy as sc
    for nn, md in [(15, 0.5), (50, 0.5), (15, 0.1)]:
        sc.pp.neighbors(adata, n_neighbors=nn)
        sc.tl.umap(adata, min_dist=md, random_state=0)
        sc.pl.umap(adata, color=['leiden', 'total_counts'], title=f'nn={nn}, min_dist={md}')
    Healthy
    Local groupings persist across settings even though the global layout rotates, shifts and rescales. Neighbors of a given cell stay the same cells.
    Red flag
    Conclusions that depend on the relative position of islands, on a bridge that vanishes in another run, or on a visible gap that closes when parameters change.
  6. For each cluster, run differential expression against the rest and read the top genes. Check whether they are lineage markers or housekeeping genes (ribosomal, mitochondrial) and stress genes such as FOS, JUN and HSPA1A. Also look at where known markers of the populations you expect appear on the plot.

    r
    markers <- FindAllMarkers(seu, only.pos = TRUE)
    head(markers[order(markers$cluster, -markers$avg_log2FC), ], 20)
    Healthy
    Each cluster has markers you can name from the literature, and cells expressing a given marker co-localize in one island.
    Red flag
    Top genes are ribosomal, mitochondrial or stress genes without lineage markers, or a marker you trust is spread across several islands.
  7. Plot a strongly cell-type-specific marker from one island and check whether it shows low-level expression across unrelated islands. Estimate background with a tool built for it (CellBender, DecontX or SoupX) and compare the embedding before and after.

    Healthy
    Markers are high in their own population and near zero elsewhere, and the bridge or halo does not depend on the decontamination step.
    Red flag
    Highly expressed genes from one abundant cell type showing up broadly in other clusters, or a thin bridge that shrinks once background counts are removed.
  8. Take cluster sizes from the metadata, not from the plot. Report counts and per-sample proportions, and check whether a cluster looks big or small on the UMAP relative to its real cell count.

    r
    table(seu$seurat_clusters)
    round(prop.table(table(seu$sample, seu$seurat_clusters), margin = 1), 3)
    Healthy
    Reported abundance comes from a table. The plot is used only to show where the cells are.
    Red flag
    Any sentence that says a population is larger, smaller, more abundant or more diverse because of how much space it occupies on the UMAP.
  9. If you want to say cells move along a smear, choose a pseudotime method that builds a graph, set the root on biological grounds (a progenitor marker, a known starting state), and ask whether the ordering is supported there. Then check that pseudotime is not just correlated with total UMI counts or genes detected.

    Healthy
    Pseudotime orders cells in a way that agrees with known marker dynamics and is not explained by depth.
    Red flag
    A trajectory that is only visible on the UMAP, a root chosen because it is at one end of the plot, or pseudotime that tracks library size.

What to do about it

Move the quantitative claims out of the plot

When: Your text says clusters are similar, distant, related or more abundant based on the embedding.

Replace each geometric statement with a computed one. For similarity, report distances between cluster centroids in PCA space or the number of differentially expressed genes between the pair. For abundance, report cell counts and per-sample proportions from the metadata. Keep the UMAP for showing where cells and markers sit.

Caveat: PCA distances depend on how features were selected and how many PCs you used, so state both. They are interpretable, not objective.

Make the plot robust before you publish it

When: The embedding will go into a figure and readers will form impressions from it.

Fix the PCs, rerun with several seeds and settings of n_neighbors and min_dist, and show the version whose local structure is the stable one. Put a technical-covariate panel (UMI counts, genes detected, sample) next to the main panel in a supplement. Label that the layout is for navigation.

Caveat: A stable local structure still says nothing about gap size or island area, so the caption has to say so.

Remove or model the technical axis, then re-embed

When: Coloring by UMI counts, genes detected or sample shows a gradient or a sample-specific island that aligns with cluster boundaries.

Handle depth and ambient RNA before clustering: use a normalization that accounts for library size, run an ambient RNA tool such as CellBender, DecontX or SoupX, and review whether the cluster still exists. For sample-specific clusters, test whether integration merges them and whether any markers remain that are biological.

Caveat: Batch correction can over- or under-correct depending on depth, and it can erase the condition effect you want to study. Check marker genes before and after.

Confirm clusters with markers and a test

When: You plan to name or describe a cluster, or a bridge suggests two clusters may be one.

Run differential expression between neighboring clusters and look at the top genes. If the difference is housekeeping genes, stress genes or a depth gradient, lower the clustering resolution or merge. If lineage markers separate them, keep them.

Caveat: Resolution has no single correct value. Scanpy's tutorial notes high resolution can look over-clustered and low resolution can lump distinct identities, so justify your choice with markers.

Use a trajectory method for trajectory claims

When: You see a smear between islands and the question is about differentiation or state transitions.

Run a pseudotime or trajectory method that builds a graph on the high-dimensional data and takes a root you can defend biologically. Report the ordering of known marker genes along pseudotime. Present the UMAP as a backdrop for that result.

Caveat: Any root choice is an assumption, and the method will return an ordering even when no real continuum exists. Treat agreement with known markers as the evidence.

When not to "fix" it

If the thing you are worried about is local, you may not need to fix anything. A small island that is pure in one marker, appears across several seeds and parameter settings, and has lineage-specific top genes is a population, whatever the shape of the plot. Rerunning the embedding until it looks tidier does harm here, because you are tuning the picture instead of testing the claim.

Do not strip out a technical covariate if it is also the biology. Depth can differ systematically between cell types, and activated or large cells legitimately carry more RNA. Regressing it out or forcing integration can erase a real difference between conditions. Look at marker genes before and after, and only correct when the structure disappears for reasons you can explain.

Five things experienced analysts do here

  1. Ask of every sentence that cites the UMAP whether it is about neighbors or about distance. Neighbor statements are mostly safe. Distance, size and direction statements need a number computed elsewhere.
  2. Keep a standard QC panel next to every embedding you show: total UMI counts, genes detected, percent mitochondrial, sample and condition. If a collaborator asks 'what is that island', the answer should be one glance away.
  3. Rerun with a new seed before you trust an arrangement. If the layout changes but the neighbors do not, you are looking at the real information.
  4. Choose the number of PCs with `ElbowPlot()` and write the number down. The embedding, the neighbor graph and the clusters should all use the same PCs, or you are comparing different spaces.
  5. Treat a clean-looking gap as a hypothesis and a thin bridge as a warning. Check the bridge cells for doublet signal, ambient RNA and depth before you give the clusters different names.

Questions people ask

Can I interpret distances between clusters on a UMAP?

No. UMAP preserves local neighborhoods, not global geometry, so two islands that look far apart may be no more different than two that look close. Compute distances in PCA space or run differential expression between the clusters, and report those instead.

Does cluster size on a UMAP reflect how many cells there are?

Not reliably. The area an island occupies depends on how the embedding spreads cells, not just on how many there are. Count cells per cluster from the metadata and report proportions per sample, never the visual area.

UMAP vs t-SNE: which is better for scRNA-seq?

Neither is a measurement tool. Both are designed to preserve local structure and both distort global distances, so treat them as navigation aids. Pick one for readability, and do the biology in the PCA space or the neighbor graph.

Can I infer a trajectory from a smear on a UMAP?

Not until a pseudotime method with a rooted graph supports it. A smear can be a depth gradient, a doublet bridge, ambient RNA or a real transition, and the picture alone cannot tell these apart. Color the smear by total UMI counts and genes detected first, then run a trajectory method with a justified root.

Do n_neighbors and min_dist change my biological conclusions?

They change the picture, and that is the point: if your conclusion changes when you change them, it was never in the data. Rerun with several settings and different seeds, and keep only the claims that survive and that are backed by a test outside the plot.

Related pages

Related reading on the blog

Sources

  1. Detecting Overclustering in UMAP — Thin bridges between clusters, housekeeping top markers, single-sample clusters above 90%, stress gene clusters.
  2. Seurat PBMC 3K Tutorial — RunUMAP on PCA dimensions, ElbowPlot, and the caution that visualization cannot fully represent the data.
  3. Scanpy Clustering Tutorial — UMAPs should not be over-interpreted; resolution changes how over- or under-clustered the plot looks.
  4. Best Practices Analysis of 10x Genomics Single Cell RNA-seq Data — Vendor best-practice guide for the upstream analysis that feeds the embedding.

Part of the UMAP misreading series.