Glossary · Single-Cell and Spatial
Overclustering
Your clustering algorithm will happily split pure noise into more pieces, so the resolution number tells you nothing about whether a cluster is real.
By Ming "Tommy" Tang, Director of Bioinformatics in Big Pharma · Reviewed October 2026 · 2 min read
Also: over-clustering, underclustering
Definition
Overclustering is the fragmentation of a biologically coherent cell population into subclusters that have no genuine biological distinction, because the clustering parameters (usually the Leiden or Louvain resolution) were set finer than the data supports. The extra clusters are driven by settings, noise, or technical variables such as sequencing depth, cell cycle phase, dissociation stress, or doublets, not by distinct cell identity. Underclustering is the opposite failure: distinct populations, often rare ones, get merged into one cluster.
You meet overclustering the moment you run FindClusters() or sc.tl.leiden(), get a number of clusters, and have to decide whether to believe it. Both tools return a tidy partition at any resolution you give them. Neither tells you whether the partition means anything.
The resolution parameter has no built-in stopping rule. Push it high enough and any point cloud, including pure noise, gets carved into more pieces. So the decision to keep a cluster, name it, and run differential expression on it is yours. Everything downstream (annotation, markers, abundance comparisons) inherits that decision.
Why it matters
Every cluster you keep becomes a cell type in your figures and a group in your differential expression tests. Overclustered groups produce markers that are not lineage genes, inflated significance, and false discoveries, because you are testing cells split apart by the same data you are testing them on. A typical case: a monocyte population is split into three clusters, and the "markers" are ribosomal genes and a depth gradient. The paper then reports three monocyte states.
Getting it right is mostly a discipline, not a tool. Pick the lowest resolution at which every cluster passes two tests: it has markers you can name, and it appears in cells from more than one sample or donor. Those two tests, not the resolution value, decide whether a cluster is real. The same discipline protects you from the opposite error. A genuine rare population has distinct markers that show up across donors, so a cluster that passes both tests stays even if it is small.
Where people get it wrong
People treat the resolution value as the answer, for example "0.8 is the right resolution" or "Scanpy defaults to 1.0, so that is fine". Published ranges, like 0.4 to 1.2 for about 3,000 cells in Seurat, were written for one dataset size and do not transfer to 100k cells or a complex tissue. The other common mistake is trusting the UMAP: two clusters that sit far apart on the plot can be indistinguishable in expression space, and two that touch can be real. A split that looks clean in 2D still has to pass the marker test and the cross-sample test.
A concrete example
You have a PBMC-like object and Seurat gives 23 clusters at one resolution and 31 at the next. Sweep resolution, look at how clusters split in clustree, then test each candidate cluster for sample composition and named markers. Count climbing linearly with no plateau, a cluster that is over 90% from one donor, or markers that are only ribosomal, mitochondrial or stress genes (FOS, JUN, HSPA1A) all say you went too far. Run this after building the neighbor graph. Adjust the clustree prefix to match your assay.
library(Seurat)
library(clustree)
# sweep resolution on the existing kNN graph
obj <- FindClusters(obj, resolution = c(0.2, 0.4, 0.6, 0.8, 1.0, 1.2, 1.4, 1.6))
# do clusters plateau, or split and reshuffle?
clustree(obj, prefix = "RNA_snn_res.")
# is any cluster dominated by one donor?
round(prop.table(table(obj$RNA_snn_res.0.8, obj$donor), margin = 1), 2)
# can you name markers for each cluster?
Idents(obj) <- "RNA_snn_res.0.8"
markers <- FindAllMarkers(obj)Related terms
Questions people ask
- What is overclustering in single-cell RNA-seq?
It is when the clustering splits a real cell population into subgroups that differ only by noise or technical variables, not by cell identity. It happens because resolution has no stopping rule. The test is whether each cluster has nameable markers and appears in more than one sample or donor.
- How many clusters is too many?
No cluster count is too many by itself, and the right number depends on tissue and dataset size. Look instead at whether every cluster has named lineage markers and exists across multiple donors, and whether the cluster count plateaus as resolution rises. If it keeps climbing linearly with resolution, you are probably splitting noise.
- Why does Seurat give me too many clusters?
Usually the resolution in
FindClusters()is higher than the data supports, or the clusters are driven by depth, cell cycle, stress genes or doublets. The kNN graph settings matter too: the default k=20 can pull other cell types into a 50-cell population, while a much smaller k can over-split a 100k-cell dataset. Check what drives each split before you lower resolution.- How do I fix overclustering?
Match the fix to the cause. Lower resolution or merge clusters that share markers, use depth-aware normalization such as SCTransform for depth-driven splits, regress out cell cycle or stress scores, and remove doublets per sample before clustering.
- Is a small cluster always overclustering?
No. A real rare population has distinct markers and shows up in several donors or samples. An overclustering artifact brings no new marker signal and tends to track depth or batch.
Related pages
- Guide · How to Choose Cell QC Thresholds in Single-Nucleus RNA-seq
- Guide · Why Your UMAP Is Misleading You in Single-Cell RNA-seq
- Guide · How to Detect Batch Effects in Spatial Transcriptomics
- Guide · How to Find and Remove Doublets in Single-Nucleus RNA-seq
- Guide · How to Detect Integration Over-Correction in Spatial Transcriptomics
Related reading on the blog
Sources
- Overclustering in scRNA-seq: How to Check and Fix It — Definition, detection checks, root causes, red flags, and remediation by cause.
- Seurat - Guided Clustering Tutorial (PBMC 3K) — FindClusters() resolution of 0.4-1.2 for about 3k cells, and FindAllMarkers() for marker testing.
- scanpy.tl.leiden documentation — Scanpy Leiden resolution parameter defaults to 1.0.