Glossary · Programming and Math Basics
Seurat object
It's the single R container holding your counts, metadata, PCA, and UMAP, and knowing exactly where each of those lives is what lets you catch a broken pipeline before you publish a UMAP.
By Ming "Tommy" Tang, Director of Bioinformatics in Big Pharma · Reviewed September 2026 · 3 min read
Also: SeuratObject, Seurat v5 object
Definition
A Seurat object is the S4-class container the Seurat R package uses to hold everything about a single-cell (or spatial, or ATAC) dataset in one variable: one or more assays, cell-level metadata, and dimensionality-reduction results like PCA and UMAP coordinates. It is organized around a fixed set of cells, and each assay it holds represents one data modality, RNA counts, ADT protein counts, ATAC peaks, rather than one assay per sample. In Seurat v5, each assay stores its matrices as named layers (counts, data, scale.data, or per-sample variants) instead of the fixed slots used in v3 and v4. You pass this one object through the entire pipeline; functions like NormalizeData() and RunPCA() read from it and write their output back into it rather than returning separate variables.
You meet the Seurat object the moment you call CreateSeuratObject() or Read10X(), before you've done a single analysis step. From then on it's the one variable you pass to every function in the pipeline: NormalizeData(), RunPCA(), RunUMAP(), FindClusters(). Each of those functions reads from the object and writes its results back into it, so the object's internal layout isn't trivia, it's the map you need to know when a result looks wrong.
That map matters most at two moments: when you have more than one assay in the object (RNA plus ADT from CITE-seq, or RNA plus an "integrated" assay after batch correction), and when you merge multiple samples in Seurat v5, where counts don't automatically pool into one matrix. If you can't say which assay, which layer, and which cells a given number came from, you can't debug a suspicious result.
Why it matters
Get the assay and layer wrong and every downstream number is computed on the wrong data without an error message telling you so. The clearest case: after batch-correction workflows, a Seurat object commonly ends up with both an "RNA" assay and an "integrated" assay. Integrated values are built for clustering and visualization, not for differential expression, running FindMarkers() against the wrong default assay gives you fold changes computed on batch-corrected values instead of the original counts, and nothing in the output flags that you did it.
The same risk shows up after merging samples in Seurat v5. Because v5 assays can hold layers split per sample rather than one pooled matrix, merge()-ing several objects doesn't automatically give you one combined counts layer. If you run FindVariableFeatures() expecting a single pooled matrix, you may be computing variance across per-sample layers instead of the full dataset. Before trusting a downstream result, check DefaultAssay(obj) and inspect the assay's layer names, that's the two-second sanity check that catches both failure modes.
Where people get it wrong
The mistake is assuming $ works the same way everywhere on a Seurat object. obj$orig.ident works because Seurat exposes meta.data columns as a shortcut, so people generalize that and expect obj$counts to reach into the expression matrix the same way. It doesn't. Slots on the S4 object itself need @ (obj@meta.data), and getting into an assay requires [[ first (obj[["RNA"]]); only once you're inside that assay object does Seurat v5 let you use $ again, as in obj[["RNA"]]$counts. Mixing these up produces a "no applicable method" error at best, or silently returns NULL at worst if you're using an operator the object doesn't define for that path. Run slotNames(obj) when you're unsure what's actually addressable before guessing at an accessor.
A concrete example
Build a Seurat object from 10x output, then inspect its structure before running any analysis, this is the two-minute check that tells you which assay and layer everything downstream will use.
library(Seurat)
counts <- Read10X("filtered_feature_bc_matrix/")
obj <- CreateSeuratObject(counts = counts, project = "pbmc",
min.cells = 3, min.features = 200)
slotNames(obj)
DefaultAssay(obj)
head(obj@meta.data)
# raw counts for the RNA assay, v5 layer syntax
obj[["RNA"]]$counts[1:5, 1:5]
obj <- NormalizeData(obj)
obj <- FindVariableFeatures(obj)
obj <- ScaleData(obj)
obj <- RunPCA(obj)
obj <- RunUMAP(obj, dims = 1:15)
Embeddings(obj, reduction = "umap")[1:5, ]Related terms
- S4 object
- AnnData
- SingleCellExperiment
- default assay
Questions people ask
- What is a Seurat object?
It's the S4-class R container the Seurat package uses to hold everything about a single-cell dataset: one or more assays (each an expression matrix for a data type like RNA or ADT), cell-level metadata, and dimensionality-reduction results like PCA and UMAP coordinates. You create one with
CreateSeuratObject()and pass that same object through the entire analysis pipeline instead of juggling separate matrices.- What is the difference between counts, data, and scale.data in a Seurat object?
countsis the raw, unnormalized expression matrix straight from your aligner.datais that matrix afterNormalizeData(), typically log-normalized.scale.datais the z-scored, variance-stabilized versionScaleData()produces for PCA and heatmaps. In Seurat v5 these are stored as named layers inside an assay rather than fixed slots, and you get them withLayerData(obj, assay = "RNA", layer = "counts")or the backwards-compatibleGetAssayData().- What's the difference between a Seurat object and a SingleCellExperiment?
Both are container objects built for the same job, organizing expression data, cell metadata, and reduced dimensions in one place, but they're built by different ecosystems with different accessor conventions. A Seurat object exposes assays,
meta.data, and reductions via@; a SingleCellExperiment exposes assays,colData,rowData, andreducedDimsvia dedicated accessor functions likecolData(sce). Converting between them is common when you need a tool that only supports one.- How do I access data inside a Seurat object?
Use
@for S4 slots, likeseurat_obj@meta.datafor cell-level metadata. Use[[to pull out an assay by name, likeseurat_obj[["RNA"]], then chain into a layer. Seurat v5 also lets you reach a layer with$, as inobj[["RNA"]]$counts, but that shortcut only works once you're inside an assay object, not directly on the Seurat object itself.- What changed with Seurat v5 layers?
Pre-v5 assays had fixed slots (
counts,data,scale.data). Seurat v5 replaces that with a named list of layers, so an assay can hold arbitrary layers, including per-sample layers after a merge, or on-disk matrices, instead of only three fixed ones. Practically, this means merged multi-sample objects can carry split counts layers rather than one pooled matrix until you deliberately combine them.
Related pages
- Guide · How to Detect Batch Effects in Single-Nucleus RNA-seq
- Guide · How to Detect Integration Over-Correction in Single-Cell RNA-seq
- Guide · How to Sanity-Check Marker Genes and Cell Type Labels in Single-Cell RNA-seq
- Guide · How to Choose a Normalization Method in Single-Cell RNA-seq
- Guide · How to Avoid Pseudoreplication in Single-Cell RNA-seq
Related reading on the blog
Sources
- Seurat v5 Essential Commands — Layer-based assay architecture in v5, LayerData(), GetAssayData(), Cells(), FetchData(), Embeddings()
- Tools for Single Cell Genomics • Seurat - Official Documentation — Assay structure, Read10X() and ReadMtx() loading functions
- Reuse the single cell data! How to create a seurat object from GEO datasets — Worked example of CreateSeuratObject() from a sparse matrix and metadata, and the standard preprocessing chain