Glossary · RNA-seq
Count matrix
The raw, unnormalized table every RNA-seq analysis starts from, and the single most common place a downstream result quietly goes wrong.
By Ming "Tommy" Tang, Director of Bioinformatics in Big Pharma · Reviewed September 2026 · 3 min read
Also: gene count matrix, expression matrix
Definition
A count matrix is a table of non-negative integers where each row is a gene (or transcript) and each column is a sample, in bulk RNA-seq, or a single cell, in scRNA-seq. Each entry is the number of sequencing reads or unique molecules assigned to that gene in that sample or cell. It is the direct output of a quantification tool (STAR+featureCounts, HTSeq, Salmon, Kallisto for bulk; Cell Ranger, STARsolo, or alevin-fry for single-cell), produced before any normalization step. The same structure recurs across omics: proteins-by-samples in proteomics, peaks-by-cells in scATAC-seq.
You meet the count matrix the moment your alignment or pseudoalignment job finishes: rows of gene IDs, columns of sample names or cell barcodes, cells full of integers. It looks finished. It is not. Every decision downstream, normalization method, which genes you filter, which test you run, depends on treating this matrix as raw material rather than a result.
In bulk RNA-seq it lands in your lap as a genes x samples table from STAR+featureCounts, HTSeq, Salmon, or Kallisto. In single-cell work it is genes x cells, produced by Cell Ranger, STARsolo, or alevin-fry, and it is mostly zeros. Knowing which kind of matrix you are holding, and what has and hasn't been done to it, decides whether your PCA, heatmap, or differential expression call means anything.
Why it matters
Raw counts are not expression levels. A gene with more reads than another in the same sample might just be longer, and a sample with more total reads will inflate every gene's count relative to a shallower-sequenced sample. Compare raw counts across samples or genes without correcting for sequencing depth, gene length, and RNA composition, and you will call genes "differentially expressed" that are only differentially sequenced.
The failure case that actually bites people: pulling two RNA-seq datasets off GEO, one quantified with STAR+HTSeq, the other with Salmon, and concatenating the count matrices before running DESeq2. Genes that show zero reads under one pipeline show real counts under the other, and reference genome version (hg19 vs hg38) changes the gene model entirely. The differential expression signal you get back is a mix of biology and pipeline artifact, and nothing in the DESeq2 output will flag which is which.
Where people get it wrong
People use "count matrix" and "normalized expression matrix" interchangeably, and it causes two opposite mistakes: feeding raw counts into a heatmap or correlation plot where they should be log-normalized or variance-stabilized first, or feeding CPM/TPM-normalized values into DESeq2 or edgeR, which expect raw integer counts because their statistical models fit a negative binomial to count data, not to pre-scaled ratios.
In single-cell work, the parallel mistake is treating the matrix as dense. Because more than 90% of a typical scRNA-seq gene-by-cell matrix is zeros, loading it as a base R matrix or NumPy array instead of a sparse format (Matrix::dgCMatrix in R, scipy.sparse.csr_matrix in Python) can blow up memory by more than 20-fold for no analytical benefit, and it also invites you to treat the zeros as missing data rather than as a mix of true non-expression and dropout, which changes how you should normalize and cluster.
A concrete example
Given raw counts from featureCounts and a DESeq2 design, generate size-factor-normalized counts for comparison, and separately compute CPM if you just need a depth-adjusted look at one gene across samples. DESeq2 wants the raw count matrix as input, not something you've already normalized.
# raw_counts: genes x samples matrix of integers, straight from featureCounts/HTSeq
dds <- DESeq2::DESeqDataSetFromMatrix(countData = raw_counts,
colData = sample_info,
design = ~condition)
dds <- DESeq2::estimateSizeFactors(dds)
normalized_counts <- DESeq2::counts(dds, normalized = TRUE)
# quick depth-only check for one gene across samples, not for DE testing
cpm_counts <- edgeR::cpm(raw_counts)Related terms
Questions people ask
- What is a count matrix in RNA-seq?
It's the table of raw read counts produced by a quantification tool: rows are genes, columns are samples (bulk) or cells (single-cell), and each value is how many reads or molecules mapped to that gene. It has not been normalized yet, so raw values aren't directly comparable across samples or genes.
- What is the difference between a count matrix and a normalized expression matrix?
The count matrix holds raw integers straight from alignment/quantification. A normalized expression matrix (CPM, TPM, or DESeq2's median-of-ratios normalized counts) has been scaled to correct for sequencing depth and, for some methods, gene length, so values are comparable across samples.
- What format is a gene count matrix usually stored in?
For bulk RNA-seq, a plain tab- or comma-delimited genes-by-samples table (or an R data.frame/matrix) is typical. For single-cell data, it's usually a sparse matrix on disk (10x-style
.mtxwith barcode and feature files, or.h5/.h5ad), loaded into memory asMatrix::dgCMatrixin R orscipy.sparse.csr_matrixin Python because most entries are zero.- Can I combine count matrices from different studies?
Only if they were generated with the same quantification pipeline and reference genome build. Mixing STAR+HTSeq counts with Salmon counts, or hg19-based counts with hg38-based counts, introduces systematic differences per gene that look like biology but are pipeline artifacts.
- Why is my single-cell count matrix mostly zeros?
Most genes are genuinely not expressed in a given cell, and on top of that, capture and amplification during library prep miss some transcripts that are actually present (dropout). Both effects combine to push sparsity above 90% in typical scRNA-seq matrices, which is why they're stored and analyzed as sparse matrices rather than dense ones.
Related pages
- Guide · How to Detect Ambient RNA Contamination in Single-Cell RNA-seq
- Guide · How to Detect Batch Effects in Single-Cell RNA-seq
- Guide · How to Detect Batch Effects in Single-Nucleus RNA-seq
- Guide · How to Choose a Normalization Method in ATAC-seq
- Guide · How to Choose a Normalization Method in Bulk RNA-seq
Related reading on the blog
Sources
- Seurat - Guided Clustering Tutorial (pbmc3k) — Definition of count matrix in scRNA-seq and sparse matrix storage efficiency
- How to convert raw counts to TPM for TCGA data and make a heatmap across cancer types — Raw count matrix definition and normalization to TPM for cross-sample comparison
- Batch Effect: To Correct or Not for Bulk RNA-seq Data — Why the raw count matrix should not be pre-corrected before differential expression testing