▸ Chatomics Field GuideWhat They Don't Teach You →

Glossary · RNA-seq

Transcript-to-gene mapping (tx2gene)

A two-column table decides whether your Salmon counts become gene-level data or a silent mess, and most failures come from version suffixes.

By Ming "Tommy" Tang, Director of Bioinformatics in Big Pharma · Reviewed October 2026 · 2 min read

Also: tx2gene, transcript gene map

Definition

A transcript-to-gene mapping, usually called tx2gene, is a table with exactly two columns: transcript ID first, gene ID second. tximport uses it to collapse transcript-level estimates from quantifiers such as Salmon, Sailfish, Kallisto and Oarfish into gene-level matrices. The transcript IDs in the table must match the IDs in the quantification files exactly. Column names do not matter, column order does.

You meet tx2gene the first time you run Salmon or kallisto and then try to load the output into DESeq2 or edgeR. The quantifier gives you one row per transcript. Your differential expression question is about genes. tx2gene is the bridge, and tximport will not build the gene-level object without it.

Where the table comes from decides how often it breaks. You can build it from an annotation package, from a TxDb made from your GTF, or with a Unix one-liner. Whichever you pick, it must come from the same annotation, and the same release, that you used to build the Salmon index. That one constraint causes most of the errors people post about.

Why it matters

A wrong tx2gene rarely fails loudly. Either tximport complains about transcripts missing from the table, or you work around it and genes quietly lose counts.

Concrete case: your quant.sf files use versioned Ensembl IDs like ENST00000xxxxxx.5, but your tx2gene was built with a tool that stripped versions. Nothing matches. The tempting fix is ignoreTxVersion = TRUE, but that is for a genuine mismatch you have inspected, not a default. Compare the IDs in both files first, then fix the table or set the flag deliberately.

A second case: a table built on gene_name loses many lncRNAs and other transcripts that have no gene_name in the GTF. Use gene_id, then convert to symbols afterwards.

Where people get it wrong

People treat tx2gene as a gene ID converter and mix it up with symbol or Entrez mapping. It is not. It maps transcripts to genes within one annotation, and the gene column should be the stable gene ID, for Ensembl the versioned gene_id. Converting to symbols is a separate step after summarization, because one ID can map to several symbols and the reverse.

The second mistake is assuming every pipeline needs it. RSEM and StringTie already write gene-level output, so there is nothing to map. Building a tx2gene for them adds a way to go wrong without adding information.

A concrete example

You quantified with Salmon against an Ensembl GTF and have no annotation package for the species. Build the table from the GTF with versions kept, transcript ID in column one and gene ID in column two. Then check that the IDs in the table look like the IDs in quant.sf before you run tximport. The field positions assume the attribute order of the Ensembl GTF in the lesson, so inspect your own file first.

bash
# gene IDs with version, from transcript rows
awk -F "\t" '$3 == "transcript" { print $9 }' myensembl.gtf | \
  tr -s ";" " " | cut -d " " -f2,4 | sed 's/\"//g' | \
  awk '{print $1"."$2}' > genes.txt

# transcript IDs with version
awk -F "\t" '$3 == "transcript" { print $9 }' myensembl.gtf | \
  tr -s ";" " " | cut -d " " -f6,8 | sed 's/\"//g' | \
  awk '{print $1"."$2}' > transcripts.txt

# transcript first, gene second
paste transcripts.txt genes.txt > tx2genes.txt

# sanity check: do the IDs look alike?
head -n 3 tx2genes.txt
cut -f1 quant.sf | head -n 4

Related terms

Questions people ask

How do I make a tx2gene file from a GTF?

Pull the transcript ID and gene ID from the transcript rows of the GTF and write them as two columns, transcript first. You can do this with awk and paste, or in R by building a TxDb with makeTxDbFromGFF() and using select(). Keep version numbers if your quantification files have them.

What format does tximport expect for tx2gene?

A data.frame with two columns: transcript ID, then gene ID. The column names are ignored but the order is not. The transcript IDs must match the ones in the quant files from your quantifier.

Do I need tx2gene for RSEM or StringTie?

No. Both write gene-level output directly, so there is nothing to summarize. You need tx2gene for transcript-level quantifiers such as Salmon, Sailfish, Kallisto and Oarfish.

When should I set ignoreTxVersion to TRUE in tximport?

Only when you have confirmed that the transcript IDs in your quant files and your tx2gene differ just by version suffix and you cannot rebuild the table. The default is FALSE. Setting it blindly can hide a real annotation mismatch.

Why are many transcripts missing from my tx2gene?

A common cause with Ensembl GTFs is building the table from gene_name, which lncRNAs and some other transcripts lack. Use gene_id for the gene column, then convert to symbols after summarization.

Related pages

Related reading on the blog

Sources

  1. Importing transcript abundance with tximport — tx2gene format, which tools need it, and TxDb and ensembldb creation methods
  2. Many transcripts missing from tx2gene created from Ensembl GTF file — Missing lncRNA transcripts when gene_name is absent; makeTxDbFromGFF
  3. Error message when importing quant.sf files: ignoreTxVersion parameter — When to use ignoreTxVersion
  4. Bioconductor - tximport — Current release version 1.40.0 and length offset behaviour
  5. Swimming downstream: statistical analysis of differential transcript usage following Salmon quantification — Transcript-level estimates are more uncertain than gene-level counts