▸ Chatomics Field GuideWhat They Don't Teach You →

Glossary · Craft and Good Practice

Gene fusion

Most fusion calls from RNA-seq are artifacts, and the number of calls per sample tells you which kind of pipeline you are running.

By Ming "Tommy" Tang, Director of Bioinformatics in Big Pharma · Reviewed October 2026 · 2 min read

Also: fusion transcript, chimeric transcript

Definition

A gene fusion is an aberrant gene formed when two separate genes are joined, through a translocation, an interstitial deletion, or a chromosomal inversion. It can produce chimeric genomic DNA, a chimeric mRNA (the fusion transcript, also called a chimeric transcript), and a truncated or chimeric protein. In RNA-seq you never see the rearrangement itself. You see reads whose two halves map to different genes, and you infer the fusion from them.

You will meet gene fusions the first time someone asks you to run a fusion caller on tumor RNA-seq. STAR-Fusion, Arriba and STAR-SEQR are the usual tools. They hand you a table of candidate gene pairs, and the table looks authoritative whether or not it is.

The decision you have to make is which rows are real. Fusion calling is a filtering problem more than a detection problem. A raw candidate list is mostly artifacts, and what you do between the caller's output and the word "fusion" in a report is the actual analysis.

Why it matters

Real fusions drive decisions in cancer: they can disrupt a gene's original function, create a new protein function, or drastically change the expression of a partner gene. A false call sends someone chasing a target that does not exist. A missed call hides one that does.

The cheapest sanity check is the count. High-quality RNA fusions number in the dozens per sample at most. If your pipeline reports 1000 fusions per sample, it is drowning in false positives, and no downstream statistic will rescue it. A statistician looking at that table sees a distribution. Someone with domain knowledge sees an impossible result and goes back to the filters.

Where people get it wrong

People treat "supported by reads" as "real". Read support only says that some reads map across two loci. The three main sources of false positives all produce exactly that: misaligned reads from repeats and segmental duplications, read-through and trans-splicing between adjacent genes, and template switching during polymerase activity. A second confusion is mixing up levels: a fusion gene (the paired loci), a fusion site (the exact junction) and a fusion isoform (a transcript variant spliced at that junction) are different objects, and counting them interchangeably inflates or deflates your numbers.

A concrete example

A caller reports a long list for one tumor sample. You check the count first, then the evidence type, then the genomic context, then look at the reads yourself. The STAR-Fusion and Arriba thresholds below are the documented defaults, so a call that barely clears them deserves suspicion rather than trust.

bash
# Triage checklist for a fusion candidate table (no tool-specific flags assumed)
#
# 1. Count: dozens per sample is plausible; ~1000 means the filters are off.
# 2. Evidence: need split (junction) reads, not only discordant pairs.
#    STAR-Fusion defaults: >=1 junction read, >=2 total supporting fragments,
#    >=3 reads if the breakpoint is not at an annotated exon junction,
#    and >=0.1 fusion fragments per million (ffpm).
#    Arriba default: min_support = 2.
# 3. Context: drop calls in repeats, pseudogenes, paralog pairs, and
#    adjacent genes (read-through).
# 4. Consensus: keep calls made by several tools (Fusion InPipe used
#    at least 3 of 5: Arriba, CICERO, deFuse, FusionCatcher, STAR-Fusion).
# 5. Look: open the breakpoint in IGV and read the alignments.

Related terms

Questions people ask

What is a gene fusion?

A gene fusion is a hybrid gene formed when parts of two separate genes are joined, usually by a translocation, deletion or inversion. It can yield a chimeric mRNA and a truncated or chimeric protein. In cancer, fusions can disrupt a gene's normal function, create a new one, or change the expression of a partner gene.

Why does RNA-seq fusion detection give so many false positives?

Three sources dominate: reads misaligned in repeats and segmental duplications, read-through or trans-splicing between neighboring genes, and template switching during polymerase activity. Paralogous genes and promiscuous partner genes add more. Each of these creates reads that span two loci without any genomic rearrangement.

How many gene fusions should I expect per sample?

High-quality RNA fusions typically number in the dozens per sample. A result near 1000 per sample means the pipeline is reporting false positives. Tighten read-support thresholds and artifact filters before interpreting anything.

Which tool should I use for fusion detection?

STAR-Fusion, Arriba and STAR-SEQR are ranked among the most accurate and fastest for cancer transcriptomes, and Arriba won the DREAM SMC-RNA Challenge. Pick one as your primary and confirm important calls with a second tool. Requiring agreement from several callers reduces false positives further.

What counts as evidence for a fusion?

Chimeric (split or junction) reads that overlap the fusion junction, and discordant read pairs that bridge or span it. Split reads pin down the exact breakpoint, so a call supported only by discordant pairs is weaker. Always inspect the alignments in a genome browser such as IGV.

Sources

  1. Accuracy assessment of fusion transcript detection via read-mapping and de novo fusion transcript assembly-based methods — Main sources of false positives, types of read evidence, and error sources.
  2. Fusion InPipe, an integrative pipeline for gene fusion detection from RNA-seq data in acute pediatric leukemia — Consensus of at least 3 of 5 tools, filtering strategies, and tool ranking.
  3. Arriba (GitHub repository) — Default min_support of 2 and the DREAM SMC-RNA Challenge result.
  4. Characterization of fusion genes and the significantly expressed fusion isoforms in breast cancer by hybrid sequencing — Definition, functional consequences, and the fusion gene, site and isoform levels.