Sanity check · Variant Calling
How to Fix Chromosome Naming Mismatches in Variant Calling
Your BED file and your VCF look fine and bedtools says they share nothing: the cause is usually one prefix, and the repair is a text edit, not a re-run.
By Ming "Tommy" Tang, Director of Bioinformatics in Big Pharma · Reviewed October 2026 · 5 min read
You intersect your exome capture BED with your VCF, or your called variants with an annotation track, and bedtools prints nothing. No error, no warning, an empty file. You check the files, they both look like genomic intervals, and you start to suspect the pipeline.
Usually nothing is wrong with the calls. One file says chr1 and the other says 1, or one says chrM and the other says MT. Tools that match contig names as exact strings find no shared chromosomes and report zero overlaps. If you miss it, you either throw away a good run or, worse, ship a result where the mitochondrial variants quietly vanished.
This page gets you from "zero overlaps" to a confirmed cause in a few minutes, tells you whether it is a naming-only problem or a real genome build mismatch, and gives you the one-liner or mapping table that fixes it without re-running alignment or calling.
What it looks like when it's happening
- `bedtools intersect -a calls.vcf.gz -b targets.bed` writes an empty file, or `-c` reports 0 for every interval, even though you know the regions overlap.
- `samtools idxstats aln.bam | grep chrM` prints nothing, while `samtools idxstats aln.bam | cut -f 1` shows a contig called `MT`.
- GATK stops with an error that a contig is not found in the reference sequence dictionary, as soon as you pass a BED, interval list or VCF built on a different naming scheme.
- `bcftools view -r chr20 sample.vcf.gz` returns no records, but `-r 20` returns plenty.
- Per-chromosome variant counts look normal for chr1 to chr22, X and Y, but the mitochondrial row is missing or zero.
- An annotation step finishes with every variant marked intergenic or unannotated, because no variant contig name matched a transcript contig name.
- `bcftools view -h` shows `##contig` lines with bare numbers while your BED file starts every line with `chr`.
Why it happens
Reference genomes are distributed by different groups with different conventions. UCSC-style references such as hg19 and hg38 use chr1, chr2 and so on. Ensembl-style references, and b37 which carries GRCh37 coordinates, use bare 1, 2. The mitochondrial contig is the least consistent: chrM is most common in UCSC hg19 and hg38, MT appears in Ensembl FASTAs and GTFs, and chrMT shows up sometimes in NCBI material. Your BAM inherits whatever names were in the FASTA used for alignment, your VCF inherits them from the BAM, and your BED or annotation file inherits them from wherever you downloaded it.
Most interval tools treat the contig name as an opaque string. chr20 is not 20. When no name in file A appears in file B, there is nothing to compare, so the overlap count is zero. The tool did exactly what it was told, which is why there is no error. bedtools behaves this way for BED and VCF inputs, and bcftools requires sequence names to match exactly without enforcing strict contig ordering.
GATK is the opposite. It validates every contig against the reference dictionary and fails loudly when one is missing. That is useful, but it makes people fix the visible error by editing one input and moving on, while another tool in the same workflow keeps silently dropping data. The mismatch you see in GATK and the mismatch you do not see in bedtools or bcftools are the same bug.
The mitochondrial case is where it hurts most. If you rename autosomes with a blanket prefix rule but leave MT as it is, or the reverse, chrM variants fail to join to the target list or annotation, and they disappear from the result with no message. Mitochondrial variants are only a few rows among millions, so nobody notices a gap that small. There is also a trap one level down: a naming fix is only safe if the build is the same. hg19 and hg38 differ in coordinates, and renaming labels on the wrong build gives you confident, wrong positions.
The checks
Run them in order. Each one tells you what healthy looks like and what the problem looks like.
0/6 checked · saved in this browser
Look at the contig names actually present in each input before running anything fancy. For a BED file take column 1, for a VCF take the first column of the records, for a BAM use the index. Compare the three lists by eye: prefix style, and the mitochondrial name.
bashcut -f 1 targets.bed | sort -u | head -30 bcftools view -H sample.vcf.gz | cut -f 1 | uniq | head -30 samtools idxstats sample.bam | cut -f 1- Healthy
- All three lists use the same style: all `chr`-prefixed or all bare, with the same mitochondrial name in each.
- Red flag
- One list has `chr1`, `chr2`, the other has `1`, `2`. Or autosomes agree and one file says `chrM` while another says `MT`.
The headers record the names and lengths the reference had when the file was made. Pull the
@SQlines from the BAM and the##contiglines from the VCF and compare them side by side.bashsamtools view -H sample.bam | grep "^@SQ" | head bcftools view -h sample.vcf.gz | grep "^##contig" | head- Healthy
- BAM and VCF headers list the same contig names in the same order, with identical lengths.
- Red flag
- Names differ between BAM and VCF, or the VCF has no `##contig` lines at all, which means you cannot trust it to carry the reference it was called against.
Search for both spellings in the BAM header and note which one exists and how long it is. A grep for only
chrMreturns nothing on anMTreference, and that empty result is the whole problem.bashsamtools view -H sample.bam | grep -E "SN:(MT|chrM)\b"- Healthy
- Exactly one line, with the name you expected from your reference and a length of 16,569 bp for the rCRS used by b37 and GRCh38.
- Red flag
- No line at all, a name that differs from your BED or annotation, or a length of 16,571 bp, which points to the GRCh37 MT sequence rather than the rCRS and tells you something about which reference you really have.
Write each file's unique contig names to a text file, sort them, and use
comm -12to list names that appear in both. Do this for each pair: BED against VCF, VCF against annotation, BAM against intervals.bashcut -f 1 targets.bed | sort -u > bed_chroms.txt bcftools view -H sample.vcf.gz | cut -f 1 | sort -u > vcf_chroms.txt comm -12 bed_chroms.txt vcf_chroms.txt | wc -l- Healthy
- The shared count is roughly the number of primary contigs you work with, covering autosomes, X, Y and the mitochondrial genome.
- Red flag
- Zero shared names, or 22 and 24 but with the mitochondrial contig missing. Zero means the whole file is on the wrong convention. A near miss means a single special-case contig is wrong.
Before you rename anything, compare contig lengths, not names. Chromosome 1 is 249,250,621 bp in hg19/GRCh37 and 248,956,422 bp in hg38/GRCh38. Read the length for chr1 or 1 from the
@SQline in the BAM and the##contigline in the VCF and match them.bashsamtools view -H sample.bam | grep -E "SN:(chr)?1\b" bcftools view -h sample.vcf.gz | grep -E "ID=(chr)?1,"- Healthy
- Same length in both files, and that length matches the build of the BED or annotation you are joining. If so, this is naming only, and relabeling is safe.
- Red flag
- Different lengths for what is nominally the same chromosome. That is a genuine build mismatch, and no amount of renaming will fix it.
After renaming, count records in one region in each file, for example a single chromosome from the VCF against the BED. Then run the intersect and confirm the count is plausible against what you expect, such as the share of variants in your exome targets.
bashbcftools view -H renamed.vcf.gz | cut -f 1 | uniq | head bedtools intersect -a renamed.vcf.gz -b targets.bed | wc -l- Healthy
- A non-empty intersect, and a first-column list that matches the other files for every contig including the mitochondrial one.
- Red flag
- Still empty, or non-empty but with the mitochondrial contig absent from the output. The second case is the one that gets past people.
What to do about it
Add or strip the prefix on a BED file
When: The BED file is the odd one out and the mismatch is only `chr` versus no `chr` on standard contigs.
To add the prefix, use the tab-aware awk form so the rest of the line is untouched: awk 'BEGIN {OFS="\t"} {$1="chr"$1; print}' my.bed > my_with_chr.bed. To strip it, use sed 's/^chr//' my.bed > my_no_chr.bed. Then fix the mitochondrial line separately (see the mapping-table fix), because a prefix rule turns MT into chrMT, which matches nothing.
Caveat: The prefix rule is blind to special contigs and to header or `track` lines. Check the first lines of the output and the mitochondrial row afterward.
Rename contigs in a VCF with a mapping table
When: You need a reliable fix for the VCF, especially when the mitochondrial contig differs or when your reference has extra contigs.
Write a two-column file of old_name new_name pairs separated by whitespace, one per line, covering 1 to 22, X, Y and the mitochondrial contig. Then run bcftools annotate --rename-chrs mapping.txt sample.vcf.gz -o renamed.vcf.gz, index the output, and re-run the checks. The same table works in the other direction if you swap the columns.
Caveat: Contigs not listed in the table keep their old name, so a partial table gives you a half-renamed file. Index the result again, because the old index no longer matches.
Rename contigs in a BAM header only after the lengths agree
When: Contig lengths in the BAM match the reference you want, and only the labels differ.
Dump the header with samtools view -H sample.bam > header.sam, edit the SN: names in that text file, and apply it with samtools reheader -h new_header.sam old.bam > new.bam. Re-index the BAM and run samtools idxstats to confirm.
Caveat: `samtools reheader` changes labels only, not coordinates. If the build differs, you will get a valid-looking file with the wrong positions. Confirm lengths first.
Fix the reference, not the inputs, when GATK refuses a contig
When: GATK reports a contig that is not in the reference dictionary and you control which reference the whole workflow uses.
Pick one reference flavor for the workflow and make every interval list, known-sites VCF and resource bundle file use that flavor's names. Rename the small inputs to match the reference rather than the other way around. If you work with GRCh38 alt contigs, note that contig names contain colons in some forms, and the -L option uses a colon as a delimiter, so you need to append :1+ to select a whole alt contig.
Caveat: Renaming a big BAM to fit a dictionary is expensive and easy to get wrong. If the BAM was aligned to a different reference flavor than the rest of the resources, switching resources is usually cheaper.
Convert coordinates or re-align when the build is actually wrong
When: The length check shows a real hg19 versus hg38 difference, or the BAM was aligned to the wrong reference flavor.
For called variants or intervals on the wrong build, use a liftover tool such as CrossMap.py to convert coordinates to the target build. If the reads themselves were aligned to the wrong flavor, go back to the raw reads and re-align.
Caveat: Liftover drops or splits some records, and it does not rescue alignment done against a mismatched reference. Count records before and after, and report what was lost.
When not to "fix" it
Do not rename contigs when the lengths show a real build difference. Relabeling chr1 on an hg19 file so it matches an hg38 BED produces a clean intersect with wrong coordinates, which is worse than an empty one. Also leave alone files that are meant to differ: if an annotation intentionally covers only primary contigs, then variants on alt, decoy or unplaced contigs having no overlap is expected, and forcing them in with a rename table adds noise. Finally, if the only unmatched records are on one contig you do not analyze, such as the mitochondrial genome in a nuclear-only exome analysis, confirm that this is a deliberate choice and write it down instead of editing files.
Five things experienced analysts do here
- Run the first-column and header checks the moment you receive any new BED, VCF or annotation file, before it goes into a pipeline. It takes seconds and the empty-output failure is silent.
- Keep one `mapping.txt` per reference pair in your project, including the mitochondrial line, and reuse it. Do not rebuild the sed expression from memory each time.
- Never grep for just `chrM`. Search for both `MT` and `chrM` in headers, and treat a missing mitochondrial contig as a finding, not a detail.
- After any rename, look at the mitochondrial row in the output specifically. Autosome counts will look right while chrM has gone missing.
- Make intersect steps fail loudly: add a check that the output is non-empty and that the shared-contig count is above zero, so a naming mismatch stops the run instead of producing a clean empty file.
Questions people ask
- Why does bedtools intersect return nothing without an error?
bedtools compares chromosome names as exact strings, so if one file says
chr1and the other says1, no interval can overlap and the result is empty. It does not treat that as an error. Print column 1 of each file and compare.- How do I convert chr1 to 1 in a VCF?
Use
bcftools annotate --rename-chrs mapping.txt sample.vcf.gz -o renamed.vcf.gzwith a file ofold_name new_namepairs. This updates the header and the records together. Piping through sed can work, but it is easy to damage lines that are not the records you meant to change.- Is chrM the same as MT in variant calling?
They label the same mitochondrial genome in different reference conventions:
chrMin UCSC hg19 and hg38,MTin Ensembl. A blanket prefix rule will not map one to the other, so include it explicitly in your rename table. Also check the length, since 16,569 bp is the rCRS and 16,571 bp is the GRCh37 MT sequence.- Why does GATK complain about contigs when bcftools does not?
GATK checks every contig against the reference sequence dictionary and stops on a mismatch. bcftools requires names to match exactly but does not enforce strict contig validation the same way, so a mismatch shows up as missing records instead of an error.
- Can I rename chromosomes in a BAM without re-aligning?
Yes, if the contig lengths already match the target reference and only the labels differ. Edit the header and apply it with
samtools reheader. If lengths differ, the build is wrong and you need liftover or re-alignment.
Related pages
- Guide · How to Catch a Genome Build Mismatch in Variant Calling
- Guide · How to Catch a Sample Swap in Variant Calling
- Guide · How to Detect Contamination in Variant Calling
- Guide · How to Avoid 0-Based vs 1-Based Coordinate Errors in Variant Calling
- Guide · How to Tell If You Sequenced Deep Enough in Variant Calling
- Glossary · Variant calling
Related reading on the blog
Sources
- Genome Build Mismatches in Variant Calling (hg19/hg38) — Contig naming differences, chr1 lengths for hg19 and hg38, MT lengths, samtools reheader, liftover.
- bcftools Manual Page — annotate --rename-chrs and exact sequence name matching.
- Reference Genome Components (GATK) — GRCh38 contig names with colons and the :1+ suffix for -L.
- Getting list of chromosome names from indexed BAM file (Biostars) — samtools idxstats piped to cut -f 1 to list contig names.
Part of the Chromosome naming series.