Sanity check · Variant Calling
How to Avoid 0-Based vs 1-Based Coordinate Errors in Variant Calling
A VCF position and a BED start are not the same number, and the mismatch silently drops the first base of every target interval.
By Ming "Tommy" Tang, Director of Bioinformatics in Big Pharma · Reviewed October 2026 · 5 min read
You have a VCF and a set of target regions, and you need to subset, annotate or intersect them. Somewhere in the pipeline a file was written from one format and read as another. Nothing crashed. The browser looks fine, because one base is invisible at screen scale.
What is at stake: a target BED built from VCF positions without subtracting one misses the first base of every interval. SNP annotation lands on the neighbouring base, motif and getfasta results come from the wrong sequence, and variants at region boundaries get dropped or kept for the wrong reason.
This page gives you checks you can run in under an hour to decide whether a given file is 0-based or 1-based, and whether your pipeline is off by one. It also gives the conversions to fix it.
What it looks like when it's happening
- A variant sits exactly on the first base of a BED interval and `bedtools intersect` or `bcftools view -R` reports it as outside the region.
- A region-restricted VCF has slightly fewer SNPs than expected, and the missing sites all sit at interval starts.
- `bedtools getfasta` returns sequences whose first base does not match the REF allele of the variant you built the interval from.
- A single-base BED interval `chr1 100 101` is annotated by one tool as position 100 and by another as position 101.
- Annotation of a known SNP returns the consequence of its neighbour: a synonymous call where you expect missense, or a motif disruption on the wrong base.
- A region of length 50 in your BED comes back as 49 or 51 bases after converting through GTF and back.
- Two tools that should agree on the same intervals differ by a handful of sites per million, concentrated at boundaries.
Why it happens
The formats were designed by different groups for different jobs. BED is 0-based and half-open: start is counted from 0 and end is exclusive, so length is end - start. VCF, GFF and GTF are 1-based and closed: both ends are included, so length is end - start + 1. SAM text is 1-based. The same first base of a chromosome is position 1 in VCF or GFF and 0-1 in BED.
The bug appears the moment you move numbers between these formats by copying columns. Take a VCF record at POS 1000 and write it to BED as 1000 1001. That interval covers 1-based position 1001, not 1000. Take a 1-based GTF exon 1000-1100 and write it as BED 1000 1100: you lose base 1000. The correct 1-based to 0-based conversion is start_0 = start_1 - 1 and end_0 = end_1. A single SNP at POS p becomes [p-1, p).
Mixing languages makes it worse. Python indexing is 0-based and matches BED; R indexing is 1-based and matches VCF and GFF. Code that slices a reference in one language using coordinates from a file meant for the other shifts by one without any error.
The damage is hard to see because it only bites at boundaries. A real bug in the eidolon tool tested 1-based VCF POS against 0-based BED intervals without conversion. The effect was about 0.009% of calls: 357 fewer SNPs and 4 extra deletions in HG002 data. That is small enough to pass a casual review and large enough to matter for a clinical exome, where the lost bases sit at the edges of exons.
The checks
Run them in order. Each one tells you what healthy looks like and what the problem looks like.
0/7 checked · saved in this browser
Decide the expected convention from the format before you look at the data. BED is 0-based half-open. VCF, GFF and GTF are 1-based closed. bcftools infers the convention from the file name:
.bedor.bed.gzis read as 0-based half-open,.vcfor.vcf.gzas 1-based. A BED file renamed to.txt, or a 1-based table named.bed, breaks that assumption.- Healthy
- Every interval file in the pipeline has an extension that matches the convention its numbers actually follow.
- Red flag
- A tab-delimited file named `.bed` that was generated from VCF or GTF columns, or a coordinate table with no clear extension that both bedtools and bcftools are asked to read.
Inspect the first lines and compute lengths. In a true BED, length is
end - startand a single base looks likechr1 99 100. If your BED contains many rows wherestart == end, it is probably 1-based closed single-base positions. A start of 0 can only exist in a 0-based file.bashhead -n 5 regions.bed awk '$2==$3' regions.bed | wc -l awk '$2==0' regions.bed | wc -l- Healthy
- No rows with `start == end`. Single-base features show as `p-1 p`.
- Red flag
- Many zero-length rows, or a BED made from variants where the start column equals VCF POS, which means POS was copied straight into the start column.
Pick one variant from the VCF whose REF base you can confirm, for example one you have looked at in IGV. Write it to BED with the documented conversion
[p-1, p). Runbedtools getfastaon that interval and compare the returned base to the VCF REF.bash# VCF line: chr1 1000 . A G printf 'chr1\t999\t1000\n' > snp.bed bedtools getfasta -fi genome.fa -bed snp.bed- Healthy
- The one-base sequence returned equals the REF allele in the VCF record.
- Red flag
- The returned base is the neighbouring one, or the REF allele matches only if you use `[p, p+1)`. A conversion somewhere in your pipeline is off by one.
Run
bedtools getfastaon the target BED and compare each output sequence length toend - start. A 50 bp interval must return 50 bases.bashbedtools getfasta -fi genome.fa -bed regions.bed -fo output.fasta awk '{print $3-$2}' regions.bed | head awk '/^>/{next}{print length($0)}' output.fasta | head- Healthy
- The two lists of lengths match row for row.
- Red flag
- Output is consistently one base longer or shorter than `end - start`, which points to a GTF-style closed interval being read as BED, or the reverse.
Write the boundary rule down and test it on three variants for a BED interval
[s, e): one atPOS = s, one atPOS = s + 1, one atPOS = e. A 1-based POSpis inside the interval whens < p <= e, equivalently whenp - 1is in[s, e). Run the same three through your tool or script and see which are included.- Healthy
- POS equal to `s` is outside, `s + 1` is inside, and POS equal to `e` is inside.
- Red flag
- POS equal to `s` is kept, or POS equal to `e` is dropped. Either means the code compares 1-based POS against 0-based bounds without conversion, as in the eidolon bug.
Count variants inside your regions two ways: with
bcftools view -R regions.bed(which treats.bedas 0-based) and withbedtools intersecton the VCF and BED. The counts should agree exactly. Then list the sites present in one set but not the other. Also compare both against your own script if you have one.bashbcftools view -H -R regions.bed calls.vcf.gz | cut -f1,2 | sort > a.txt bedtools intersect -u -a calls.vcf.gz -b regions.bed | grep -v '^#' | cut -f1,2 | sort > b.txt wc -l a.txt b.txt comm -3 a.txt b.txt | head- Healthy
- Identical counts and an empty `comm -3` output.
- Red flag
- A small discrepancy where every differing site has POS equal to an interval start or start + 1. That clustering at boundaries is the signature of an off-by-one.
For a larger test, build
[POS-1, POS)BED intervals for a few thousand SNPs, extract the bases withbedtools getfasta, and compare each to the VCF REF. Count mismatches. This tests the whole chain, including the reference build.- Healthy
- Essentially all SNP REF alleles match the reference base, ignoring soft-masking case.
- Red flag
- A large fraction of REF alleles mismatch, which is what a systematic one-base shift gives you on near-random sequence. A small fraction mismatching is more likely a build or contig-name problem.
What to do about it
Convert VCF positions to BED as [POS-1, POS)
When: You are building a target or mask BED from VCF records, for example to restrict calling or to run bedtools on a variant set.
Subtract one from the start and keep POS as the end, then sort and merge:
awk '!/^#/ {print $1"\t"$2-1"\t"$2}' input.vcf | sort -k1,1 -k2,2n | bedtools merge -i stdin
The IARC VCF-tricks one-liner uses the same structure, but as published it prints $2 and $2+1. That is [p, p+1), which shifts every site by one. Adjust it to $2-1 and $2, and confirm with the single-SNP round trip.
Caveat: Indels and multi-base REF alleles span `len(REF)` bases. The interval is `[POS-1, POS-1+len(REF))`, not `[POS-1, POS)`.
Convert 1-based closed intervals (GTF, GFF) to BED with start minus one
When: You are turning exon or gene features from GTF or GFF into a BED target file.
Use start_0 = start_1 - 1 and end_0 = end_1. In awk, for a GTF with start in column 4 and end in column 5:
awk 'BEGIN{OFS="\t"} !/^#/ {print $1, $4-1, $5}' annotation.gtf > exons.bed
Then sort and merge before using it as a capture or calling target.
Caveat: If the annotation came from a tool that already wrote BED-style starts, you will now shift by one in the other direction. Run the round-trip check on a known exon boundary first.
Convert BED to 1-based closed when a downstream tool needs it
When: A tool or collaborator expects 1-based closed coordinates, such as GFF-style tables or R code indexing a sequence.
Use start_1 = start_0 + 1 and end_1 = end_0:
awk '{print $1,$2+1,$3}' OFS='\t' input.bed > converted.tsv
Give the output a name that does not end in .bed, so bcftools and bedtools do not read it with the wrong convention.
Caveat: The result is no longer a valid BED. Anything that auto-detects by extension will misread it if you keep the `.bed` name.
Let the tool handle the conversion and stop doing it by hand
When: You only need to subset or intersect and do not need to write intermediate coordinate files.
Pass the original files with their correct extensions. bcftools treats .bed and .bed.gz as 0-based half-open and .vcf and .vcf.gz as 1-based. bedtools handles BED and GFF/GTF internally, so the same bedtools intersect command works with either. Delete any hand-written awk conversion that sits between two tools that already do this.
Caveat: This only holds when the extension and the content agree. A 1-based table renamed to `.bed` defeats it, and your own Python or R code still needs explicit conversion.
Fix language-level slicing explicitly
When: Your own Python or R code reads coordinates from a file and slices a sequence or array.
Python is 0-based and matches BED: seq[start:end] for BED intervals, and seq[pos-1:pos-1+len(ref)] for VCF POS. R is 1-based and matches VCF and GFF: substr(seq, start, end) for those, and start + 1 for BED starts. Write one small conversion function per project and test it against the single-SNP round trip.
Caveat: Libraries differ in what they return. Check the documentation of any pysam, pybedtools or Bioconductor object before assuming its convention.
When not to "fix" it
Do not convert coordinates when the files already agree and the numbers are right. If the round-trip check returns the VCF REF base and your boundary test passes, adding a further minus one will break a working pipeline. Also leave arithmetic alone when the shift has a different cause: a genome build mismatch (hg19 versus hg38) or a differently normalised indel representation can look like an off-by-one but is not fixed by editing the start column. And for display, genome browsers handle the conversion for you, so a one-base visual offset between a BED track and a VCF track is not by itself evidence of an error.
Five things experienced analysts do here
- Write the convention into the filename or a README for every interval file you create, for example `targets.0based.bed`, so the next person does not have to guess.
- Keep one canonical test SNP with a known REF base in your pipeline's tests and run the `[POS-1, POS)` getfasta round trip after any change to a step that writes coordinates.
- Look at boundary sites first. An off-by-one never shows in the bulk of a region, so count variants at `POS == start` and `POS == end` separately when you compare two outputs.
- Prefer a tool's native reader over awk. bcftools and bedtools already know the conventions, and every hand-written conversion is one more place to be wrong.
- Treat a discrepancy of a few hundred sites in millions as a real signal, not noise. The eidolon bug was about 0.009% of calls and was still a genuine boundary error.
Questions people ask
- Is BED 0-based or 1-based?
BED is 0-based and half-open: start is counted from 0 and end is exclusive, so length is
end - start. The first base of a chromosome ischr1 0 1. This is the opposite of VCF, GFF and GTF, which are 1-based and closed.- Is VCF 0-based or 1-based?
VCF POS is 1-based. A SNP at POS p becomes the BED interval
[p-1, p). For indels and multi-base REF alleles, the interval spans the length of the REF allele.- How do I convert a GTF to BED without an off-by-one error?
GTF is 1-based closed and BED is 0-based half-open, so subtract one from the start and keep the end:
start_0 = start_1 - 1,end_0 = end_1. In awk that isprint $1, $4-1, $5. Check one known exon boundary withbedtools getfastabefore trusting the file.- Does bcftools handle 0-based and 1-based coordinates automatically?
For region and target files it decides from the file extension:
.bedor.bed.gzis read as 0-based half-open,.vcfor.vcf.gzas 1-based. That only helps if the extension matches the real content of the file.- Is SAM 0-based or 1-based, and what about BAM?
SAM text is 1-based. The book lesson on samtools describes BAM as 0-based, reflecting its binary storage. Tools such as samtools and pysam present positions according to their own documented conventions, so check the documentation of the library you use.
Related pages
Sources
- 0-Based vs. 1-Based Genomic Coordinates — Conversion formulas and the boundary condition s < p <= e for a VCF position inside a BED interval.
- bedtools 2.31.0 Documentation - Overview — BED is 0-based half-open; bedtools handles BED and GFF coordinates internally.
- bcftools Manual Page — bcftools picks the coordinate convention from the .bed or .vcf extension.
- BedRecord::contains is given 1-based VCF POS against 0-based BED intervals (Issue #770) — Real off-by-one bug: 357 fewer SNPs and 4 additional deletions in HG002 data.
- VCF-tricks Repository — awk and bedtools one-liner for turning VCF records into merged BED intervals.
Part of the 0-based vs 1-based coordinates series.