▸ Chatomics Field GuideWhat They Don't Teach You →

Glossary · Genomics and Variants

ChIP-seq

A ChIP-seq peak file always looks like an answer, so the real skill is checking it against biology you already know before you believe it.

By Ming "Tommy" Tang, Director of Bioinformatics in Big Pharma · Reviewed October 2026 · 3 min read

Also: chromatin immunoprecipitation

Definition

ChIP-seq (chromatin immunoprecipitation followed by sequencing) uses an antibody to pull down a specific protein or histone modification together with the DNA it is bound to, then sequences that DNA. Reads pile up at the genomic locations where the target was bound, and a peak caller such as MACS3 turns those pileups into a list of binding sites or modified regions. It gives genome-wide binding sites for transcription factors, histone marks, and other DNA-binding proteins in one experiment. Think of the antibody as a fishing hook: you only catch what it grabs, so the result is only as good as the hook.

You will meet ChIP-seq the first time someone hands you a FASTQ folder and says "call the peaks", or the first time you want to reuse a published dataset from GEO. The pipeline is well worn. Align, call peaks with MACS, make a bigWig, look at it in a genome browser. Every step will run and produce a file, even when the experiment failed.

That is the problem. A peak caller has no idea whether your antibody pulled down the right protein. It reports enrichment over background, and a bad antibody produces enrichment too. What decides your analysis is not the pipeline. It is whether the peaks land where the biology says they must.

Why it matters

The cost of skipping validation is a clean-looking peak set built on nothing. Take an estrogen receptor ChIP-seq. You must see peaks at known targets such as TFF1 or GREB1. If they are missing, there are three explanations: the ChIP failed, the antibody is not specific, or the sample had no ER binding. The pipeline cannot tell you which one. Your knowledge of the biology can.

Antibody quality is a real source of this failure. Many antibodies cross-react, and some H3K9me3 antibodies recognize H3K27me3. One study tested 14 anti-NLRP3 antibodies against knockout and wild-type cells and only a few passed, the rest being nonspecific noise. Without a knockout control, you would not have known.

Mark type also changes how you read results. Point-source marks such as H3K4me3 and H3K27ac give consistent calls across peak callers and saturate at about 2.5 million reads or fewer. Broad marks give highly variable calls and need about 10 million reads or more. A caller choice or depth that is fine for one mark can quietly fail for another.

Where people get it wrong

People treat a successful pipeline run as evidence of a successful experiment. MACS will call peaks on a failed ChIP, an off-target antibody, or a noisy input, and a bigWig will render either way. The fix is to check known positive loci first, then look at QC metrics (FRiP, NSC and RSC, NRF, IDR across replicates). A second confusion is using one caller and one set of settings for every mark: broad marks need a broad-peak-aware approach such as SICER or MACS in broad mode, not the point-source defaults. The packet does not give threshold values for FRiP, NSC or RSC, so use the ENCODE guidelines for your assay rather than a number from memory.

A concrete example

You have an aligned, indexed ChIP BAM and a matched input BAM from a transcription factor experiment. Before calling anything final, make normalized tracks and open them at positive control loci. For an ER ChIP, load TFF1 and GREB1 and confirm the ChIP track is enriched over input there. If the signal is flat, stop and work out whether the ChIP, the antibody, or the sample is at fault before you interpret any other peak.

bash
samtools index sample.bam

# normalized ChIP signal, 20 bp bins for higher resolution
bamCoverage -b IP.bam -o IP.bw --binSize 20 --normalizeUsing RPKM

# ChIP vs input, log2 ratio
bamCompare -b1 ChIP.bam -b2 Input.bam -o ChIP_vs_Input.bw

Related terms

Questions people ask

What is ChIP-seq used for?

It maps where a protein or histone modification sits on the genome. Typical uses are transcription factor binding sites and histone marks such as H3K4me3 and H3K27ac. You can also pair it with RNA-seq, for example with BETA, to predict which genes a transcription factor directly regulates.

What are the main steps of a ChIP-seq analysis pipeline?

Download and QC the reads, align to a reference genome such as hg38, call peaks, then normalize and generate tracks for visualization. Outputs are BAM files (with a .bai index), BED files for peaks, and bigWig files for coverage. nf-core/chipseq packages these steps, and a Snakemake pipeline can reprocess GEO data.

How do I know my ChIP-seq worked?

Check for peaks at known targets first, such as TFF1 or GREB1 for estrogen receptor. Then look at FRiP, NSC and RSC, NRF, and IDR between replicates. Antibody specificity should be validated with knockout versus wild-type controls before the experiment.

What is the difference between ChIP-seq and CUT&Tag?

Both map protein-DNA interactions, but they use different chemistry. Crosslink ChIP-seq uses formaldehyde and sonication, while CUT&Tag cuts and tags DNA in place using an antibody-guided tagmentation step. The packet lists CUT&Tag as one of four approaches, with native ChIP-seq and ChIPmentation, but does not give a head-to-head benchmark, so check the literature for your target.

Which peak caller should I use for ChIP-seq?

MACS3 is the most highly cited and a sensible default for point-source marks and transcription factors. For broad marks, results vary widely between callers, so consider SICER or run more than one and compare. ChIP-AP combines MACS2, HOMER, SICER and SPP for consensus calls.

Related pages

Related reading on the blog

Sources

  1. Comparative analysis of commonly used peak calling programs for ChIP-Seq — Peak caller behavior for point versus broad marks, read depth saturation, nf-core/chipseq and ChIP-AP
  2. File formats for Peak Visualization — BAM, BED and bigWig formats; bamCoverage and bamCompare commands and normalization options
  3. Reviving BETA for Python 3: Integrating ChIP-seq and RNA-seq — BETA integrates ChIP-seq and RNA-seq to predict direct TF targets