ChIP-seq guidelines and practices of the ENCODE and modENCODE consortia.
The main results reproduced, with only marginal, non-material deviations.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- Nothing in this column.
- 🟡Could not use the authors’ exact input data
- 🟡Reported values were only indirectly comparable
- 🟡A deviation arose in the data or preprocessing
- 🟡A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The deviation was non-trivial in magnitude
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Reproduced the paper's core computational pipeline (aligner -> MACS peak calling of ChIP vs Input) end-to-end on a real ENCODE dataset (PolII ChIP-seq vs Input, K562 cells) drawn from the paper's own PRJNA63441 accession tree, using MACS3 (the current incarnation of the paper's cited MACS repo). The pipeline ran without errors and produced 31,419 called peaks; FRiP (25.7%) and NRF (0.90/0.95) both independently satisfy the paper's own stated QC guideline thresholds, providing meaningful (if qualitative/order-of-magnitude) support for the paper's central methodological claims. Dataset profiling of PRJNA63441 shows it is a live, still-growing umbrella BioProject: the paper's reported 478 ChIP-seq datasets (2012 snapshot) has grown to 14,516 unique ChIP-seq experiments today, expected given ENCODE's continued production for over a decade, not a reproducibility defect. NSC/RSC cross-correlation QC and IDR replicate-concordance analysis -- the paper's other flagship QC metrics -- were not attempted in this pass and are flagged as missing/out of scope, not as failures.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-07-28
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-07-31no human curator yet
- Last updated
- 2026-07-31
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusThe paper asks what experimental and analytical standards are needed to make ChIP-seq experiments reliable and comparable, and reports the working guidelines (antibody validation, replication, sequencing depth, quality metrics, data reporting) developed by the ENCODE and modENCODE consortia to address the substantial variability in how such experiments are performed, scored, and archived.
- ★ ENCODE/modENCODE define a set of working standards and guidelines for ChIP-seq covering antibody validation, experimental replication, sequencing depth, data/metadata reporting, and data quality assessment. resource
- ★ Antibodies must pass a primary characterization assay (immunoblot for transcription factors, with immunostaining as alternative) plus at least one of five secondary assays (knockdown/RNAi, second antibody or complex member, epitope tag, affinity enrichment plus mass spectrometry, or motif analysis). method
- ★ Only a minority of commercially available transcription-factor antibodies both meet the characterization guidelines and work in ChIP-seq. finding
- ★ Two independent biological replicates are set as the consortium standard, with the irreproducible discovery rate (IDR) method used to assess replicate agreement and set peak thresholds. method
- ★ The number of called peaks for a typical point-source factor continues to increase with sequencing depth rather than saturating, because ChIP signal strength is a continuum rather than a discrete set of positive sites. finding
- ★ A minimum of 20 million mapped reads is set for ENCODE point-source transcription-factor ChIP experiments, since signal enrichment plateaus and peaks discovered beyond that depth are progressively weaker. method
- Chromatin-associated proteins fall into point-source, broad-source, and mixed-source classes that require different analytical approaches. mechanism
- Epitope tagging with large clones (fosmids/BACs) provides near-physiological expression and is an effective alternative when suitable antibodies are unavailable, though overexpression can cause occupancy of non-physiological sites. method
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| ChIP-seq (chromatin immunoprecipitation followed by high-throughput DNA sequencing) | Human and mouse cell lines/tissues (e.g., K562, GM12878, HepG2), D. melanogaster embryos, C. elegans; >100 cell types across four organisms | none (native factor immunoprecipitation); epitope-tagged constructs in some cases | Genome-wide binding sites/peaks of transcription factors and histone modifications; peak counts and fold-enrichment | — |
| ChIP-chip (ChIP followed by DNA microarray hybridization) | Small-genome organisms used by modENCODE (D. melanogaster, C. elegans) | none | Enriched genomic regions relative to differentially labeled reference DNA | DNA microarray |
| Immunoblot (Western blot) — primary antibody characterization assay | Nuclear extract from GM12878 and K562 cells | none | Presence/absence of band at expected molecular weight and detection of cross-reacting bands (e.g., SIN3B at 133 kDa) | Santa Cruz sc13145 (pass) and sc996 (fail) anti-SIN3B antibodies |
| Immunoprecipitation followed by immunoblot | Nuclear lysates of K562 cells | none | Efficiency and specificity of immunoprecipitation of the expected-size band (TBLR1, 56 kDa) across input, IP, and depleted lanes | Abcam ab24550 anti-TBLR1 antibody |
| Immunofluorescence / immunostaining (alternative primary assay) | Cultured cells | none | Nuclear localization pattern; pass/fail quality-control call | — |
| Immunoprecipitation followed by mass spectrometry (secondary characterization) | Whole-cell lysates of K562, GM12878, and HepG2 | none | Peptide identification confirming the immunoprecipitated protein (SP1, ~106 kDa) from a Coomassie-stained excised gel band | Santa Cruz sc-17824 anti-SP1 antibody; MASCOT (Matrix Science); Scaffold (Proteome Software, Inc.) |
| Binding-site motif enrichment analysis (secondary characterization) | ENCODE human ChIP-seq data sets for 85 transcription factors | none | Motif fold-enrichment relative to all DNase-accessible sites (bias-corrected with shuffle motifs) and motif representation as percentage of analyzed peaks | Motif search stringency 4–6; peaks from IDR analysis at 0.01 cut-off |
| Peak calling as a function of sequencing depth (saturation analysis) | 11 deeply sequenced human ENCODE ChIP-seq data sets (e.g., MAFK in HepG2) | none (in silico read subsampling in 2.5 million-read increments) | Number of called peaks and median fold-enrichment of newly called peaks versus number of uniquely mapped reads | PeakSeq (0.01% FDR cut-off) |
- – Only about one-fifth of tested commercially available transcription-factor antibodies met the characterization guidelines and also functioned in ChIP-seq. ~20% (44 of 227)
- ▲ Peak counts continued to increase with sequencing depth for nearly all data sets; clear saturation was observed for only one factor with few binding sites, and one data set yielded >150,000 peaks at 100 million mapped reads. >150,000 peaks at 100 million mapped reads
- – Signal enrichment plateaus with greater sequencing depth; at 20 million mapped reads, median enrichments are typically five- to 13-fold. five- to 13-fold median enrichment
- ▼ Peaks newly identified beyond 20 million reads are much weaker, with enrichment about 20% of that of the strongest peaks; additional peaks at three- to sevenfold enrichment can still be found at much greater depth, likely low-affinity or open-chromatin sites. ~20% of strongest-peak enrichment; three- to sevenfold
- – Of 85 transcription factors with ENCODE ChIP-seq data, 60% had a data set meeting the fourfold motif enrichment standard, and 96% of those data sets met the >10% motif representation standard. 60% of 85 factors; 96% of qualifying data sets
- – Secondary antibody characterization data submitted to the consortia were dominated by mass spectrometry, followed by second-antibody/epitope-tag/complex-member ChIP, motif analysis, and siRNA knockdown. 55% / 28% / 10% / 7%
- – All six GFP-tagged C. elegans factors tested to date complemented null mutants, supporting epitope tagging as a viable alternative to native antibodies. 6 of 6
- – Initial RNA polymerase II ChIP-seq experiments showed that more than two replicates did not significantly improve site discovery, motivating the two-biological-replicate standard.
- count more than a thousand individual ChIP-seq experiments for more than 140 different factors and histone modifications in more than 100 cell types in four organisms (Scope of ENCODE/modENCODE ChIP-seq production)
- count 145 polyclonal and 43 monoclonal antibodies (Antibodies used to successfully generate ChIP-seq data as of October 2011)
- count ~20% (44 of 227) (Tested commercial transcription-factor antibodies meeting guidelines and working in ChIP-seq)
- other 55% mass spectrometry, 28% second antibody/epitope tag/complex member, 10% motif analysis, 7% siRNA knockdown (Secondary characterization data type submitted with consortia antibodies)
- fold_change five- to 13-fold median enrichment (Median peak enrichment at 20 million mapped reads across deeply sequenced data sets)
- fold_change three- to sevenfold (Enrichment of additional peaks found only at much greater sequencing depths)
- other 60% of factors meet fourfold motif enrichment; 96% meet >10% motif representation (Motif-based secondary validation across ENCODE transcription factors (IDR 0.01 peaks))
- pvalue P < 0.05; 0.0% protein FDR and 0.0% peptide FDR (MASCOT probability-based peptide matching and Scaffold analysis for SP1 IP–mass spectrometry)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This is a methods/guidelines resource paper (not a hypothesis-testing study) describing ENCODE/modENCODE consortium standards for ChIP-seq experimental design, antibody validation, and data quality assessment. Reproducibility across biological replicates was assessed using the irreducible/irreproducible discovery rate (IDR) framework, peak calling was evaluated with FDR-based thresholds, and antibody/target identification via mass spectrometry used probability-based peptide matching with FDR control. Results are largely reported as summary statistics (percentages of data sets/antibodies meeting a criterion, fold-enrichment values) across many aggregated data sets rather than as single pairwise significance tests.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Irreproducible discovery rate (IDR) analysis (Li et al. 2011) | assessing agreement/reproducibility between biological replicate ChIP-seq experiments and setting peak thresholds | two biological replicates per ChIP measurement (consortium standard) | not stated |
| PeakSeq peak calling with FDR cut-off | calling ChIP-seq peaks as a function of sequencing depth for 11 ENCODE data sets (Fig. 3) | 0.01% FDR cut-off stated | not stated |
| MASCOT probability-based peptide matching | identifying immunoprecipitated proteins by mass spectrometry (e.g., SP1 antibody validation, Fig. 2D) | — | P < 0.05 threshold stated |
| Scaffold protein/peptide FDR analysis | post-processing of mass spectrometry identifications for antibody characterization | — | 0.0% protein FDR and 0.0% peptide FDR stated as thresholds used |
| Motif fold-enrichment analysis with sequence-bias correction (shuffled motifs) | evaluating transcription-factor binding motif enrichment at ChIP-seq peaks relative to DNase-accessible sites (Fig. 2E, 85 factors) | 85 transcription factors; peaks defined by IDR analysis at 0.01 cutoff | not stated |
-
Replicate reproducibility was assessed using the IDR framework applied to ranked peak lists from paired biological replicates.↳ Could also: Other reproducibility metrics such as the Pearson/Spearman correlation of signal tracks between replicates, or Jaccard/overlap statistics on called peak sets, could also be used — These would provide complementary views of replicate concordance (e.g., genome-wide signal correlation vs. rank-based peak agreement) and are commonly reported alongside or instead of IDR in ChIP-seq quality assessments
-
Peak significance was determined using an FDR-based cut-off from PeakSeq.↳ Could also: Alternative peak callers with different statistical models (e.g., MACS with a Poisson/local-lambda background model, or SPP) could also be applied — Different peak-calling statistical models make different assumptions about background read distribution, so comparing across callers can illustrate how peak set composition depends on the chosen model
-
Motif enrichment was summarized as fold-enrichment relative to DNase-accessible sites, corrected using shuffled-sequence controls.↳ Could also: A formal statistical test (e.g., a hypergeometric or Fisher's exact test comparing motif occurrence in peaks vs. a matched background) could also be used to accompany the fold-enrichment value — A formal test statistic with a p-value would complement the fold-enrichment metric by quantifying how unlikely the observed enrichment is under a null background model
-
Antibody and data set compliance with guidelines was summarized as percentages of the total tested (e.g., '20% of tested antibodies met characterization guidelines').↳ Could also: Reporting these proportions with binomial confidence intervals could also be used — A confidence interval around each proportion would convey the precision of the estimate, which is particularly informative when the denominator (number of antibodies or data sets) varies across categories
-
Mass-spectrometry-based protein identification used a fixed FDR threshold (0.0% protein/peptide FDR) together with a MASCOT probability cutoff (P < 0.05).↳ Could also: Reporting a range of FDR thresholds or a q-value distribution across identified peptides could also be used — Showing the number of identifications retained across a range of FDR cut-offs can help readers judge the sensitivity of the mass-spec-based antibody validation to the specific threshold chosen
-
Sequencing-depth effects on peak discovery were shown descriptively as peak counts and fold-enrichment plotted against reads sequenced (Fig. 3).↳ Could also: A saturation/rarefaction curve model (e.g., fitting a saturation function to estimate asymptotic peak numbers) could also be used — A fitted saturation curve would let readers extrapolate the expected gain in peak discovery from additional sequencing depth beyond what was directly observed
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
The reproduction executed the paper's prescribed ChIP-seq pipeline (bowtie2 -> MACS3 callpeak of ChIP vs matched Input) end-to-end on a real ENCODE K562 PolII dataset from the paper's own accession tree and confirmed both computed QC guidelines: FRiP = 0.2572 sits inside the paper's stated 0.2-0.5 band for successful point-source experiments, and redundant rates of 0.10/0.05 give NRF 0.90 / 0.95, clearing the >0.8 recommendation. The one graded mismatch — the paper's 478 ChIP-seq datasets under PRJNA63441 vs 14,516 unique ChIP-Seq experiments / 45,081 runs on a 2026-07-28 ENA query — is a 14-year snapshot effect on a living umbrella BioProject, correctly diagnosed by the room as expected consortium growth and not an authors' defect. The residual weaknesses are on our side and are about scope, not correctness: the test dataset, aligner, genome build and q-cutoff were all self-chosen because the paper specifies none; tool versions drifted (MACS3 vs MACS-1.x, GRCh38 vs hg18); and the paper's two other flagship QC contributions (NSC>1.05 / RSC>0.8 cross-correlation and IDR replicate concordance) were never computed, so 1 of 14,516 experiments carries the whole verdict. Overall a credible, transparently-scoped partial confirmation with explainable deviations and no fabrication signal — yellow across the board rather than green (incomplete coverage, interval-only endpoints) or red (nothing contradicts the paper).
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.