Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

ChIP-seq guidelines and practices of the ENCODE and modENCODE consortia.

Genome Res · 2012
L1 74/100 3/4
Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +8
✓ What held up
  • Nothing in this column.
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🟡Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
74/100
Reproducibility score
at the mean
vs. all fields · 1173 studies
🎯 Scores higher than 43% of all assessed papers rank 644 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Reproduced the paper's core computational pipeline (aligner -> MACS peak calling of ChIP vs Input) end-to-end on a real ENCODE dataset (PolII ChIP-seq vs Input, K562 cells) drawn from the paper's own PRJNA63441 accession tree, using MACS3 (the current incarnation of the paper's cited MACS repo). The pipeline ran without errors and produced 31,419 called peaks; FRiP (25.7%) and NRF (0.90/0.95) both independently satisfy the paper's own stated QC guideline thresholds, providing meaningful (if qualitative/order-of-magnitude) support for the paper's central methodological claims. Dataset profiling of PRJNA63441 shows it is a live, still-growing umbrella BioProject: the paper's reported 478 ChIP-seq datasets (2012 snapshot) has grown to 14,516 unique ChIP-seq experiments today, expected given ENCODE's continued production for over a decade, not a reproducibility defect. NSC/RSC cross-correlation QC and IDR replicate-concordance analysis -- the paper's other flagship QC metrics -- were not attempted in this pass and are flagged as missing/out of scope, not as failures.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-07-28
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-07-31
no human curator yet
Last updated
2026-07-31

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

The paper asks what experimental and analytical standards are needed to make ChIP-seq experiments reliable and comparable, and reports the working guidelines (antibody validation, replication, sequencing depth, quality metrics, data reporting) developed by the ENCODE and modENCODE consortia to address the substantial variability in how such experiments are performed, scored, and archived.

Core claims
  • ENCODE/modENCODE define a set of working standards and guidelines for ChIP-seq covering antibody validation, experimental replication, sequencing depth, data/metadata reporting, and data quality assessment. resource
  • Antibodies must pass a primary characterization assay (immunoblot for transcription factors, with immunostaining as alternative) plus at least one of five secondary assays (knockdown/RNAi, second antibody or complex member, epitope tag, affinity enrichment plus mass spectrometry, or motif analysis). method
  • Only a minority of commercially available transcription-factor antibodies both meet the characterization guidelines and work in ChIP-seq. finding
  • Two independent biological replicates are set as the consortium standard, with the irreproducible discovery rate (IDR) method used to assess replicate agreement and set peak thresholds. method
  • The number of called peaks for a typical point-source factor continues to increase with sequencing depth rather than saturating, because ChIP signal strength is a continuum rather than a discrete set of positive sites. finding
  • A minimum of 20 million mapped reads is set for ENCODE point-source transcription-factor ChIP experiments, since signal enrichment plateaus and peaks discovered beyond that depth are progressively weaker. method
  • Chromatin-associated proteins fall into point-source, broad-source, and mixed-source classes that require different analytical approaches. mechanism
  • Epitope tagging with large clones (fosmids/BACs) provides near-physiological expression and is an effective alternative when suitable antibodies are unavailable, though overexpression can cause occupancy of non-physiological sites. method
Experimental setups
Assay System Perturbation Readout Platform
ChIP-seq (chromatin immunoprecipitation followed by high-throughput DNA sequencing) Human and mouse cell lines/tissues (e.g., K562, GM12878, HepG2), D. melanogaster embryos, C. elegans; >100 cell types across four organisms none (native factor immunoprecipitation); epitope-tagged constructs in some cases Genome-wide binding sites/peaks of transcription factors and histone modifications; peak counts and fold-enrichment
ChIP-chip (ChIP followed by DNA microarray hybridization) Small-genome organisms used by modENCODE (D. melanogaster, C. elegans) none Enriched genomic regions relative to differentially labeled reference DNA DNA microarray
Immunoblot (Western blot) — primary antibody characterization assay Nuclear extract from GM12878 and K562 cells none Presence/absence of band at expected molecular weight and detection of cross-reacting bands (e.g., SIN3B at 133 kDa) Santa Cruz sc13145 (pass) and sc996 (fail) anti-SIN3B antibodies
Immunoprecipitation followed by immunoblot Nuclear lysates of K562 cells none Efficiency and specificity of immunoprecipitation of the expected-size band (TBLR1, 56 kDa) across input, IP, and depleted lanes Abcam ab24550 anti-TBLR1 antibody
Immunofluorescence / immunostaining (alternative primary assay) Cultured cells none Nuclear localization pattern; pass/fail quality-control call
Immunoprecipitation followed by mass spectrometry (secondary characterization) Whole-cell lysates of K562, GM12878, and HepG2 none Peptide identification confirming the immunoprecipitated protein (SP1, ~106 kDa) from a Coomassie-stained excised gel band Santa Cruz sc-17824 anti-SP1 antibody; MASCOT (Matrix Science); Scaffold (Proteome Software, Inc.)
Binding-site motif enrichment analysis (secondary characterization) ENCODE human ChIP-seq data sets for 85 transcription factors none Motif fold-enrichment relative to all DNase-accessible sites (bias-corrected with shuffle motifs) and motif representation as percentage of analyzed peaks Motif search stringency 4–6; peaks from IDR analysis at 0.01 cut-off
Peak calling as a function of sequencing depth (saturation analysis) 11 deeply sequenced human ENCODE ChIP-seq data sets (e.g., MAFK in HepG2) none (in silico read subsampling in 2.5 million-read increments) Number of called peaks and median fold-enrichment of newly called peaks versus number of uniquely mapped reads PeakSeq (0.01% FDR cut-off)
Key results
  • Only about one-fifth of tested commercially available transcription-factor antibodies met the characterization guidelines and also functioned in ChIP-seq. ~20% (44 of 227)
  • Peak counts continued to increase with sequencing depth for nearly all data sets; clear saturation was observed for only one factor with few binding sites, and one data set yielded >150,000 peaks at 100 million mapped reads. >150,000 peaks at 100 million mapped reads
  • Signal enrichment plateaus with greater sequencing depth; at 20 million mapped reads, median enrichments are typically five- to 13-fold. five- to 13-fold median enrichment
  • Peaks newly identified beyond 20 million reads are much weaker, with enrichment about 20% of that of the strongest peaks; additional peaks at three- to sevenfold enrichment can still be found at much greater depth, likely low-affinity or open-chromatin sites. ~20% of strongest-peak enrichment; three- to sevenfold
  • Of 85 transcription factors with ENCODE ChIP-seq data, 60% had a data set meeting the fourfold motif enrichment standard, and 96% of those data sets met the >10% motif representation standard. 60% of 85 factors; 96% of qualifying data sets
  • Secondary antibody characterization data submitted to the consortia were dominated by mass spectrometry, followed by second-antibody/epitope-tag/complex-member ChIP, motif analysis, and siRNA knockdown. 55% / 28% / 10% / 7%
  • All six GFP-tagged C. elegans factors tested to date complemented null mutants, supporting epitope tagging as a viable alternative to native antibodies. 6 of 6
  • Initial RNA polymerase II ChIP-seq experiments showed that more than two replicates did not significantly improve site discovery, motivating the two-biological-replicate standard.
Key statistics
  • count more than a thousand individual ChIP-seq experiments for more than 140 different factors and histone modifications in more than 100 cell types in four organisms (Scope of ENCODE/modENCODE ChIP-seq production)
  • count 145 polyclonal and 43 monoclonal antibodies (Antibodies used to successfully generate ChIP-seq data as of October 2011)
  • count ~20% (44 of 227) (Tested commercial transcription-factor antibodies meeting guidelines and working in ChIP-seq)
  • other 55% mass spectrometry, 28% second antibody/epitope tag/complex member, 10% motif analysis, 7% siRNA knockdown (Secondary characterization data type submitted with consortia antibodies)
  • fold_change five- to 13-fold median enrichment (Median peak enrichment at 20 million mapped reads across deeply sequenced data sets)
  • fold_change three- to sevenfold (Enrichment of additional peaks found only at much greater sequencing depths)
  • other 60% of factors meet fourfold motif enrichment; 96% meet >10% motif representation (Motif-based secondary validation across ENCODE transcription factors (IDR 0.01 peaks))
  • pvalue P < 0.05; 0.0% protein FDR and 0.0% peptide FDR (MASCOT probability-based peptide matching and Scaffold analysis for SP1 IP–mass spectrometry)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This is a methods/guidelines resource paper (not a hypothesis-testing study) describing ENCODE/modENCODE consortium standards for ChIP-seq experimental design, antibody validation, and data quality assessment. Reproducibility across biological replicates was assessed using the irreducible/irreproducible discovery rate (IDR) framework, peak calling was evaluated with FDR-based thresholds, and antibody/target identification via mass spectrometry used probability-based peptide matching with FDR control. Results are largely reported as summary statistics (percentages of data sets/antibodies meeting a criterion, fold-enrichment values) across many aggregated data sets rather than as single pairwise significance tests.

Replicationbiological Sample sizeStandard set at two independent biological replicates per ChIP measurement, with additional replicates generated when quality metrics were poor; minimum sequencing depth standard of 20 million mapped reads for point-source factors GroupsNot a classical two/multi-group comparison; the paper aggregates results across many ChIP-seq data sets (e.g., 85 transcription factors, 11 deeply sequenced data sets, 227 tested antibodies) to characterize guideline compliance and data quality Pairingna Randomization/blindingnot stated Dispersionunclear Exact p-valuesno Effect sizesyes Confidence intervalsno Multiplicity correctionFalse discovery rate (FDR) control, applied separately in peak calling (PeakSeq FDR cut-off) and in mass-spectrometry protein/peptide identification (Scaffold FDR)
Statistical tests used
Test Applied to n Assumptions
Irreproducible discovery rate (IDR) analysis (Li et al. 2011) assessing agreement/reproducibility between biological replicate ChIP-seq experiments and setting peak thresholds two biological replicates per ChIP measurement (consortium standard) not stated
PeakSeq peak calling with FDR cut-off calling ChIP-seq peaks as a function of sequencing depth for 11 ENCODE data sets (Fig. 3) 0.01% FDR cut-off stated not stated
MASCOT probability-based peptide matching identifying immunoprecipitated proteins by mass spectrometry (e.g., SP1 antibody validation, Fig. 2D) P < 0.05 threshold stated
Scaffold protein/peptide FDR analysis post-processing of mass spectrometry identifications for antibody characterization 0.0% protein FDR and 0.0% peptide FDR stated as thresholds used
Motif fold-enrichment analysis with sequence-bias correction (shuffled motifs) evaluating transcription-factor binding motif enrichment at ChIP-seq peaks relative to DNase-accessible sites (Fig. 2E, 85 factors) 85 transcription factors; peaks defined by IDR analysis at 0.01 cutoff not stated
Approaches that could also have been used
  • Replicate reproducibility was assessed using the IDR framework applied to ranked peak lists from paired biological replicates.
    Could also: Other reproducibility metrics such as the Pearson/Spearman correlation of signal tracks between replicates, or Jaccard/overlap statistics on called peak sets, could also be used — These would provide complementary views of replicate concordance (e.g., genome-wide signal correlation vs. rank-based peak agreement) and are commonly reported alongside or instead of IDR in ChIP-seq quality assessments
  • Peak significance was determined using an FDR-based cut-off from PeakSeq.
    Could also: Alternative peak callers with different statistical models (e.g., MACS with a Poisson/local-lambda background model, or SPP) could also be applied — Different peak-calling statistical models make different assumptions about background read distribution, so comparing across callers can illustrate how peak set composition depends on the chosen model
  • Motif enrichment was summarized as fold-enrichment relative to DNase-accessible sites, corrected using shuffled-sequence controls.
    Could also: A formal statistical test (e.g., a hypergeometric or Fisher's exact test comparing motif occurrence in peaks vs. a matched background) could also be used to accompany the fold-enrichment value — A formal test statistic with a p-value would complement the fold-enrichment metric by quantifying how unlikely the observed enrichment is under a null background model
  • Antibody and data set compliance with guidelines was summarized as percentages of the total tested (e.g., '20% of tested antibodies met characterization guidelines').
    Could also: Reporting these proportions with binomial confidence intervals could also be used — A confidence interval around each proportion would convey the precision of the estimate, which is particularly informative when the denominator (number of antibodies or data sets) varies across categories
  • Mass-spectrometry-based protein identification used a fixed FDR threshold (0.0% protein/peptide FDR) together with a MASCOT probability cutoff (P < 0.05).
    Could also: Reporting a range of FDR thresholds or a q-value distribution across identified peptides could also be used — Showing the number of identifications retained across a range of FDR cut-offs can help readers judge the sensitivity of the mass-spec-based antibody validation to the specific threshold chosen
  • Sequencing-depth effects on peak discovery were shown descriptively as peak counts and fold-enrichment plotted against reads sequenced (Fig. 3).
    Could also: A saturation/rarefaction curve model (e.g., fitting a saturation function to estimate asymptotic peak numbers) could also be used — A fitted saturation curve would let readers extrapolate the expected gain in peak discovery from additional sequencing depth beyond what was directly observed
Software: PeakSeq · MASCOT (Matrix Science) · Scaffold (Proteome Software, Inc.) · IDR analysis framework (Li et al. 2011)

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

claim-pipeline-executes
Reported
Paper prescribes a standard ChIP-seq processing pipeline (short-read aligner -> uniquely-mapped reads -> MACS peak calling of ChIP signal vs matched Input/control) as the consortium-wide method for calling enriched regions from raw reads, using the taoliu/MACS tool the paper cites.
Reproduced
Ran the full pipeline end-to-end on a real ENCODE dataset from the paper's own accession (bowtie2 2.4.1 alignment to GRCh38 no-alt-analysis-set -> samtools 1.9 sort/index -> MACS3 3.0.4 callpeak, -g hs -q 0.05) using PolII ChIP-seq vs Input in K562 cells (SRX001941 vs SRX001942, PRJNA63447/GSE13008, a sub-project of the paper's cited PRJNA63441). Produced 31,419 called peaks (narrowPeak), predicted fragment length d=107bp, without errors.
exact
claim-frip-guideline-range
Reported
Paper states FRiP (Fraction of Reads in Peaks) is a key ChIP-seq QC metric, and that the most successful point-source-factor ChIP-seq experiments typically show FRiP values in the 0.2-0.5 range.
Reproduced
Computed FRiP directly from reproduced alignment+peak-calling output: 18,523,644 mapped ChIP reads total; 4,764,780 of them fall inside the 31,419 called peaks -> FRiP = 0.2572 (25.7%).
within tolerance
claim-nrf-threshold
Reported
Paper defines NRF (Non-Redundant Fraction = unique/total mapped reads after removing PCR duplicates) as a library-complexity QC metric and recommends NRF > 0.8 for a usable ChIP-seq/Input library.
Reproduced
MACS3 reports redundant rate 0.10 for the ChIP (ChIP treatment) track and 0.05 for the Input (control) track -> NRF_chip = 0.90, NRF_input = 0.95. Both clear the paper's stated 0.8 threshold.
exact
claim-prjna63441-chipseq-dataset-count
Reported
Paper text states that, at time of writing (~2011-2012), 478 ChIP-seq data sets had been submitted to GEO/SRA under accession PRJNA63441.
Reproduced
Queried the live ENA Portal API for PRJNA63441 on 2026-07-28: the umbrella now spans 4 sub-projects (PRJNA63441 itself: 9 runs; PRJNA63443: 58,200 runs; PRJNA30709: 7,591 runs; PRJNA63447: 2,208 runs), 68,008 total sequencing runs, of which 45,081 runs / 14,516 unique experiments are library_strategy=ChIP-Seq -- roughly 30x the paper's reported count.
did not match

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 74/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟡2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +8

The reproduction executed the paper's prescribed ChIP-seq pipeline (bowtie2 -> MACS3 callpeak of ChIP vs matched Input) end-to-end on a real ENCODE K562 PolII dataset from the paper's own accession tree and confirmed both computed QC guidelines: FRiP = 0.2572 sits inside the paper's stated 0.2-0.5 band for successful point-source experiments, and redundant rates of 0.10/0.05 give NRF 0.90 / 0.95, clearing the >0.8 recommendation. The one graded mismatch — the paper's 478 ChIP-seq datasets under PRJNA63441 vs 14,516 unique ChIP-Seq experiments / 45,081 runs on a 2026-07-28 ENA query — is a 14-year snapshot effect on a living umbrella BioProject, correctly diagnosed by the room as expected consortium growth and not an authors' defect. The residual weaknesses are on our side and are about scope, not correctness: the test dataset, aligner, genome build and q-cutoff were all self-chosen because the paper specifies none; tool versions drifted (MACS3 vs MACS-1.x, GRCh38 vs hg18); and the paper's two other flagship QC contributions (NSC>1.05 / RSC>0.8 cross-correlation and IDR replicate concordance) were never computed, so 1 of 14,516 experiments carries the whole verdict. Overall a credible, transparently-scoped partial confirmation with explainable deviations and no fabrication signal — yellow across the board rather than green (incomplete coverage, interval-only endpoints) or red (nothing contradicts the paper).

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.