A ChIP-exo screen of 887 Protein Capture Reagents Program transcription factor antibodies in human cells.
The main results reproduced, with only marginal, non-material deviations.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Reported values are derivable from the shared data
- 🟡Could not use the authors’ exact input data
- 🟡Reported values were only indirectly comparable
- 🟡A deviation arose in the data or preprocessing
- 🟡A deviation was attributed to the published material
- 🟡The deviation was non-trivial in magnitude
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Described well enough to reproduce the easy 80%, via a THIRD-PARTY standard chain on the paper's own data (valid per P16). The BRIEF's code pointer (ebartom/NGSbartom) was WRONG; the paper's real artifacts are CEGRcode/PCRPpipeline (analysis code) + CEGRcode/2021-Lai_PCRP (deposited figure peak BEDs). I reproduced the ChIP-exo positive-control results (R3): for 3 USF1 + 3 NRF1 K562 samples from GSE151287 I ran bowtie2->hg19, samtools dedup, MACS3 callpeak, on «our HPC» («job», COMPLETED 6:47). RESULT 1:1: alignment rates 79-93% reproduce the paper's 70-90% standard (within-tol). RESULT (core): independently-called peaks land on the authors' deposited motif-bound peak loci FAR above chance - pooled USF1 439/2425 (18.1%) and NRF1 15/1009 (1.5%) recovered vs 0.00% from size-matched random shuffles (>1000x enrichment), reproducing the positive-control specificity claim directionally and quantitatively-above-background. NOT a coordinate-identical peak reproduction: I used MACS3 (the authors used GeneTrack/ChExMix), single low-depth replicates vs their 43-replicate aggregate, and strict q<0.01. PCR-duplicate standard NOT met by most screen samples (consistent with the paper's own note that high dups are common; our dedup definition may also differ). NOT ATTEMPTED (hard 20%): the authors' PCRPpipeline verbatim (python2.7 + ChExMix JAR + unshipped REF bundle), the AUROC/GENRE motif 'Direct Binding' analysis (R6), the 245-clone replicate-concordance (R4), and ChIP-seq/CUT&RUN/STORM/PBM (separate series / wet-lab). No fabrication indicators: deposited peak sets are internally consistent and recoverable from raw data. Note: reported per-cell-type dataset counts are smaller than the GEO deposit (analyzed subset vs full deposit) - benign denominator difference, flagged for human audit.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 57assessed: 2026-06-15 ⛓ 56135877ec3b
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-15
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusCan the renewable monoclonal antibodies generated by the NIH Protein Capture Reagents Program (PCRP) against human transcription factors reliably enrich their cognate genomic targets in chromatin-based assays such as ChIP-exo? The study field-tests 887 PCRP antibodies to assess their utility and the metrics/challenges of antibody validation.
- ★ About 5% of tested PCRP antibodies showed high-confidence cognate-target enrichment in at least one assay and are strong candidates for further validation. finding
- ★ An additional 34% produced ChIP-exo data distinct from background warranting further testing, while 61% were not substantially different from background. finding
- ★ Massively parallel ChIP-exo provides a high-throughput, ultra-high-resolution platform suitable for screening and evaluating large antibody collections. method
- ★ ChIP-exo peaks enriched at a precise distance from cognate motifs provide strong support for antibody target specificity. mechanism
- ★ Independent hybridoma clones against the same target can differ markedly in ChIP performance, so failure of one clone warrants testing others. finding
- DSHB-derived hybridoma culture supernatants detected more cognate-motif binding events than CDI concentrates at equal reported antibody amounts. finding
- RNAi knockdown of NRF1 diminished NRF1 ChIP-seq signal, demonstrating specificity of PCRP mAbs 3D4 and 3H1. finding
- ★ The screen generated a large public resource of ~1200 ChIP-exo data sets covering 887 antibodies against 681 unique human transcription factors. resource
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| ChIP-exo/seq | K562 (human bone marrow lymphoblast); also MCF-7, HepG2, and human tissues (liver, kidney, placenta, breast) | none (antibody screening) | genome-wide near-base-pair protein-DNA binding enrichment at cognate motifs | — |
| ChIP-seq | HCT116 cells | none (137 PCRP hybridomas, 70 targets) | significantly enriched peaks | — |
| ChIP-seq with RNAi knockdown | HCT116 cells | shRNA knockdown of NRF1 (two targeting shRNAs vs nontargeting) | loss of NRF1 ChIP-seq signal | — |
| Western blot / immunoblot | HCT116 cells | shRNA knockdown of NRF1 (SH1, SH2 vs NonT) | NRF1 protein level normalized to tubulin beta (quantified by ImageJ) | SDS-PAGE; ImageJ quantification |
| CUT&RUN | human cells (subset of mAbs) | none | genomic enrichment | — |
| STORM super-resolution microscopy | human cells (subset of mAbs) | none | subcellular protein localization | — |
| Protein binding microarray (PBM) | in vitro | none | DNA-binding specificity | — |
| Vendor source comparison ChIP-exo | K562 cells | NRF1, USF1, YY1 mAbs from DSHB vs CDI (3 µg) | relative ChIP yield / binding events at cognate motifs | protein A/G magnetic beads |
- – ~5% of tested antibodies showed high-confidence cognate-target enrichment in at least one assay 5%
- – 34% produced ChIP-exo data distinct from background; 61% not substantially different from background 34% / 61%
- – 43 of 45 (>95%) USF1 data sets produced a USF1-specific ChIP-exo pattern around E-boxes >95% (43/45)
- – Of 245 replicated hybridoma clones, 36 (14.7%) showed reproducible enrichment of the same genomic feature class 14.7% (36/245)
- – 102 (41.6%) of replicated clones gave no enrichment in both replicates; 107 (43.7%) enriched in one replicate but not the other 41.6% / 43.7%
- – 19 of 137 (14%) PCRP hybridomas produced significantly enriched ChIP-seq peaks in HCT116 14% (19/137)
- ▼ NRF1 ChIP-seq peaks diminished by two specific shRNAs but not by untargeted oligo, confirming specificity of mAbs 3D4 and 3H1
- – Both HSF1 clones 1A10 and 1A8 gave near-identical ChIP-exo patterns while clones 1C1 and 1D11 failed
- count 887 unique antibodies against 681 unique human transcription factors (antibodies assayed by ChIP-exo)
- count ~1200 ChIP-exo data sets (total ChIP-exo data generated, primarily in K562)
- count 943 unique mAb clones tested (887 ChIP-exo, 59 other assays); 642 targeted putative ssTFs (overall reagents tested)
- count ~164,000 E-box motif instances with significant ChIP-exo peak-pair (USF1 Q<0.01 peak-pairs)
- pvalue Q < 0.01 (significance threshold for ChIP-exo peak-pairs)
- count 43 independent USF1 replicates as positive control (technical reproducibility assessment)
- count 1009 K562, 134 MCF-7, 96 HepG2 data sets (distribution of hybridoma supernatant testing across cell lines)
- other AUROC > 0.6 (top-motif enrichment threshold across 259 ChIP-exo data sets with >500 peaks)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This large-scale antibody validation resource screened 887 monoclonal antibodies against 681 human transcription factors primarily by ChIP-exo in K562 cells (~1200 data sets), with subsets further tested by ChIP-seq, CUT&RUN, STORM microscopy, immunoblots, and protein binding microarrays. Antibody performance was evaluated using FDR-based peak-pair calling (Q < 0.01), motif enrichment scoring (AUROC), and cross-replicate Pearson correlation. The majority of antibodies were assayed in a single pass without replication, and results were reported descriptively as percentages of antibodies falling into performance tiers rather than as formal inferential statistics.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| FDR-based ChIP-exo peak-pair calling (Q-value threshold Q < 0.01) | Identification of significant ChIP-exo peak-pairs at ~164,000 E-box motif instances for USF1, and used across the broader antibody screen for peak detection | ~164,000 E-box motif instances for USF1 analysis; 1261 total data sets in the full screen | not stated |
| Pearson's pairwise correlation | Reproducibility assessment of USF1 ChIP-exo occupancy at putative USF1-bound E-boxes across all 43 USF1 replicates and IgG/no-antibody negative controls | 43 USF1 independent replicates plus negative controls | not stated |
| AUROC (Area Under the Receiver Operating Characteristic Curve) for motif enrichment | Scoring motif enrichment in 259 ChIP-exo data sets with >500 peaks; top PWM assigned per data set by highest AUROC; AUROC > 0.6 used as enrichment threshold | 259 ChIP-exo data sets; 100 putative ssTF binding PWMs evaluated per data set | not stated |
| ChIP-seq peak enrichment analysis (method details deferred to Methods section) | 137 PCRP hybridomas tested by ChIP-seq in HCT116 cells; 19/137 (14%) produced significantly enriched peaks | 137 hybridoma clones corresponding to 70 targets | not stated |
| ImageJ band-intensity quantification with loading-control normalization | NRF1 knockdown efficiency quantification by Western blot, normalized to tubulin beta | — | not stated |
-
Category proportions (~5%, ~34%, ~61%) describing antibody performance tiers were reported as point estimates without uncertainty quantification↳ Could also: Exact binomial confidence intervals or bootstrap CIs could also be reported alongside these proportions — Uncertainty estimates on screening proportions help readers assess stability of the observed rates and inform decisions about how many antibodies to expect to succeed in follow-up validation, particularly given the large but finite sample of 887 reagents
-
Reproducibility across ChIP-exo replicates was assessed using Pearson's pairwise correlation of occupancy values at USF1-bound E-boxes↳ Could also: The Irreproducibility Discovery Rate (IDR) framework, designed specifically for ranked peak lists from ChIP experiments, could also be applied to quantify replicate concordance — IDR models the reproducibility of each peak individually using a mixture model of reproducible and irreproducible signals, providing a principled peak-level reproducibility threshold that complements global correlation summaries and is recommended by ENCODE for ChIP data quality assessment
-
Pearson correlation was used to compare occupancy values across replicates, which assumes a linear relationship between continuous signal values↳ Could also: Spearman rank correlation could also be computed as a non-parametric alternative — Spearman correlation is less influenced by extreme high-occupancy outlier sites common in ChIP data at highly occupied regulatory elements and does not presuppose normality of occupancy values
-
Motif enrichment was scored by AUROC with a fixed threshold of >0.6 to classify data sets as showing cognate-motif enrichment↳ Could also: Permutation-based significance testing for AUROC — shuffling peak summit-to-motif-match assignments within each data set — could also be used to derive data-adaptive null distributions and thresholds — Permutation-based thresholds account for dataset-specific variation in peak count, motif frequency, and peak-width distributions, potentially reducing sensitivity to a single universal cutoff value across experiments of very different scale
-
Cell type for ChIP testing was selected using a less-than-twofold FPKM difference threshold in target mRNA expression, defaulting to K562 when no difference was detected↳ Could also: Protein-level abundance data (e.g., from mass spectrometry-based proteomics or quantitative immunoblot across cell lines) could also be used to guide cell type selection — mRNA and protein abundances are not always concordant due to translational regulation and protein stability differences; protein-level evidence of target expression might more directly predict ChIP enrichment success, particularly for post-transcriptionally regulated transcription factors
-
Most antibodies were screened in a single ChIP-exo experiment, with replication performed only on the subset of 245 clones showing initial enrichment or selected as negative controls↳ Could also: A minimum of two biological replicates for all antibodies, analyzed with IDR-based peak filtering, is also used in large-scale ChIP screening programs such as ENCODE — Two replicates per antibody enable formal reproducibility-based peak filtering and provide an empirical estimate of false-positive rates from single-experiment sampling variation; the authors acknowledge this trade-off explicitly in framing the study as a first-pass screen
Result convergence & founder nodes
Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.
-
NRF1 ChIP-seq peak signal is reduced by two NRF1-targeting shRNAs but not by non-targeting control, confirming specificity of anti-NRF1 monoclonal antibodiesChIP-seq hct116 down 2021×1papers★ This paper is the founder (earliest)
-
HSF1 ChIP-exo enrichment is clone-dependent: clones 1A10 and 1A8 produce concordant binding patterns while clones 1C1 and 1D11 fail to enrichChIP-seq k562 mixed 2021×1papers★ This paper is the founder (earliest)
-
USF1 produces reproducible cognate E-box ChIP-exo enrichment in >95% of antibody datasets tested, demonstrating high antibody specificityChIP-seq k562 none 2021×1papers★ This paper is the founder (earliest)
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-34426512
Paper: Lai WKM et al. (2021) A ChIP-exo screen of 887 Protein Capture Reagents Program transcription factor antibodies in human cells. Genome Res. PMID 34426512, PMCID PMC8415381, DOI 10.1101/gr.275472.121.
Corrected code/data pointers (BRIEF.md was wrong / generic). The paper's Data access section names the actual artifacts:
- ChIP-exo analysis pipeline (authors' own):
github.com/CEGRcode/PCRPpipeline(BRIEF.md listedebartom/NGSbartom, which is a different, generic lab toolkit — not the pipeline used here). - Deposited figure/peak files (the reported results):
github.com/CEGRcode/2021-Lai_PCRP. - Data: GEO
GSE151287(ChIP-exo, 1666 GEO samples),GSE152144(ChIP-seq),GSE151326(CUT&RUN). PBM in UniPROBE (LAI20A).
What the paper reports (candidate results)
| # | Reported result | Source | Pipeline-derived? | In scope? |
|---|---|---|---|---|
| R1 | 887 unique mAbs vs 681 unique targets (605 putative ssTFs) assayed by ChIP-exo | Abstract/Results | No (experimental design count) | partial (metadata cross-check only) |
| R2 | Datasets per cell type: K562 1009, MCF-7 134, HepG2 96 (1261 total in replicate) | Results §"Screening" | No (count of deposited datasets) | partial (metadata cross-check) |
| R3 | USF1 positive control: 43/45 (>95%) gave a USF1-specific ChIP-exo pattern at E-boxes; ~164,000 relaxed E-box instances | Results + Suppl Fig 1 | Yes — alignment + peak-pair calling (GeneTrack) | YES (core target) |
| R4 | Replicated clones (245): 36 (14.7%) concordant feature enrichment, 102 (41.6%) none, 107 (43.7%) one-of-two | Results + Suppl Table 2 | Yes (ChExMix feature-enrichment pipeline) | hard-20% (not attempted, see below) |
| R5 | ChIP-seq HCT116: 19/137 (14%) significantly enriched | Results + Suppl Table 3 | Yes | out (different series GSE152144; deferred) |
| R6 | Motif analysis: 259 ChIP-exo + 19 ChIP-seq peak files; 20 mAbs / 16 ssTFs "Direct Binding" (motif enriched + centered, AUROC) | Results + Suppl Table 4 | Yes (AUROC + GENRE + 100 PWMs) | hard-20% (not attempted) |
| R7 | ~5% of antibodies validated for cognate target across ≥1 assay; +34% distinct-from-background; 61% background | Abstract | Yes (composite, multi-assay) | out (whole-study composite) |
In-scope target chosen (80/20): R3 — ChIP-exo positive controls
Pipeline named by the paper: reads aligned to hg19; ChIP-exo peaks called with
GeneTrack (manuscript figures) / ChExMix; peak-pairs intersected with cognate motifs.
The full PCRPpipeline needs an unshipped REF/ bundle (hg19 genome+background model,
JASPAR2020 motif clusters, chromHMM/segway/repeat tracks, ChExMix JAR, python2.7) →
that is the hard 20% and is not reproduced verbatim.
Reproduction (third-party tools on the paper's own data — explicitly valid per
BRIEF P16/§2): for a small panel of positive-control ChIP-exo samples
(USF1 ×3, NRF1 ×3, K562), from raw GEO/SRA reads we run a standard, fully-specified
chain — bwa mem → hg19, dedup, MACS2 narrowPeak — on «our HPC», and check:
- Technical performance (Methods §"Technical performance"): % aligned (paper standard 70–90%) and % PCR duplicates (paper standard <40%).
- Specificity / recovery of deposited results: overlap of our called peaks
with the authors' deposited USF1 / NRF1 motif-bound peak-pair BEDs
(
2021-Lai_PCRP/Figure1— USF1 2425, NRF1 1009 peak-pairs) vs a size-matched random-background expectation. A working positive control must recover the deposited motif-bound loci far above chance.
This reproduces the paper's central qualitative+quantitative claim that the USF1/NRF1 positive controls yield specific, motif-concordant ChIP-exo enrichment, from raw data, with an independent standard pipeline.
Explicitly NOT attempted (and why)
- R4/R6 full ChExMix + AUROC motif pipeline — unshipped REF bundle, python2.7, custom GENRE/PWM tooling; underspecified for exact reproduction (hard 20%).
- R5/R7 ChIP-s
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
Using an independent standard chain (bowtie2→hg19, samtools dedup, MACS3) on the paper's own GSE151287 reads, the alignment QC (79-93%, mean 84.4%) reproduces the 70-90% standard within tolerance and independently-called USF1/NRF1 peaks land on the deposited motif-bound loci far above a 0.00% random shuffle (>1000x enrichment), so the deposited values are derivable from the shared data with no fabrication signal. The deviations — modest absolute recovery (439/2425 USF1, 15/1009 NRF1), dup rates above the <40% standard, and dataset-count denominators — sit on our methodology side (single low-depth replicates, MACS3 vs ChExMix, stricter dedup, GEO superset vs analyzed subset), not the authors'. Severity is moderate: magnitude and direction of the positive-control specificity hold, but only 3/45 datasets were tested and the broader R4/R6 analyses were not attempted, so the central conclusion is confirmed only in a limited form.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.