Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

TP53 engagement with the genome occurs in distinct local chromatin environments via pioneer factor activity.

Genome Res · 2014
L1 80/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +4
✓ What held up
  • Reported values are derivable from the shared data
  • The central claim held under reproduction
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🟡Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡The deviation was non-trivial in magnitude
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
80/100
Reproducibility score
0.3 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 56% of all assessed papers rank 484 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Reproduced the paper's poverlap-based TP53/chromatin-overlap claims (Fig. 2H-J / Supplemental Fig. S2G-I, P=0.01; and the ATAC-seq overlap percentages by chromatin category in Fig. 4C) by applying the assigned third-party tool poverlap (github.com/brentp/poverlap) to the paper's own pre-processed peak files from GSE58740 -- no re-alignment or re-peak-calling was needed since GEO already hosts the exact category-partitioned peak BEDs (TSS/Enhancers/Protoenhancers) and ATAC-seq MACS peaks. All 8 poverlap permutation tests (4 chromatin categories x {genome-wide, local 10kb} shuffle, n=1000 shuffles each) completed and are directionally and statistically consistent with the paper's P=0.01 claim, though every test hit the n=1000 permutation floor (p=0.000999) rather than matching the paper's exact reported P-value -- expected given our smaller shuffle count and the use of ATAC-seq peaks as a proxy for the paper's unverified-in-detail 'regions of enriched chromatin' definition. The more precise, independently checkable Fig. 4C overlap-percentage claims reproduced with a mix of grades: distal/proto-enhancer overlap (9.62% vs reported '<10%') graded exact; TSS/promoter overlap (81.83% vs reported ~80%) graded within-tol; enhancer overlap (59.49% vs reported ~50%) graded partial (same qualitative ordering -- enhancer overlap is clearly intermediate between distal and promoter -- but ~9.5 percentage points above the reported figure). Getting a working Python-2/bedtools/poverlap environment required 6 SLURM job iterations to work around several genuine, documented bugs: a conda Terms-of-Service plugin crash, a solver-plugin regression from disabling all conda plugins, poverlap.py's own py2/py3 print-syntax mixing (unparseable under real Python 2 without a print_function future-import patch), the commandr dependency's use of a py3-only inspect API, and an internal poverlap inconsistency where --genome is required by an assert even in --shuffle_loc mode despite the docstring saying it's ignored there -- all patches are minimal, narrowly scoped, and documented in poverlap_job.sh with rationale comments. Also profiled all 3 GEO accessions the paper relies on: GSE58740 (primary, own data, complete, quality A), GSE21823 (external MNase-seq, profiled at manifest level only since out of scope for the assigned poverlap pipeline, quality B, delivers_promised uncheckable since content wasn't parsed), and GSE56640 (external TP53/TP63 keratinocyte ChIP-seq, only 4 of its 25 samples are actually cited by this paper, quality B, one series-wide completeness gap: no processed file for the Input control sample). NOT attempted: independent RNA-seq re-quantification, the separate H4K16ac random-matched-peak enrichment test (different method, not poverlap), and content-level parsing of GSE21823/GSE56640 raw data (their assigned use in the paper is unrelated to the TP53/chromatin-overlap statistic this room's assigned repo addresses). No completeness claim is made beyond what is stated here.

💻 Code ↗ 🗄 Data: GSE21823

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-08-03
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-08-03
no human curator yet
Last updated
2026-08-03

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

The paper asks how TP53 engages the genome within the context of chromatin to activate transcription, testing whether TP53 binding sites can be distinguished by the dynamics of their local chromatin environment genome-wide, including at gene-distal regulatory elements.

Core claims
  • TP53 binding events fall into three distinct categories defined by the local chromatin environment: TSS (H3K4me3+), enhancer (H3K4me1+/H3K4me3-), and distal (H3K4me1-/H3K4me3-) peaks. finding
  • Distal TP53 binding sites lack classic chromatin modifications and lie within inaccessible chromatin, indicating TP53 has intrinsic pioneer factor activity. mechanism
  • TP53 binds preestablished enhancers marked by H3K4me1/H3K27ac and eRNA transcription, and further activates them by inducing H3K27ac/H4K16ac deposition and RNA pol II recruitment rather than commissioning new enhancers within 6 h. finding
  • H4K16ac is the most dynamic chromatin modification at TP53 binding sites upon TP53 activation, especially at gene-distal elements. finding
  • Inaccessible TP53 binding sites display enhancer-like properties in epithelial-lineage cell types, representing 'proto-enhancers' that become active enhancers in the appropriate cellular context. finding
  • TP53, along with TP63, may act as a pioneer factor to specify epithelial enhancers, so that TP53 tunes its stress response to the lineage-specific epigenomic landscape rather than following a cell-type invariant program. mechanism
  • A compendium of new and existing genome-wide data sets (TP53, RNA pol II, histone modification ChIP-seq, RNA-seq, ATAC-seq, MNase-seq in IMR90) provides a resource for analyzing TP53–chromatin relationships. resource
  • Classification of TP53 peaks by overlap with MACS-defined H3K4me3/H3K4me1 enriched regions is a method for resolving TP53 binding-site classes. method
Experimental setups
Assay System Perturbation Readout Platform
ChIP-seq (TP53) IMR90 primary human lung fibroblasts 5 μM nutlin (MDM2 inhibitor) vs DMSO, 6 h Genome-wide TP53 binding sites / significantly enriched peaks versus condition-specific input
ChIP-seq (histone modifications: H3K4me1, H3K4me2, H3K4me3, H3K27ac, H4K16ac) IMR90 primary human lung fibroblasts 5 μM nutlin vs DMSO, 6 h Enrichment/occupancy of each modification at TP53 peaks (peak center ±750 bp or ±2500 bp windows)
ChIP-seq (RNA polymerase II) IMR90 primary human lung fibroblasts 5 μM nutlin vs DMSO, 6 h RNA pol II occupancy at TP53 peaks, enhancers and gene bodies
RNA-seq (poly(A)+ selected RNA) IMR90 primary human lung fibroblasts 5 μM nutlin vs DMSO Differential transcript expression; TP53-induced and down-regulated genes; GO enrichment
ATAC-seq IMR90 primary human lung fibroblasts nutlin vs DMSO Chromatin accessibility (MACS-defined peaks and tag enrichment) at each TP53 peak class
MNase-seq IMR90 primary human lung fibroblasts none/nutlin-induced TP53 sites analyzed Average nucleosome tag density ±750 bp at TP53 and CTCF binding sites
GRO-seq (published data) HCT116 colon carcinoma cells nutlin vs DMSO Bidirectional nascent transcription (eRNA) tags at TP53-bound enhancers, normalized to 1 × 10^-7 reads
Meta-analysis of published TP53 genome-wide binding data sets / motif analysis Multiple human cell types (published ChIP data sets); IMR90 peaks none (computational) Representation of each TP53 peak class across data sets; percentage of peaks containing consensus TP53 response element motif
Key results
  • Nutlin treatment stabilized TP53 and increased the number of significantly enriched TP53 peaks (FDR < 1) relative to DMSO
  • The majority of induced TP53 binding events occur >5 kb from the TSS of a protein-coding gene, with the modal group 5–50 kb from the nearest TSS
  • H3K27ac and especially H4K16ac increase at TP53 binding sites after TP53 induction, while average H3K4 methylation enrichment does not change
  • Distal (H3K4me1-/H3K4me3-) peaks are the largest group of TP53 binding sites and show nutlin-induced H4K16ac enrichment over random distance- and size-matched peaks 44% of peaks; 2.6-fold H4K16ac induction
  • Bidirectional GRO-seq tags at putative TP53 enhancers increase upon nutlin treatment, though bidirectional eRNA transcription already occurs at most of these enhancers in DMSO twofold
  • TP53-bound enhancers are more likely to be transcribed (RNA pol II+) than genome-wide enhancers and less likely to be poised >60% vs <20%
  • TP53-bound enhancers are far more often dual-marked by both H3K27ac and H4K16ac than genome-wide enhancers ~70% vs 25%
  • H3K4me1, H3K4me2, and H3K27ac peaks are present at essentially all TP53-bound enhancers before nutlin, whereas RNA pol II and H4K16ac are gained de novo at about half of them ~50% de novo RNA pol II and H4K16ac
Key statistics
  • count 73% of identified TP53 peaks contain a consensus TP53 response element (RE) motif (Motif analysis across all TP53 peaks)
  • count <50% of H3K4me3+ TP53 sites contain the RE (Motif content of TSS-class TP53 peaks)
  • count 44% (Fraction of TP53 binding sites lacking H3K4 methylation (distal class), the largest group)
  • fold_change 2.6-fold induction, P < 2.2 × 10^-16 (H4K16ac enrichment at distal TP53 sites versus random distance- and size-matched peaks)
  • pvalue P < 2.2 × 10^-16 (H3K4me3 enrichment at TSS TP53 peaks versus enhancer or distal peaks; also for lower nutlin-inducible TP53 binding at peaks lacking RE motifs)
  • pvalue P = 0.01 (TP53 intersection with regions of enriched chromatin versus locally or genome-wide random peak locations)
  • fold_change twofold increase in bidirectional GRO-seq tags (eRNA transcription at TP53-bound enhancers, nutlin vs DMSO, HCT116)
  • count Fewer than 5% of transcripts show significant differential expression between DMSO and nutlin (RNA-seq in IMR90 after nutlin treatment)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The study identified genome-wide TP53 and histone-modification ChIP-seq binding sites in DMSO- versus nutlin-treated IMR90 fibroblasts using MACS peak calling against input controls (FDR < 1%), then classified TP53 peaks into three groups (TSS, enhancer, distal) based on overlap with H3K4me3/H3K4me1 domains. Group-level differences in enrichment, motif content, and chromatin overlap were assessed with reported p-values (e.g., P < 2.2 × 10⁻¹⁶, P = 0.01), and results were also visualized with box plots, heatmaps, and average enrichment profiles across ChIP-seq, ATAC-seq, GRO-seq, RNA-seq, and MNase-seq datasets. The specific statistical test names, exact sample sizes, and software/version details for most comparisons are not stated in the provided text.

Replicationunclear GroupsDMSO (vehicle) vs. nutlin (TP53-activating, MDM2 inhibitor) treatment in IMR90 fibroblasts; also comparisons among TSS, enhancer, and distal TP53 peak classes Pairingunclear Randomization/blindingnot stated DispersionIQR Exact p-valuesyes Effect sizesyes Confidence intervalsno Multiplicity correctionFDR thresholding used within MACS peak calling (FDR < 1%)
Statistical tests used
Test Applied to n Assumptions
MACS peak enrichment calling vs. input (FDR-based) identification of significant TP53 and histone-modification ChIP-seq peaks (Fig. 1A; Supplemental Tables S1, S2) not stated
unspecified statistical comparison (reported as P < 2.2 × 10⁻¹⁶) H3K4me1/H3K4me3 enrichment differences across TSS, enhancer, and distal TP53 peak classes (Fig. 2B,C) not stated
unspecified statistical comparison (reported as P < 2.2 × 10⁻¹⁶) nutlin-induced TP53 binding and H4K16ac enrichment in motif-containing vs. motif-lacking TP53 peaks (Supplemental Fig. S2E,F) not stated
comparison of observed vs. randomized/expected peak overlap (reported as P = 0.01) TP53 intersection with enriched chromatin regions vs. locally/genome-wide random peak locations (Supplemental Fig. S2G–I) not stated
differential expression analysis (method not named) DMSO vs. nutlin transcript comparison, reported as "fewer than 5% of transcripts show significant differential expression" (Supplemental Fig. S1C) not stated
Approaches that could also have been used
  • Peak significance was determined via MACS with an FDR threshold, while several other group comparisons are reported only as p-values (e.g., P < 2.2 × 10⁻¹⁶) without naming the specific statistical test.
    Could also: explicitly naming the test used for each comparison, such as a Wilcoxon rank-sum (Mann-Whitney U) test for the box-plot enrichment comparisons — ChIP-seq enrichment scores are often skewed/non-normal, so a named non-parametric test would clarify what assumptions the reported p-values rest on and let readers judge their appropriateness.
  • Differences in enrichment across the three TP53 peak classes (TSS, enhancer, distal) were reported as individual p-values rather than through a single omnibus comparison across all three groups.
    Could also: a Kruskal-Wallis test (or one-way ANOVA) across the three classes followed by post-hoc pairwise tests with a multiple-comparison correction (e.g., Dunn's test with Benjamini-Hochberg adjustment) — an omnibus test establishes whether the groups differ overall before pairwise comparisons are drawn, and a post-hoc correction explicitly controls the false-discovery rate across the pairwise tests.
  • No multiple-testing correction is described for the several pairwise comparisons reported across the results, beyond the FDR used internally for MACS peak calling.
    Could also: applying a family-wise correction (e.g., Benjamini-Hochberg FDR or Bonferroni) across the full set of reported comparisons — since many comparisons draw on the same underlying ChIP-seq/RNA-seq datasets, a shared correction would bound the overall false-positive rate across all reported tests.
  • Sample sizes (n) underlying the statistical comparisons (e.g., number of peaks per class) are not explicitly stated alongside the reported p-values.
    Could also: reporting the n of peaks/regions contributing to each comparison directly in figure legends or the text — this would allow readers to gauge the statistical power and precision behind each comparison, which is particularly relevant for the smaller TSS peak category described as the least numerous group.
  • Differential expression between DMSO and nutlin conditions was described descriptively ("fewer than 5% of transcripts show significant differential expression") without naming the underlying model.
    Could also: a model-based RNA-seq differential expression tool such as DESeq2 or edgeR, using a negative binomial generalized linear model with a stated FDR-adjusted significance cutoff — these tools explicitly model the mean-variance relationship of RNA-seq counts and document their multiple-testing correction, which can aid reproducibility of the differential expression calls.
  • Enrichment and expression changes are frequently reported as fold-change point estimates (e.g., "twofold increase," "2.6-fold induction") without accompanying confidence intervals.
    Could also: reporting bootstrap-based or model-based confidence intervals around the fold-change estimates — confidence intervals would convey the precision of each fold-change estimate alongside the point estimate and the associated p-value.
Software: MACS

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

fig4c_distal_protoenhancer_atac_overlap
Reported
Fig. 4C: <10% of distal/proto-enhancer TP53 peaks overlap ATAC-seq-defined accessible chromatin regions (Nutlin-treated IMR90).
Reproduced
180/1871 (9.62%) proto-enhancer TP53 peaks (GSE58740_p53Nut_Protoenhancers_hg19.bed.gz) overlap >=1bp with ATAC-seq Nutlin MACS peaks (bedtools intersect -wo).
exact
fig4c_enhancer_atac_overlap
Reported
Fig. 4C: ~50% of enhancer-associated TP53 peaks overlap ATAC-seq accessible chromatin (Nutlin-treated IMR90).
Reproduced
925/1555 (59.49%) enhancer TP53 peaks (GSE58740_p53Nut_Enhancers_hg19.bed.gz) overlap ATAC-seq Nutlin MACS peaks.
partial
fig4c_tss_promoter_atac_overlap
Reported
Fig. 4C: ~80% of TSS/promoter-associated TP53 peaks overlap ATAC-seq accessible chromatin (Nutlin-treated IMR90).
Reproduced
743/908 (81.83%) TSS TP53 peaks (GSE58740_p53Nut_TSS_hg19.bed.gz) overlap ATAC-seq Nutlin MACS peaks.
within tolerance
fig2h_j_suppfig_s2g_i_permutation_enrichment
Reported
Fig. 2H-J / Supplemental Fig. S2G-I; Methods: TP53 peak overlap with regions of enriched chromatin is significantly higher (P=0.01) than expected for either locally or genome-wide random peak locations, tested via poverlap with local peak shuffling (github.com/brentp/poverlap).
Reproduced
poverlap permutation test (n=1000 shuffles, --ncpus 8), all TP53 Nutlin FDR1 peaks (n=4334, GSE58740_p53_Nutlin_Peaks_hg19_FDR1.bed.gz) vs ATAC-seq Nutlin MACS peaks (n=57830, used as proxy for 'regions of enriched chromatin') via bedtools intersect -wo: genome-wide bedtools-shuffle control observed=1848 vs simulated mean=187.6, simulated_p=0.000999 (permutation floor at n=1000); local 10kb poverlap local-shuffle control observed=1848 vs simulated mean=579.2, simulated_p=0.000999 (floor). Repeated identically per chromatin subcategory (Enhancers/Protoenhancers/TSS, both shuffle modes) -- all 8 combinations hit the same p=0.000999 floor with observed overlap 3-18x the simulated mean.
within tolerance

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 80/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟡2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟢5. Derivability / plausibility
🟡6. Severity of the deviation
🟢7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +4

Reproduces well. Using the authors' own deposited hg19 peak BEDs (GSE58740), two of three Fig. 4C values land essentially on target (proto-enhancer 180/1871 = 9.62% vs '<10%'; TSS 743/908 = 81.83% vs '~80%'), and the poverlap permutation test confirms significant enrichment in all 8 category x shuffle-mode combinations (observed=1848 vs simulated mean 187.6 genome-wide / 579.2 local-10kb, p at the 0.000999 floor vs the reported P=0.01). The one real gap is the enhancer category at 59.49% vs the paper's '~50%' — a ~9.5-percentage-point deviation that sits on our side of the ledger: the paper states no overlap criterion and no numeric Fig. 4C values, so the >=1bp bedtools intersect rule and the ATAC MACS peak set were our choices, and the same substitution was needed to stand in for the paper's undefined 'regions of enriched chromatin'. Nothing here suggests an authors' defect — the values are derivable from the shared data and the pioneer-factor conclusion (accessibility rising monotonically proto-enhancer -> enhancer -> TSS, all significantly above shuffled expectation) is fully confirmed. Overall: yellow, solid reproduction with explainable, method-choice-driven deviation.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.