Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

DNA Methylation Directs Polycomb-Dependent 3D Genome Re-organization in Naive Pluripotency.

Cell Rep · 2019
50/100 3/4
Reproduced on the brainbox compute brainarbeit.com
How its reproducibility compares
50/100
Reproducibility score
1.4 SD below mean
vs. all fields · 1173 studies
🎯 Scores higher than 8% of all assessed papers rank 1026 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

DESCRIBED WELL ENOUGH: YES. The Methods fully specify the downstream Hi-C pipeline (distiller-nf mapping -> cooler matrices -> cooltools insulation/eigs + coolpup.py pileups), and GEO GSE124342 ships the balanced multi-resolution .mcool matrices (serum + 2i mESC) that those exact third-party tools consume. This makes the paper clearly reproducible by running the named third-party tools on the paper's own shipped data (brief P16), without raw re-mapping. NOT A DROP. Work completed before forced finalization: paper+Methods+figures read; scope.md written (in-scope downstream analyses vs out-of-scope wet-lab/raw-mapping); data confirmed and staged on «infra» (serum .mcool 2.24 GB complete, 2i .mcool + mm9.fa fetching); conda env (cooler/cooltools/coolpuppy/bioframe) building on «our HPC»; self-contained analyses (insulation 25kb, compartments 200kb, P(s)) scripted and submitted as «job»; headline CGI-CGI decompaction pileup (coolpup.py) scoped with CGI source (UCSC mm9 cpgIslandExt) and optional RING1B source (GSE69955). The operator requested finalization while «job» was ~10 min in and still solving/downloading the conda env, so NO numeric claim was graded yet (no values fabricated). To complete: let «job» finish (outputs land in «path»), then author+run the coolpup.py pileup job and grade against Fig 3B/4C direction. NOT ATTEMPTED (deliberate last-20%/out-of-scope): raw FASTQ re-mapping with distiller-nf; H3K27me3-quantile obs/exp (Fig 3C, needs Marks 2012 ChIP); HiCCUPS loop calling + embryo pileups (Fig 4D, needs deep Bonev 2017 + many external Hi-C); all FISH/mass-spec/RNA-seq (wet-lab, not pipeline-derived from this GEO).

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-14
Rubric version
not recorded
Assessed by
Last updated
2026-07-29

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

Does the DNA hypomethylation that occurs when mouse ESCs transition to the naive ground state (2i) alter 3D genome organization, and is any such re-organization driven by DNA methylation loss acting through polycomb redistribution rather than by the altered cell state itself?

Core claims
  • The altered 3D genome of 2i-cultured ESCs (chromatin decompaction and loss of polycomb interactions at polycomb targets) is due to redistribution of polycomb away from its targets. finding
  • DNA hypomethylation is the direct cause of polycomb redistribution and the consequent 3D genome re-organization in 2i ESCs. mechanism
  • Preventing DNA hypomethylation during the transition to 2i restores primed-like H3K27me3 distribution and polycomb-mediated 3D genome organization, yet cells retain 2i ground-state functional characteristics. finding
  • Hypomethylated mouse pre-implantation blastocysts show 3D chromatin (HoxD decompaction) organization similar to 2i ESCs. finding
  • Restoring the epigenome/3D genome has limited impact on gene expression, cautioning against assuming causal roles for the epigenome and 3D genome in gene regulation in ESCs. finding
  • 3D FISH and in situ Hi-C were used to assay local chromatin compaction and genome-wide chromatin interactions across culture conditions and genotypes. method
  • The 3B3L mESC line (Dnmt3a−/−;Dnmt3b−/− with CAG-driven Dnmt3b and Dnmt3l transgenes) maintains high DNA methylation under 2i, providing a resource to uncouple methylation from cell state. resource
Experimental setups
Assay System Perturbation Readout Platform
3D fluorescence in situ hybridization (FISH) inter-probe distance measurement WT, Ring1B−/−, Eed−/−, and E14 mESCs in serum/LIF or 2i/LIF culture condition (serum vs 2i) and polycomb knockouts (Ring1B, Eed) inter-probe distances (μm) across HoxD, HoxB, HoxC and a control locus (chromatin compaction)
3D FISH in whole embryos E3.5 mouse blastocysts (13 individual blastocysts) none (in vivo developmental state) HoxD and control locus inter-probe distances
in situ Hi-C E14 mESCs grown in serum/LIF and 2i/LIF (two independent datasets per condition) culture condition (serum vs 2i) genome-wide normalized chromatin contact frequencies, local insulation, compartment eigenvectors, pileups at RING1B/CGI/loop sites
HPLC / mass spectrometry of DNA methylation WT, 3B3L, and TKO mESCs in serum/LIF and 2i/LIF genotype (Dnmt transgene-rescued 3B3L; TKO) and culture condition global percentage of methylated cytosine normalized to total guanine HPLC / mass spectrometry
ChIP-seq (referenced/integrated) mESCs in serum or 2i culture condition H3K27me3, RING1B, Suz12 occupancy profiles
Transcript level measurement WT and 3B3L mESCs in 2i genotype/culture condition Uhrf1 transcript levels
Key results
  • HoxD locus significantly decompacts in 2i/LIF relative to serum/LIF mESCs median inter-probe distance ~300 to ~400 nm
  • Decompaction in serum occurs to the same extent in Ring1B−/− or Eed−/− cells, with no further decompaction under 2i, indicating polycomb titration accounts for it
  • Control locus inter-probe distances are not significantly different between conditions or genotypes despite hypomethylation in 2i
  • HoxD in E3.5 blastocysts is decompact relative to serum mESCs and resembles 2i compaction state
  • Hi-C contact frequencies are depleted at HoxA, HoxB, HoxC, HoxD in 2i; significant for HoxA/B/C but not HoxD
  • Long-range contacts between polycomb loci (e.g., Skida1-Bmi1, En2-Shh-Mnx1) are depleted/lost in 2i
  • RING1B-associated loops are depleted in 2i while CTCF-associated interactions are preserved or enhanced
  • 3B3L cells retain high CpG DNA methylation in 2i, similar to serum, unlike WT 2i cells
Key statistics
  • other up to 75% reduction of H3K27me3 at polycomb targets in 2i (H3K27me3 loss at polycomb targets including Hox clusters in 2i (Marks et al.))
  • pvalue p < 0.0001 (HoxD inter-probe distance increase serum vs 2i (highly significant))
  • mean median ~300 nm (serum) to ~400 nm (2i) (HoxD inter-probe distances)
  • count n = 181 (long (>10 kb) regions of RING1B binding used in rescaled pileups)
  • count 13 individual blastocysts (E3.5 blastocysts measured for HoxD compaction)
  • other 650 kb (distance separating Skida1 and Bmi1 polycomb loci on chromosome 2 with long-range contacts)
  • other 1.3 Mb (span across En2, Shh, Mnx1 loci on chromosome 5 with long-range contacts)
  • other 100-kb cluster (size of HoxD polycomb domain demarked by H3K27me3, PRC2, PRC1)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This study used 3D FISH inter-probe distance measurements and in situ Hi-C to characterize chromatin compaction and long-range interactions at polycomb target loci across multiple mESC conditions (serum/LIF vs 2i/LIF, WT vs polycomb mutants vs methylation-preserved 3B3L cells) and in E3.5 blastocysts. FISH distance distributions were compared using tests not named in the main text but deferred to supplementary Tables S1 and S2, with results displayed as violin plots and p-value thresholds. Hi-C data were analyzed via Z-score tests for local contact depletion and pileup averaging for long-range interactions, with local interactions reported as mean ± 95% CI across H3K27me3 quantiles. DNA methylation was quantified by mass spectrometry and reported as means of two technical replicates with SD.

Replicationmixed Sample sizeTwo independent Hi-C datasets per condition explicitly stated; FISH n deferred to supplementary Tables S1 and S2; 13 blastocysts for one comparison; two technical replicates for mass spectrometry (Figure 5A) GroupsSerum/LIF vs 2i/LIF mESCs; WT vs Ring1B−/− vs Eed−/−; WT vs 3B3L (methylation-preserved in 2i); E3.5 blastocysts vs cultured cells; published embryo Hi-C datasets (ICM, E3.5, E6.5 epiblast/visceral endoderm, E7.5 ectoderm) Pairingunpaired Randomization/blindingnot stated Dispersionmixed Exact p-valuesno Effect sizesno Confidence intervalsyes Multiplicity correctionnot stated
Statistical tests used
Test Applied to n Assumptions
Not named in main text; deferred to supplementary Tables S1 and S2 FISH inter-probe distances at HoxD, HoxB, HoxC, and control loci across serum, 2i, Ring1B−/−, Eed−/−, 3B3L conditions, and E3.5 blastocysts Not stated in main text; 13 individual blastocysts noted for one comparison not stated
Z score analysis of Hi-C contact frequency (observed vs expected) Local Hi-C contact depletion at HoxA, HoxB, HoxC, HoxD in 2i vs serum (Figure S2C) Two independent Hi-C datasets per condition not stated
Rescaled pileup averaging of Hi-C contact frequencies (Flyamer et al., 2019) Genome-wide aggregated interactions at RING1B binding sites (>10 kb), CGI-to-CGI contacts, and annotated RING1B- and CTCF-associated loops n = 181 long RING1B binding sites for Figure 3B; loop counts not stated in main text na
Approaches that could also have been used
  • Statistical test names for FISH inter-probe distance comparisons are absent from the main text and figure legends, deferred entirely to supplementary Tables S1 and S2
    Could also: Name the test (e.g., Mann-Whitney U or Kolmogorov-Smirnov, both common for non-normal distance distributions) directly in the figure legend or Methods — Reporting the test name alongside results allows readers to immediately assess its appropriateness for the data type without locating supplementary material; it also facilitates reproducibility
  • Multiple pairwise comparisons are made across loci (HoxA–D, control), genotypes (WT, Ring1B−/−, Eed−/−, 3B3L), and conditions (serum, 2i) with threshold-based p values
    Could also: A factorial or two-way ANOVA with a post-hoc correction (e.g., Tukey HSD or Benjamini-Hochberg FDR) applied across the family of contrasts — An omnibus test followed by corrected pairwise comparisons is one standard approach for controlling error rate across many simultaneous contrasts within a structured experimental design
  • FISH results are described only via p-value thresholds; median inter-probe distances are mentioned narratively (~300 to ~400 nm) but no standardized effect size is reported
    Could also: Report a standardized effect size (e.g., rank-biserial correlation for a non-parametric test, or Cohen's d) alongside each p value — Effect sizes quantify the magnitude of the chromatin compaction difference independently of sample size, facilitating cross-study comparison and power estimation for future work
  • DNA methylation quantification in Figure 5A is based on the mean of two technical replicates with SD
    Could also: Use biological replicates (independent cell cultures or passages) as the unit of variation rather than, or in addition to, technical replicates — Biological replicates capture between-experiment variability, which is typically of greater interest than within-run measurement precision when reporting global methylation levels across cell lines
  • Hi-C pileup enrichment at RING1B and CTCF loop anchors is summarized by the center-pixel value displayed on heatmaps, with conditions compared by visual inspection
    Could also: Bootstrap or permutation-based confidence intervals on pileup scores, enabling formal statistical comparison of enrichment between serum and 2i conditions — Uncertainty estimates on aggregate contact scores would allow quantitative comparison of enrichment across conditions, complementing the visual heatmap comparison
  • Differential Hi-C contact analysis at Hox loci uses Z scores on contact frequency matrices
    Could also: Negative-binomial or Poisson model-based differential analysis (e.g., diffHic, multiHiCcompare) applied to the count data across replicate datasets — Model-based approaches account for the count nature and overdispersion of Hi-C data and provide FDR-adjusted p values genome-wide, which is a standard alternative in the field for comparing contact frequencies between conditions
Software: Not named in the provided text (Methods section not included)

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-31722211

Paper: McLaughlin et al. 2019, Cell Reports 29(7):1974-1985.e6. "DNA Methylation Directs Polycomb-Dependent 3D Genome Re-organization in Naive Pluripotency." DOI 10.1016/j.celrep.2019.10.031 · PMCID PMC6856714 · Data GSE124342.

What the paper computes (pipeline map)

Result Pipeline / tool Inputs In scope?
Hi-C read mapping → contact matrices distiller-nfcooler (mm9, mapq≥30, PCR-dup removed), balanced raw FASTQ (SRP174419) No — heavy raw re-mapping. The output (balanced .mcool) is shipped on GEO and is what downstream tools consume → we reproduce the downstream steps on the shipped coolers (brief P16: third-party tool on the paper's data is equally valid).
Pile-up of CGI–CGI cis interactions, RING1B-bound vs not, serum vs 2i (Fig 4C, Fig 3B) coolpup.py (Flyamer 2019), 5 kb, 205 kb window, obs/exp shipped coolers + CGI list (mm9) + RING1B peaks (Illingworth 2015) Yes (headline) — needs external RING1B peaks + CGI annotation.
Insulation scores, serum vs 2i (Fig S3C; "globally similar") cooltools diamond-insulation, 25 kb, window 100 kb shipped coolers only Yes (self-contained)
Compartments / eigenvector, serum vs 2i (Fig S2B; "larger effect on compartmentalization") cooltools call-compartments/eigs-cis, 200 kb, GC phasing shipped coolers + mm9 GC Yes (self-contained)
Contact decay P(s) / loss of local interactions in 2i (Fig 3A) cooltools expected-cis shipped coolers only Yes (supporting)
obs/exp in 25 kb windows split by H3K27me3 quantile (Fig 3C) custom obs/exp + H3K27me3 ChIP (Marks 2012) No — external ChIP; last-20%.
Loop calling + pileups across embryo Hi-C (Fig 4D) cooltools call-dots (HiCCUPS) on deep Bonev 2017 data external deep Hi-C + many embryo datasets No — external/heavy; last-20%.
FISH inter-probe distances, mass-spec methylation, RNA-seq rlog (Figs 1,2,5,6) wet-lab / imaging No — not pipeline-derived from GSE124342.

Reproduction targets (this room)

  1. Insulation similarity — Pearson r of insulation scores 2i vs serum (claim: high / "globally similar").
  2. Compartment effect — Pearson r of eigenvector 2i vs serum (claim: compartments more affected than insulation ⇒ r_eig < r_insul).
  3. CGI/RING1B decompaction (headline) — center-pixel obs/exp enrichment of CGI–CGI pileups, RING1B-bound vs not, serum vs 2i (claim: enriched in serum at RING1B CGIs, depleted in 2i).
  4. (supporting) P(s) 2i/serum short-range contact ratio.

Reported values to compare against

  • Pile-up enrichment values are printed in figure corners (Fig 3B, 4C) — direction is text-stated: RING1B-bound CGI contacts enriched in serum, depleted in 2i.
  • RING1B regions >10 kb used for rescaled pileup: n = 181 (Fig 3B legend).
  • Insulation: 2i vs serum "globally similar" (qualitative, Fig S3C).
  • Compartments: "larger effect of culture conditions on compartmentalization" than insulation (Fig S2B).

Data

  • GSE124342_serum.mm9.mapq_30.1000.mcool (balanced, base 1 kb, multires)
  • GSE124342_2i.mm9.mapq_30.1000.mcool
  • GSE124342_statistics.tsv.gz (distiller QC stats)
  • All staged on «infra» .../reproductions/pmid-31722211/; never on «host».

External annotations (for headline pileup)

  • CGI list: UCSC mm9 cpgIslandExt (canonical, unambiguous) — primary.
  • RING1B peaks (paper cites Illingworth et al. 2015): GSE69955 (Ring1B ChIP-seq in WT mESCs). Used only for the optional RING1B-split bonus; peak-calling choices make it the "last 20%".
Figures / tables: Fig S3CFig S2BFig 3BFig 4CFig 3A
insulation-similarity
Reported
Insulation scores are globally similar between serum and 2i mESCs (Fig S3C; 'local Hi-C interactions globally are similar... as assessed by correlation of their insulation scores').
Reproduced
NOT YET COMPUTED — SLURM «job» still building conda env at forced finalization. cooltools diamond-insulation (25kb, 100kb window) on shipped coolers was scripted+submitted; output -> «infra» outputs/insulation_25kb_serum_vs_2i.tsv.
partial
compartment-effect
Reported
Culture conditions affect compartmentalization more than insulation (Fig S2B, 'a larger effect of culture conditions on compartmentalization').
Reproduced
NOT YET COMPUTED — cooltools eigs-cis (200kb, GC phasing) scripted+submitted; expected testable ordering r_eigenvector < r_insulation; output -> «infra» outputs/eigenvector_200kb_serum_vs_2i.tsv.
partial
cgi-ring1b-decompaction-headline
Reported
Contacts between RING1B-bound CGIs are enriched in serum and greatly depleted (decompacted) in 2i (Fig 3B rescaled pileups, n=181 RING1B regions >10kb; Fig 4C CGI pileups, coolpup.py).
Reproduced
NOT YET COMPUTED — coolpup.py CGI-CGI pileup job (5kb, 205kb window) prepared but not run; pending env completion. Primary uses UCSC mm9 CGIs; optional RING1B split via GSE69955.
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

No assessment has been recorded yet.
🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

297.1 k
tokens (I/O) · 15.7 M incl. cache
78 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.