Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Quantitative epigenetic co-variation in CpG islands and co-regulation of developmental genes.

Sci Rep · 2013
L1 95/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +6
✓ What held up
  • Any deviation was negligible
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🟡Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🔴A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
95/100
Reproducibility score
1.2 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 89% of all assessed papers rank 105 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Liu et al. 2013 (PMID 23999385) ships a working, self-contained Java tool (QDCMR) whose core Shannon-entropy DEM-CGI classification algorithm was independently reproduced exactly (threshold=0.962 for n=3, SD=0.07) by running the authors' own binary on synthetic test data. The paper's central result -- 5,194 DEM-CGIs -- was verified as internally consistent: the authors' own supplementary Table S1 contains exactly 5,194 well-formed data rows with plausible genomic/epigenomic values. All 8 supplementary tables were confirmed genuine and well-structured (an earlier suspicion of file corruption, based on an incomplete manual byte-scan, was corrected after proper parsing). Real GEO series-matrix metadata was downloaded and inspected for all 4 declared/discovered accessions (GSE12241, GSE11172, GSE8024, GSE10246), confirming a coherent (though not fully paper-documented) data-provenance story across ESC/NPC/Brain and 4 epigenetic marks (H3K4me2, H3K4me3, H3K27me3, DNAm). A full, independent, raw-data-to-final-table reproduction was NOT attempted: the paper's own code repository does not ship the raw-ChIP-seq/RRBS-to-matrix pipeline (only the downstream classification step), making that a much larger undertaking than reproducing the shipped, documented tool. One of five underlying data sources (a legacy, pre-GEO Broad Institute RRBS archive for DNA methylation) could not be located from the Methods' description and returned an external HTTP 502 error on the one plausible archive path tried. This is NOT a claim of full reproduction -- it is a faithful, auditable account of what was and was not verified, and why.

💻 Code ↗ 🗄 Data: GSE12241

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-07-30
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-07-31
no human curator yet
Last updated
2026-07-31

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

The paper asks whether multiple epigenetic modifications (DNA methylation, H3K4me2, H3K4me3, H3K27me3) in CpG islands vary in a quantitatively coordinated (co-varying) manner during mammalian neural differentiation, and whether such co-variation contributes to co-regulation of developmental genes.

Core claims
  • Four epigenetic modifications (DNA methylation, H3K4me2, H3K4me3, H3K27me3) in mouse CGIs undergo combinatorial variation (co-variation) across ESCs, NPCs and adult brain during neuron differentiation. finding
  • DNA methylation variation is significantly positively correlated with H3K27me3 variation and negatively correlated with H3K4me2/3 variation. finding
  • An optimized entropy-based QDMR strategy quantifies epigenetic variation across multiple samples and identifies CGIs differentially modified by epigenetic modifications (DEM-CGIs), overcoming the two-sample limitation of methods such as ChIPDiff and DIME. method
  • 5,194 DEM-CGIs were identified, 92% of which lie near 4,508 known genes (DEMGs) enriched for embryonic development and neuron differentiation processes. finding
  • Differentially DNA methylated CGIs overlap significantly with H3K27me3-differentially modified CGIs, and the two repressive marks behave antagonistically across development stages (methylation up / H3K27me3 down in NPCs). mechanism
  • Combinations of several epigenetic modifications explain gene expression better than any single modification, though the optimal combination differs between developmental stages. finding
  • Differentially expressed genes are preferentially differentially epigenetically modified, and imprinted genes are enriched among DEMGs, suggesting dynamic CGI modification marks imprinted genes. finding
  • Dynamic CGI epigenetic modification affects core reprogramming transcription factors (Klf4, Oct4/Pou5f1, Sox2 dynamic; c-Myc stable) during ESC-to-brain differentiation. mechanism
Experimental setups
Assay System Perturbation Readout Platform
DNA methylation profiling (genome-wide, CGI-level; reanalyzed public data) Mouse embryonic stem cells, neural precursor cells and adult brain none (developmental stage comparison) DNA methylation level per CpG island
ChIP-seq for H3K4me2 Mouse ESCs, NPCs and adult brain none H3K4me2 modification level per CpG island
ChIP-seq for H3K4me3 Mouse ESCs, NPCs and adult brain none H3K4me3 modification level per CpG island
ChIP-seq for H3K27me3 Mouse ESCs, NPCs and adult brain none H3K27me3 modification level per CpG island
Gene expression profiling Mouse ESCs, NPCs and adult brain none Expression levels of 6,026 genes; identification of 429 differentially expressed genes
Computational entropy-based quantification (optimized QDMR) of epigenetic variation 8,337 mouse CGIs annotated to seven genome regions (Up2kb, 5'UTR, CodingExon, Intron, 3'UTR, Down2kb, Intergenic) none Per-CGI entropy value per modification; DEM-CGI calls at entropy threshold 0.962 UCSC Table Browser CGI annotation; Circos for visualization
Correlation and best subsets regression analysis of modifications versus expression 3,916 DEM-CGIs paired with 3,699 related DEMGs across three stages none Correlation coefficients between modifications and between modification and expression; regression models of expression
Gene ontology functional enrichment and overlap/enrichment statistics DEMG, DEG, DEMG&DEG and imprinted gene sets (mouse) none Enriched GO biological process terms; overlap significance (chi-square test)
Key results
  • Of 8,337 CGIs with all four modifications across three stages, 5,194 were differentially modified by at least one modification (DEM-CGIs) 62% (5,194/8,337)
  • 4,778 DEM-CGIs were located near 4,508 known genes (DEMGs) 92% (4,778/5,194)
  • CGIs with low methylation entropy had low H3K27me3 entropy but high H3K4me2/3 entropy in all genome regions; methylation variation positively correlated with H3K27me3 variation and negatively with H3K4me2/3 variation
  • DNA methylation and H3K27me3 entropy showed similar bimodal distributions, while H3K4me2 and H3K4me3 shared a similar unimodal distribution
  • 504 CGIs were differentially modified by both DNA methylation and H3K27me3, and were enriched in CodingExon regions relative to other CGIs 40% (199/504) vs 21% (1,710/8,337)
  • Most DNAm&H3K27me3-DEM-CGIs increased methylation in NPCs (decreased in brain and ESCs) while decreasing H3K27me3 in NPCs and increasing it in brain
  • 80% of differentially expressed genes were also DEMGs, more than expected by chance 80% (341/429) vs 61% (3,699/6,026) expected; p < 0.0001
  • Imprinted genes were related to DEM-CGIs more often than non-imprinted genes 83% (25/30) vs 62% (4,449/7,214); chi-square p < 0.05
Key statistics
  • count 15,948 mouse CGIs obtained from UCSC Table Browser; 8,337 selected with all four modifications in three stages (Dataset construction)
  • count 5,194/8,337 (>62%) DEM-CGIs; 4,778 near 4,508 genes (92%); six DEM-CGIs modified by all four marks, five near transcriptional start sites (DEM-CGI identification)
  • other entropy threshold 0.962 from a probability model for three samples (DEM-CGI calling cutoff)
  • count 504 DNAm&H3K27me3-DEM-CGIs; 199/504 (40%) in CodingExon vs 1,710/8,337 (21%) of other CGIs (Overlap of two repressive marks)
  • pvalue p < 0.0001 (Enrichment of DEMGs among 429 DEGs (341/429, 80%) vs 61% expected)
  • pvalue p < 0.05 (Chi-square test) (Imprinted genes (25/30, 83%) overlapping DEMGs vs non-imprinted (4,449/7,214, 62%))
  • count expression levels of 6,026 genes; 3,916 DEM-CGIs analyzed against 3,699 DEMGs (82%, 3,699/4,508) (Expression-epigenome correlation analysis)
  • count 30 imprinted genes among 7,244 genes related to 8,337 CGIs; ~50% of mouse CGIs located in promoters of known genes (Genome annotation context)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The paper used an entropy-based quantification strategy (an extension of their previously published QDMR method) to score epigenetic variation across three developmental stages (ESCs, NPCs, adult brain) in CpG islands, then applied a probability-model-derived threshold to call differentially modified CGIs (DEM-CGIs). Relationships between variation in different epigenetic marks, and between marks and gene expression, were assessed with correlation analysis and a best subsets regression, while categorical overlaps (e.g., imprinted genes vs. DEMGs, DEGs vs. DEMGs) were assessed with a chi-square test and chance-expectation comparisons. Results are reported mainly as counts, percentages, and p-value thresholds rather than as measures of dispersion or confidence intervals.

Replicationunclear Sample sizeDescribed as counts of genomic features (e.g., 8,337 CGIs, 15,948 total CGIs, 6,026 genes with expression data) rather than a biological-replicate sample size or power calculation GroupsThree developmental stages (mouse ESCs, neural precursor cells, adult brain), compared genome-wide across CGIs/genes Pairingna Randomization/blindingna Dispersionnone Exact p-valuesno Confidence intervalsno Multiplicity correctionnot stated
Statistical tests used
Test Applied to n Assumptions
Correlation analysis (type not specified) methylation entropy vs. H3K27me3/H3K4me2/H3K4me3 entropy (Figure 1d, Supplementary Figure S3) 8,337 CGIs not stated
Probability-model-based threshold (entropy cutoff = 0.962) identification of DEM-CGIs across ESC/NPC/brain (Figure 2a) 8,337 CGIs (three samples) not stated
Chi-square test overlap of imprinted genes with DEMGs vs. non-imprinted genes (Supplementary Table S7) 30 imprinted genes among 7,244 genes not stated
Chance-expectation comparison (test type not named), reported as p < 0.0001 overlap of differentially expressed genes (DEGs) with DEMGs (Supplementary Table S6) 429 DEGs among 6,026 genes not stated
Best subsets regression modeling gene expression levels from combinations of epigenetic modifications (Results, 'co-regulate the developmental genes' section) 3,699 DEMGs / 3,916 DEM-CGIs not stated
Functional/GO enrichment analysis DEMGs and DEGs gene ontology biological process terms (Figure 2d, Table 1) 4,508 DEMGs; 429 DEGs not stated
Approaches that could also have been used
  • Correlation between entropy-based variation scores (e.g., methylation entropy vs. H3K27me3 entropy) was assessed without specifying whether a Pearson or Spearman coefficient was used.
    Could also: A rank-based Spearman correlation could also be used — Entropy scores are bounded and can have non-normal or bimodal distributions (as the paper itself notes for methylation/H3K27me3 entropy), so a rank-based measure can be a robust complement to a linear correlation coefficient in that setting.
  • Many correlation, enrichment, and overlap comparisons were performed across multiple epigenetic marks, genome regions, and developmental stages.
    Could also: A multiple-testing correction such as Benjamini-Hochberg FDR could also be applied across the family of comparisons — When many tests are run in parallel, an FDR or Bonferroni adjustment is a standard way to control the overall false-positive rate across the full set of comparisons.
  • The threshold for calling DEM-CGIs (entropy = 0.962) was derived from a probability model described in a prior publication.
    Could also: A permutation-based empirical null distribution could also be used to set the cutoff — An empirical/permutation approach can complement a parametric probability model by directly estimating the null distribution from the observed data, which some readers find easier to relate to a chosen false-discovery threshold.
  • A best subsets regression was used to relate combinations of epigenetic modifications to gene expression levels.
    Could also: A regularized regression approach such as LASSO or elastic net could also be used — With several correlated epigenetic predictors, regularized regression can handle collinearity and perform variable selection in a way that scales well as the number of candidate predictors grows.
  • The overlap between imprinted genes and DEMGs was tested with a chi-square test on a 30-gene imprinted-gene set.
    Could also: Fisher's exact test could also be used for this comparison — Fisher's exact test is often preferred for contingency tables with small expected cell counts, which can arise with a category as small as 30 imprinted genes.
  • Gene ontology enrichment results for DEMGs and DEGs were reported as enriched terms without a stated correction method.
    Could also: A hypergeometric or Fisher's exact test with FDR correction (as implemented in tools like DAVID or clusterProfiler) could also be used — This is a widely used standard for GO enrichment that explicitly reports adjusted p-values across the many gene sets tested, which can aid comparison with other enrichment studies.
Software: Circos · QDMR (entropy-based method, authors' own tool, ref. 23)

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

C1_dem_cgi_threshold_algorithm
Reported
Paper's methods state the Shannon-entropy-based DEM-CGI classification threshold is 0.962 for n=3 samples (SD=0.07), computed via the authors' QDCMR tool.
Reproduced
Tool output (Statisitcs.txt): 'The Threshold for DCMR is: 0.962' - exact match to paper's reported value, obtained independently from the shipped binary rather than by reading the paper.
exact
C2_dem_cgi_count_5194
Reported
Paper reports 5,194 DEM-CGIs (CpG islands differentially modified across ESC/NPC/Brain) as its central catalogue.
Reproduced
Table S1 'Information of DEM-CGIs and related DEMGs' contains exactly 5,194 data rows (5196 total rows minus 1 title row minus 1 header row), with plausible per-CGI genomic coordinates, CpG/GC content, obsExp ratios, and per-mark classification columns (H3K4me2-DEM-CGI, H3K4me3-DEM-CGI, H3K27me3-DEM-CGI, DNAm-DEM-CGI).
exact
C3_data_provenance_coherence
Reported
Paper's stated data sources (GEO accessions for ChIP-seq/expression, non-GEO Broad FTP for RRBS methylation) actually contain what the paper implies.
Reproduced
GSE12241 = Mikkelsen et al. chromatin-state compendium (21 samples: ESC/NPC/MEF/ESHyb x H3K4me3/H3K9me3/H3K27me3/H3K36me3/H4K20me3/RPol2/H3/WCE). GSE11172 = a purpose-built 7-sample series explicitly described as being 'to examine the correlation between histone and DNA methylation during lineage-commitment' (NP/ES H3K4me2+H3K4me1, Brain H3K4me3+H3K4me2+H3K27me3) -- this is clearly the source of H3K4me2 for all 3 cell types AND Brain H3K4me3/H3K27me3 (absent from GSE12241). GSE8024 = 8-sample Affymetrix GPL1261 expression set for ES/NPC/MEF (isogenic 129SvJae x C57BL/6) -- covers ESC/NPC expression but has MEF, not Brain. GSE10246 = GNF Mouse GeneAtlas V3, 189-sample GPL1261 tissue atlas containing many brain SUB-regions (cerebral cortex, cerebellum, hippocampus, hypothalamus, amygdala, striatum, olfactory bulb, spinal cord, pituitary, dorsal root ganglia) but no single generic 'whole brain' sample.
within tolerance

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 95/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟡2. Endpoint comparability
🟡3. Location of the main deviation
🔴4. Cause of the deviation
🟡5. Derivability / plausibility
🟢6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +6

The one directly comparable scalar reproduced exactly — running the authors' own QDCMR.jar on synthetic 5-CGI x 3-sample input yielded The Threshold for DCMR is: 0.962, matching the paper's stated n=3/SD=0.07 threshold. The headline product, 5,194 DEM-CGIs, was only verified as internally consistent (Table S1 of srep02576-s2.xls holds exactly 5,194 well-formed data rows); it could not be re-derived because the repo ships no raw-ChIP-seq/RRBS-to-GCT pipeline (and not even its own documented DATA/example.gct), the Methods omit which GSE10246 brain region stands for 'Brain', and the legacy Broad RRBS methylation archive has no accession and returned HTTP 502. The gap therefore sits on the authors'/availability side, not in our computation: there are zero numeric discrepancies anywhere, no magnitude or significance flip, and no fabrication signal — but also no independent path to the central number. Hence yellow across derivability/core-claim/overall rather than green (nothing independently confirms the biology) or red (nothing contradicts it).

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.