DNA demethylation is associated with malignant progression of lower-grade gliomas.
Part of the results reproduced; minor but material deviations remained.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- Nothing in this column.
- 🔴Could not use the authors’ exact input data
- 🔴Reported values were only indirectly comparable
- 🟡A deviation arose in the data or preprocessing
- 🟡A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The deviation was non-trivial in magnitude
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
This paper has a computational component, but its primary data is legally or ethically access-restricted — identifiable patient cohorts, rare-disease genomes, or controlled-access biobanks that cannot be openly shared. The reproduction therefore could not be attempted. That is a neutral verdict: it does not mean the result is wrong or that the authors fell short — only that, for legitimate privacy reasons, it cannot be independently checked from public data. We deliberately do NOT assign a 0–100 score here, because a low number would wrongly read as a failed reproduction.
▸Reproduction agent’s raw note
PARTIAL reproduction with compute executed in-room as a SLURM batch job («our HPC» «job», compute node n155, COMPLETED). The paper's RAW data (122-sample Infinium 450K methylation, 36-sample whole-exome, 31-sample RNA-seq) is deposited in the JGA controlled-access archive JGAS00000000146, so the upstream pipelines (unsupervised methylation clustering, DMP t-test, RNA-seq differential expression, the karkinos somatic genotyper, and Genomon-fusion) cannot be executed from raw input; those headline numbers (C2 DMP counts, C3b down-genes, C6 replication-timing enrichment p-value, C7 fusions) are therefore honestly data_restricted and were graded without inventing any value. GSE63428 in the manifest is NOT the paper's own data: it is the external Rivera-Mulia replication-timing reference (PMID 26055160), which was profiled (1.9 GB tar, 120 raw NimbleGen .pair across 4 platforms, intensities only, no processed RT track). Using the OPEN-access Sci Rep supplement (Tables S1-S6 = the deposited pipeline outputs + per-sample annotation), the job re-derived the reported summary statistics: cohort sizes (122 methylation, 31 RNA-seq) reproduce EXACTLY; the C.3 demethylated class (9 tumors, 8/9 grade IV) reproduces EXACTLY from S1/S2; the 116 upregulated-gene count reproduces EXACTLY from S3; the WES mutation burden, recomputed from the deposited 4,524-mutation karkinos call set (S6), lands within ~8-17% of the reported 26.2 (initial) / 56.9 (recurrent non-hypermutator) using the same metric, with the two temozolomide hypermutators (MT2-3, MT40-2) correctly identified, so the reported burdens are corroborated as real; and the '54% shared mutations' figure reproduces to 53.5% once the paper's denominator is identified as the fraction of the initial tumor's mutations retained at recurrence (intersection/primary), rather than a symmetric Jaccard (which gives 26.6%). The karkinos repo (the somatic genotyper behind C4) was cloned at the pinned commit d54c0ddf28252f69b448f1ce30d8a62d27792fdf and its calling parameters documented; it cannot be run because its tumor/normal WES BAM inputs are exactly the JGA-restricted data. Out of scope (non-pipeline): RT-PCR/Sanger/IHC validation, survival description, the DAVID GO web tool. No value was fabricated; every grade is provisional and human-checkable via AUDIT.md, original/claims.tsv, reproduction/agreement.json, and reproduction/outputs/repro_outputs.json.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessmentassessed: 2026-06-16 ⛓ 5ccac385ad9a
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-22
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-16no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetThe paper investigates the molecular mechanisms underlying malignant progression of lower-grade gliomas, hypothesizing that partial loss of the G-CIMP DNA hypermethylation phenotype (demethylation) during recurrence drives malignant transformation, potentially via passive demethylation in late-replicating chromatin during accelerated cell division.
- ★ Nearly half of IDH-mutant glioblastomas that progressed from lower-grade gliomas show characteristic partial DNA demethylation in previously methylated (G-CIMP) genomic regions of their initial tumor. finding
- ★ Demethylated regions in progressed tumors are significantly enriched in late-replicating chromatin domains, suggesting passive demethylation due to delayed maintenance methylation during accelerated cell division. mechanism
- ★ Cell cycle-related genes, RB pathway and PI3K-AKT pathway genes are frequently altered in the G-CIMP-demethylated glioblastomas. finding
- ★ IGF2BP3 is upregulated via promoter demethylation in G-CIMP-demethylated tumors, potentially drives cell proliferation, and its high expression is associated with worse patient survival. finding
- ★ Most demethylated regions in G-CIMP-demethylated (C.3) tumors are located outside CpG islands, in non-regulatory ('open sea') regions. finding
- ★ Unsupervised clustering of methylation profiles from 122 gliomas identifies five clusters (C.1-C.5) reflecting histology and genetics, including a distinct G-CIMP-demethylated subgroup (C.3) resembling TCGA's 'G-CIMP-low'. method
- IDH1-mutant glioma cell lines retaining both mutant and wildtype alleles cluster together with G-CIMP-demethylated tumors, suggesting rapid proliferation in culture drives a similar demethylation profile. finding
- TTK, CDK2, and NCAPG, all cell-cycle-associated genes, also show demethylated promoters among genes upregulated in C.3 tumors. finding
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| DNA methylation array | 122 gliomas (including 24 matched primary/recurrent pairs) and 3 normal brain samples | none | genome-wide DNA methylation (beta values), unsupervised clustering | Infinium HumanMethylation450K BeadChip |
| whole-exome sequencing | 36 gliomas | none | somatic mutations | — |
| RNA sequencing | 31 gliomas (C.3 n=8 vs C.1 non-codel n=23) | none | gene expression levels (FPKM), differential expression | — |
| Repli-seq replication timing analysis (reanalysis of prior report) | neural progenitor cells (NPCs) | none | genome-wide replication timing correlated with methylation change | — |
| gene ontology enrichment analysis | in silico analysis of differentially expressed genes from RNA-seq | none | functional category enrichment | DAVID (http://david.abcc.ncifcrf.gov/) |
| methylation and expression reanalysis of public dataset | TCGA G-CIMP-high and G-CIMP-low glioma tumors | none | IGF2BP3 promoter methylation and gene expression | — |
| DNA methylation clustering of public cell line data | two IDH1-mutant glioma cell lines retaining mutant and wildtype alleles | IDH1 mutant/wildtype allele retention | methylation profile similarity to tumor clusters | — |
| Kaplan-Meier survival analysis | TCGA G-CIMP-high and G-CIMP-low astrocytoma patients | none | overall survival stratified by IGF2BP3 expression | — |
- ▼ 33,695 probes were significantly hypomethylated versus only 635 hypermethylated in G-CIMP-demethylated (C.3) tumors compared to C.1 non-codel tumors. 33,695 vs 635 probes (q<0.05, diff>0.2)
- ▼ Probes in late-replicating regions were significantly enriched among demethylated probes in C.3 tumors. p<2.2×10^-16, Chi-square test
- ▼ Enrichment of late-replicating regions among demethylated probes was verified using independent TCGA G-CIMP-high vs G-CIMP-low data. p<2.2×10^-16
- – 116 genes were upregulated and 383 genes downregulated in C.3 versus C.1 non-codel tumors by RNA-seq. q<0.05, FC>2 or <0.5
- ▲ Upregulated genes in C.3 tumors were enriched for cell division, mitotic nuclear division, and chromatin/chromosome segregation.
- ▲ Only three genes showed both promoter demethylation and upregulated expression in C.3 tumors, including IGF2BP3. 3 genes identified
- ▼ In TCGA data, the G-CIMP-low group showed significantly lower IGF2BP3 promoter methylation and higher IGF2BP3 expression than the G-CIMP-high group. p<0.001 (methylation); p<0.001 (expression)
- ▲ Higher IGF2BP3 expression was significantly correlated with worse overall survival among TCGA G-CIMP-high and G-CIMP-low astrocytoma patients.
- count 33,695 hypomethylated probes vs 635 hypermethylated probes (C.3 vs C.1 non-codel gliomas, q<0.05, methylation difference>0.2)
- pvalue p<2.2×10^-16 (Chi-square test) (enrichment of late-replicating regions among demethylated probes, C.3 vs C.1)
- pvalue p<2.2×10^-16 (TCGA validation of replication timing vs methylation, G-CIMP-high vs -low)
- count 116 upregulated genes, 383 downregulated genes (RNA-seq comparison, C.3 (n=8) vs C.1 non-codel (n=23), q<0.05)
- pvalue p<0.001 (IGF2BP3 promoter methylation, TCGA G-CIMP-low vs G-CIMP-high)
- pvalue p<0.001 (IGF2BP3 gene expression, TCGA G-CIMP-low vs G-CIMP-high)
- count 122 gliomas profiled by methylation array (24 matched pairs); 36 by WES; 31 by RNA-seq (overall cohort size for molecular profiling)
- mean 5.3 years (lower-grade glioma to GBM); 1.4 years (anaplastic astrocytoma to GBM) (population-based mean time to malignant progression)
Statistical methods review
Model: opusA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This is an integrated multi-omics observational study of 122 gliomas (with whole-exome sequencing on 36 and RNA-seq on 31), including 24 matched primary/recurrent pairs, that profiles DNA methylation, mutations, and gene expression. Methylation subgroups were defined by unsupervised hierarchical clustering, and differential methylation/expression between predefined clusters (G-CIMP-demethylated C.3 vs. C.1 non-codel) was tested with paired two-sided moderated Welch's t-tests with Benjamini-Hochberg FDR correction. Enrichment of late-replicating domains was assessed by Chi-square test, and TCGA validation used Wilcoxon rank-sum and log-rank (Kaplan-Meier) tests. Results were reported with q-values/p-values and visualized via heatmaps, volcano plots, starburst plots, and box plots.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Unsupervised hierarchical clustering (top 10,000 variant probes) | Grouping of 122 gliomas into clusters C.1–C.5 (Fig. 1) and subset clustering (Supplementary Fig. S1, S3) | 122 tumors (subset analyses on C.1–C.4 and C.1/C.2) | na |
| Paired two-sided moderated Welch's t-test with Benjamini-Hochberg FDR | Differential methylation of probes between C.3 and C.1 non-codel tumors (Supplementary Fig. S2a, Fig. 2b) | C.3 n=9 vs. C.1 non-codel n=44 (419,382 probes examined) | not stated |
| Paired two-sided moderated Welch's t-test with Benjamini-Hochberg FDR | Differential gene expression between C.3 and C.1 non-codel tumors (Fig. 3a, Fig. 3c) | C.3 n=8 vs. C.1 non-codel n=23 | not stated |
| Paired two-sided moderated Welch's t-test with Benjamini-Hochberg FDR | Combined promoter methylation vs. expression (starburst plot, Fig. 4a) | — | not stated |
| Chi-square test | Enrichment of late-replicating regions among demethylated probes (Fig. 2b, Fig. 2c for TCGA) | — | not stated |
| Wilcoxon rank-sum test | IGF2BP3 promoter methylation and expression between TCGA G-CIMP-low vs. G-CIMP-high tumors (Fig. 4c) | — | na |
| Log-rank test (Kaplan-Meier) | Overall survival by IGF2BP3 expression in TCGA G-CIMP tumors (Fig. 4e, Supplementary Fig. S4) | — | na |
-
Methylation subgroups were defined by unsupervised hierarchical clustering on the top 10,000 variant probes, then used as fixed groups for downstream differential testing.↳ Could also: Cluster stability could also be assessed with consensus clustering or bootstrap/silhouette resampling, and downstream tests interpreted with awareness that groups were data-derived. — Quantifying cluster robustness conveys how reproducible the subgroup boundaries are, which is helpful when the same data both define and compare groups.
-
Group comparisons of methylation and expression used a paired two-sided moderated Welch's t-test with Benjamini-Hochberg FDR.↳ Could also: For RNA-seq counts, count-based models such as DESeq2 or edgeR (negative binomial) could also be used, and for methylation β-values an M-value transformation or beta-regression is an option. — These models are tailored to the mean–variance structure of count and bounded-proportion data and can improve sensitivity and calibration, especially with small n.
-
Late-replicating-domain enrichment was evaluated with a Chi-square test treating probes as independent observations.↳ Could also: A permutation/block-resampling test that accounts for spatial correlation along the genome, or reporting an odds ratio with a confidence interval, could also be used. — Genomically adjacent probes are correlated, so a resampling approach and an interval estimate convey both significance and effect magnitude while respecting that structure.
-
Differential thresholds combined a q-value cutoff with fold-change/methylation-difference cutoffs to define significant features.↳ Could also: A formal effect-size estimate with confidence intervals (e.g., log fold-change shrinkage) alongside FDR could also accompany each feature. — Interval estimates communicate the precision of each change and complement the binary significance call.
-
Survival association of IGF2BP3 in TCGA data was assessed with a univariate log-rank test on dichotomized high/low expression.↳ Could also: A Cox proportional-hazards model treating expression continuously and adjusting for covariates (e.g., grade, IDH/codel status) could also be used. — A multivariable model yields a hazard ratio with a confidence interval and helps describe the association independent of known prognostic factors, while avoiding information loss from dichotomization.
-
Dispersion in the box plots is shown graphically without an explicitly stated summary statistic (SD/SEM/IQR) in the text.↳ Could also: Explicitly stating the dispersion measure (e.g., IQR for box plots, or 95% CI of group medians) could also accompany the figures. — Naming the spread statistic makes the visualized variability unambiguous for readers, particularly given the modest group sizes.
Result convergence & founder nodes
Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.
-
Global CpG methylation is predominantly reduced in G-CIMP-demethylated gliomas vs non-codel lower-grade gliomas (33,695 hypomethylated vs 635 hypermethylated probes).microarray human glioma down 2019×1papers★ This paper is the founder (earliest)
-
G-CIMP-demethylated gliomas predominantly comprise IDH-mutant recurrent GBMs without 1p/19q codeletion, consistent with malignant progression.other human glioma 2019×1papers★ This paper is the founder (earliest)
-
Higher IGF2BP3 expression is significantly associated with worse overall survival in TCGA G-CIMP gliomas.other human glioma 2019×1papers★ This paper is the founder (earliest)
-
IGF2BP3 shows concurrent promoter demethylation and transcriptional upregulation in G-CIMP-demethylated gliomas.other human glioma up 2019×1papers★ This paper is the founder (earliest)
-
Late-replicating genomic regions are significantly enriched among demethylated CpG sites in G-CIMP-demethylated gliomas.other human glioma 2019×1papers★ This paper is the founder (earliest)
-
Cell division and chromosome segregation gene sets are enriched among genes upregulated in G-CIMP-demethylated vs non-codel lower-grade gliomas.RNA-seq human glioma up 2019×1papers★ This paper is the founder (earliest)
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope analysis — pmid-30760837
Paper: Nomura M, et al. "DNA demethylation is associated with malignant progression of lower-grade gliomas." Sci Rep 2019;9:1903. PMID 30760837 · PMCID PMC6374451 · DOI 10.1038/s41598-019-38510-0.
Pointers: code = github.com/genome-rcast/karkinos (HEAD
d54c0ddf28252f69b448f1ce30d8a62d27792fdf); manifest data = geo:GSE63428.
Data situation (the central constraint)
- Primary data (450K methylation 122 tumors, WES 36 tumors+normal, RNA-seq 31) is deposited in the Japanese Genotype-phenotype Archive (JGA) under controlled-access accession JGAS00000000146 → cannot be downloaded without a Data Access Committee application. Raw-input re-execution of the clustering / DMP / DE / karkinos / fusion pipelines is therefore blocked.
- GSE63428 (the manifest accession) is NOT the paper's data: it is the
external Rivera-Mulia et al. replication-timing reference (PMID 26055160,
60 samples / 29 cell types, raw NimbleGen
.pair). Public, but ships RAW intensities only (120.pair+ manifests), no processed RT track. - HOWEVER, the open-access Sci Rep Supplementary Tables S1–S6 ship the deposited pipeline OUTPUTS + per-sample annotation: S1 master table (per- sample cluster assignment, grade, IDH, RNA-seq/WES/methylation flags), S2 the 9 C.3 demethylated samples, S3 the 116 up-genes, S4 WES read stats, S5 tumor content, S6 the full 4,524 somatic-mutation call set (karkinos output). This enables re-derivation of the reported summary statistics from the deposited data — a genuine internal-consistency / fabrication check — even though the upstream pipelines cannot be re-run from raw input.
Reproduction strategy adopted
- Re-derive every reported number that the deposited supplement tables make computable (cohort Ns, C.3 class composition, up-gene count, mutation burden, shared-mutation %). Run as a logged SLURM job on «our HPC» («job», node n155, COMPLETED). → graded in claims.tsv / agreement.json.
- Profile GSE63428 fully (data type, file inventory, N, QC) in the same job.
- Mark restricted every claim whose only path is the JGA raw data
(C2 DMP list, C3 down-genes, C6 enrichment p-value, C7 fusions) — honest
data_restricted, no fabricated value.
| # | Reported result | Pipeline | Input | Verdict |
|---|---|---|---|---|
| C0a/b/c | cohort 122 meth / 31 RNA / 36 WES | sample QC | S1/S4 | exact / within-tol (deposited annotation) |
| C1 | 9 C.3 tumors, 8/9 grade IV | 450K clustering | S1+S2 | exact (output corroborated; clustering not rerun – JGA) |
| C2 | 33,695 hypo / 635 hyper probes | t-test on 450K beta | JGA 450K | uncheckable (probe list not deposited) |
| C3a | 116 up-genes | RNA-seq DE | S3 | exact |
| C3b | 383 down-genes | RNA-seq DE | JGA | uncheckable (not deposited) |
| C4a/b | burden 26.2 / 56.9 | karkinos → aggregate | S6 | partial (recomputed 30.6/52.3; karkinos not rerun – JGA) |
| C4c | 54% shared | pair compare | S6 | within-tol (53.5% = inter/primary; definition pinned) |
| C5 | 3 promoter-demeth+upreg genes | join C2×C3 | S3(+JGA) | partial (expression side only) |
| C6 | RT enrichment p<2.2e-16 | RT classify × DMP | GSE63428(raw)+JGA | uncheckable (no RT track; DMP restricted) |
| C7 | gene fusions | Genomon-fusion | JGA RNA-seq | uncheckable (restricted + wet-lab) |
Out of scope (non-pipeline): RT-PCR/Sanger/IHC validation, survival description, DAVID GO web tool.
Outcome
Partial reproduction. Compute ran on «our HPC»; the cleanly-specified deposited outputs (cohort Ns, C.3 composition, 116 up-genes) reproduce exactly, the mutation-burden statistics reproduce within ~8–17% from the deposited call set, the 54%-shared figure reproduces (53.5% once the denominator = inter/primary is identified), and the remaining headline numbers (C2 DMPs, C3b down-genes, C6 RT-enrichment, C7 fusions) are honest
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
This is a clean data_restricted drop: all seven reported targets (e.g. 33,695/635 hypo/hypermethylated probes, 116/383 DE genes, WES burden 26.2/56.9 via karkinos) derive from controlled-access JGA data (JGAS00000000146) that cannot be obtained without a Data Access Committee application. The deviation/cause therefore lies entirely on the data-availability axis, not on the authors' side and not in our methodology — the karkinos genotyper is public and builds, the values are presumably derivable for a JGA-credentialed holder, and no value was fabricated. Nothing was confirmed or refuted, so the work is solid but unverifiable; criticality is yellow, not red.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.