DMN-seq enriches DNA hypomethylated regions for biomarker discovery using 5-methylcytosine glycosylase.
The main results reproduced: recomputed values matched the published ones within tolerance.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Reported values are derivable from the shared data
- ✓Any deviation was negligible
- ✓The central claim held under reproduction
- 🟡Could not use the authors’ exact input data
- 🟡Reported values were only indirectly comparable
- 🟡A deviation arose in the data or preprocessing
- 🟡A deviation was attributed to the published material
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Described well enough to run: the authors ship a complete, MIT-licensed Snakemake pipeline (DMN-seq @4d8f1d3) with pinned environment.yml and example data; their pipeline ran end-to-end on «our HPC». The central single-base-resolution mechanism claim (DME nicks DNA at 5mC) reproduces cleanly and quantitatively: on lambda-DNA the consistently methylated cytosines are the Dcm CCWGG motifs, and read-start nick sites are 15.9x enriched at those cytosines in DME-treated vs input, with input sitting exactly at the random background (142/48502=0.29%) and the detected-site consensus motif = CCWGG. This is partial rather than full 1:1 because the RAW GSM reads are NOT publicly released (BioProject PRJNA1336081 has no SRA links; only processed *.spikein.sites.bed.gz are public), so the Dcm/motif numbers were derived from the authors' DEPOSITED lambda sites; the repo's example FASTQs are a minimal generic subset (only ~0.07% map to lambda) and cannot regenerate the lambda statistics from raw. NOT attempted (80/20): mESC >97%-CpG (needs mm10 + GSE312186 raw reads), CRC/cfDNA hypomethylation + 2,319 DMRs + 0.1ng sensitivity (human GSE312185, restricted; DMR-calling code not in repo), replicate Pearson R=0.99 (raw replicate reads unavailable). No fabrication concern: deposited data robustly supports the reproduced mechanism claim.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 80assessed: 2026-06-14 ⛓ 048a5de2e38c
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-14
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusCan the 5mC-specific bifunctional DNA glycosylase DEMETER (DME), which nicks DNA at 5mC sites, be harnessed to detect 5mC at single-base resolution and, conversely, to deplete hypermethylated regions so as to enrich and profile hypomethylated DNA regions for cancer biomarker discovery, including in low-input cell-free DNA?
- ★ DME-mediated nicking enables DMN-seq (DMN+) to detect 5mC at single-base resolution by ligating adaptors only to 5mC-containing fragments generated by DME excision method
- ★ DME specifically excises 5mC but shows no reactivity toward 5hmC, allowing DMN-seq to distinguish 5mC from 5hmC finding
- ★ A modified design (DMN–) depletes hypermethylated regions and enriches hypomethylated DNA regions, including promoters/enhancers and open chromatin method
- ★ DMN– robustly maps hypomethylated regions and tumor-specific motifs in colorectal cancer, identifying functionally relevant biomarker candidates resource
- ★ DMN– can be applied to cell-free DNA with inputs as low as 0.1 ng with high sensitivity and reproducibility finding
- DMN– shows higher mapping ratio and greater specificity for unmethylated regions than MRE-seq and UBS-seq in repetitive regions and LINE-1 elements finding
- DME generates single-strand breaks specific to 5mC locations and proportional to 5mC density within a methylated region by cleaving only one strand of symmetrically methylated dsDNA mechanism
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| in vitro DNA excision/cleavage assay | synthetic 82 bp dsDNA oligo with single 5mC site and identical oligo with single 5hmC site | DME enzyme titration (5mC vs 5hmC substrate) | DNA excision efficiency at the modified site | — |
| DMN+ sequencing (DMN-seq, methylation enrichment) | partially methylated λ-DNA with synthetic dsDNA spike-in oligos | DME excision | reads starting at 5mC sites (enrichment intensity), DCM motif methylation | Illumina NGS |
| UBS-seq (Ultrafast Bisulfite Sequencing) | partially methylated λ-DNA | bisulfite/none (reference standard) | methylated sites and methylation levels | — |
| DMN+ sequencing | mouse embryonic stem cell (mESC) genomic DNA with λ-DNA and 200-bp synthetic oligo internal controls | DME excision | 5mC site detection, CpG/CHG/CHH motif distribution, replicate correlation | Illumina NGS (~10 million reads shallow) |
| DMN– sequencing (hypomethylation enrichment) | mESC genomic DNA with internal controls | DME treatment vs untreated/input | enrichment of hypomethylated regions around TSS/TES, metagene profiles | Illumina NGS |
| MRE-seq (methylation-sensitive restriction enzyme digestion sequencing) | mESC genomic DNA repetitive regions / LINE-1 | restriction enzyme digestion | mapping ratio, unmethylated CpG enrichment vs 5mC level | — |
| DMN– sequencing | genomic DNA from 6 paired primary CRC tumor and adjacent normal colonic mucosa (UCMC patients) | DME treatment vs input | differentially methylated/hypomethylated regions flanking TSS, HOMER motif enrichment | Illumina NGS (avg 55.2 million reads/sample) |
| DMN– sequencing on cell-free DNA | clinical cfDNA (low input) | DME treatment | hypomethylated region detection sensitivity and reproducibility | — |
- – DME excised the synthetic 5mC dsDNA in a dose-dependent manner but showed no reactivity toward 5hmC at any enzyme level
- ▲ DMN+ enrichment intensity correlated highly between two technical replicates in partially methylated λ-DNA R=0.99
- ▲ DMN+ accurately identified methylated DCM sites across methylation levels, correlating with UBS-seq
- – More than 97% of 5mC sites detected by DMN+ in mESC gDNA were CpG motifs, comparable to BS-seq >97%
- ▲ DMN– showed methylation enrichment at TSS ~twofold higher than at TES region, consistent with UBS-seq ~2-fold
- ▲ DMN– signals overlapped strongly with open chromatin (ATAC-seq, DNaseI-seq) and enhancer/promoter histone marks (H3K27ac, H3K4me3)
- ▲ DMN– detected hypomethylated regions enriched in CRC tumor vs healthy tissue and CRC-specific TF motifs (Nrf1, ELF3, EGR1, IKZF2)
- ▲ DMN– had higher mapping ratio and steeper decline of enrichment with increasing 5mC than MRE-seq and UBS-seq in repetitive/LINE-1 regions
- correlation R=0.99 (Pearson correlation of DMN+ enrichment intensity between two technical replicates in partially methylated λ-DNA)
- count >97% (Fraction of DMN+ detected 5mC sites that are CpG motifs in mESC gDNA (coverage threshold ≥5))
- fold_change ~2-fold (DMN– methylation enrichment at TSS relative to TES region)
- count 55.2 million reads per sample (Average sequencing depth for DMN– applied to 6 paired CRC/healthy tissue samples)
- count ~10 million reads (Shallow sequencing depth for DMN+ on mESC gDNA with internal controls)
- count 0.1 ng (Lowest cfDNA input for DMN– with high sensitivity and reproducibility)
- count N=3937 (Number of methylation sites used for consensus motif analysis in mESC DMN-seq)
- count 82 bp (Length of synthetic dsDNA oligos used to test DME activity on 5mC and 5hmC)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
DMN-seq is validated through controlled experiments using synthetic oligonucleotides, partially methylated λ-DNA, and mouse embryonic stem cell genomic DNA, with reproducibility assessed primarily by Pearson correlation between technical replicates. For clinical proof-of-concept, DMN– was applied to paired colorectal cancer (CRC) tumor and adjacent healthy tissue from 6 patients; differentially hypomethylated 2000 bp genomic windows were identified and visualized as scatter and volcano plots, and sequence motif enrichment in those regions was performed using HOMER. Methodological benchmarking against UBS-seq, RRBS, and MRE-seq was presented as metagene profiles and mapping ratios rather than formal statistical tests.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Pearson correlation | Comparison of enrichment intensity between two technical replicates of partially methylated λ-DNA (Fig. 1f) | All CpG/DCM/non-CpG sites detected in λ-DNA; exact N not stated | not stated |
| Differential enrichment analysis for hypomethylated regions (specific statistical test not named) | Identification of 2000 bp segments significantly enriched in CRC vs healthy tissue (Fig. 3a scatter plot) | 6 paired CRC tumor and adjacent normal tissue samples | not stated |
| Differential methylation region analysis visualized as volcano plot (specific test not named) | DMRs within 2000 bp upstream of TSS in paired male samples (Fig. 3e; full description truncated in provided text) | Subset of male patients from the 6-pair cohort; exact n not stated | not stated |
| HOMER de novo and known motif enrichment analysis | Sequence motifs enriched in hypomethylated regions of CRC tumor vs healthy tissues (Fig. 3c, d) | Not stated | not stated |
-
The statistical test underlying the identification of 'significantly enriched' 2000 bp hypomethylated windows in CRC vs healthy tissue (Fig. 3a) is not named in the text↳ Could also: Established count-based differential analysis tools such as DESeq2, edgeR, or limma-voom applied to per-window read counts would also identify differentially enriched regions — Named tools provide a fully transparent, reproducible statistical framework with well-characterized dispersion modeling and built-in multiple-testing correction (e.g., Benjamini-Hochberg FDR), which facilitates cross-study comparisons and reviewer scrutiny
-
No multiple-testing correction method is stated for the genome-wide DMR discovery across thousands of 2000 bp windows↳ Could also: Benjamini-Hochberg false discovery rate (FDR) control applied across all tested genomic windows is the standard approach for genome-wide epigenomic differential analyses — Genome-wide analyses test thousands to millions of windows simultaneously; explicit FDR control is a widely expected convention in epigenomics that limits the expected proportion of spurious discoveries among reported findings
-
Technical reproducibility was assessed using Pearson correlation (R = 0.99) between two replicates of enrichment intensity at λ-DNA methylation sites↳ Could also: Spearman rank correlation or intraclass correlation coefficient (ICC) could also quantify replicate agreement — Spearman correlation is robust to non-normal distributions and outliers common in count-based enrichment data; ICC additionally captures absolute agreement (not just rank order), a distinction often preferred in method-validation contexts
-
The paired clinical design (tumor vs adjacent healthy tissue from the same patient, n = 6 pairs) is described but no power analysis or sample size justification is provided↳ Could also: A prospective or post-hoc power analysis based on effect sizes from published CRC methylation datasets could also accompany the cohort description — A power analysis contextualizes the study's sensitivity to detect DMRs at a given effect size and supports interpretation of negative or marginal findings, which is particularly informative for a proof-of-concept clinical cohort
-
Methodological benchmarking of DMN– against MRE-seq and UBS-seq was presented as metagene enrichment profiles and mapping ratios (visual comparison)↳ Could also: Quantitative performance metrics such as area under the receiver operating characteristic curve (AUROC), precision-recall curves, or F1-score computed against a gold-standard methylation call set could also be reported — Single summary metrics allow direct numerical comparisons between methods across thresholds and enable readers to assess sensitivity-specificity trade-offs in a way that visual profile overlaps alone do not
-
The paired structure of the clinical cohort (same patient as the unit of pairing) is described, but the statistical model used to exploit this pairing in the DMR analysis is not stated↳ Could also: A paired Wilcoxon signed-rank test, paired t-test, or a linear mixed-effects model with patient as a random effect could also be applied to explicitly leverage the paired design — Formally modeling the paired structure removes between-patient variability from the error term, increasing power to detect true tumor-vs-normal methylation differences—particularly valuable when the cohort is small (n = 6 pairs)
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Reproduction scope — pmid-41673887 (DMN-seq)
Paper: Wang Y, Li Y, Ye C, et al. DMN-seq enriches DNA hypomethylated regions for biomarker discovery using 5-methylcytosine glycosylase. Genome Biol 2026. PMID 41673887 · PMCID PMC13097799 · DOI 10.1186/s13059-026-03991-6. Code: https://github.com/yangli04/DMN-seq (MIT, commit 4d8f1d3, 2026-02-15) — authors' own Snakemake pipeline. Data: GEO GSE309434 (+ GSE312185, GSE312186).
Pipelines in the repo (both Snakemake)
5mC_detection/— cutadapt → UMI trim → AAA-start filter → HISAT2 (spike-in / hg38 / mm10) →count_sites.py(per-read 5′ "signal" sites) → collapse to TSV.hypomethylation_enrichment/— cutadapt → HISAT2 (spike-in / hg38) → Picard dedup → featureCounts over TSS±2 kb → deepTools heatmaps.
Data inventory (what is actually obtainable)
| accession | content | organism | raw reads (SRA) | processed deposited |
|---|---|---|---|---|
| GSE309434 | DMN+ on partially-methylated λ-DNA + synthetic spike-ins; 8 DMN+ libs + 1 UBS-seq | Lambda phage | NOT in SRA (BioProject PRJNA1336081 has no SRA links) | per-sample *.spikein.sites.bed.gz (e.g. c_A1 = 10.4 M λ reads) ✔ |
| GSE312186 | mESC gDNA + 0.1% λ spike-in | Mus musculus | not resolved | *.spikein.sites.bed.gz (p_A0/1/2) |
| GSE312185 | CRC tissue + cfDNA (1/0.5/0.1 ng) | Human | not resolved (human, likely restricted) | — |
Raw FASTQs for the specific GSMs are not publicly released (no SRA records).
The repo ships small example λ-DNA FASTQs (5mC_detection/raw_data/DME_*,
Input_*) that correspond to the DMN+ λ design → used as the run input.
λ-DNA reference (GenBank J02459.1, 48,502 bp) is public and tiny.
IN SCOPE (attempted) — 5mC_detection on λ-DNA (spikein path)
Run the authors' 5mC_detection Snakemake pipeline on the shipped example λ data,
mapping to λ-DNA J02459.1, producing per-read 5mC "signal" sites (count_sites).
Reproduced claims compared against the authors' own deposited sites + paper text:
- R1 concordance — reproduced λ signal positions match the authors' deposited
GSM9267342_c_A1.spikein.sites.bed.gz(λ subset). Pipeline-derived. - R2 DCM motif (Fig 1c / "All 5mC sites detected in λ-DNA also presented strong DCM motifs"): fraction of detected λ sites at a CCWGG (Dcm) motif. Computed on both our run and the authors' deposited bed.
- R3 Fig 1d example — 5mC sites at λ positions 43,404 and 43,406 detected.
- R4 enrichment — DME-treated vs Input enrichment of read-starts at Dcm sites.
OUT OF SCOPE (not attempted — reason)
- mESC ">97 % of detected 5mC sites are CpG (cov≥5)" (Fig 2) — needs full mESC reads (GSE312186) + mm10 HISAT2 index (heavy build); raw reads not resolved. Skipped (80/20).
- CRC / cfDNA hypomethylation, 2,319 DMRs, 55.2 M / 17.3 M reads, 0.1 ng
sensitivity (Figs 3–5) — human data (GSE312185), not publicly obtainable
(restricted); DMR-calling code not in repo.
data_restricted/no_codefor that part. - Replicate Pearson R = 0.99 — needs multiple full replicate libraries (raw reads unavailable); example data is one DME + one Input. Not attempted.
- Wet-lab claims (protein purification, excision efficiency gels) — non-pipeline.
Expected honest outcome: partial — pipeline mechanics + λ DCM-motif claim reproducible on the authors' code+example data and checkable against deposited outputs; the headline biological numbers depend on raw reads that the authors did not release to SRA, so they are not 1:1 reproducible.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
The foundational DME-nicks-5mC mechanism (Fig 1) reproduces cleanly and quantitatively from the authors' deposited GSE309434 sites: treated read-starts hit Dcm CCWGG cytosines at 4.66% vs input 0.29% (= 15.9x), input sits exactly at the 142/48502=0.29% random background, and the detected-site consensus is CCWGG — strong internal validation with no fabrication concern. It is partial only because the raw GSM reads were never released (no SRA links) and the repo example FASTQs map ~0.07% to lambda, so a from-raw 1:1 and the downstream mESC/CRC/cfDNA/DMR claims could not be run. The limitation is therefore data-availability (authors'/repository side), not a derivability or correctness defect in the claim that was tested.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.