Systematic benchmarking of tools for CpG methylation detection from nanopore sequencing.
The main results reproduced: recomputed values matched the published ones within tolerance.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Same input data as the authors
- ✓Reported values were directly comparable
- ✓No relevant deviation in data/preprocessing
- ✓No authors-side cause for any deviation
- ✓Reported values are derivable from the shared data
- ✓Any deviation was negligible
- ✓The central claim held under reproduction
- ✓Overall, the reproduction was clean
- Every checked point held up.
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
REPRODUCED (deposited-data recomputation). Recomputed the paper's central tool-benchmark statistics directly from the authors' deposited Supplementary Data tables and compared against the exact printed figure-panel values. Fig 2a (mixture dataset 1, per-site Pearson r / r2 / RMSE for all 6 tools): 18/18 numbers EXACT. Fig 3a ROC-AUC and Fig 3b PR-AUC (5 tools, per-read): 9/10 exact, Megalodon ROC-AUC 0.977 vs 0.978 (within-tol, trapezoid integration of deposited ROC points). Fig 3e (mixture dataset 2, per-site, 4 tools + METEORE RF + REG): all exact except METEORE REG within 0.0003. The paper's headline claim that the METEORE RF consensus (Megalodon+DeepSignal) achieves lower RMSE (0.0687) than every individual tool (best Megalodon 0.0773) is confirmed. This is strong anti-fabrication evidence: the figures are faithfully derivable from the deposited data with no discrepancy. SCOPE/LIMITS: this verifies the deposited derived tables reproduce the figures and that METEORE's reported consensus values are internally consistent; it does NOT independently re-run the six methylation callers from raw fast5 (Fig 3c per-read 10-fold-CV AUCs and the end-to-end basecalling are NOT attempted — they require GPU/heavy compute and per-read score tables not deposited as small files). Figs 1 (schematic), 4-5 (WGBS recapitulation on NA12878) and the nCATS analyses were not attempted. INFRA NOTE: «our HPC» was fully storage-blocked during this run («infra» user quota exhausted, home quota exhausted, /tmp 50M at 100%), so no SLURM job could write anywhere; because the reproducible core here is small-data downstream statistics (<10 MB inputs, pandas/scipy in seconds, no compute node needed), it was run locally on «host» in a scratch dir and only small derived results are stored. Operator was notified of the quota outage.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 99assessed: 2026-06-18 ⛓ 3be5ce191b7c
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-18
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusHow do existing computational tools for detecting CpG (5mC) DNA methylation from Nanopore sequencing compare in accuracy, and can a consensus approach combining multiple tools improve detection accuracy at the single-read and per-site levels?
- ★ Nanopore methylation detection tools exhibit a tradeoff between false positives and false negatives and high dispersion relative to expected per-site methylation frequencies. finding
- ★ A consensus approach (METEORE) combining predictions from two or more tools improves accuracy over individual tools. method
- ★ Megalodon achieves the highest correlation and lowest RMSE among individual tools on control mixtures and the highest AUC/AUC-PR at the per-read level. finding
- ★ Varying single score cutoffs and discarding reads of uncertain methylation state improve prediction accuracy over default cutoffs. method
- ★ Nanopore methylation predictions recapitulate WGBS data, especially at sites of low and high methylation, with all tools overpredicting at intermediate methylation. finding
- METEORE and Snakemake reproducibility pipelines are provided as an open resource at https://github.com/comprna/METEORE. resource
- Combining methylation predictions from both strands at CpG sites improves correlation with WGBS over per-strand predictions. finding
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Nanopore sequencing methylation calling on control mixtures | E. coli reference genome (PCR-amplified negative control and M.SssI-treated positive control DNA) | M.SssI methyltransferase treatment vs PCR amplification; defined methylated/unmethylated read mixtures (0–100%) | Per-site and per-read CpG methylation frequency/score | Nanopolish, Megalodon, DeepSignal, Guppy, Tombo, DeepMod (fast5 input) |
| Cas9-targeted nanopore sequencing (nCATS) | Human lymphoblastoid cell line NA12878 native nuclear DNA, ten forensically relevant regions | Cas9-targeted enrichment; none (native DNA) | CpG methylation frequency at targeted regions | Oxford Nanopore MinION flowcell |
| Whole-genome bisulfite sequencing (WGBS) comparison | Human NA12878 | bisulfite conversion | Per-CpG-site percentage methylation (reference standard) | Illumina (ENCODE data) |
| Consensus model training/evaluation (METEORE RF and REG) | Control mixture datasets 1 and 2 (E. coli-derived reads) | none (computational combination of Megalodon + DeepSignal or all five tools) | Combined per-read and per-site methylation prediction accuracy (AUC, RMSE, correlation) | Random forest (max_depth=3, n_estimator=10) and multiple linear regression |
- – All tools except Tombo achieved Pearson correlation above 0.8 for per-site methylation frequency on control mixture dataset 1 r>0.8 (p<2.2e-16)
- – All five per-read-capable tools achieved ROC AUC and PR AUC above 0.8, with Megalodon highest AUC and AUC-PR > 0.8
- ▲ METEORE (REG) combining Megalodon and DeepSignal achieved the highest Pearson correlation and lowest RMSE vs WGBS r=0.9262, RMSE=0.1607
- ▼ METEORE RF combining Megalodon and DeepSignal achieved lower RMSE than individual tools on mixture dataset 2
- ▼ Applying optimized single score cutoffs gave all tools lower RMSE than default cutoffs on mixture dataset 2
- ▲ All tools showed positive correlation with WGBS using both-strand combined methylation; Tombo and DeepMod lowest r ranging 0.7401 (DeepMod) to 0.9262 (METEORE REG)
- ▼ Guppy systematically underpredicted per-site methylation and failed to predict any 100% methylated sites at >0.8 cutoff
- ▲ Nanopolish and Tombo systematically overpredicted methylation with high dispersion
- correlation r=0.9262, r2=0.8579, ρ=0.8885, RMSE=0.1607 (METEORE (REG) Megalodon+DeepSignal vs WGBS)
- correlation r=0.9177, r2=0.8423, ρ=0.8765, RMSE=0.1708 (DeepSignal vs WGBS)
- correlation r=0.9117, r2=0.8312, ρ=0.8801, RMSE=0.1772 (Megalodon vs WGBS)
- correlation r=0.7401, r2=0.5477, ρ=0.7264, RMSE=0.2874 (DeepMod vs WGBS (lowest))
- pvalue <2.2e-16 (Pearson correlations of per-site methylation for all tools on mixture dataset 1)
- count 346,793 CpG sites; 100 sites selected; ~2400 reads per mixture set; 11 benchmarking datasets (E. coli control mixture design)
- count median coverage 85×, minimum 50× (read coverage at selected control sites)
- count 1743 methylation calls; 793 low (0.0–0.3), 264 intermediate (0.3–0.7), 686 high (0.7–1.0) (WGBS methylation bins for nanopore comparison)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This benchmarking study compared six nanopore CpG methylation detection tools (Nanopolish, Megalodon, DeepSignal, Guppy, Tombo, DeepMod) using controlled read mixtures spanning 0–100% methylation, per-read level analysis, and comparison with whole-genome bisulfite sequencing (WGBS). Accuracy was quantified primarily via Pearson and Spearman correlations, RMSE, and area under ROC and precision-recall curves across multiple datasets and coverage levels. A consensus ensemble approach (METEORE) combining tool outputs via random forest or linear regression was evaluated with tenfold cross-validation. Results were visualized using violin plots, boxplots, and ECDF curves, with p-values for correlation tests reported as inequality bounds.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Pearson correlation (r) and coefficient of determination (r²) | Per-site methylation frequency prediction vs. expected value across all mixture datasets and WGBS comparison (Figs 2a, 3e, Table 1) | 100 CpG sites per mixture subset; 1661–1739 sites for WGBS comparison | not stated |
| Root mean square error (RMSE) | Per-site methylation frequency prediction accuracy across all tools and all datasets (Figs 2a, 3e, Table 1, Supplementary Figs 5–6) | 100 CpG sites per mixture subset; 1661–1739 sites for WGBS comparison | na |
| Area under the ROC curve (AUC) | Per-read binary methylation classification accuracy for five tools on 0% and 100% methylated read sets (Fig 3a) | Reads from 0% and 100% methylated sets (~2400 reads per set in mixture dataset 1) | not stated |
| Area under the precision-recall curve (AUC-PR) | Per-read precision-recall tradeoff for five tools on 0% and 100% methylated read sets (Fig 3b) | Reads from 0% and 100% methylated sets (~2400 reads per set in mixture dataset 1) | not stated |
| Tenfold cross-validation (random forest and multiple linear regression) | METEORE consensus model evaluation for all two-tool and five-tool combinations (Fig 3c, 3d, Supplementary Fig 3) | 100 CpG sites across 11 methylation mixtures in dataset 1 | not stated |
| Spearman rank correlation (ρ) | Nanopore vs. WGBS methylation frequency agreement for all tools (Table 1) | 1661–1739 CpG sites | not stated |
| Empirical cumulative distribution function (ECDF) | Classification of fully unmethylated and fully methylated sites across varying score thresholds (Figs 2c, 2d) | 100 CpG sites in the 0% and 100% methylation sets | na |
-
Method agreement was assessed primarily with Pearson correlation and RMSE↳ Could also: Bland-Altman analysis (limits of agreement) could also be used for method comparison against the WGBS reference — Bland-Altman plots visualize systematic bias and proportional error across the full measurement range, which is a standard complement to correlation in method-comparison studies; high Pearson r does not preclude meaningful fixed or proportional biases between the nanopore tool and the WGBS reference
-
Pearson correlation was the primary agreement metric, measuring linear association↳ Could also: Lin's concordance correlation coefficient (CCC) could also be used for method agreement — Lin's CCC jointly quantifies precision (Pearson r) and accuracy (proximity to the identity line of perfect agreement) in a single index, which is well suited to benchmarking where one method serves as a reference standard and where predictions that are linearly related but offset would still yield high Pearson r
-
Multiple tools and metrics were compared across multiple datasets without a stated multiple testing correction↳ Could also: Benjamini-Hochberg FDR or Bonferroni correction could also be applied to the family of correlation tests — When many simultaneous comparisons are made across tools and datasets, applying a multiplicity correction is one standard approach to controlling the expected false-positive rate across the comparison family; this would place bounds on the risk that any single reported significant result is a false positive
-
METEORE model performance was estimated with tenfold cross-validation on mixture dataset 1 and then applied to the separate mixture dataset 2↳ Could also: Repeated k-fold cross-validation or bootstrap confidence intervals around AUC and RMSE could also be used — Repeating the cross-validation over multiple random splits, or using bootstrap resampling, yields variance estimates around the performance metrics, allowing formal quantification of uncertainty and more rigorous pairwise comparison of model performance differences
-
Per-site prediction accuracy was summarized with the proportion of sites falling outside a fixed 10% window around the expected methylation value↳ Could also: Calibration curves (reliability diagrams) binning predicted vs. observed methylation proportions across the 0–1 range could also be used — Calibration analysis explicitly tests whether predicted methylation frequencies are systematically over- or under-estimated at each level of the scale, separating discrimination from calibration and making it easier to characterize the directional pattern of deviations across tools
-
The WGBS comparison was conducted using a single human cell line (NA12878)↳ Could also: Validation across multiple cell lines or primary tissue samples with diverse methylation landscapes could also be incorporated — A single biological source constrains generalizability; replication across samples with different baseline methylation distributions would allow assessment of whether tool performance rankings are stable across varied genomic contexts
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — PMID 34103501 (METEORE)
Paper: Yuen, Jack, Eyras et al., Systematic benchmarking of tools for CpG methylation detection from nanopore sequencing, Nat Commun 2021;12:3438. DOI 10.1038/s41467-021-23778-6 · Code: https://github.com/comprna/METEORE (MIT).
What kind of paper
A benchmark/methods paper. It runs six existing nanopore 5mC-calling tools (Nanopolish, DeepSignal, Megalodon, Tombo, Guppy, DeepMod) on controlled methylation mixtures and on human data, quantifies their accuracy against ground truth (control mixtures + WGBS), and introduces METEORE, a consensus of ≥2 tools (random forest [RF] and ridge regression [REG]) that improves accuracy. All headline results are pipeline-derived (in scope per P16: applying these third-party tools + the authors' combiner to the data).
In scope (pipeline-derived) — and what was attempted
| Result | Pipeline | Reproduced here? |
|---|---|---|
| Fig 2a per-site Pearson r / r² / RMSE vs expected, 6 tools, mixture dataset 1 | tool freq calls → aggregate per site → cor/RMSE | YES — from deposited Supp Data 5 |
| Fig 3a per-read ROC AUC (5 tools) | per-read scores → ROC | YES — integrated from deposited Supp Data 8 |
| Fig 3b per-read PR AUC (5 tools) | per-read scores → PR | YES — integrated from deposited Supp Data 9 |
| Fig 3e per-site r/r²/RMSE, mixture dataset 2, 4 tools + METEORE RF & REG | tool/METEORE freq → cor/RMSE | YES — from deposited Supp Data 10 |
| Headline: METEORE RF RMSE < every individual tool | RF consensus | YES (0.0687 < 0.0773) |
| Fig 3c/d METEORE RF ROC/PR AUC by 10-fold CV on per-read scores | RF train + CV on per-read score matrix | NOT attempted — needs per-read score tables for all tools on mixture set 1, which are NOT deposited as small files; requires re-running tools on raw fast5 (heavy/GPU) |
| Fig 2b/2c/2d, Supp Tables (cutoff tuning) | thresholds | not attempted (secondary; partly derivable from Supp 6/7) |
| Fig 4–5 nanopore vs WGBS recapitulation (NA12878) | nanopore + WGBS correlation | NOT attempted (additional analysis; heavy) |
| End-to-end basecalling + 6-tool methylation calling from raw fast5 | Guppy/Megalodon (GPU), Nanopolish, DeepSignal, Tombo, DeepMod | NOT attempted (heavy GPU compute; + infra outage, see below) |
Out of scope
- Fig 1 (schematic). Wet-lab: gRNA/RNP assembly, library prep, sequencing (Methods).
Reproduction level (important nuance)
This reproduction recomputes the reported summary statistics from the authors' deposited derived tables (Supplementary Data 5/8/9/10) and compares them to the exact figure-panel values. It therefore verifies (a) my statistics match theirs and (b) the figures are faithfully derivable from the deposited data — i.e. strong anti-fabrication evidence. It does not independently re-run the six methylation callers from raw fast5; that end-to-end re-run (and Fig 3c per-read CV) is the remaining harder tier, blocked here by a «our HPC» storage outage («infra» + home quotas exhausted, /tmp full) — see AUDIT.md. Because the reproduced core is small-data downstream statistics (<10 MB inputs, seconds of pandas/scipy, no compute node needed), it was run locally on «host»; only small derived results are stored.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
This is a clean, essentially 1:1 reproduction: the central benchmark statistics for all six tools (Fig 2a) plus Fig 3a ROC-AUC, Fig 3b PR-AUC, and Fig 3e mixture-2 metrics were recomputed directly from the authors' deposited Supplementary Data tables and matched the printed values exactly (34/36 claims delta=0.0). The only deviations — Megalodon ROC-AUC 0.977 vs 0.978 and METEORE REG RMSE 0.0956 vs 0.0953 — are within tolerance and attributable to trapezoid integration/rounding, our side, not the authors'. The headline claim that the METEORE RF consensus (RMSE 0.0687) beats every individual tool (best Megalodon 0.0773) is confirmed. Scope limit (raw-fast5 caller re-runs and Figs 4–5 WGBS not attempted) is a coverage caveat, not a discrepancy; values are fully derivable, so no fabrication concern.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.