Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Systematic benchmarking of tools for CpG methylation detection from nanopore sequencing.

Nat Commun · 2021
L1 99/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Every question reproduced
-1 pts
From: “every question reproduced”
Total score -7
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
  • No relevant deviation in data/preprocessing
  • No authors-side cause for any deviation
  • Reported values are derivable from the shared data
  • Any deviation was negligible
  • The central claim held under reproduction
  • Overall, the reproduction was clean
What did not (or only partly)
  • Every checked point held up.
How its reproducibility compares
99/100
Reproducibility score
1.4 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 95% of all assessed papers rank 55 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

REPRODUCED (deposited-data recomputation). Recomputed the paper's central tool-benchmark statistics directly from the authors' deposited Supplementary Data tables and compared against the exact printed figure-panel values. Fig 2a (mixture dataset 1, per-site Pearson r / r2 / RMSE for all 6 tools): 18/18 numbers EXACT. Fig 3a ROC-AUC and Fig 3b PR-AUC (5 tools, per-read): 9/10 exact, Megalodon ROC-AUC 0.977 vs 0.978 (within-tol, trapezoid integration of deposited ROC points). Fig 3e (mixture dataset 2, per-site, 4 tools + METEORE RF + REG): all exact except METEORE REG within 0.0003. The paper's headline claim that the METEORE RF consensus (Megalodon+DeepSignal) achieves lower RMSE (0.0687) than every individual tool (best Megalodon 0.0773) is confirmed. This is strong anti-fabrication evidence: the figures are faithfully derivable from the deposited data with no discrepancy. SCOPE/LIMITS: this verifies the deposited derived tables reproduce the figures and that METEORE's reported consensus values are internally consistent; it does NOT independently re-run the six methylation callers from raw fast5 (Fig 3c per-read 10-fold-CV AUCs and the end-to-end basecalling are NOT attempted — they require GPU/heavy compute and per-read score tables not deposited as small files). Figs 1 (schematic), 4-5 (WGBS recapitulation on NA12878) and the nCATS analyses were not attempted. INFRA NOTE: «our HPC» was fully storage-blocked during this run («infra» user quota exhausted, home quota exhausted, /tmp 50M at 100%), so no SLURM job could write anywhere; because the reproducible core here is small-data downstream statistics (<10 MB inputs, pandas/scipy in seconds, no compute node needed), it was run locally on «host» in a scratch dir and only small derived results are stored. Operator was notified of the quota outage.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 99
    assessed: 2026-06-18 ⛓ 3be5ce191b7c
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-18
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

How do existing computational tools for detecting CpG (5mC) DNA methylation from Nanopore sequencing compare in accuracy, and can a consensus approach combining multiple tools improve detection accuracy at the single-read and per-site levels?

Core claims
  • Nanopore methylation detection tools exhibit a tradeoff between false positives and false negatives and high dispersion relative to expected per-site methylation frequencies. finding
  • A consensus approach (METEORE) combining predictions from two or more tools improves accuracy over individual tools. method
  • Megalodon achieves the highest correlation and lowest RMSE among individual tools on control mixtures and the highest AUC/AUC-PR at the per-read level. finding
  • Varying single score cutoffs and discarding reads of uncertain methylation state improve prediction accuracy over default cutoffs. method
  • Nanopore methylation predictions recapitulate WGBS data, especially at sites of low and high methylation, with all tools overpredicting at intermediate methylation. finding
  • METEORE and Snakemake reproducibility pipelines are provided as an open resource at https://github.com/comprna/METEORE. resource
  • Combining methylation predictions from both strands at CpG sites improves correlation with WGBS over per-strand predictions. finding
Experimental setups
Assay System Perturbation Readout Platform
Nanopore sequencing methylation calling on control mixtures E. coli reference genome (PCR-amplified negative control and M.SssI-treated positive control DNA) M.SssI methyltransferase treatment vs PCR amplification; defined methylated/unmethylated read mixtures (0–100%) Per-site and per-read CpG methylation frequency/score Nanopolish, Megalodon, DeepSignal, Guppy, Tombo, DeepMod (fast5 input)
Cas9-targeted nanopore sequencing (nCATS) Human lymphoblastoid cell line NA12878 native nuclear DNA, ten forensically relevant regions Cas9-targeted enrichment; none (native DNA) CpG methylation frequency at targeted regions Oxford Nanopore MinION flowcell
Whole-genome bisulfite sequencing (WGBS) comparison Human NA12878 bisulfite conversion Per-CpG-site percentage methylation (reference standard) Illumina (ENCODE data)
Consensus model training/evaluation (METEORE RF and REG) Control mixture datasets 1 and 2 (E. coli-derived reads) none (computational combination of Megalodon + DeepSignal or all five tools) Combined per-read and per-site methylation prediction accuracy (AUC, RMSE, correlation) Random forest (max_depth=3, n_estimator=10) and multiple linear regression
Key results
  • All tools except Tombo achieved Pearson correlation above 0.8 for per-site methylation frequency on control mixture dataset 1 r>0.8 (p<2.2e-16)
  • All five per-read-capable tools achieved ROC AUC and PR AUC above 0.8, with Megalodon highest AUC and AUC-PR > 0.8
  • METEORE (REG) combining Megalodon and DeepSignal achieved the highest Pearson correlation and lowest RMSE vs WGBS r=0.9262, RMSE=0.1607
  • METEORE RF combining Megalodon and DeepSignal achieved lower RMSE than individual tools on mixture dataset 2
  • Applying optimized single score cutoffs gave all tools lower RMSE than default cutoffs on mixture dataset 2
  • All tools showed positive correlation with WGBS using both-strand combined methylation; Tombo and DeepMod lowest r ranging 0.7401 (DeepMod) to 0.9262 (METEORE REG)
  • Guppy systematically underpredicted per-site methylation and failed to predict any 100% methylated sites at >0.8 cutoff
  • Nanopolish and Tombo systematically overpredicted methylation with high dispersion
Key statistics
  • correlation r=0.9262, r2=0.8579, ρ=0.8885, RMSE=0.1607 (METEORE (REG) Megalodon+DeepSignal vs WGBS)
  • correlation r=0.9177, r2=0.8423, ρ=0.8765, RMSE=0.1708 (DeepSignal vs WGBS)
  • correlation r=0.9117, r2=0.8312, ρ=0.8801, RMSE=0.1772 (Megalodon vs WGBS)
  • correlation r=0.7401, r2=0.5477, ρ=0.7264, RMSE=0.2874 (DeepMod vs WGBS (lowest))
  • pvalue <2.2e-16 (Pearson correlations of per-site methylation for all tools on mixture dataset 1)
  • count 346,793 CpG sites; 100 sites selected; ~2400 reads per mixture set; 11 benchmarking datasets (E. coli control mixture design)
  • count median coverage 85×, minimum 50× (read coverage at selected control sites)
  • count 1743 methylation calls; 793 low (0.0–0.3), 264 intermediate (0.3–0.7), 686 high (0.7–1.0) (WGBS methylation bins for nanopore comparison)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This benchmarking study compared six nanopore CpG methylation detection tools (Nanopolish, Megalodon, DeepSignal, Guppy, Tombo, DeepMod) using controlled read mixtures spanning 0–100% methylation, per-read level analysis, and comparison with whole-genome bisulfite sequencing (WGBS). Accuracy was quantified primarily via Pearson and Spearman correlations, RMSE, and area under ROC and precision-recall curves across multiple datasets and coverage levels. A consensus ensemble approach (METEORE) combining tool outputs via random forest or linear regression was evaluated with tenfold cross-validation. Results were visualized using violin plots, boxplots, and ECDF curves, with p-values for correlation tests reported as inequality bounds.

Replicationmixed Sample size100 CpG sites selected from 346,793 in the E. coli reference genome; ~2400 reads per mixture dataset at minimum 50× coverage (median 85×); single human lymphoblastoid cell line (NA12878) used for WGBS comparison yielding 1661–1739 analyzable CpG sites GroupsSix methylation detection tools plus METEORE consensus (RF and REG variants) compared across 11 controlled methylation mixtures (0–100%) and against WGBS reference data Pairingna Randomization/blindingnot stated Dispersionmixed Exact p-valuesno Effect sizesno Confidence intervalsno Multiplicity correctionnone stated
Statistical tests used
Test Applied to n Assumptions
Pearson correlation (r) and coefficient of determination (r²) Per-site methylation frequency prediction vs. expected value across all mixture datasets and WGBS comparison (Figs 2a, 3e, Table 1) 100 CpG sites per mixture subset; 1661–1739 sites for WGBS comparison not stated
Root mean square error (RMSE) Per-site methylation frequency prediction accuracy across all tools and all datasets (Figs 2a, 3e, Table 1, Supplementary Figs 5–6) 100 CpG sites per mixture subset; 1661–1739 sites for WGBS comparison na
Area under the ROC curve (AUC) Per-read binary methylation classification accuracy for five tools on 0% and 100% methylated read sets (Fig 3a) Reads from 0% and 100% methylated sets (~2400 reads per set in mixture dataset 1) not stated
Area under the precision-recall curve (AUC-PR) Per-read precision-recall tradeoff for five tools on 0% and 100% methylated read sets (Fig 3b) Reads from 0% and 100% methylated sets (~2400 reads per set in mixture dataset 1) not stated
Tenfold cross-validation (random forest and multiple linear regression) METEORE consensus model evaluation for all two-tool and five-tool combinations (Fig 3c, 3d, Supplementary Fig 3) 100 CpG sites across 11 methylation mixtures in dataset 1 not stated
Spearman rank correlation (ρ) Nanopore vs. WGBS methylation frequency agreement for all tools (Table 1) 1661–1739 CpG sites not stated
Empirical cumulative distribution function (ECDF) Classification of fully unmethylated and fully methylated sites across varying score thresholds (Figs 2c, 2d) 100 CpG sites in the 0% and 100% methylation sets na
Approaches that could also have been used
  • Method agreement was assessed primarily with Pearson correlation and RMSE
    Could also: Bland-Altman analysis (limits of agreement) could also be used for method comparison against the WGBS reference — Bland-Altman plots visualize systematic bias and proportional error across the full measurement range, which is a standard complement to correlation in method-comparison studies; high Pearson r does not preclude meaningful fixed or proportional biases between the nanopore tool and the WGBS reference
  • Pearson correlation was the primary agreement metric, measuring linear association
    Could also: Lin's concordance correlation coefficient (CCC) could also be used for method agreement — Lin's CCC jointly quantifies precision (Pearson r) and accuracy (proximity to the identity line of perfect agreement) in a single index, which is well suited to benchmarking where one method serves as a reference standard and where predictions that are linearly related but offset would still yield high Pearson r
  • Multiple tools and metrics were compared across multiple datasets without a stated multiple testing correction
    Could also: Benjamini-Hochberg FDR or Bonferroni correction could also be applied to the family of correlation tests — When many simultaneous comparisons are made across tools and datasets, applying a multiplicity correction is one standard approach to controlling the expected false-positive rate across the comparison family; this would place bounds on the risk that any single reported significant result is a false positive
  • METEORE model performance was estimated with tenfold cross-validation on mixture dataset 1 and then applied to the separate mixture dataset 2
    Could also: Repeated k-fold cross-validation or bootstrap confidence intervals around AUC and RMSE could also be used — Repeating the cross-validation over multiple random splits, or using bootstrap resampling, yields variance estimates around the performance metrics, allowing formal quantification of uncertainty and more rigorous pairwise comparison of model performance differences
  • Per-site prediction accuracy was summarized with the proportion of sites falling outside a fixed 10% window around the expected methylation value
    Could also: Calibration curves (reliability diagrams) binning predicted vs. observed methylation proportions across the 0–1 range could also be used — Calibration analysis explicitly tests whether predicted methylation frequencies are systematically over- or under-estimated at each level of the scale, separating discrimination from calibration and making it easier to characterize the directional pattern of deviations across tools
  • The WGBS comparison was conducted using a single human cell line (NA12878)
    Could also: Validation across multiple cell lines or primary tissue samples with diverse methylation landscapes could also be incorporated — A single biological source constrains generalizability; replication across samples with different baseline methylation distributions would allow assessment of whether tool performance rankings are stable across varied genomic contexts
Software: Snakemake (pipeline orchestration) · METEORE (random forest and linear regression consensus; RF parameters max_depth=3, n_estimator=10 are consistent with scikit-learn but not explicitly named)

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — PMID 34103501 (METEORE)

Paper: Yuen, Jack, Eyras et al., Systematic benchmarking of tools for CpG methylation detection from nanopore sequencing, Nat Commun 2021;12:3438. DOI 10.1038/s41467-021-23778-6 · Code: https://github.com/comprna/METEORE (MIT).

What kind of paper

A benchmark/methods paper. It runs six existing nanopore 5mC-calling tools (Nanopolish, DeepSignal, Megalodon, Tombo, Guppy, DeepMod) on controlled methylation mixtures and on human data, quantifies their accuracy against ground truth (control mixtures + WGBS), and introduces METEORE, a consensus of ≥2 tools (random forest [RF] and ridge regression [REG]) that improves accuracy. All headline results are pipeline-derived (in scope per P16: applying these third-party tools + the authors' combiner to the data).

In scope (pipeline-derived) — and what was attempted

Result Pipeline Reproduced here?
Fig 2a per-site Pearson r / r² / RMSE vs expected, 6 tools, mixture dataset 1 tool freq calls → aggregate per site → cor/RMSE YES — from deposited Supp Data 5
Fig 3a per-read ROC AUC (5 tools) per-read scores → ROC YES — integrated from deposited Supp Data 8
Fig 3b per-read PR AUC (5 tools) per-read scores → PR YES — integrated from deposited Supp Data 9
Fig 3e per-site r/r²/RMSE, mixture dataset 2, 4 tools + METEORE RF & REG tool/METEORE freq → cor/RMSE YES — from deposited Supp Data 10
Headline: METEORE RF RMSE < every individual tool RF consensus YES (0.0687 < 0.0773)
Fig 3c/d METEORE RF ROC/PR AUC by 10-fold CV on per-read scores RF train + CV on per-read score matrix NOT attempted — needs per-read score tables for all tools on mixture set 1, which are NOT deposited as small files; requires re-running tools on raw fast5 (heavy/GPU)
Fig 2b/2c/2d, Supp Tables (cutoff tuning) thresholds not attempted (secondary; partly derivable from Supp 6/7)
Fig 4–5 nanopore vs WGBS recapitulation (NA12878) nanopore + WGBS correlation NOT attempted (additional analysis; heavy)
End-to-end basecalling + 6-tool methylation calling from raw fast5 Guppy/Megalodon (GPU), Nanopolish, DeepSignal, Tombo, DeepMod NOT attempted (heavy GPU compute; + infra outage, see below)

Out of scope

  • Fig 1 (schematic). Wet-lab: gRNA/RNP assembly, library prep, sequencing (Methods).

Reproduction level (important nuance)

This reproduction recomputes the reported summary statistics from the authors' deposited derived tables (Supplementary Data 5/8/9/10) and compares them to the exact figure-panel values. It therefore verifies (a) my statistics match theirs and (b) the figures are faithfully derivable from the deposited data — i.e. strong anti-fabrication evidence. It does not independently re-run the six methylation callers from raw fast5; that end-to-end re-run (and Fig 3c per-read CV) is the remaining harder tier, blocked here by a «our HPC» storage outage («infra» + home quotas exhausted, /tmp full) — see AUDIT.md. Because the reproduced core is small-data downstream statistics (<10 MB inputs, seconds of pandas/scipy, no compute node needed), it was run locally on «host»; only small derived results are stored.

Figures / tables: Fig 2aFig 3eFig 3aFig 3b
fig2a_megalodon_RMSE
Reported
RMSE=0.0758 (r=0.9860, r2=0.9723)
Reproduced
RMSE=0.0758 (r=0.9860, r2=0.9723)
exact
fig2a_nanopolish_r
Reported
r=0.8733, r2=0.7626, RMSE=0.1665
Reproduced
r=0.8733, r2=0.7626, RMSE=0.1665
exact
fig2a_deepsignal
Reported
r=0.9420, r2=0.8875, RMSE=0.1275
Reproduced
r=0.9420, r2=0.8875, RMSE=0.1275
exact
fig2a_tombo
Reported
r=0.7945, r2=0.6312, RMSE=0.2282
Reproduced
r=0.7945, r2=0.6312, RMSE=0.2282
exact
fig2a_guppy
Reported
r=0.8993, r2=0.8088, RMSE=0.3135
Reproduced
r=0.8993, r2=0.8088, RMSE=0.3135
exact
fig2a_deepmod
Reported
r=0.9467, r2=0.8962, RMSE=0.1235
Reproduced
r=0.9467, r2=0.8962, RMSE=0.1235
exact
fig3a_roc_auc_megalodon
Reported
AUC=0.978
Reproduced
AUC=0.977
within tolerance
fig3a_roc_auc_deepsignal
Reported
AUC=0.937
Reproduced
AUC=0.937
exact
fig3a_roc_auc_nanopolish
Reported
AUC=0.921
Reproduced
AUC=0.921
exact
fig3a_roc_auc_tombo
Reported
AUC=0.898
Reproduced
AUC=0.898
exact
fig3a_roc_auc_guppy
Reported
AUC=0.833
Reproduced
AUC=0.833
exact
fig3b_pr_auc_megalodon
Reported
AUC=0.984
Reproduced
AUC=0.984
exact
fig3b_pr_auc_deepsignal
Reported
AUC=0.940
Reproduced
AUC=0.940
exact
fig3b_pr_auc_nanopolish
Reported
AUC=0.924
Reproduced
AUC=0.924
exact
fig3b_pr_auc_tombo
Reported
AUC=0.887
Reproduced
AUC=0.887
exact
fig3b_pr_auc_guppy
Reported
AUC=0.885
Reproduced
AUC=0.885
exact
fig3e_meteore_rf_RMSE
Reported
RMSE=0.0687 (r=0.9806, r2=0.9616)
Reproduced
RMSE=0.0687 (r=0.9806, r2=0.9616)
exact
fig3e_meteore_reg_RMSE
Reported
RMSE=0.0953 (r=0.9796, r2=0.9596)
Reproduced
RMSE=0.0956 (r=0.9797, r2=0.9598)
within tolerance
fig3e_megalodon
Reported
r=0.9886, r2=0.9773, RMSE=0.0773
Reproduced
r=0.9886, r2=0.9773, RMSE=0.0773
exact
fig3e_claim_meteore_rf_lower_rmse_than_all_individual
Reported
METEORE RF (Megalodon+DeepSignal) RMSE 0.0687 < best individual Megalodon 0.0773 (mix dataset 2)
Reproduced
0.0687 < 0.0773 confirmed
exact

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 99/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟢3. Location of the main deviation
🟢4. Cause of the deviation
🟢5. Derivability / plausibility
🟢6. Severity of the deviation
🟢7. Core claim
🟢8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Every question reproduced
-1 pts
From: “every question reproduced”
Total score -7

This is a clean, essentially 1:1 reproduction: the central benchmark statistics for all six tools (Fig 2a) plus Fig 3a ROC-AUC, Fig 3b PR-AUC, and Fig 3e mixture-2 metrics were recomputed directly from the authors' deposited Supplementary Data tables and matched the printed values exactly (34/36 claims delta=0.0). The only deviations — Megalodon ROC-AUC 0.977 vs 0.978 and METEORE REG RMSE 0.0956 vs 0.0953 — are within tolerance and attributable to trapezoid integration/rounding, our side, not the authors'. The headline claim that the METEORE RF consensus (RMSE 0.0687) beats every individual tool (best Megalodon 0.0773) is confirmed. Scope limit (raw-fast5 caller re-runs and Figs 4–5 WGBS not attempted) is a coverage caveat, not a discrepancy; values are fully derivable, so no fabrication concern.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

161.5 k
tokens (I/O) · 10 M incl. cache
19 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.