Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

RNA modifications detection by comparative Nanopore direct RNA sequencing.

Nat Commun · 2021
L1 50/100 3/4
Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
✓ What held up
  • Nothing in this column.
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🟡Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
50/100
Reproducibility score
1.4 SD below mean
vs. all fields · 1173 studies
🎯 Scores higher than 8% of all assessed papers rank 1026 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

REPRODUCED (qualitative) the central synthetic-oligo benchmark of Nanocompore (Leger 2021, Fig. 2): running the authors' own pipeline (minimap2 2.17 -> f5c eventalign -> nanocompore 1.0.4 eventalign_collapse -> sampcomp) on PRJEB44511 Oligo_1 (3x m6A) vs unmodified control, Nanocompore calls all 3 engineered m6A sites significant (p<0.01) as contiguous kmer clusters, INCLUDING the two non-DRACH contexts (CUAGC@39, CGACC@60) the paper highlights. KS_intensity is the discriminating statistic; GMM-logit is weak with a single replicate. Grade is 'partial' only because the paper exposes no numeric value for these oligos to match 1:1 -- the qualitative detection claim reproduces fully (3/3). A non-obvious blocker was diagnosed and fixed: nanocompore's Whitelist requires ref length STRICTLY > min_ref_length (default 100), and the backbone is exactly 100 nt, so the reference was silently dropped (empty DB) until --min_ref_length 50 was passed. NOT attempted: yeast IME4 (very large FAST5); human METTL3-KD/7SK (FASTQ-only, no signal, honestly out of scope).

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 50
    assessed: 2026-06-18 ⛓ 06e0ccf26342
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-21
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator headless) · v1.0 · run #1 2026-06-19
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The paper tests whether a model-free, comparative analytical framework (Nanocompore) applied to Oxford Nanopore direct RNA sequencing signal data (current intensity and dwell time) can accurately detect and localize RNA post-transcriptional modifications without requiring a trained model, by comparing a modification-containing sample to a control sample with fewer or no modifications.

Core claims
  • Nanocompore is a model-free comparative method that uses a 2-component Gaussian mixture model (GMM) and univariate statistical tests on signal intensity/dwell time to detect RNA modifications in Nanopore direct RNA sequencing data without needing a training set method
  • Nanocompore detects multiple distinct RNA modification types (m6A, I, m5C, Ψ, m6,2A, m1G, 2'-OMeA) in synthetic oligonucleotides with positional accuracy finding
  • The GMM-logit test gives the best balance of specificity and sensitivity (highest F1 score) among Nanocompore's tests, particularly at high read coverage finding
  • At 512x coverage, the GMM-logit test achieved a mean accuracy of 94.48% for detecting m6A and 89.8% for detecting other modifications finding
  • Nanocompore detected known m7G and pseudouridine modification sites in E. coli 16S rRNA methyltransferase knockout controls with extremely high statistical significance finding
  • In yeast, Nanocompore identified m6A-associated peaks enriched near stop codons and within DRACH motifs, with partial overlap to an orthogonal reference set of known m6A sites finding
  • Metacompore, a Snakemake pipeline comparing six modification-detection tools, showed Eligos2 had higher sensitivity than Nanocompore's GMM methods for m6A detection in yeast resource
Experimental setups
Assay System Perturbation Readout Platform
Nanopore direct RNA sequencing (in silico simulation) simulated data none detection of simulated shifts in current intensity/dwell time
Nanopore direct RNA sequencing synthetic oligonucleotides (3 oligos, 100nt) chemical incorporation of modifications (m6A, I, m5C, Ψ, m6,2A, m1G, 2'-OMeA) Nanocompore p-values and peak calls at modified positions Oxford Nanopore R9 pores
Nanopore direct RNA sequencing Escherichia coli strain MRE600, 16S rRNA knock-out of RsmG or RsuA detection of m7G at position G527 and Ψ at position 516 Oxford Nanopore
Nanopore direct RNA sequencing (polyA+ transcriptome) Saccharomyces cerevisiae, WT vs ime4Δ knock-out of IME4 (m6A methyltransferase) significant kmers/peaks indicating m6A sites; enrichment near stop codon and DRACH motif Oxford Nanopore; basecalled with Guppy, resquiggled with Nanopolish eventalign
Computational benchmarking (Metacompore pipeline) Saccharomyces cerevisiae DRS dataset vs orthogonal m6A reference set none sensitivity, specificity, precision of 6 modification-detection algorithms (Nanocompore, Tombo, Eligos, Diff_err, Epinano, MINES) Snakemake pipeline
Key results
  • Nanocompore detected all tested modifications (m6A in 3 contexts, I, m5C, Ψ, m6,2A, m1G, 2'-OMeA) in synthetic oligos
  • m1G was detected as a significant signal one kmer downstream rather than at the modified kmer itself
  • GMM-logit test had lower sensitivity but higher specificity than KS tests on intensity or dwell time, with best F1 score above 512x coverage
  • Mean accuracy at 512x coverage and p=0.05 threshold using GMM-logit test 94.48% (m6A), 89.8% (other modifications)
  • At 4096x coverage, 75% of m6A sites detected when only 20% of reads were modified 75%
  • Nanocompore detected m7G (RsmG KO) and Ψ (RsuA KO) sites in E. coli 16S rRNA as highly significant p-value<10^-300
  • Nanocompore analysis of yeast WT vs ime4Δ identified significant kmers and peaks, refined by peak calling 15,961 significant kmers in 1,510 transcripts; 10,217 peaks
  • Eligos2 had higher sensitivity for m6A detection than Nanocompore's GMM methods in yeast benchmarking 45.8% (Eligos2) vs 16% (GMM) vs 5.5% (GMM context2)
Key statistics
  • pvalue <10^-300 (Nanocompore detection of m7G (G527) and Ψ (516) in E. coli 16S rRNA knockout controls)
  • mean 94.48% accuracy (m6A), 89.8% accuracy (other modifications) (GMM-logit test accuracy at 512x coverage, p-value cutoff 0.05, synthetic oligonucleotides)
  • count 14,554,547 reads total; coverage above 30x for 2,523 genes (40% of annotated transcriptome) (Yeast WT and ime4Δ DRS sequencing, three biological replicates per condition)
  • count 15,961 significant kmers in 1,510 distinct transcripts (FDR 1%) (Nanocompore analysis of yeast m6A detection)
  • count 10,217 peaks, median 3 peaks per transcript (Peak calling refinement of Nanocompore yeast kmer results)
  • other 21% (124/602) of known m6A sites overlapped Nanocompore peaks; 8% (124/1549) of Nanocompore peaks supported by orthogonal reference (Validation of yeast Nanocompore m6A predictions against orthogonal reference set)
  • other Sensitivity: Eligos2 45.8%, Nanocompore GMM 16%, Nanocompore GMM context2 5.5% (Metacompore benchmarking of 6 tools, nominal FDR threshold 1%, log odds ratio threshold 0.5, all kmers considered)
  • count 80,000 total simulation runs (100 datasets per combination of 3 factors) (In silico simulation of modification stoichiometry and knock-down efficiency across coverage levels)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

Nanocompore identifies RNA modification sites by comparing Oxford Nanopore direct-RNA sequencing signal data (median current intensity and log10 dwell time) from a modified sample against a non-modified control, position by position across transcripts. The primary statistical test is a 2-component Gaussian mixture model followed by logistic regression (GMM-logit); Kolmogorov–Smirnov (KS) tests on intensity or dwell time are offered as univariate alternatives. P-values from adjacent kmers can be combined using Hou's method, and all position-level p-values are corrected for multiple comparisons via Benjamini–Hochberg FDR. Method performance was characterised via ROC curves, F1 scores, and sensitivity/specificity on 100 subsampled and in silico mixed datasets spanning a factorial grid of coverage, modification stoichiometry, and knockdown efficiency, as well as against known m6A sites in yeast.

Replicationbiological Sample sizeThree biological replicates per condition for yeast in vivo experiment (WT vs ime4Δ), each sequenced in an individual flowcell; one flowcell per oligonucleotide for in vitro experiments; 100 independently generated datasets per factorial condition combination (coverage × stoichiometry × knockdown efficiency) for simulation benchmarks GroupsModified experimental RNA vs non-modified control RNA (KO/KD of modification enzyme, or unmodified synthetic/in vitro-transcribed RNA) Pairingunpaired Randomization/blindingnot stated DispersionCI Exact p-valuesyes Effect sizesno Confidence intervalsyes Multiplicity correctionBenjamini–Hochberg FDR
Statistical tests used
Test Applied to n Assumptions
2-component Gaussian mixture model (GMM) followed by logistic regression (GMM-logit) Primary method for identifying modified positions transcriptome-wide; used for all in vivo analyses (yeast, human, E. coli rRNA re-analysis) and main in vitro benchmarks Average 648,543.5 reads per oligonucleotide after quality filtering (in vitro); 14,554,547 total reads with >30x coverage for 2,523 genes across 6 flowcells (yeast in vivo, 3 biological replicates per condition); 32–4096 reads per subsampled dataset in simulation benchmarks not stated
Kolmogorov–Smirnov (KS) test on median current intensity Univariate pairwise test on per-position current intensity; benchmarked against GMM-logit in ROC, F1, TPR, and FPR analyses (Figs. 2B–E, 3B) Same subsampled datasets as GMM-logit benchmarks (32–4096 reads, n = 100 datasets per condition combination) not stated
Kolmogorov–Smirnov (KS) test on log10(dwell time) Univariate pairwise test on per-position dwell time; benchmarked alongside KS-intensity and GMM-logit (Figs. 2B–E, 3B) Same subsampled datasets as GMM-logit benchmarks not stated
Hou's method for combining non-independent p-values of neighbouring kmers Applied to aggregate evidence across proximal kmer contexts to account for modification-induced signal spread to adjacent positions; offered as an option within Nanocompore not stated
Mann–Whitney (MW) test and t-tests Listed as additional supported univariate tests in Nanocompore (Fig. 1B caption); not the focus of any reported analyses in the text not stated
Approaches that could also have been used
  • The paper uses a 2-component Gaussian mixture model followed by logistic regression (GMM-logit) as the primary bivariate test for comparing read distributions between conditions
    Could also: A multivariate permutation test on the joint (intensity, dwell time) distribution, or a likelihood ratio test comparing fitted mixture models, could also be used — Permutation tests make no parametric distributional assumptions and remain valid when the two-component Gaussian structure is an imperfect fit to the empirical signal; a likelihood ratio test offers a principled framework for comparing nested mixture models without requiring a separate logistic regression step
  • Univariate KS tests on intensity and on dwell time are run as separate tests, and the bivariate GMM-logit combines both dimensions through clustering then regression
    Could also: A single multivariate test such as Hotelling's T² or a two-sample energy statistic applied jointly to intensity and dwell time could also be used from the outset — A jointly multivariate test accounts for covariance between intensity and dwell time without first requiring a clustering step, potentially offering a more direct and assumption-free comparison of the bivariate distributions
  • P-values from adjacent kmers are combined using Hou's method for correlated tests
    Could also: Fisher's combined probability test, the Lancaster method for weighted dependent tests, or Brown's method for correlated p-values could also be applied to aggregate kmer-level evidence — These are established alternatives for combining dependent p-values and may perform differently depending on the actual correlation structure of neighbouring kmer signals; reporting sensitivity to the choice of combination method would allow readers to assess robustness
  • Benchmarking used 100 independently generated subsampled or in silico mixed datasets per factorial condition combination
    Could also: Stratified cross-validation or bootstrap resampling on a single full-size dataset could also estimate method performance and its variance — Bootstrap or k-fold cross-validation approaches provide variance estimates without requiring repeated data simulation and are standard in bioinformatics tool evaluation; they may also better reflect variance under real experimental conditions
  • Performance is summarised using sensitivity, specificity, precision, F1 score, and ROC curves
    Could also: Matthews Correlation Coefficient (MCC) or area under the precision–recall curve (AUPRC) could also be reported as complementary summary metrics — MCC and AUPRC are considered more informative when the positive class (modified sites) is rare relative to all tested positions, a common scenario in transcriptome-wide modification detection; they are less sensitive to class imbalance than F1 or ROC-AUC
  • Each position is summarised by the median current intensity and log10(dwell time) aggregated across reads before statistical testing
    Could also: Additional distributional features per position (e.g., signal variance, interquartile range, or quantiles of the intensity distribution) could also be extracted and incorporated as additional dimensions in the GMM or as further univariate test statistics — Richer per-position feature vectors may capture modification-induced changes in signal shape that are not fully reflected in the median, potentially improving sensitivity for modifications with heterogeneous or tail-heavy effects on pore current
Software: Nanocompore · Nextflow (nanocompore_pipeline) · Nanopolish eventalign / NanopolishComp Eventalign_collapse · Guppy (basecalling) · Samtools · Snakemake (Metacompore pipeline)

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Figures / tables: Fig. 2
oligo_m6A_all
Reported
Fig. 2: Nanocompore detects the engineered m6A modifications on the synthetic curlcake oligo, including non-DRACH contexts
Reproduced
3/3 engineered m6A sites called significant (p<0.01) as contiguous kmer clusters: GGACU@A18, CUAGC@A39, CGACC@A60
partial
oligo_m6A_GGACU
Reported
m6A at DRACH site GGACU detected (Fig. 2)
Reproduced
KS_intensity p=1.1e-12 (kmer16); context_2 p=4.4e-6; cluster pos 14-18
partial
oligo_m6A_CUAGC
Reported
m6A at non-DRACH site CUAGC detected (Fig. 2)
Reproduced
KS_intensity context_2 p=1.3e-3; cluster pos 36-38
partial
oligo_m6A_CGACC
Reported
m6A at non-DRACH site CGACC detected (Fig. 2)
Reproduced
KS_intensity p=4.2e-22 (kmer59); context_2 p=2.1e-10; cluster pos 57-60
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator headless) · v1.0 L1 50/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟡2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

714.1 k
tokens (I/O) · 55.9 M incl. cache
134 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.