Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

RADAR: differential analysis of MeRIP-seq data with a random effect model.

Genome Biol · 2019
L1 10/100 3/4
⚑ Flagged for review — a reproduced result did not match the reported value

Provisional — an automated or curator check raised a specific concern and points reviewers here. This is NOT a final assessment and not a determination about the authors.

Why this verdict

The main result did not reproduce in this reproduction attempt. Where our recomputation produced values that differ from the published ones, those discrepancies are listed below. This is a single automated attempt — not peer review and not a finding of error or misconduct — and differences can also arise from data access, undocumented parameters or the computing environment. The verdict can be contested via “report an error”.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +7
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
What did not (or only partly)
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
10/100
Reproducibility score
3.6 SD below mean
vs. all fields · 1173 studies
🎯 Scores higher than 0% of all assessed papers rank 1169 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

RADAR's quick-start example workflow was reproduced end-to-end on «our HPC» («job», COMPLETED on node n096): RADAR @ b3c0cd7 compiled from source, conda R4.3.3 + Bioconductor, iGenomes hg38 UCSC genes.gtf (the manual-recommended annotation), and the exact documented default parameters (binSize=50, fragmentLength=150, minCountsCutOff=15, FDR<0.1, Beta_cutoff=0.5). The shipped toy BAMs are hg38, chr1-only. Pipeline: 1,535,061 bins -> 1,111 after filterBins -> deterministic PoissonGamma test. RESULT: 0 differential loci, vs the manual's stated 15 -> MISMATCH. The discrepancy is NOT a near-miss (min FDR=0.27): real but sub-threshold signal exists (94 bins p<0.05, top genes incl. RBM15, an m6A writer-complex subunit). To yield 15 loci the gene model would need ~270-300 filtered bins (~4x sparser) -> the dominant cause is the exact-but-under-specified GTF/gene-model build ('hg38_UCSC.gtf' in the manual). The result was NOT tuned toward 15 (no GTF-shopping). PROVISIONAL possible-overstatement flag raised for human review (not a confirmed fabrication). NOT attempted: GSE119168 full real-data analysis, paper main figures, simulation/method-comparison panels. Pipeline determinism confirmed by source inspection (no RNG). All artifacts under reproduction/.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 50
    assessed: 2026-06-15 ⛓ 36ebdc986f85
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-23
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The paper hypothesizes that a random-effect model using gene-level INPUT read counts (rather than peak-level counts and negative-binomial models used by prior tools) can more accurately and robustly detect differentially methylated (m6A) loci from MeRIP-seq data, including in complex study designs with covariates.

Core claims
  • RADAR is a novel analytical tool for differential methylation analysis of MeRIP-seq data combining gene-level INPUT normalization with a Poisson random effect model. method
  • RADAR uses gene-level read counts rather than peak/bin-level counts as a more robust measurement of pre-IP RNA expression level, reducing unwanted variance. method
  • RADAR models post-IP read counts with a flexible Poisson random effect model instead of a negative binomial distribution, better capturing the observed mean-variance relationship of m6A-seq data. mechanism
  • RADAR's generalized linear model framework allows incorporation of covariates, enabling analysis of complex study designs. method
  • RADAR achieves higher sensitivity and better-calibrated FDR than exomePeak, Fisher's exact test, MeTDiff/MeTPeak, and QNB on simulated MeRIP-seq data, especially when covariates are present. finding
  • Applying RADAR to four real m6A-seq datasets (ovarian cancer, T2D islets, mouse liver METTL14 KO, mouse brain stress) demonstrates it can accommodate diverse study designs and confounders. finding
  • RADAR is implemented as open-source software available on GitHub. resource
  • Confounding covariates (e.g., age in the ovarian cancer dataset, batch in the T2D dataset) can be regressed out to better separate samples by disease status. finding
Experimental setups
Assay System Perturbation Readout Platform
MeRIP-seq (m6A-seq) simulation simulated data derived from T2D islet INPUT library simulated true-site effect sizes (0.5, 0.75, 1) with/without covariate sensitivity and FDR of differential methylation detection
m6A-MeRIP-seq human fallopian tube (normal) vs metastatic omental tumor tissue disease status (ovarian cancer) differential m6A methylation sites GSE119168
m6A-MeRIP-seq human pancreatic islets, T2D patients vs healthy controls disease status (T2D) confounded with batch differential m6A methylation sites GSE120024
m6A-MeRIP-seq mouse liver, wild type vs METTL14 knockout METTL14 knockout differential m6A methylation sites GSE119490
m6A-MeRIP-seq mouse cortex/brain, stress-exposed vs control stress exposure differential m6A methylation sites GSE113781
PCA of m6A enrichment (IP counts adjusted for INPUT) ovarian cancer and T2D datasets covariate regression (age, batch) principal components / sample separation by disease status
Key results
  • RADAR achieved 95.7% sensitivity with 10.1% FDR at effect size 1 in the simple simulation case (FDR cutoff 10%) 95.7% sensitivity, 10.1% FDR
  • exomePeak and Fisher's exact test achieved high sensitivity (72-96%) but unsatisfactory FDR (>46-52%) across simulated effect sizes FDR >46%
  • QNB had little to no power for small effect sizes and only 13.9% sensitivity at effect size 1 13.9% sensitivity, 0.5% FDR
  • With covariates included (difficult case), RADAR maintained sensitivity (38.4-95.7%) and calibrated FDR (13.7-18.2%), while MeTDiff lost power sensitivity 38.4-95.7%, FDR 13.7-18.2%
  • Gene-level read counts show lower cross-sample variance than bin-level (local) read counts across four m6A-seq datasets
  • Median peak read coverage is 18 reads (7 reads per 50-bp bin) versus 272 reads on average at gene level in a typical 20M-read INPUT sample 272 vs 18 reads
  • After regressing out batch/age covariates, samples separate by disease status on PCA in the T2D and ovarian cancer datasets
  • All methods failed to detect enriched low p-values in the mouse brain dataset, indicating little to no true signal, consistent with prior publication
Key statistics
  • count 272 reads (gene-level) vs 18 reads/peak (7 reads/50bp bin) (read coverage for pre-IP expression estimation)
  • count 26,324 simulated sites, 20% true positives (simulation design for benchmarking)
  • other sensitivity 29.1%, FDR 12.0% (RADAR, simple case, effect size 0.5)
  • other sensitivity 95.7%, FDR 10.1% (RADAR, simple case, effect size 1)
  • other sensitivity 38.4%, FDR 18.2% (RADAR, difficult case (with covariate), effect size 0.5)
  • count 7 normal fallopian tube vs 6 metastatic omental tumor samples (ovarian cancer dataset (GSE119168))
  • count 8 T2D vs 7 healthy control islet samples (T2D dataset (GSE120024))
  • count 4 wild type vs 4 METTL14 knockout mouse liver samples (mouse liver dataset (GSE119490))

Statistical methods review

Model: opus

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This is a methods paper presenting RADAR, a tool that models MeRIP-seq read counts with a Poisson random-effect generalized linear model to detect differentially methylated loci, using gene-level INPUT counts as the pre-IP expression estimate and DESeq2 median-of-ratios normalization. Performance is assessed by simulation studies (data generated under both a random-effect model and a quad-negative-binomial model, across effect sizes 0.5/0.75/1, with and without covariates) and on four real m6A-seq datasets. Results are reported primarily as sensitivity and empirical false discovery rate at a fixed FDR cutoff, plus p-value histogram distributions, rather than via classic hypothesis-test summaries.

Replicationbiological Sample sizesample sizes given per real dataset (ovarian 7 normal vs 6 tumor; T2D 8 vs 7 in three batches; mouse liver 4 vs 4; mouse brain 7 vs 7) and simulations used 8 samples with 10 simulated copies each; no formal power analysis described Groupsdisease/condition vs control (e.g., tumor vs normal, T2D vs control, KO vs WT, stress vs control) Pairingpaired Randomization/blindingnot stated Dispersionnone Exact p-valuesno Effect sizesyes Confidence intervalsno Multiplicity correctionFDR control (results reported and thresholded at an FDR cutoff of 0.1); empirical FDR computed against known true sites in simulations
Statistical tests used
Test Applied to n Assumptions
Poisson random-effect generalized linear model (RADAR's own differential methylation test) main differential methylation calls across simulated and real m6A-seq datasets 8 simulated samples/replicates per simulation scenario; real datasets: ovarian cancer 7 vs 6, T2D 8 vs 7, mouse liver 4 vs 4, mouse brain 7 vs 7 stated
Fisher's exact test (exomePeak and as a comparator method) benchmark comparison of differential methylation detection on simulated and real data not stated
Likelihood ratio test based on binomial distribution ('bltest', exomePeak later version) described as an existing comparator approach not stated
Beta-binomial model (MeTPeak/MeTDiff) comparator method in benchmarks not stated
Negative binomial model (DRME/QNB) comparator method in benchmarks stated (assumes quadratic mean-variance relationship)
Approaches that could also have been used
  • Method performance was summarized using sensitivity and empirical FDR at a single FDR cutoff of 0.1, with supplementary curves at varying cutoffs.
    Could also: Full ROC or precision-recall curves with summary statistics such as AUC or partial AUC could also be reported. — A single-threshold summary captures one operating point; threshold-free curves would additionally convey performance across the whole sensitivity-specificity trade-off and give a single comparable scalar across methods.
  • Benchmarking relied on simulated data generated under the random-effect model and the quad-negative-binomial model.
    Could also: A purely resampling- or plasmode-based simulation drawn from real data (e.g., permuting/spiking known signals into observed counts) could also be used alongside parametric generators. — Parametric simulation can favor methods whose assumptions match the generating model; complementary data-driven simulation would add evidence that conclusions are not tied to a particular parametric form.
  • The Poisson random-effect GLM was chosen over a negative binomial model, motivated by an observed non-quadratic mean-variance relationship in post-IP counts.
    Could also: A negative binomial or quasi-Poisson GLM with random effects, or a flexible mean-variance trend (e.g., voom-style precision weighting), could also accommodate over-dispersion. — These are widely used over-dispersion frameworks; presenting them as alternatives clarifies which aspects of variability the random-effect Poisson choice captures and would let readers compare model-fit trade-offs.
  • Covariates (e.g., age, batch) were incorporated within the GLM, illustrated partly via PCA before/after regressing out covariates.
    Could also: Explicit interaction or sensitivity analyses, or surrogate-variable/RUV approaches for unmeasured confounders, could also be reported. — These would additionally characterize robustness to unmodeled or partially confounded covariates beyond the known ones included in the design.
  • FDR control at a 0.1 cutoff was used to define discoveries.
    Could also: Reporting the specific multiple-testing procedure (e.g., Benjamini-Hochberg vs. q-value/local FDR) and effect-size thresholds jointly could also be stated. — Naming the exact procedure and any effect-size filter makes the discovery rule fully reproducible and clarifies how family-wise versus false-discovery error is being controlled.
  • Real-data comparisons used p-value histogram shape as a proxy for sensitivity in the absence of ground truth.
    Could also: Reproducibility-based metrics such as IDR across replicates or concordance with orthogonal validation could also serve as evaluation criteria. — Without true labels in real data, reproducibility or orthogonal-assay agreement offers an additional, label-free way to compare methods that complements p-value distribution inspection.
Software: R/RADAR (the presented tool) · DESeq2 (median-of-ratios normalization) · exomePeak · MeTDiff/MeTPeak · QNB

Result convergence & founder nodes

Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
100
Impact: high
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

C1
Reported
15 differential m6A loci (FDR<0.1 & |logFC|>0.5) on the shipped 17-sample (9 Ctl + 8 Case) RADAR example data (RADAR manual / workflow.html)
Reproduced
0 differential loci (0 bins pass FDR<0.1 & |beta|>0.5; min FDR observed = 0.27 over 1,111 tested bins; reportResult stops with 'no bin passing threshold')
did not match

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 10/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +7

This is an incomplete reproduction, not a discrepancy: the RADAR example data ships inside the package (1:1 input) and the target is a single deterministic tutorial value (15 differential loci, no RNG in the code), but SLURM «job» was finalized while still in the multi-hour countReads phase, so no reproduced number was emitted. The shortfall is squarely on our side (compute/runtime), not the authors' — nothing about the 15-loci figure looks unsupported or fabricated. The only genuine open uncertainty is the under-specified GTF build (an input/preprocessing lever). Verdict: yellow across the judgement axes (pending, derivable-in-principle, no measured deviation), criticality yellow.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

266.5 k
tokens (I/O) · 16.1 M incl. cache
107 min
runtime · 7.73 CPU-h
34.6 GB
peak RAM
2 (1 failed)
HPC jobs
hummel
machine