Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Comprehensive analysis of m6A methylome alterations after azacytidine plus venetoclax treatment for acute myeloid leukemia by nanopore sequencing.

Comput Struct Biotechnol J · 2024
L1 No data access 3/4
Why this verdict

The main result did not reproduce in this reproduction attempt. Where our recomputation produced values that differ from the published ones, those discrepancies are listed below. This is a single automated attempt — not peer review and not a finding of error or misconduct — and differences can also arise from data access, undocumented parameters or the computing environment. The verdict can be contested via “report an error”.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Main result did not reproduce
Decisive
From: Q5 · Derivability / plausibility 🔴
Main result did not reproduce
Decisive
From: Q7 · Core claim 🔴
Main result did not reproduce
Decisive
From: Q8 · Severity of the miss (overall human judgment) 🔴
✓ What held up
  • Nothing in this column.
What did not (or only partly)
  • 🔴Could not use the authors’ exact input data
  • 🔴Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🔴A deviation was attributed to the published material
  • 🔴Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🔴The central claim did not (fully) hold under reproduction
  • 🔴Overall, the reproduction showed a material discrepancy
No data access Data access not granted

This paper has a computational component, but its primary data is legally or ethically access-restricted — identifiable patient cohorts, rare-disease genomes, or controlled-access biobanks that cannot be openly shared. The reproduction therefore could not be attempted. That is a neutral verdict: it does not mean the result is wrong or that the authors fell short — only that, for legitimate privacy reasons, it cannot be independently checked from public data. We deliberately do NOT assign a 0–100 score here, because a low number would wrongly read as a failed reproduction.

Reproduction agent’s raw note

DROP (no_data_accession). The paper is described well enough at the method level (Guppy 3.3.0 -> Tombo 1.5.1 resquiggle --rna -> Minimap2 2.22 map-ont GRCh38 -> DENA LSTM m6A calling -> NanoCount DEGs; thresholds given: m6A high if ratio>0.5, diff-m6A paired t-test P<0.1, DEG FC>2 & P<0.05), and all code is public third-party (tombo, DENA both resolve). But it is NOT reproducible because the SOLE input to every pipeline-derived result — the authors' Oxford Nanopore direct-RNA FAST5 raw signals from 3 AML/CR bone-marrow pairs + HL-60 — was never deposited: exhaustive full-text search (Europe PMC PMC10950754 fullTextXML, 125913 bytes) found NO SRA/ENA/BioProject/GSA/NGDC/CNGB/figshare/zenodo/dryad accession and NO 'available on request' statement. The only accessions present (GSE1159/6891/8970/12417/37642) are external Affymetrix microarray cohorts used only via the kmplot.com online survival tool, NOT the paper's data; the registry's data_accession=GSE1159 is a confirmed text-mining false positive (Valk 2004 microarray). Applying a third-party tool to the paper's data (brief P16) is impossible with no data to apply it to. NOT ATTEMPTED, by design: any «our HPC» compute (no input exists); the kmplot.com online survival curves (web tool over third-party microarrays, out of scope); wet-lab claims (IC50=1.86uM, qPCR, polyA length). No compute spent; nothing fabricated. Flagged for human review: the paper's specific pipeline numbers are currently unverifiable from any shipped data/code (possible-fabrication class).

💻 Code ↗ 🗄 Data: GSE1159

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment
    assessed: 2026-06-14 ⛓ 554357368022
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-14
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

Whether azacytidine (AZA) plus venetoclax (VEN) treatment alters the global m6A RNA methylome in AML patients, and whether these alterations reveal biomarkers linked to AML prognosis.

Core claims
  • m6A site number and m6A levels are significantly lower in post-treatment complete remission (CR) bone marrow than in pre-treatment AML bone marrow finding
  • AZA exerts a significant demethylation effect at the RNA level, confirmed both in patient BM pairs and in AZA-treated HL-60 AML cells finding
  • 13 genes show concordant decreases in both m6A modification and expression level between AML and CR BMs finding
  • Nanopore direct RNA sequencing with the DENA pipeline enables antibody-free, single-nucleotide-resolution quantification of m6A levels method
  • Down-regulated m6A sites are concentrated in exonic regions, predominantly 3' UTRs, of protein-coding genes finding
  • The genomic and RRACH-motif distribution pattern of m6A sites (e.g., AAACA most abundant) is largely unchanged between AML and CR BMs despite overall level differences finding
  • Genes with down-regulated m6A are enriched in translation regulation, apoptosis, mRNA splicing, protein methylation, nucleotide metabolism, and AML-related pathways finding
  • HPRT1, SNRPC, and ANP32B are prognostic genes for AML, with decreased expression associated with favorable overall survival finding
Experimental setups
Assay System Perturbation Readout Platform
Nanopore direct RNA sequencing (dRNA-seq) for m6A mapping (DENA pipeline) paired bone marrow samples (pre-treatment AML vs post-treatment CR) from 3 AML patients AZA+VEN induction chemotherapy genome-wide m6A site number, m6A ratio/level, motif distribution Oxford Nanopore GridION, R9.4.1 flow cell, SQK-RNA002 kit
Nanopore dRNA-seq for m6A quantification HL-60 AML cell line AZA treatment at IC50 concentration m6A levels Oxford Nanopore GridION
Nanopore RNA-seq quantification (NanoCount) paired AML vs CR bone marrow (3 patients) AZA+VEN treatment differentially expressed genes (fold change, P-value) Oxford Nanopore GridION
Functional enrichment analysis (GO: BP/MF/CC; KEGG pathway) gene lists (m6A down-regulated genes) derived from BM samples none (bioinformatic analysis) enriched biological processes, molecular functions, cellular components, pathways DAVID; Gene Ontology tools; Enrichr
Kaplan-Meier survival analysis AML patients from 5 GEO datasets (GSE1159, GSE12417, GSE37642, GSE6891, GSE8970) none (retrospective expression-survival correlation) overall survival stratified by HPRT1/SNRPC/ANP32B expression Kaplan-Meier Plotter (kmplot.com)
Site-specific m6A quantification qPCR (MazF/BstI-based RT-qPCR) paired AML vs CR bone marrow samples AZA+VEN treatment relative m6A ratio at specific sites in HPRT1, SNRPC, ANP32B SYBR green qPCR (BioRad)
Key results
  • Relative amount of m6A sites in CR BMs was about a quarter of that in AML BMs ~4-fold lower, P<0.0001
  • m6A levels significantly higher in AML BMs than CR BMs P<0.001
  • Majority of differentially m6A sites and genes between AML and CR BMs were down-regulated rather than up-regulated
  • 478 DEGs identified between AML and CR BMs (209 up-regulated, 269 down-regulated) 209 up / 269 down, P<0.05
  • 110 m6A up-regulated genes and 604 m6A down-regulated genes identified genome-wide 110 up / 604 down
  • Intersection of differentially m6A-modified genes and DEGs yielded 13 genes with concordant decreases in m6A and expression 13 genes
  • Of the 13 overlapping genes, HPRT1, SNRPC, and ANP32B were significantly associated with AML prognosis; decreased expression linked to favorable overall survival 3 genes, n=1608 patients
  • 86.55% of genes with reduced m6A levels in CR BMs were protein-coding genes 86.55%
Key statistics
  • pvalue P<0.0001 (relative m6A site proportion, AML vs CR BM)
  • pvalue P<0.001 (m6A level difference, AML vs CR BM)
  • pvalue P<0.1 (threshold for differentially m6A sites (paired t-test across 3 sample pairs))
  • fold_change FC>2 (also stated as |logFC|>1) (threshold for significant DEGs)
  • pvalue P<0.05 (threshold for DEG significance)
  • count 478 DEGs (209 up, 269 down) (CR vs AML BM gene expression comparison)
  • count 110 m6A up-regulated genes; 604 m6A down-regulated genes (differentially m6A-modified genes, CR vs AML BM)
  • count 1608 AML patients (pooled cohort from 5 GEO datasets used for Kaplan-Meier survival analysis)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This paired pre-post observational study used Nanopore direct RNA sequencing on bone marrow samples from three AML patients before and after AZA/VEN treatment. Differential m6A sites were identified by paired t-test across three patient pairs using a P<0.1 threshold, and differentially expressed genes (DEGs) were filtered by fold-change and nominal P-value. Prognostic significance of three selected genes was evaluated via log-rank tests on Kaplan-Meier survival curves drawn from five external GEO cohorts (n=1608). Results were reported primarily with threshold P-value indicators and volcano/Manhattan-style visualizations.

Replicationbiological Sample sizeThree AML patients who achieved complete remission enrolled; no formal power calculation or sample size justification stated GroupsPre-treatment AML bone marrow vs post-treatment complete remission bone marrow (paired); AZA-treated HL-60 cells vs untreated (secondary validation) Pairingpaired Randomization/blindingnot stated Dispersionnone Exact p-valuesno Effect sizesno Confidence intervalsno Multiplicity correctionnone stated
Statistical tests used
Test Applied to n Assumptions
Paired t-test Comparison of m6A ratios at each genomic site between pre-treatment (AML) and post-treatment (CR) bone marrow pairs; used to identify differentially m6A-modified sites (P<0.1 threshold, Figs 2A-B) 3 matched patient pairs (6 bone marrow samples) not stated
Independent samples t-test or Mann-Whitney U test (either/or, selection criterion unstated) Two-group comparisons described generically in statistical methods section 2.8; applied to group-level comparisons including m6A site counts and m6A levels (Figs 1D-E) 3 per group (AML vs CR bone marrow) not stated
Log-rank test Overall survival analysis comparing high vs low expression of HPRT1, SNRPC, and ANP32B in AML patients via Kaplan-Meier Plotter (Figs 3D-F) 1608 AML patients from 5 GEO databases (GSE1159, GSE12417, GSE37642, GSE6891, GSE8970) not stated
Fold-change threshold plus nominal P-value filter Identification of DEGs between AML and CR bone marrow (|log2FC|>1 and P<0.05; Fig 3A volcano plot) 3 matched patient pairs not stated
Approaches that could also have been used
  • Genome-wide m6A site comparisons were conducted by paired t-test at thousands of sites simultaneously, with a nominal P<0.1 threshold and no multiple-testing correction
    Could also: A Benjamini-Hochberg FDR adjustment across all tested sites could also be applied, yielding q-values in addition to or instead of nominal P-values — When testing thousands of genomic positions simultaneously, FDR control is the standard approach in methylation and sequencing studies to limit the expected proportion of false discoveries; this is the convention in analogous MeRIP-seq pipelines
  • A parametric paired t-test was used to compare m6A ratios across three matched pairs
    Could also: A non-parametric Wilcoxon signed-rank test could also be used for the same three-pair matched design — With only n=3 pairs, the normality assumption underlying the paired t-test cannot be empirically verified; the Wilcoxon signed-rank test is a standard distribution-free alternative for small matched-pair designs
  • DEGs were filtered by a fold-change threshold (|log2FC|>1) combined with a nominal P-value cutoff (P<0.05) without a stated multiple-testing correction
    Could also: An adjusted P-value (e.g., Benjamini-Hochberg FDR q<0.05 or q<0.1) could also serve as the significance criterion for DEG calling, as implemented in tools such as DESeq2 or edgeR — RNA-seq DEG analyses conventionally apply FDR adjustment to account for simultaneous testing across thousands of genes; with n=3 per group, quasi-likelihood negative-binomial models (e.g., edgeR QLF) are also designed for very small replicates and include dispersion shrinkage
  • Group-level comparisons of m6A site counts and m6A levels (Figs 1D-E) were reported with threshold P-value symbols but without any measure of dispersion
    Could also: Reporting mean ± SD or mean ± SEM, or displaying individual paired data points with connecting lines, could also accompany the significance indicators — With n=3 per group, individual-point plots (strip or dot plots with pair connectors) give readers complete visibility of the underlying data and highlight consistency across pairs; dispersion statistics help contextualize effect magnitude relative to variability
  • Three separate log-rank tests were performed for HPRT1, SNRPC, and ANP32B survival analyses without mention of correction for the three comparisons
    Could also: A Bonferroni correction (adjusted alpha=0.017) or Benjamini-Hochberg FDR across the three tests could also be applied and reported — Testing three genes for prognostic association constitutes a family of comparisons; acknowledging and adjusting for this is standard practice in prognostic biomarker studies and helps calibrate confidence in each individual finding
  • No sample size justification or power calculation was reported for the primary cohort of three patients
    Could also: A formal power analysis, or an explicit characterization of the study as a hypothesis-generating pilot with stated effect size assumptions, could also accompany the study design description — Documenting the power of the primary comparison—or explicitly framing the n=3 design as exploratory—helps readers interpret the scope, precision, and generalizability of the sequencing-based findings before external validation
Software: SPSS 26.0 · Guppy (ONT basecaller) 3.3.0 · Tombo (re-squiggle) 1.5.1 · Minimap2 2.22 · DENA (LSTM-based m6A site caller) · NanoCount 1.0.0 · BEDTools 2.26.0 · R/CMplot · R/ggseqlogo · DAVID · Enrichr · Kaplan-Meier Plotter (online tool)

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
7
Impact: low
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

GSE12417 GEO in Results (http://purl.org/orb/Results)
also used by 1 paper:
GSE37642 GEO in Results (http://purl.org/orb/Results)
also used by 1 paper:
GSE1159 GEO in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GSE6891 GEO in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GSE8970 GEO in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet

Downstream reach in the literature

298 downstream papers · 5 datasets

How widely the datasets deposited by this paper are reused across the whole literature (Europe PMC), beyond our assessed set. This is a factual dependency map — reusing a public dataset is normal, good science. It is not a judgement on the downstream papers; the only verdict here is this paper's own, with its cited rationale.

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — PMID 38510975 (PMC10950754, doi:10.1016/j.csbj.2024.02.029)

Zhang et al. 2024, Comput Struct Biotechnol J 23:1144-1153. "Comprehensive analysis of m6A methylome alterations after azacytidine plus venetoclax treatment for acute myeloid leukemia by nanopore sequencing."

Pipeline-derived results (would be IN scope)

Every quantitative result in this paper is derived from one primary input: Oxford Nanopore direct-RNA sequencing (SQK-RNA002, R9.4.1, GridION) of bone marrow from 3 AML patient pairs (pre-treatment AML vs post-treatment CR) plus HL-60 cells (DMSO vs AZA). The described pipeline:

step tool (paper) resolves?
basecalling Guppy 3.3.0 (--flowcell FLO-MIN106 --kit SQK-RNA002) yes (ONT login-walled)
signal re-squiggle Tombo 1.5.1 preprocess annotate_raw_with_fastqs + resquiggle --rna yes — github.com/nanoporetech/tombo (public, not archived, last push 2023-05, deprecated by ONT)
genome alignment Minimap2 2.22 -ax map-ont → GRCh38 yes
m6A calling DENA LSTM_extract.py / LSTM_predict.py (pre-trained LSTM) yes — github.com/weir12/DENA (public)
m6A ratio modified reads / total reads; site "high" if ratio>0.5; diff-m6A if paired t-test P<0.1 n/a
gene annotation BEDTools 2.26.0 intersect yes
DEGs NanoCount 1.0.0; FC>2 & P<0.05 yes
visualisation CMplot, ggseqlogo (R) yes

Pinnable numeric claims that this pipeline produces (see original/claims.tsv): relative m6A-site amount in CR ≈ ¼ of AML (Fig 1D); 478 DEGs (209 up / 269 down, Fig 3); 110 m6A-up / 604 m6A-down genes (Fig 3); 13 genes with decreased m6A + expression; 3 prognosis genes HPRT1/SNRPC/ANP32B; "AAACA" most / "AGACA" least abundant RRACH motif (Fig 1).

Out of scope

  • IC50 of AZA for HL-60 = 1.86 µM (Fig S2) — wet-lab assay.
  • qPCR m6A-site validation, polyA-tail length (Fig 4E/F) — wet-lab.
  • Overall-survival KM curves of the 3 genes — produced by the kmplot.com online tool over 5 external Affymetrix microarray cohorts (GSE1159, GSE6891, GSE8970, GSE12417, GSE37642; 1608 patients). A web tool over third-party data, not the paper's own pipeline; not «our HPC»-reproducible and tangential to the core claim.

DROP — data_accession blocker (no «our HPC» compute spent)

The paper deposits NO raw nanopore data and gives NO availability statement. Exhaustive full-text search (Europe PMC PMC10950754/fullTextXML, 125 913 bytes) found:

  • the only sequence-archive-style accessions are the 5 GSE microarray cohorts above, and they appear only in the Fig-3 survival-curve caption as inputs to the kmplot.com online tool — NOT the paper's nanopore data;
  • no SRA / ENA / BioProject (PRJ*) / GSA / NGDC-OMIX / CNGB / figshare / zenodo / dryad accession;
  • no "data available on request" / "deposited in" / repository sentence anywhere (the single literal "FAST5" hit is in the Methods Tombo command);
  • the registry's recorded data_accession=GSE1159 is a text-mining false positive — confirmed: GSE1159 = Valk et al. 2004 "Expression profiles of AML patient samples" Affymetrix microarray, unrelated to this study's nanopore data.

Consequence: the input to every pipeline-derived result is unobtainable. Applying the third-party tools (Tombo/DENA, fully valid per P16) is impossible without the paper's own FAST5 signals, which were never released. No public substitute dataset reproduces this paper's reported numbers.

drop_reason = no_data_accession (code resolves & is public; no resolvable public dataset for the pipeline input). Nothing fabricated; no compute submitted.

Figures / tables: Fig 1DFig 3Fig 1
C1
Reported
m6A-site amount in CR ~1/4 of AML (Fig 1D)
Reproduced
not reproduced — input data not deposited
partial
C2
Reported
478 DEGs = 209 up + 269 down (Fig 3)
Reproduced
not reproduced — input data not deposited
partial
C3
Reported
110 m6A-up + 604 m6A-down genes (Fig 3)
Reproduced
not reproduced — input data not deposited
partial
C4
Reported
13 genes with decreased m6A + expression
Reproduced
not reproduced — input data not deposited
partial
C5
Reported
3 prognosis genes: HPRT1, SNRPC, ANP32B
Reproduced
not reproduced — input data not deposited
partial
C6
Reported
RRACH motif: AAACA most / AGACA least (Fig 1)
Reproduced
not reproduced — input data not deposited
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 13/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🔴1. Data identity
🔴2. Endpoint comparability
🟡3. Location of the main deviation
🔴4. Cause of the deviation
🔴5. Derivability / plausibility
🟡6. Severity of the deviation
🔴7. Core claim
🔴8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Main result did not reproduce
Decisive
From: Q5 · Derivability / plausibility 🔴
Main result did not reproduce
Decisive
From: Q7 · Core claim 🔴
Main result did not reproduce
Decisive
From: Q8 · Severity of the miss (overall human judgment) 🔴

This RU is a clean drop (no_data_accession): the only input to every quantitative claim — the authors' Nanopore direct-RNA FAST5 signals for 3 AML/CR pairs + HL-60 — was never deposited, and PMC10950754 carries no data-availability statement (the registry's GSE1159 is a text-mining false positive). The third-party code (tombo, DENA) resolves, so the blocker is purely data on the authors' side, not our method. Consequently none of C1–C6 (478 DEGs; 110/604 diff-m6A genes; 13 overlap; HPRT1/SNRPC/ANP32B; AAACA/AGACA motif) are derivable or testable, leaving the central conclusion unconfirmed and the values unverifiable — a possible-fabrication/unverifiable class rather than demonstrated error.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

71.2 k
tokens (I/O) · 3.4 M incl. cache
5 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.