Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

5-methylcytosine RNA modification regulators-based patterns and features of immune microenvironment in acute myeloid leukemia.

Aging (Albany NY) · 2024
L1 93/100 PQI 88
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score -6
✓ What held up
  • Reported values were directly comparable
  • No relevant deviation in data/preprocessing
  • No authors-side cause for any deviation
  • Reported values are derivable from the shared data
  • Any deviation was negligible
  • The central claim held under reproduction
  • Overall, the reproduction was clean
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
How its reproducibility compares
93/100
Reproducibility score
1.1 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 85% of all assessed papers rank 154 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough to reproduce, and 1:1 on the headline validation claim. The paper's named code is the generic survival R package (third-party tool; the authors shipped no pipeline) and the assigned data is GSE12417. I took the paper's fully-published 11-gene m5C-regulator LASSO-Cox risk-score formula (coefficients verbatim from Results), applied it to GSE12417 (per-gene z-score signature transfer, mean-collapse multi-probe, median split), and recomputed survival statistics with the survival package + timeROC on «our HPC». GSE12417 has three sub-series; only GPL570 (HG-U133 Plus 2.0, n=79) carries all 11 genes, where the 1-year time-dependent AUC = 0.726 vs the reported 0.720 (delta 0.006, essentially exact) and the high/low-risk Kaplan-Meier separation is significant (log-rank p=0.0044, HR 2.31, CI excludes 1, high-risk worse OS) — the central GSE12417 validation claim reproduces. The GPL96 sub-series (n=163) gives a null result, fully explained by the older HG-U133A array lacking 2 of the 11 genes (NSUN4, TET2), not a contradiction. No fabrication indicated: the headline AUC is independently recovered from public data + published coefficients. NOT attempted (out of this RU's data/code scope, per scope.md): LASSO model TRAINING (done on TCGA-LAML+GTEx, not GSE12417 -> we validated, did not refit); consensus clustering / 17-regulator differential expression (TCGA+GTEx); immune deconvolution EPIC/ESTIMATE, GSVA pathways, scRNA cell-type maps, ENCORI miRNA network, GDSC drug-IC50, cBioPortal mutations (each uses other datasets); RT-qPCR validation (wet-lab, n=4+4); and GSE37642 (second validation cohort, not the assigned accession). Caveats: the paper does not state the exact AUC timepoint, which GSE12417 sub-series, or the signature-transfer preprocessing; standard conventions were used and disclosed.

💻 Code ↗ 🗄 Data: GSE12417

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 93
    assessed: 2026-06-15 ⛓ a7d8934e1e3f
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-15
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

The role of 5-methylcytosine (m5C) RNA modification regulators in acute myeloid leukemia (AML) is unknown; this study tests whether the expression, genetic/epigenetic alterations, and patterns of 17 m5C regulators define molecular subtypes, predict prognosis, and shape the immune microenvironment in AML.

Core claims
  • Most of the 17 m5C regulators are differentially expressed between AML and normal samples and correlate with AML prognosis. finding
  • Two m5C modification subtypes (C1, C2) and high-/low-risk subgroups based on m5C regulator expression differ significantly in prognosis and immune cell infiltration. finding
  • A 6-gene (m5C-regulator) prognostic risk model predicts AML survival and was validated in GSE12417 and GSE37642 datasets. resource
  • Methylation status of certain m5C regulators (e.g., DNMT3A, DNMT3B, TRDMT1) affects AML patient survival. finding
  • m5C regulators (notably DNMT3A) correlate with miRNA expression and drug IC50 values in AML. finding
  • DNMT3A is the most frequently mutated m5C regulator in AML (19%), mainly missense mutations, but the mutation does not significantly affect survival. finding
  • qPCR validation was used to confirm differential expression of m5C regulators between normal and AML samples. method
  • m5C regulators correlate with immune cell infiltration (CD4/CD8 T-cells, macrophages) and stromal score in the AML microenvironment. mechanism
Experimental setups
Assay System Perturbation Readout Platform
Bulk RNA-seq differential expression (merged GTEx+TCGA, batch-corrected) AML patient and normal samples (TCGA, GTEx) none mRNA expression of 17 m5C regulators
Mutation annotation analysis AML patients (cBioPortal, 4 datasets) none mutation type and frequency of m5C regulators cBioPortal
Copy number variation analysis AML patients (TCGA) none homozygous/heterozygous CNV and CNV-mRNA correlation
DNA methylation analysis AML patients none methylation level vs expression and survival of m5C regulators
Consensus clustering / LASSO Cox regression / ROC modeling AML patient cohort (TCGA), validation in GSE12417 and GSE37642 none molecular subtypes, risk score, survival, AUC
Immune cell infiltration deconvolution (EPIC, ESTIMATE) AML patient tumor microenvironment none immune cell proportions, stromal score and correlation with m5C regulators EPIC, ESTIMATE algorithms
Drug sensitivity, RNAss, and miRNA correlation analysis; GSVA pathway analysis AML patient data none drug IC50 correlation, RNA stemness score, miRNA expression, KEGG/hallmark pathway scores GSVA
scRNA-seq cell-type-specific expression; qPCR validation; peripheral blood smear/genome sequencing AML scRNA-seq datasets (GSE116256, GSE135851, GSE147989, GSE154109); normal vs AML samples; two AML patients none cell-type-specific expression; m5C regulator expression by qPCR; morphology and DNMT3A mutation qPCR
Key results
  • Nine m5C regulators (NOP2, NSUN3, NSUN6, NSUN7, DNMT1, DNMT3A, TRDMT1, TET2, TET3) highly expressed in AML tumor group
  • Low-risk score group survived longer than high-risk group; prognostic model AUC = 0.744 AUC=0.744
  • Model validated with AUC 0.720 (GSE12417) and 0.757 (GSE37642), low-risk group better prognosis AUC=0.720 and 0.757
  • DNMT3A mutation rate highest at 19% (mostly missense); TET2 ~9% (mainly truncating) 19%; ~9%
  • High DNA methylation of DNMT3A and TRDMT1 shortened overall survival; high DNMT3B methylation lengthened survival
  • DNMT3A positively correlated with Nelarabine drug sensitivity (IC50) R=0.638, p<0.001
  • DNMT3A positively correlated with hsa-mir-146a expression R=0.65, P=1.6e-19
  • CD4 T-cells enriched in low-risk group; CD8 T-cells and macrophages enriched in high-risk group
Key statistics
  • other AUC=0.744 (ROC of m5C prognostic model in TCGA)
  • other AUC=0.720 (model validation in GSE12417)
  • other AUC=0.757 (model validation in GSE37642)
  • count 19% (DNMT3A mutation frequency in AML)
  • correlation R=0.638, p<0.001 (DNMT3A vs Nelarabine IC50)
  • correlation R=−0.51, FDR=1e-12 (NSUN7 methylation vs expression)
  • correlation R=−0.36, FDR=1.9e-06 (DNMT3A methylation vs expression)
  • correlation R=0.65, P=1.6e-19 (DNMT3A vs hsa-mir-146a)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This retrospective bioinformatics study characterized 17 m5C RNA modification regulators in AML using merged TCGA and GTEx expression data, mutation/CNV data from cBioPortal, and GEO cohorts. The primary analytical strategy combined unsupervised consensus clustering (ConsensusClusterPlus) to define patient subtypes, LASSO-penalized Cox regression to build a prognostic risk score, and Kaplan-Meier analysis to compare survival between risk groups. Secondary analyses used Pearson and Spearman correlations to link regulator expression with immune infiltration (EPIC/ESTIMATE), drug sensitivity (IC50), and miRNA networks; external GEO cohorts (GSE12417, GSE37642) and qPCR served as validation steps.

Replicationunclear Sample sizeCohort sizes not stated in the available text; discovery cohort is TCGA AML; independent validation uses GSE12417 and GSE37642; cell-type-specific expression uses four scRNA-seq GEO datasets (GSE116256, GSE135851, GSE147989, GSE154109); qPCR validation in unspecified number of clinical samples GroupsAML vs. normal (TCGA+GTEx merged); consensus clusters C1 vs. C2; high-risk vs. low-risk (median prognostic score); clinical subgroups by age, CR status, FAB subtype, sex, survival status, treatment status; methylation-high vs. methylation-low per regulator Pairingunpaired Randomization/blindingnot stated Dispersionnone Confidence intervalsno Multiplicity correctionFDR (method not named; Benjamini-Hochberg is conventional but not stated) applied to Spearman methylation-expression correlations (Figure 5B); no correction stated for all other comparison families
Statistical tests used
Test Applied to n Assumptions
Differential expression test (specific test not named; significance reported as threshold asterisks via boxplots) AML vs. normal sample expression for all 17 m5C regulators after ComBat batch correction (Figure 2A) and clinical subgroup comparisons across age, CR, FAB, sex, survival status, treatment status (Figure 3A–F) not stated
Pearson correlation Inter-regulator expression correlation matrix (Figure 2B); drug IC50 vs. regulator expression with exact R and p reported for selected pairs (Figure 9B,C); miRNA vs. regulator expression and immune cell vs. miRNA (Figure 9D–G); regulator vs. stromal score (Figure 8G); pathway scores vs. regulators (Figure 10B,D) not stated
Univariate Cox proportional hazards regression Association of each individual m5C regulator expression level with overall survival (Figure 2C) not stated
Kaplan-Meier survival analysis (log-rank test implied by p-value annotations on KM curves) DNMT3A mutation vs. overall survival (Figure 4D); CNV deletions vs. survival (Figure 4G); methylation-high vs. -low for DNMT3A, DNMT3B, TRDMT1 (Figure 5C,D); high- vs. low-risk group in TCGA discovery set, GSE12417, and GSE37642 validation sets (Figures 6E, 7A,B) not stated
Spearman correlation Methylation status vs. mRNA expression of m5C regulators (Figure 5A); exact R and FDR reported for four representative genes: NSUN4 (R=−0.31, FDR=4.4e-05), NSUN7 (R=−0.51, FDR=1e-12), DNMT3A (R=−0.36, FDR=1.9e-06), DNMT1 (R=−0.20, FDR=9.8e-03) (Figure 5B) not stated
LASSO-penalized Cox regression Feature selection and coefficient estimation for the m5C prognostic risk model; minimum-criteria used to identify 11-gene signature from candidate differentially expressed genes in the two consensus clusters (Figure 6C,D) not stated
Time-dependent ROC curve / AUC Prognostic model discrimination: AUC=0.744 in TCGA training set (Figure 6F), AUC=0.720 in GSE12417 (Figure 7C), AUC=0.757 in GSE37642 (Figure 7D) na
GSVA (gene set variation analysis) Enrichment of KEGG and Hallmark pathway gene sets compared between high- and low-risk groups (Figure 10A,C) not stated
Approaches that could also have been used
  • Multiple differential expression comparisons across clinical subgroups (CR, FAB, treatment, sex) were performed with significance denoted only by threshold asterisks, without naming the underlying statistical test
    Could also: For multi-level group comparisons, a Kruskal-Wallis test (non-parametric) or one-way ANOVA (parametric) followed by a post-hoc correction such as Dunn's test or Tukey HSD could also be applied and would make the inference structure explicit — Specifying the omnibus test and post-hoc procedure clarifies what null hypothesis is being rejected and how the family-wise error rate is controlled across the many regulators and subgroup combinations tested
  • The prognostic risk score was derived from LASSO Cox regression in the TCGA discovery cohort and then evaluated directly in external GEO validation cohorts without an internal cross-validation step
    Could also: Internal k-fold cross-validation (e.g., 10-fold CV) or bootstrap optimism-correction within the TCGA cohort could also be applied before proceeding to external validation — Internal validation quantifies overfitting introduced during the penalized selection step and complements external validation by distinguishing optimism from transportability
  • Survival differences between high- and low-risk groups were assessed with Kaplan-Meier curves (unadjusted log-rank comparison)
    Could also: A multivariable Cox regression including established AML prognostic covariates (age, cytogenetic risk, FAB subtype) alongside the m5C risk score could also evaluate whether the score provides independent prognostic information — Unadjusted log-rank tests do not account for confounders; a multivariable model would clarify the incremental prognostic contribution of the m5C risk score beyond known clinical factors
  • Immune cell infiltration proportions were estimated using a single deconvolution algorithm (EPIC)
    Could also: Applying additional deconvolution methods such as CIBERSORT-ABS, MCP-counter, or xCell to the same expression data and comparing concordance across methods is also common practice — Different deconvolution algorithms make different modeling assumptions and show varying performance across tumor types and data platforms; cross-algorithm concordance strengthens confidence in the reported immune-infiltration patterns
  • Methylation-expression correlations are reported with FDR-adjusted values for four representative regulators (Figure 5B), while most other large families of statistical tests do not report a correction method
    Could also: Applying a consistent FDR or Bonferroni correction across the full families of miRNA-regulator correlations, drug-IC50 correlations, and per-regulator differential expression tests could also be used to control false-discovery rates across these multiple comparisons — Consistent multiplicity adjustment across test families aids in interpreting which associations are likely to replicate and aligns with current reporting standards for large-scale omics correlation analyses
  • Results of group comparisons (e.g., immune cell proportions between C1/C2 or high-/low-risk) are reported with p-value thresholds but without quantitative effect size measures
    Could also: Reporting fold-differences in estimated immune cell proportions, Cohen's d, or hazard ratios with 95% confidence intervals alongside p-values would also characterize the magnitude of observed differences — Effect size estimates allow readers to assess clinical or biological relevance independently of sample size, and confidence intervals convey estimation uncertainty that p-value thresholds alone do not capture
Software: R/inSilicoMerging · R/SVA (ComBat batch correction) · R/ConsensusClusterPlus · cBioPortal (mutation/CNV visualization) · EPIC (immune cell deconvolution) · ESTIMATE (stromal/immune scoring) · R/GSVA · LASSO Cox regression (likely R/glmnet; not explicitly named)

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
5
Impact: low
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

GSE116256 GEO in Methods (http://purl.org/orb/Methods)
also used by 1 paper:
GSE12417 GEO in Methods (http://purl.org/orb/Methods)
also used by 1 paper:
GSE37642 GEO in Methods (http://purl.org/orb/Methods)
also used by 1 paper:
GSE135851 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE147989 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE154109 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet

Downstream reach in the literature

179 downstream papers · 2 datasets

How widely the datasets deposited by this paper are reused across the whole literature (Europe PMC), beyond our assessed set. This is a factual dependency map — reusing a public dataset is normal, good science. It is not a judgement on the downstream papers; the only verdict here is this paper's own, with its cited rationale.

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-38277218

Paper: Ding Y et al. (2024) 5-methylcytosine RNA modification regulators-based patterns and features of immune microenvironment in acute myeloid leukemia. Aging (Albany NY). PMID 38277218 · PMCID PMC10911375 · DOI 10.18632/aging.205484

Assigned code artifact: https://github.com/therneau/survival — the generic survival R package (Terry Therneau). This is NOT the authors' own analysis code; it is the third-party statistics tool the paper used for Cox / Kaplan–Meier / log-rank survival analysis. Per the brief's P16 rule, applying this named tool to the paper's named data is an equally valid reproduction. The authors did not release their own pipeline code (Aging papers typically ship none).

Assigned data: geo:GSE12417 — a public AML microarray cohort (Metzeler et al., Blood 2008) used in the paper as an independent validation cohort for the prognostic risk model. Subseries: GPL96 (HG-U133A, 163 CN-AML, training in the original Metzeler paper), GPL97 (HG-U133B, 163), GPL570 (HG-U133 Plus 2.0, 79).

In scope (pipeline-derived, reproducible with GSE12417 + survival)

# Reported result Where Reproducible?
C1 11-gene m5C risk-score validation AUC = 0.720 in GSE12417 Results / model-validation figure YES — formula + coefficients fully published; GSE12417 public. PRIMARY TARGET.
C2 High- vs low-risk Kaplan–Meier separation (median split) is prognostic in GSE12417 same YES — KM + log-rank via survival; the paper's central claim for this cohort.

The published risk-score formula (11 genes, LASSO-Cox coefficients):

RiskScore = -0.3202*NSUN3 -0.2526*NSUN4 +0.0372*NSUN5 -0.3514*NSUN6 +0.3160*NSUN7
            +0.3244*DNMT1 -0.3648*DNMT3A +0.3476*DNMT3B -0.0334*TET2 -0.0897*TET3
            +0.2137*ALKBH1

Apply to GSE12417 expression (per-gene z-score, standard cross-platform signature transfer), median-split into high/low risk, then survival/timeROC: KM + log-rank p, Cox HR, C-index, and time-dependent ROC-AUC at 1/3/5 yr → compare the AUC against the reported 0.720.

Out of scope (not attempted, and why)

  • Model training (LASSO-Cox coefficient estimation): done on TCGA-LAML (+GTEx merge via inSilicoMerging/ComBat), not GSE12417. We take the published coefficients as given and validate them on the assigned cohort — we do not refit. Out of this RU's data scope.
  • Consensus clustering (C1/C2 m5C patterns), 17-regulator DE vs normal: needs TCGA+GTEx merged matrix (not GSE12417). Out of data scope.
  • Immune deconvolution (EPIC/ESTIMATE), GSVA pathways, scRNA cell-type maps, miRNA(ENCORI) network, GDSC drug-IC50 correlations, cBioPortal mutations: each uses other datasets (TCGA, 4 scRNA GEOs, GDSC, cBioPortal). Out of data scope.
  • RT-qPCR validation (ALYREF/DNMT3B/NSUN2/NSUN5/TET1/TET3, n=4+4): wet-lab, non-computational. Out of scope by definition.
  • GSE37642 (second validation cohort, AUC 0.757): not the assigned accession; same method as C1 but skipped to stay within the assigned data (80/20).

Honesty notes

  • The paper does not state the exact preprocessing for signature transfer to the microarray cohort (z-score? raw? which probe-collapse?) nor the exact timepoint of the 0.720 AUC. We use the standard convention (per-gene z-score, mean-collapse multi-probe, report 1/3/5-yr AUC) and disclose this. Any AUC discrepancy may reflect that ambiguity, not a fabrication per se.
  • Whether the paper used GSE12417 GPL96 (163) or GPL570 (79) is unspecified; we run every platform with survival + the genes and report each.
Figures / tables: figure
C1
Reported
GSE12417 validation AUC = 0.720
Reproduced
1-yr timeROC AUC = 0.726 on the GPL570 sub-series (HG-U133 Plus 2.0, n=79, all 11 signature genes present); 0.515 on GPL96 (n=163, only 9/11 genes on array)
within tolerance
C2
Reported
11-gene risk score (median split) significantly stratifies overall survival in GSE12417 (high-risk worse)
Reproduced
GPL570: log-rank p=0.0044, HR(high vs low)=2.31 (95% CI 1.28-4.18), C-index 0.620, correct direction
exact

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 93/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟢2. Endpoint comparability
🟢3. Location of the main deviation
🟢4. Cause of the deviation
🟢5. Derivability / plausibility
🟢6. Severity of the deviation
🟢7. Core claim
🟢8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score -6

On the only GSE12417 sub-series carrying all 11 signature genes (GPL570, n=79), the published risk-score validation reproduces essentially 1:1: 1-yr timeROC AUC=0.726 vs reported 0.720 (Δ=0.006) and a significant prognostic split (log-rank p=0.0044, HR=2.31, high-risk worse OS). The deviation is negligible and on our methodology side (self-chosen sub-series, timepoint, and z-score transfer the paper left unstated), not the authors'. The GPL96 null is fully explained by 2 missing genes on the older array and is not a contradiction. Both numbers are derivable from public data + published coefficients, so there is no fabrication concern.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

170.4 k
tokens (I/O) · 10.8 M incl. cache
20 min
runtime · 0 CPU-h
0.4 GB
peak RAM
2
HPC jobs
hummel
machine