Corpus 1,286 assessed · 1,187 scored · 648 reproduced ≥75 · 174 flagged ·∅ 73.9/100
← New search

5-methylcytosine RNA modification regulators-based patterns and features of immune microenvironment in acute myeloid leukemia.

Aging (Albany NY) · 2024
L1 93/100 PQI 88
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score -6
✓ What held up
  • Reported values were directly comparable
  • No relevant deviation in data/preprocessing
  • No authors-side cause for any deviation
  • Reported values are derivable from the shared data
  • Any deviation was negligible
  • The central claim held under reproduction
  • Overall, the reproduction was clean
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
How its reproducibility compares
93/100
Reproducibility score
1.1 SD above mean
vs. all fields · 1187 studies
🎯 Scores higher than 85% of all assessed papers rank 155 of 1187 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough to reproduce, and 1:1 on the headline validation claim. The paper's named code is the generic survival R package (third-party tool; the authors shipped no pipeline) and the assigned data is GSE12417. I took the paper's fully-published 11-gene m5C-regulator LASSO-Cox risk-score formula (coefficients verbatim from Results), applied it to GSE12417 (per-gene z-score signature transfer, mean-collapse multi-probe, median split), and recomputed survival statistics with the survival package + timeROC on «our HPC». GSE12417 has three sub-series; only GPL570 (HG-U133 Plus 2.0, n=79) carries all 11 genes, where the 1-year time-dependent AUC = 0.726 vs the reported 0.720 (delta 0.006, essentially exact) and the high/low-risk Kaplan-Meier separation is significant (log-rank p=0.0044, HR 2.31, CI excludes 1, high-risk worse OS) — the central GSE12417 validation claim reproduces. The GPL96 sub-series (n=163) gives a null result, fully explained by the older HG-U133A array lacking 2 of the 11 genes (NSUN4, TET2), not a contradiction. No fabrication indicated: the headline AUC is independently recovered from public data + published coefficients. NOT attempted (out of this RU's data/code scope, per scope.md): LASSO model TRAINING (done on TCGA-LAML+GTEx, not GSE12417 -> we validated, did not refit); consensus clustering / 17-regulator differential expression (TCGA+GTEx); immune deconvolution EPIC/ESTIMATE, GSVA pathways, scRNA cell-type maps, ENCORI miRNA network, GDSC drug-IC50, cBioPortal mutations (each uses other datasets); RT-qPCR validation (wet-lab, n=4+4); and GSE37642 (second validation cohort, not the assigned accession). Caveats: the paper does not state the exact AUC timepoint, which GSE12417 sub-series, or the signature-transfer preprocessing; standard conventions were used and disclosed.

💻 Code ↗ 🗄 Data: GSE12417

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 93
    assessed: 2026-06-15 ⛓ a7d8934e1e3f
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-15
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-09-19

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The paper investigates whether 5-methylcytosine (m5C) RNA modification regulators are differentially expressed and functionally/prognostically relevant in acute myeloid leukemia (AML), and whether their expression patterns can define molecular subtypes and a risk-stratifying prognostic model.

Core claims
  • m5C regulators are differentially expressed between AML and normal samples and correlate with AML prognosis finding
  • Methylation status of specific m5C regulators (e.g., DNMT3A, DNMT3B, TRDMT1) affects overall survival of AML patients finding
  • Two m5C modification subtypes and high-/low-risk subgroups defined by m5C regulator expression show significant differences in prognosis and immune cell infiltration finding
  • m5C regulators correlate with miRNA expression and drug IC50 sensitivity in AML finding
  • An 11-gene m5C regulator-based prognostic risk model predicts AML patient survival and was validated in independent GEO cohorts method
  • qPCR confirmed differential expression of m5C regulators between AML and normal samples method
  • DNMT3A has the highest mutation rate (19%) among m5C regulators in AML, predominantly missense mutations finding
  • DNMT3A mutation status does not significantly affect overall survival of AML patients finding
Experimental setups
Assay System Perturbation Readout Platform
bulk RNA-seq/microarray expression profiling TCGA AML patient samples vs GTEx normal samples none mRNA expression levels of 17 m5C regulators TCGA/GTEx databases (inSilicoMerging, SVA combat)
mutation and copy number variation analysis AML patients (4 integrated cBioPortal datasets) none mutation frequency, mutation type, CNV status of m5C regulators cBioPortal
DNA methylation array analysis TCGA AML patient samples none correlation of methylation beta values with mRNA expression and overall survival TCGA methylation data
consensus clustering and LASSO Cox regression TCGA AML cohort none m5C modification subtypes (C1/C2) and risk-score prognostic gene signature ConsensusClusterPlus, R LASSO Cox
prognostic model validation (survival/ROC analysis) GSE12417 and GSE37642 AML cohorts none overall survival and AUC of risk model GEO microarray datasets
immune cell deconvolution (EPIC) and stromal scoring (ESTIMATE) TCGA AML cohort none immune cell type proportions and stromal score correlated with m5C regulators and risk groups EPIC algorithm, ESTIMATE
single-cell RNA-seq analysis AML patient samples (GSE116256, GSE135851, GSE147989, GSE154109) none cell-type-specific expression of m5C regulators across immune/hematopoietic cell subsets GEO scRNA-seq datasets
qPCR AML vs normal patient/clinical samples none mRNA expression levels of m5C regulators
Key results
  • Nine m5C regulators (NOP2, NSUN3, NSUN6, NSUN7, DNMT1, DNMT3A, TRDMT1, TET2, TET3) were unexpectedly highly expressed in AML tumor samples vs normal
  • Higher expression of most m5C regulators (except TET1, TET2, TET3, ALYREF) was associated with worse prognosis by univariate Cox regression
  • DNMT3A and TRDMT1 methylation negatively correlated with their expression; high DNMT3A methylation was associated with shorter overall survival R=-0.36, FDR=1.9e-06 (DNMT3A)
  • High DNA methylation of DNMT3B was associated with significantly longer overall survival, opposite to DNMT3A/TRDMT1
  • DNMT3A had the highest mutation rate among m5C regulators in AML, mostly missense mutations 19%
  • DNMT3A mutation status did not significantly affect AML patient survival by Kaplan-Meier analysis
  • The 11-gene m5C regulator prognostic model showed good predictive performance in the TCGA cohort and was validated in two independent GEO datasets AUC=0.744 (TCGA); AUC=0.720 (GSE12417); AUC=0.757 (GSE37642)
  • CD4 T-cells were enriched in the low-risk group while CD8 T-cells and macrophages were enriched in the high-risk group
Key statistics
  • correlation R=-0.36, FDR=1.9e-06 (DNMT3A methylation vs mRNA expression)
  • correlation R=-0.51, FDR=1e-12 (NSUN7 methylation vs mRNA expression)
  • other 19% (DNMT3A mutation rate among AML patients, highest of all m5C regulators)
  • other AUC=0.744 (ROC AUC of m5C regulator prognostic risk model in TCGA cohort)
  • correlation R=0.638, p<0.001 (DNMT3A expression vs Nelarabine IC50)
  • correlation R=0.65, P=1.6e-19 (DNMT3A expression vs hsa-mir-146a expression)
  • correlation R=0.75, P=1.1e-23 (CD8 T-cell infiltration vs hsa-mir-6503 expression)
  • count 33 miRNAs significantly correlated with m5C regulators; DNMT3A positively regulated 13 (P<0.0001) and negatively regulated 7 (P<0.0001) (gene-miRNA correlation network analysis)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This retrospective bioinformatics study characterized 17 m5C RNA modification regulators in AML using merged TCGA and GTEx expression data, mutation/CNV data from cBioPortal, and GEO cohorts. The primary analytical strategy combined unsupervised consensus clustering (ConsensusClusterPlus) to define patient subtypes, LASSO-penalized Cox regression to build a prognostic risk score, and Kaplan-Meier analysis to compare survival between risk groups. Secondary analyses used Pearson and Spearman correlations to link regulator expression with immune infiltration (EPIC/ESTIMATE), drug sensitivity (IC50), and miRNA networks; external GEO cohorts (GSE12417, GSE37642) and qPCR served as validation steps.

Replicationunclear Sample sizeCohort sizes not stated in the available text; discovery cohort is TCGA AML; independent validation uses GSE12417 and GSE37642; cell-type-specific expression uses four scRNA-seq GEO datasets (GSE116256, GSE135851, GSE147989, GSE154109); qPCR validation in unspecified number of clinical samples GroupsAML vs. normal (TCGA+GTEx merged); consensus clusters C1 vs. C2; high-risk vs. low-risk (median prognostic score); clinical subgroups by age, CR status, FAB subtype, sex, survival status, treatment status; methylation-high vs. methylation-low per regulator Pairingunpaired Randomization/blindingnot stated Dispersionnone Confidence intervalsno Multiplicity correctionFDR (method not named; Benjamini-Hochberg is conventional but not stated) applied to Spearman methylation-expression correlations (Figure 5B); no correction stated for all other comparison families
Statistical tests used
Test Applied to n Assumptions
Differential expression test (specific test not named; significance reported as threshold asterisks via boxplots) AML vs. normal sample expression for all 17 m5C regulators after ComBat batch correction (Figure 2A) and clinical subgroup comparisons across age, CR, FAB, sex, survival status, treatment status (Figure 3A–F) not stated
Pearson correlation Inter-regulator expression correlation matrix (Figure 2B); drug IC50 vs. regulator expression with exact R and p reported for selected pairs (Figure 9B,C); miRNA vs. regulator expression and immune cell vs. miRNA (Figure 9D–G); regulator vs. stromal score (Figure 8G); pathway scores vs. regulators (Figure 10B,D) not stated
Univariate Cox proportional hazards regression Association of each individual m5C regulator expression level with overall survival (Figure 2C) not stated
Kaplan-Meier survival analysis (log-rank test implied by p-value annotations on KM curves) DNMT3A mutation vs. overall survival (Figure 4D); CNV deletions vs. survival (Figure 4G); methylation-high vs. -low for DNMT3A, DNMT3B, TRDMT1 (Figure 5C,D); high- vs. low-risk group in TCGA discovery set, GSE12417, and GSE37642 validation sets (Figures 6E, 7A,B) not stated
Spearman correlation Methylation status vs. mRNA expression of m5C regulators (Figure 5A); exact R and FDR reported for four representative genes: NSUN4 (R=−0.31, FDR=4.4e-05), NSUN7 (R=−0.51, FDR=1e-12), DNMT3A (R=−0.36, FDR=1.9e-06), DNMT1 (R=−0.20, FDR=9.8e-03) (Figure 5B) not stated
LASSO-penalized Cox regression Feature selection and coefficient estimation for the m5C prognostic risk model; minimum-criteria used to identify 11-gene signature from candidate differentially expressed genes in the two consensus clusters (Figure 6C,D) not stated
Time-dependent ROC curve / AUC Prognostic model discrimination: AUC=0.744 in TCGA training set (Figure 6F), AUC=0.720 in GSE12417 (Figure 7C), AUC=0.757 in GSE37642 (Figure 7D) na
GSVA (gene set variation analysis) Enrichment of KEGG and Hallmark pathway gene sets compared between high- and low-risk groups (Figure 10A,C) not stated
Approaches that could also have been used
  • Multiple differential expression comparisons across clinical subgroups (CR, FAB, treatment, sex) were performed with significance denoted only by threshold asterisks, without naming the underlying statistical test
    Could also: For multi-level group comparisons, a Kruskal-Wallis test (non-parametric) or one-way ANOVA (parametric) followed by a post-hoc correction such as Dunn's test or Tukey HSD could also be applied and would make the inference structure explicit — Specifying the omnibus test and post-hoc procedure clarifies what null hypothesis is being rejected and how the family-wise error rate is controlled across the many regulators and subgroup combinations tested
  • The prognostic risk score was derived from LASSO Cox regression in the TCGA discovery cohort and then evaluated directly in external GEO validation cohorts without an internal cross-validation step
    Could also: Internal k-fold cross-validation (e.g., 10-fold CV) or bootstrap optimism-correction within the TCGA cohort could also be applied before proceeding to external validation — Internal validation quantifies overfitting introduced during the penalized selection step and complements external validation by distinguishing optimism from transportability
  • Survival differences between high- and low-risk groups were assessed with Kaplan-Meier curves (unadjusted log-rank comparison)
    Could also: A multivariable Cox regression including established AML prognostic covariates (age, cytogenetic risk, FAB subtype) alongside the m5C risk score could also evaluate whether the score provides independent prognostic information — Unadjusted log-rank tests do not account for confounders; a multivariable model would clarify the incremental prognostic contribution of the m5C risk score beyond known clinical factors
  • Immune cell infiltration proportions were estimated using a single deconvolution algorithm (EPIC)
    Could also: Applying additional deconvolution methods such as CIBERSORT-ABS, MCP-counter, or xCell to the same expression data and comparing concordance across methods is also common practice — Different deconvolution algorithms make different modeling assumptions and show varying performance across tumor types and data platforms; cross-algorithm concordance strengthens confidence in the reported immune-infiltration patterns
  • Methylation-expression correlations are reported with FDR-adjusted values for four representative regulators (Figure 5B), while most other large families of statistical tests do not report a correction method
    Could also: Applying a consistent FDR or Bonferroni correction across the full families of miRNA-regulator correlations, drug-IC50 correlations, and per-regulator differential expression tests could also be used to control false-discovery rates across these multiple comparisons — Consistent multiplicity adjustment across test families aids in interpreting which associations are likely to replicate and aligns with current reporting standards for large-scale omics correlation analyses
  • Results of group comparisons (e.g., immune cell proportions between C1/C2 or high-/low-risk) are reported with p-value thresholds but without quantitative effect size measures
    Could also: Reporting fold-differences in estimated immune cell proportions, Cohen's d, or hazard ratios with 95% confidence intervals alongside p-values would also characterize the magnitude of observed differences — Effect size estimates allow readers to assess clinical or biological relevance independently of sample size, and confidence intervals convey estimation uncertainty that p-value thresholds alone do not capture
Software: R/inSilicoMerging · R/SVA (ComBat batch correction) · R/ConsensusClusterPlus · cBioPortal (mutation/CNV visualization) · EPIC (immune cell deconvolution) · ESTIMATE (stromal/immune scoring) · R/GSVA · LASSO Cox regression (likely R/glmnet; not explicitly named)

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
5
Impact: low
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

GSE116256 GEO in Methods (http://purl.org/orb/Methods)
also used by 1 paper:
GSE12417 GEO in Methods (http://purl.org/orb/Methods)
also used by 1 paper:
GSE37642 GEO in Methods (http://purl.org/orb/Methods)
also used by 1 paper:
GSE135851 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE147989 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE154109 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet

Downstream reach in the literature

179 downstream papers · 2 datasets

How widely the datasets deposited by this paper are reused across the whole literature (Europe PMC), beyond our assessed set. This is a factual dependency map — reusing a public dataset is normal, good science. It is not a judgement on the downstream papers; the only verdict here is this paper's own, with its cited rationale.

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-38277218

Paper: Ding Y et al. (2024) 5-methylcytosine RNA modification regulators-based patterns and features of immune microenvironment in acute myeloid leukemia. Aging (Albany NY). PMID 38277218 · PMCID PMC10911375 · DOI 10.18632/aging.205484

Assigned code artifact: https://github.com/therneau/survival — the generic survival R package (Terry Therneau). This is NOT the authors' own analysis code; it is the third-party statistics tool the paper used for Cox / Kaplan–Meier / log-rank survival analysis. Per the brief's P16 rule, applying this named tool to the paper's named data is an equally valid reproduction. The authors did not release their own pipeline code (Aging papers typically ship none).

Assigned data: geo:GSE12417 — a public AML microarray cohort (Metzeler et al., Blood 2008) used in the paper as an independent validation cohort for the prognostic risk model. Subseries: GPL96 (HG-U133A, 163 CN-AML, training in the original Metzeler paper), GPL97 (HG-U133B, 163), GPL570 (HG-U133 Plus 2.0, 79).

In scope (pipeline-derived, reproducible with GSE12417 + survival)

# Reported result Where Reproducible?
C1 11-gene m5C risk-score validation AUC = 0.720 in GSE12417 Results / model-validation figure YES — formula + coefficients fully published; GSE12417 public. PRIMARY TARGET.
C2 High- vs low-risk Kaplan–Meier separation (median split) is prognostic in GSE12417 same YES — KM + log-rank via survival; the paper's central claim for this cohort.

The published risk-score formula (11 genes, LASSO-Cox coefficients):

RiskScore = -0.3202*NSUN3 -0.2526*NSUN4 +0.0372*NSUN5 -0.3514*NSUN6 +0.3160*NSUN7
            +0.3244*DNMT1 -0.3648*DNMT3A +0.3476*DNMT3B -0.0334*TET2 -0.0897*TET3
            +0.2137*ALKBH1

Apply to GSE12417 expression (per-gene z-score, standard cross-platform signature transfer), median-split into high/low risk, then survival/timeROC: KM + log-rank p, Cox HR, C-index, and time-dependent ROC-AUC at 1/3/5 yr → compare the AUC against the reported 0.720.

Out of scope (not attempted, and why)

  • Model training (LASSO-Cox coefficient estimation): done on TCGA-LAML (+GTEx merge via inSilicoMerging/ComBat), not GSE12417. We take the published coefficients as given and validate them on the assigned cohort — we do not refit. Out of this RU's data scope.
  • Consensus clustering (C1/C2 m5C patterns), 17-regulator DE vs normal: needs TCGA+GTEx merged matrix (not GSE12417). Out of data scope.
  • Immune deconvolution (EPIC/ESTIMATE), GSVA pathways, scRNA cell-type maps, miRNA(ENCORI) network, GDSC drug-IC50 correlations, cBioPortal mutations: each uses other datasets (TCGA, 4 scRNA GEOs, GDSC, cBioPortal). Out of data scope.
  • RT-qPCR validation (ALYREF/DNMT3B/NSUN2/NSUN5/TET1/TET3, n=4+4): wet-lab, non-computational. Out of scope by definition.
  • GSE37642 (second validation cohort, AUC 0.757): not the assigned accession; same method as C1 but skipped to stay within the assigned data (80/20).

Honesty notes

  • The paper does not state the exact preprocessing for signature transfer to the microarray cohort (z-score? raw? which probe-collapse?) nor the exact timepoint of the 0.720 AUC. We use the standard convention (per-gene z-score, mean-collapse multi-probe, report 1/3/5-yr AUC) and disclose this. Any AUC discrepancy may reflect that ambiguity, not a fabrication per se.
  • Whether the paper used GSE12417 GPL96 (163) or GPL570 (79) is unspecified; we run every platform with survival + the genes and report each.
Figures / tables: figure
C1
Reported
GSE12417 validation AUC = 0.720
Reproduced
1-yr timeROC AUC = 0.726 on the GPL570 sub-series (HG-U133 Plus 2.0, n=79, all 11 signature genes present); 0.515 on GPL96 (n=163, only 9/11 genes on array)
within tolerance
C2
Reported
11-gene risk score (median split) significantly stratifies overall survival in GSE12417 (high-risk worse)
Reproduced
GPL570: log-rank p=0.0044, HR(high vs low)=2.31 (95% CI 1.28-4.18), C-index 0.620, correct direction
exact

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 93/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟢2. Endpoint comparability
🟢3. Location of the main deviation
🟢4. Cause of the deviation
🟢5. Derivability / plausibility
🟢6. Severity of the deviation
🟢7. Core claim
🟢8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score -6

On the only GSE12417 sub-series carrying all 11 signature genes (GPL570, n=79), the published risk-score validation reproduces essentially 1:1: 1-yr timeROC AUC=0.726 vs reported 0.720 (Δ=0.006) and a significant prognostic split (log-rank p=0.0044, HR=2.31, high-risk worse OS). The deviation is negligible and on our methodology side (self-chosen sub-series, timepoint, and z-score transfer the paper left unstated), not the authors'. The GPL96 null is fully explained by 2 missing genes on the older array and is not a contradiction. Both numbers are derivable from public data + published coefficients, so there is no fabrication concern.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at [email protected].

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

170.4 k
tokens (I/O) · 10.8 M incl. cache
20 min
runtime · 0 CPU-h
0.4 GB
peak RAM
2
HPC jobs
hummel
machine