5-methylcytosine RNA modification regulators-based patterns and features of immune microenvironment in acute myeloid leukemia.
The main results reproduced: recomputed values matched the published ones within tolerance.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Reported values were directly comparable
- ✓No relevant deviation in data/preprocessing
- ✓No authors-side cause for any deviation
- ✓Reported values are derivable from the shared data
- ✓Any deviation was negligible
- ✓The central claim held under reproduction
- ✓Overall, the reproduction was clean
- 🟡Could not use the authors’ exact input data
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Described well enough to reproduce, and 1:1 on the headline validation claim. The paper's named code is the generic survival R package (third-party tool; the authors shipped no pipeline) and the assigned data is GSE12417. I took the paper's fully-published 11-gene m5C-regulator LASSO-Cox risk-score formula (coefficients verbatim from Results), applied it to GSE12417 (per-gene z-score signature transfer, mean-collapse multi-probe, median split), and recomputed survival statistics with the survival package + timeROC on «our HPC». GSE12417 has three sub-series; only GPL570 (HG-U133 Plus 2.0, n=79) carries all 11 genes, where the 1-year time-dependent AUC = 0.726 vs the reported 0.720 (delta 0.006, essentially exact) and the high/low-risk Kaplan-Meier separation is significant (log-rank p=0.0044, HR 2.31, CI excludes 1, high-risk worse OS) — the central GSE12417 validation claim reproduces. The GPL96 sub-series (n=163) gives a null result, fully explained by the older HG-U133A array lacking 2 of the 11 genes (NSUN4, TET2), not a contradiction. No fabrication indicated: the headline AUC is independently recovered from public data + published coefficients. NOT attempted (out of this RU's data/code scope, per scope.md): LASSO model TRAINING (done on TCGA-LAML+GTEx, not GSE12417 -> we validated, did not refit); consensus clustering / 17-regulator differential expression (TCGA+GTEx); immune deconvolution EPIC/ESTIMATE, GSVA pathways, scRNA cell-type maps, ENCORI miRNA network, GDSC drug-IC50, cBioPortal mutations (each uses other datasets); RT-qPCR validation (wet-lab, n=4+4); and GSE37642 (second validation cohort, not the assigned accession). Caveats: the paper does not state the exact AUC timepoint, which GSE12417 sub-series, or the signature-transfer preprocessing; standard conventions were used and disclosed.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 93assessed: 2026-06-15 ⛓ a7d8934e1e3f
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-15
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusThe role of 5-methylcytosine (m5C) RNA modification regulators in acute myeloid leukemia (AML) is unknown; this study tests whether the expression, genetic/epigenetic alterations, and patterns of 17 m5C regulators define molecular subtypes, predict prognosis, and shape the immune microenvironment in AML.
- ★ Most of the 17 m5C regulators are differentially expressed between AML and normal samples and correlate with AML prognosis. finding
- ★ Two m5C modification subtypes (C1, C2) and high-/low-risk subgroups based on m5C regulator expression differ significantly in prognosis and immune cell infiltration. finding
- ★ A 6-gene (m5C-regulator) prognostic risk model predicts AML survival and was validated in GSE12417 and GSE37642 datasets. resource
- ★ Methylation status of certain m5C regulators (e.g., DNMT3A, DNMT3B, TRDMT1) affects AML patient survival. finding
- ★ m5C regulators (notably DNMT3A) correlate with miRNA expression and drug IC50 values in AML. finding
- DNMT3A is the most frequently mutated m5C regulator in AML (19%), mainly missense mutations, but the mutation does not significantly affect survival. finding
- qPCR validation was used to confirm differential expression of m5C regulators between normal and AML samples. method
- m5C regulators correlate with immune cell infiltration (CD4/CD8 T-cells, macrophages) and stromal score in the AML microenvironment. mechanism
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Bulk RNA-seq differential expression (merged GTEx+TCGA, batch-corrected) | AML patient and normal samples (TCGA, GTEx) | none | mRNA expression of 17 m5C regulators | — |
| Mutation annotation analysis | AML patients (cBioPortal, 4 datasets) | none | mutation type and frequency of m5C regulators | cBioPortal |
| Copy number variation analysis | AML patients (TCGA) | none | homozygous/heterozygous CNV and CNV-mRNA correlation | — |
| DNA methylation analysis | AML patients | none | methylation level vs expression and survival of m5C regulators | — |
| Consensus clustering / LASSO Cox regression / ROC modeling | AML patient cohort (TCGA), validation in GSE12417 and GSE37642 | none | molecular subtypes, risk score, survival, AUC | — |
| Immune cell infiltration deconvolution (EPIC, ESTIMATE) | AML patient tumor microenvironment | none | immune cell proportions, stromal score and correlation with m5C regulators | EPIC, ESTIMATE algorithms |
| Drug sensitivity, RNAss, and miRNA correlation analysis; GSVA pathway analysis | AML patient data | none | drug IC50 correlation, RNA stemness score, miRNA expression, KEGG/hallmark pathway scores | GSVA |
| scRNA-seq cell-type-specific expression; qPCR validation; peripheral blood smear/genome sequencing | AML scRNA-seq datasets (GSE116256, GSE135851, GSE147989, GSE154109); normal vs AML samples; two AML patients | none | cell-type-specific expression; m5C regulator expression by qPCR; morphology and DNMT3A mutation | qPCR |
- ▲ Nine m5C regulators (NOP2, NSUN3, NSUN6, NSUN7, DNMT1, DNMT3A, TRDMT1, TET2, TET3) highly expressed in AML tumor group
- – Low-risk score group survived longer than high-risk group; prognostic model AUC = 0.744 AUC=0.744
- – Model validated with AUC 0.720 (GSE12417) and 0.757 (GSE37642), low-risk group better prognosis AUC=0.720 and 0.757
- – DNMT3A mutation rate highest at 19% (mostly missense); TET2 ~9% (mainly truncating) 19%; ~9%
- – High DNA methylation of DNMT3A and TRDMT1 shortened overall survival; high DNMT3B methylation lengthened survival
- ▲ DNMT3A positively correlated with Nelarabine drug sensitivity (IC50) R=0.638, p<0.001
- ▲ DNMT3A positively correlated with hsa-mir-146a expression R=0.65, P=1.6e-19
- – CD4 T-cells enriched in low-risk group; CD8 T-cells and macrophages enriched in high-risk group
- other AUC=0.744 (ROC of m5C prognostic model in TCGA)
- other AUC=0.720 (model validation in GSE12417)
- other AUC=0.757 (model validation in GSE37642)
- count 19% (DNMT3A mutation frequency in AML)
- correlation R=0.638, p<0.001 (DNMT3A vs Nelarabine IC50)
- correlation R=−0.51, FDR=1e-12 (NSUN7 methylation vs expression)
- correlation R=−0.36, FDR=1.9e-06 (DNMT3A methylation vs expression)
- correlation R=0.65, P=1.6e-19 (DNMT3A vs hsa-mir-146a)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This retrospective bioinformatics study characterized 17 m5C RNA modification regulators in AML using merged TCGA and GTEx expression data, mutation/CNV data from cBioPortal, and GEO cohorts. The primary analytical strategy combined unsupervised consensus clustering (ConsensusClusterPlus) to define patient subtypes, LASSO-penalized Cox regression to build a prognostic risk score, and Kaplan-Meier analysis to compare survival between risk groups. Secondary analyses used Pearson and Spearman correlations to link regulator expression with immune infiltration (EPIC/ESTIMATE), drug sensitivity (IC50), and miRNA networks; external GEO cohorts (GSE12417, GSE37642) and qPCR served as validation steps.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Differential expression test (specific test not named; significance reported as threshold asterisks via boxplots) | AML vs. normal sample expression for all 17 m5C regulators after ComBat batch correction (Figure 2A) and clinical subgroup comparisons across age, CR, FAB, sex, survival status, treatment status (Figure 3A–F) | — | not stated |
| Pearson correlation | Inter-regulator expression correlation matrix (Figure 2B); drug IC50 vs. regulator expression with exact R and p reported for selected pairs (Figure 9B,C); miRNA vs. regulator expression and immune cell vs. miRNA (Figure 9D–G); regulator vs. stromal score (Figure 8G); pathway scores vs. regulators (Figure 10B,D) | — | not stated |
| Univariate Cox proportional hazards regression | Association of each individual m5C regulator expression level with overall survival (Figure 2C) | — | not stated |
| Kaplan-Meier survival analysis (log-rank test implied by p-value annotations on KM curves) | DNMT3A mutation vs. overall survival (Figure 4D); CNV deletions vs. survival (Figure 4G); methylation-high vs. -low for DNMT3A, DNMT3B, TRDMT1 (Figure 5C,D); high- vs. low-risk group in TCGA discovery set, GSE12417, and GSE37642 validation sets (Figures 6E, 7A,B) | — | not stated |
| Spearman correlation | Methylation status vs. mRNA expression of m5C regulators (Figure 5A); exact R and FDR reported for four representative genes: NSUN4 (R=−0.31, FDR=4.4e-05), NSUN7 (R=−0.51, FDR=1e-12), DNMT3A (R=−0.36, FDR=1.9e-06), DNMT1 (R=−0.20, FDR=9.8e-03) (Figure 5B) | — | not stated |
| LASSO-penalized Cox regression | Feature selection and coefficient estimation for the m5C prognostic risk model; minimum-criteria used to identify 11-gene signature from candidate differentially expressed genes in the two consensus clusters (Figure 6C,D) | — | not stated |
| Time-dependent ROC curve / AUC | Prognostic model discrimination: AUC=0.744 in TCGA training set (Figure 6F), AUC=0.720 in GSE12417 (Figure 7C), AUC=0.757 in GSE37642 (Figure 7D) | — | na |
| GSVA (gene set variation analysis) | Enrichment of KEGG and Hallmark pathway gene sets compared between high- and low-risk groups (Figure 10A,C) | — | not stated |
-
Multiple differential expression comparisons across clinical subgroups (CR, FAB, treatment, sex) were performed with significance denoted only by threshold asterisks, without naming the underlying statistical test↳ Could also: For multi-level group comparisons, a Kruskal-Wallis test (non-parametric) or one-way ANOVA (parametric) followed by a post-hoc correction such as Dunn's test or Tukey HSD could also be applied and would make the inference structure explicit — Specifying the omnibus test and post-hoc procedure clarifies what null hypothesis is being rejected and how the family-wise error rate is controlled across the many regulators and subgroup combinations tested
-
The prognostic risk score was derived from LASSO Cox regression in the TCGA discovery cohort and then evaluated directly in external GEO validation cohorts without an internal cross-validation step↳ Could also: Internal k-fold cross-validation (e.g., 10-fold CV) or bootstrap optimism-correction within the TCGA cohort could also be applied before proceeding to external validation — Internal validation quantifies overfitting introduced during the penalized selection step and complements external validation by distinguishing optimism from transportability
-
Survival differences between high- and low-risk groups were assessed with Kaplan-Meier curves (unadjusted log-rank comparison)↳ Could also: A multivariable Cox regression including established AML prognostic covariates (age, cytogenetic risk, FAB subtype) alongside the m5C risk score could also evaluate whether the score provides independent prognostic information — Unadjusted log-rank tests do not account for confounders; a multivariable model would clarify the incremental prognostic contribution of the m5C risk score beyond known clinical factors
-
Immune cell infiltration proportions were estimated using a single deconvolution algorithm (EPIC)↳ Could also: Applying additional deconvolution methods such as CIBERSORT-ABS, MCP-counter, or xCell to the same expression data and comparing concordance across methods is also common practice — Different deconvolution algorithms make different modeling assumptions and show varying performance across tumor types and data platforms; cross-algorithm concordance strengthens confidence in the reported immune-infiltration patterns
-
Methylation-expression correlations are reported with FDR-adjusted values for four representative regulators (Figure 5B), while most other large families of statistical tests do not report a correction method↳ Could also: Applying a consistent FDR or Bonferroni correction across the full families of miRNA-regulator correlations, drug-IC50 correlations, and per-regulator differential expression tests could also be used to control false-discovery rates across these multiple comparisons — Consistent multiplicity adjustment across test families aids in interpreting which associations are likely to replicate and aligns with current reporting standards for large-scale omics correlation analyses
-
Results of group comparisons (e.g., immune cell proportions between C1/C2 or high-/low-risk) are reported with p-value thresholds but without quantitative effect size measures↳ Could also: Reporting fold-differences in estimated immune cell proportions, Cohen's d, or hazard ratios with 95% confidence intervals alongside p-values would also characterize the magnitude of observed differences — Effect size estimates allow readers to assess clinical or biological relevance independently of sample size, and confidence intervals convey estimation uncertainty that p-value thresholds alone do not capture
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
- Comprehensive analysis of m<sup>6</sup>A methy... L1 No data access
- Comprehensive analysis of m<sup>6</sup>A methy... L1 No data access
Downstream reach in the literature
179 downstream papers · 2 datasetsHow widely the datasets deposited by this paper are reused across the whole literature (Europe PMC), beyond our assessed set. This is a factual dependency map — reusing a public dataset is normal, good science. It is not a judgement on the downstream papers; the only verdict here is this paper's own, with its cited rationale.
- Data-Driven Phenotypic Dissection of AML Reveals Pro... 2015 · 1,971 cites
- PrognoScan: a new database for meta-analysis of the... 2009 · 772 cites
- Chemotherapy-Resistant Human Acute Myeloid Leukemia... 2017 · 728 cites
- Targeting the RNA m<sup>6</sup>A Reader... 2019 · 400 cites
- An 86-probe-set gene-expression signature predicts s... 2008 · 335 cites
- Association of a leukemic stem cell gene expression... 2010 · 311 cites
- The M2 macrophage marker <i>CD206</i>: a novel progn... 2020 · 171 cites
- Identification of a 24-gene prognostic signature tha... 2013 · 165 cites
- miR-22 has a potent anti-tumour role with therapeuti... 2016 · 118 cites
- IL-8 as mediator in the microenvironment-leukaemia n... 2015 · 59 cites
- Longitudinal single-cell profiling of chemotherapy r... 2023 · 59 cites
- TIGIT blockade repolarizes AML-associated TIGIT<sup>... 2022 · 55 cites
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-38277218
Paper: Ding Y et al. (2024) 5-methylcytosine RNA modification regulators-based patterns and features of immune microenvironment in acute myeloid leukemia. Aging (Albany NY). PMID 38277218 · PMCID PMC10911375 · DOI 10.18632/aging.205484
Assigned code artifact: https://github.com/therneau/survival — the generic
survival R package (Terry Therneau). This is NOT the authors' own analysis
code; it is the third-party statistics tool the paper used for Cox / Kaplan–Meier /
log-rank survival analysis. Per the brief's P16 rule, applying this named tool to
the paper's named data is an equally valid reproduction. The authors did not
release their own pipeline code (Aging papers typically ship none).
Assigned data: geo:GSE12417 — a public AML microarray cohort (Metzeler et al.,
Blood 2008) used in the paper as an independent validation cohort for the
prognostic risk model. Subseries: GPL96 (HG-U133A, 163 CN-AML, training in the
original Metzeler paper), GPL97 (HG-U133B, 163), GPL570 (HG-U133 Plus 2.0, 79).
In scope (pipeline-derived, reproducible with GSE12417 + survival)
| # | Reported result | Where | Reproducible? |
|---|---|---|---|
| C1 | 11-gene m5C risk-score validation AUC = 0.720 in GSE12417 | Results / model-validation figure | YES — formula + coefficients fully published; GSE12417 public. PRIMARY TARGET. |
| C2 | High- vs low-risk Kaplan–Meier separation (median split) is prognostic in GSE12417 | same | YES — KM + log-rank via survival; the paper's central claim for this cohort. |
The published risk-score formula (11 genes, LASSO-Cox coefficients):
RiskScore = -0.3202*NSUN3 -0.2526*NSUN4 +0.0372*NSUN5 -0.3514*NSUN6 +0.3160*NSUN7
+0.3244*DNMT1 -0.3648*DNMT3A +0.3476*DNMT3B -0.0334*TET2 -0.0897*TET3
+0.2137*ALKBH1
Apply to GSE12417 expression (per-gene z-score, standard cross-platform signature
transfer), median-split into high/low risk, then survival/timeROC:
KM + log-rank p, Cox HR, C-index, and time-dependent ROC-AUC at 1/3/5 yr → compare
the AUC against the reported 0.720.
Out of scope (not attempted, and why)
- Model training (LASSO-Cox coefficient estimation): done on TCGA-LAML (+GTEx merge via inSilicoMerging/ComBat), not GSE12417. We take the published coefficients as given and validate them on the assigned cohort — we do not refit. Out of this RU's data scope.
- Consensus clustering (C1/C2 m5C patterns), 17-regulator DE vs normal: needs TCGA+GTEx merged matrix (not GSE12417). Out of data scope.
- Immune deconvolution (EPIC/ESTIMATE), GSVA pathways, scRNA cell-type maps, miRNA(ENCORI) network, GDSC drug-IC50 correlations, cBioPortal mutations: each uses other datasets (TCGA, 4 scRNA GEOs, GDSC, cBioPortal). Out of data scope.
- RT-qPCR validation (ALYREF/DNMT3B/NSUN2/NSUN5/TET1/TET3, n=4+4): wet-lab, non-computational. Out of scope by definition.
- GSE37642 (second validation cohort, AUC 0.757): not the assigned accession; same method as C1 but skipped to stay within the assigned data (80/20).
Honesty notes
- The paper does not state the exact preprocessing for signature transfer to the microarray cohort (z-score? raw? which probe-collapse?) nor the exact timepoint of the 0.720 AUC. We use the standard convention (per-gene z-score, mean-collapse multi-probe, report 1/3/5-yr AUC) and disclose this. Any AUC discrepancy may reflect that ambiguity, not a fabrication per se.
- Whether the paper used GSE12417 GPL96 (163) or GPL570 (79) is unspecified; we run every platform with survival + the genes and report each.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
On the only GSE12417 sub-series carrying all 11 signature genes (GPL570, n=79), the published risk-score validation reproduces essentially 1:1: 1-yr timeROC AUC=0.726 vs reported 0.720 (Δ=0.006) and a significant prognostic split (log-rank p=0.0044, HR=2.31, high-risk worse OS). The deviation is negligible and on our methodology side (self-chosen sub-series, timepoint, and z-score transfer the paper left unstated), not the authors'. The GPL96 null is fully explained by 2 missing genes on the older array and is not a contradiction. Both numbers are derivable from public data + published coefficients, so there is no fabrication concern.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.