Integrated drug resistance and leukemic stemness gene-expression scores predict outcomes in large cohort of over 3500 AML patients from 10 trials.
The main results reproduced, with only marginal, non-material deviations.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Same input data as the authors
- 🟡Reported values were only indirectly comparable
- 🟡A deviation arose in the data or preprocessing
- 🟡A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The deviation was non-trivial in magnitude
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
DESCRIBED WELL ENOUGH: YES for the in-scope piece. The code repo (Abdelrahman-Elsayed/kit-nfold-cv-glmnet @09e3513) is a generic third-party LASSO/Cox 1000x-CV tool (P16 application case). The ADE-RS5 score equation is fully specified (Eq.1, 5 genes) and the assigned public validation cohort GSE71014 (104 CN-AML, Illumina HumanHT-12 V4.0/GPL10558) has per-sample OS-months+event in GEO. OUTCOME: PARTIAL. On «our HPC» («job», R 4.2.3 + GEOquery 2.66.0 + survival 3.7), the published ADE-RS5 was computed on all 104 GSE71014 samples (36 OS events) and tested vs OS. All 5 genes mapped to probes; ABCC1 is absent from GEO's curated AnnotGPL for GPL10558 and was recovered from the Illumina manufacturer platform table (ABCC1->ILMN_1802404). RESULT: higher ADE-RS5 associates with WORSE OS (HR>1) in all three dichotomizations (continuous HR=1.75 p=0.33; 40/60 HR=1.61 p=0.15; median HR=1.83 p=0.078/logrank 0.074) — DIRECTION matches the paper, but the association is NOT statistically significant within this single small cohort. This is consistent with (not contradictory to) the paper's significance claim, which is for the POOLED meta-analysis across 10 cohorts; the GSE71014-specific HR is only inside a forest-plot figure, so we grade directionality+significance rather than a numeric HR match. No fabrication flags. NOT ATTEMPTED (and why): ADE-RS5 derivation (AML02 discovery expression+EFS restricted); pLSC6 / integrated 4-group score (pLSC6 coefficients are external, ref.7, not in this paper); the other 9 validation cohorts (mix of dbGaP/on-request restricted + other GEO). CAVEAT: dichotomization here is within-cohort percentile (paper used rpart cutpoints fit on the restricted discovery cohort), and one representative probe per gene was used (highest-mean collapse).
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 50assessed: 2026-06-16 ⛓ dd34048f7f72
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-16
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-16no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetComprehensive transcriptomic evaluation of genes with pharmacological relevance to ara-C, daunorubicin, and etoposide (ADE chemotherapy) can be used to derive a drug-resistance gene-expression score that, alone or combined with a leukemic stemness score (pLSC6), predicts treatment outcomes in AML patients.
- ★ A 5-gene ADE-Resistance Score (ADE-RS5), derived via LASSO regression from 67 pharmacologically relevant genes, predicts MRD positivity, EFS and OS in pediatric AML. finding
- ★ ADE-RS5 was developed using LASSO penalized Cox regression on mRNA expression of 67 candidate genes in the AML02 discovery cohort. method
- ★ Integrating ADE-RS5 with the previously defined pLSC6 leukemic stemness score into four patient groups improves prognostic stratification. finding
- ★ The integrated pLSC6/ADE-RS5 score group remains an independent predictor of EFS and OS after multivariable adjustment for risk group, WBC, FLT3-ITD status and age. finding
- ★ ADE-RS5 and pLSC6/ADE-RS5 integrated scores were validated in >3500 pediatric and adult AML patients across 10 independent cohorts. resource
- pLSC6 stemness score is significantly associated with established high-risk AML features (risk group, cytogenetics, FLT3 status). finding
- High ADE-RS5 score is associated with higher induction I/II MRD positivity in validation cohorts. finding
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| LASSO penalized Cox regression on mRNA expression | pediatric AML patients, AML02 trial (discovery cohort, n=163) | none | event-free survival (EFS) association with expression of 67 candidate genes | — |
| Gene-expression-based risk scoring (ADE-RS5) with recursive partitioning | pediatric AML, AML02 discovery cohort (n=163) | none | MRD1 positivity, EFS, OS | — |
| Integrated transcriptomic score analysis (pLSC6 + ADE-RS5, four-group classification) | pediatric AML, AML02 discovery cohort (n=163) | none | MRD1 positivity, EFS, OS across four score groups | — |
| Multivariable Cox regression analysis | pediatric AML, AML02 discovery cohort (n=163) | none | EFS/OS adjusted for risk group, diagnostic WBC count, FLT3 status, age | — |
| Transcriptomic score validation (ADE-RS5, pLSC6, integrated score) | combined pediatric AML validation cohorts (4 trials, n=1861) | none | EFS, OS | — |
| Transcriptomic score validation (ADE-RS5, pLSC6, integrated score) | combined adult AML validation cohorts (5 trials, n=1669) | none | EFS, OS | — |
- ▲ Each unit increase in ADE-RS5 associated with increased EFS event rate in single-predictor Cox model (discovery cohort) 7.32-fold, p<0.00001, 95% CI 3.75-14.28
- ▲ High ADE-RS5 predicts higher MRD1 positivity in discovery cohort OR=2.39, 95% CI 1.23-4.63, p=0.013
- ▼ High ADE-RS5 associated with lower EFS and OS in discovery cohort EFS HR=4.07 p<0.0001; OS HR=4.54 p<0.0001
- ▼ Integrated Group 4 (high pLSC6 + high ADE-RS5) had worst EFS/OS vs Group 1 (both low) in discovery cohort EFS HR=8.89 p<0.0001; OS HR=12.68 p<0.0001
- ▲ Integrated score group remained independent predictor in multivariable analysis (discovery cohort) EFS: Group2 vs1 HR=4.68 p<0.001; Group3 vs1 HR=3.22 p=0.01; Group4 vs1 HR=7.26 p<0.001
- ▲ ADE-RS5 and pLSC6 showed consistent significant EFS/OS association in combined pediatric validation cohort ADE-RS5 EFS HR=1.38, OS HR=1.6; pLSC6 EFS HR=1.9, OS HR=2.1 (all p<0.001)
- ▼ 5-year EFS and OS were markedly lower in high vs low pLSC6/ADE-RS5 groups in both pediatric and adult validation cohorts e.g. pediatric Group1 vs Group4 5-yr EFS 57.76% vs 29.27%, p<0.0001
- fold_change 7.32-fold increase in EFS event rate per unit ADE-RS5 (single-predictor Cox regression, discovery cohort)
- pvalue p=0.013 (ADE-RS5 high group vs MRD1 positivity, discovery cohort)
- pvalue p<0.0001 (ADE-RS5 high group vs EFS and OS, discovery cohort)
- fold_change HR=7.26, p<0.001 (multivariable analysis, integrated Group4 vs Group1, EFS)
- fold_change HR=9.72, p<0.001 (multivariable analysis, integrated Group4 vs Group1, OS)
- count n=163 (AML02 discovery/model-development cohort)
- count ~3634 patients (1861 pediatric, 1773 adult) across 10 cohorts (combined validation cohorts)
- fold_change ADE-RS5 EFS HR=1.38, OS HR=1.6 (p<0.001) (combined pediatric validation cohort)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
The paper developed a five-gene drug resistance score (ADE-RS5) using LASSO-penalized Cox regression with extensive leave-10%-out cross-validation in a discovery cohort of 163 pediatric AML patients, then validated both ADE-RS5 and an integrated four-group score (combining ADE-RS5 with a pre-existing stemness score, pLSC6) across 10 independent cohorts totaling ~3,634 patients. Survival outcomes (EFS, OS) were compared between score groups using Cox proportional hazards models (univariate and multivariable), with MRD1 positivity assessed via logistic regression. Results were reported as hazard ratios or odds ratios with 95% confidence intervals and p-values; categorical covariate distributions across groups were summarized with p-values consistent with chi-square testing. No explicit multiplicity correction across the many comparisons, endpoints, and cohorts was described.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| LASSO-penalized Cox proportional hazards regression | Model development: 67 pharmacological genes vs. EFS in AML02 discovery cohort; gene stability threshold set at ≥950/1000 leave-10%-out cross-validation replications | 163 | not stated |
| Leave-10%-out cross-validation (1000 replications) | Internal validation of LASSO gene selection stability in AML02 | 163 | na |
| Univariate Cox proportional hazards regression | ADE-RS5 (continuous) vs. EFS in AML02 discovery cohort (HR=7.32, 95% CI=3.75–14.28); ADE-RS5 group vs. EFS and OS; pLSC6 group vs. EFS and OS in each cohort | 163 (discovery); 1861 pediatric and 1669 adult (validation combined) | not stated |
| Recursive partitioning | Dichotomization of continuous ADE-RS5 score into low (n=98, 60%) vs. high (n=65, 40%) groups in AML02 | 163 | na |
| Logistic regression | ADE-RS5 score group vs. MRD1 positivity in AML02 (OR=2.39, 95% CI=1.23–4.63, p=0.013) | 163 | not stated |
| Multivariable Cox proportional hazards regression | ADE-RS5 score group and integrated pLSC6/ADE-RS5 four-group score vs. EFS and OS, adjusting for diagnostic risk group, WBC count, FLT3 status, and age in AML02 and validation cohorts | 163 (discovery); combined validation cohorts (n not explicitly restated for MVA) | not stated |
| Log-rank test (implied by Kaplan-Meier comparisons with p-values) | Group-by-group survival comparisons (EFS and OS) across all discovery and validation cohorts | null | not stated |
| Chi-square or Fisher's exact test (implied) | Categorical patient characteristics (gender, age group, risk group, cytogenetics, FLT3, WBC, MRD, induction response) compared across pLSC6, ADE-RS5, and integrated score groups in Table 1 | 1861 pediatric and 1669 adult | not stated |
-
Continuous ADE-RS5 scores were dichotomized into two groups using recursive partitioning↳ Could also: The continuous score could be retained as-is in Cox models, or patients could be divided into tertiles or quartiles — Dichotomization discards within-group variation and can inflate apparent effect sizes; using the score as a continuous predictor or in ordered quantile groups preserves the full dose-response relationship and avoids the multiple-testing concern inherent in data-driven cutpoint selection
-
Ten validation cohorts were each analyzed separately, with results summarized narratively across cohorts↳ Could also: A random-effects meta-analysis pooling cohort-specific HRs could be applied across the 10 validation datasets — Meta-analytic pooling yields a single quantitative summary estimate with a formal heterogeneity statistic (I²/Q-test), making between-cohort consistency directly visible and quantified rather than assessed visually or narratively
-
Multiple endpoints (EFS, OS, MRD1 positivity), multiple score comparisons, and multiple cohorts were tested without a described multiplicity correction↳ Could also: A Benjamini-Hochberg false discovery rate (FDR) correction or Bonferroni adjustment could be applied within families of related tests (e.g., all pairwise group contrasts within a cohort for a given endpoint) — When many null hypotheses are tested simultaneously, the family-wise type-I error rate rises; a pre-specified correction procedure helps readers calibrate the expected proportion of false positives among the reported findings
-
Variable selection for ADE-RS5 used LASSO-penalized Cox regression↳ Could also: Elastic net (combining L1 and L2 penalties) or random survival forests could also be used for variable selection from the 67-gene panel — LASSO tends to select one variable from correlated groups arbitrarily; elastic net can retain grouped correlated predictors, and random survival forests are nonparametric and capture nonlinear or interaction effects — both are standard alternatives when gene-expression predictors are expected to be correlated
-
Five-year EFS and OS rates in Table 1 are presented with values in parentheses consistent with standard errors of Kaplan-Meier estimates↳ Could also: 95% confidence intervals around the Kaplan-Meier point estimates could be reported instead of or alongside standard errors — CIs are more directly interpretable as the plausible range for the true population rate and are the conventional format recommended by reporting guidelines such as STROBE and CONSORT; SEs require an additional mental step (±1.96 × SE) to interpret as an uncertainty range
-
MRD1 positivity was analyzed with logistic regression treating it as a binary outcome↳ Could also: Competing-risks regression (e.g., Fine-Gray subdistribution hazard model) or time-to-event analysis of MRD could also be used if time-to-MRD-assessment varied across patients — If time to MRD measurement differed across patients or if early death precluded MRD assessment, a competing-risks or survival framework for MRD would account for informative censoring that binary logistic regression does not address
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
The published ADE-RS5 equation was applied verbatim to the fully public CN-AML cohort GSE71014 (n=104, 36 events) and reproduces the paper's worse-OS direction (HR>1 in all three cutpoints: continuous 1.752, 40/60 1.613, median 1.827). The single deviation — non-significance in this cohort — sits on our methodology / data-availability side: the discovery rpart cutpoints are restricted (forcing self-chosen percentile splits) and the cohort-specific HR exists only inside a forest-plot figure, so an exact numeric match is impossible. This is not an authors' defect and shows no fabrication signal; a non-significant single cohort within a significant pooled 10-cohort meta-analysis is normal. Overall a solid partial reproduction with explainable deviations.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.