Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Integrated drug resistance and leukemic stemness gene-expression scores predict outcomes in large cohort of over 3500 AML patients from 10 trials.

NPJ Precis Oncol · 2024
L1 50/100 PQI 83
Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +8
✓ What held up
  • Same input data as the authors
What did not (or only partly)
  • 🟡Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
50/100
Reproducibility score
1.4 SD below mean
vs. all fields · 1173 studies
🎯 Scores higher than 8% of all assessed papers rank 1026 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

DESCRIBED WELL ENOUGH: YES for the in-scope piece. The code repo (Abdelrahman-Elsayed/kit-nfold-cv-glmnet @09e3513) is a generic third-party LASSO/Cox 1000x-CV tool (P16 application case). The ADE-RS5 score equation is fully specified (Eq.1, 5 genes) and the assigned public validation cohort GSE71014 (104 CN-AML, Illumina HumanHT-12 V4.0/GPL10558) has per-sample OS-months+event in GEO. OUTCOME: PARTIAL. On «our HPC» («job», R 4.2.3 + GEOquery 2.66.0 + survival 3.7), the published ADE-RS5 was computed on all 104 GSE71014 samples (36 OS events) and tested vs OS. All 5 genes mapped to probes; ABCC1 is absent from GEO's curated AnnotGPL for GPL10558 and was recovered from the Illumina manufacturer platform table (ABCC1->ILMN_1802404). RESULT: higher ADE-RS5 associates with WORSE OS (HR>1) in all three dichotomizations (continuous HR=1.75 p=0.33; 40/60 HR=1.61 p=0.15; median HR=1.83 p=0.078/logrank 0.074) — DIRECTION matches the paper, but the association is NOT statistically significant within this single small cohort. This is consistent with (not contradictory to) the paper's significance claim, which is for the POOLED meta-analysis across 10 cohorts; the GSE71014-specific HR is only inside a forest-plot figure, so we grade directionality+significance rather than a numeric HR match. No fabrication flags. NOT ATTEMPTED (and why): ADE-RS5 derivation (AML02 discovery expression+EFS restricted); pLSC6 / integrated 4-group score (pLSC6 coefficients are external, ref.7, not in this paper); the other 9 validation cohorts (mix of dbGaP/on-request restricted + other GEO). CAVEAT: dichotomization here is within-cohort percentile (paper used rpart cutpoints fit on the restricted discovery cohort), and one representative probe per gene was used (highest-mean collapse).

💻 Code ↗ 🗄 Data: GSE71014

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 50
    assessed: 2026-06-16 ⛓ dd34048f7f72
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-16
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-16
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

Comprehensive transcriptomic evaluation of genes with pharmacological relevance to ara-C, daunorubicin, and etoposide (ADE chemotherapy) can be used to derive a drug-resistance gene-expression score that, alone or combined with a leukemic stemness score (pLSC6), predicts treatment outcomes in AML patients.

Core claims
  • A 5-gene ADE-Resistance Score (ADE-RS5), derived via LASSO regression from 67 pharmacologically relevant genes, predicts MRD positivity, EFS and OS in pediatric AML. finding
  • ADE-RS5 was developed using LASSO penalized Cox regression on mRNA expression of 67 candidate genes in the AML02 discovery cohort. method
  • Integrating ADE-RS5 with the previously defined pLSC6 leukemic stemness score into four patient groups improves prognostic stratification. finding
  • The integrated pLSC6/ADE-RS5 score group remains an independent predictor of EFS and OS after multivariable adjustment for risk group, WBC, FLT3-ITD status and age. finding
  • ADE-RS5 and pLSC6/ADE-RS5 integrated scores were validated in >3500 pediatric and adult AML patients across 10 independent cohorts. resource
  • pLSC6 stemness score is significantly associated with established high-risk AML features (risk group, cytogenetics, FLT3 status). finding
  • High ADE-RS5 score is associated with higher induction I/II MRD positivity in validation cohorts. finding
Experimental setups
Assay System Perturbation Readout Platform
LASSO penalized Cox regression on mRNA expression pediatric AML patients, AML02 trial (discovery cohort, n=163) none event-free survival (EFS) association with expression of 67 candidate genes
Gene-expression-based risk scoring (ADE-RS5) with recursive partitioning pediatric AML, AML02 discovery cohort (n=163) none MRD1 positivity, EFS, OS
Integrated transcriptomic score analysis (pLSC6 + ADE-RS5, four-group classification) pediatric AML, AML02 discovery cohort (n=163) none MRD1 positivity, EFS, OS across four score groups
Multivariable Cox regression analysis pediatric AML, AML02 discovery cohort (n=163) none EFS/OS adjusted for risk group, diagnostic WBC count, FLT3 status, age
Transcriptomic score validation (ADE-RS5, pLSC6, integrated score) combined pediatric AML validation cohorts (4 trials, n=1861) none EFS, OS
Transcriptomic score validation (ADE-RS5, pLSC6, integrated score) combined adult AML validation cohorts (5 trials, n=1669) none EFS, OS
Key results
  • Each unit increase in ADE-RS5 associated with increased EFS event rate in single-predictor Cox model (discovery cohort) 7.32-fold, p<0.00001, 95% CI 3.75-14.28
  • High ADE-RS5 predicts higher MRD1 positivity in discovery cohort OR=2.39, 95% CI 1.23-4.63, p=0.013
  • High ADE-RS5 associated with lower EFS and OS in discovery cohort EFS HR=4.07 p<0.0001; OS HR=4.54 p<0.0001
  • Integrated Group 4 (high pLSC6 + high ADE-RS5) had worst EFS/OS vs Group 1 (both low) in discovery cohort EFS HR=8.89 p<0.0001; OS HR=12.68 p<0.0001
  • Integrated score group remained independent predictor in multivariable analysis (discovery cohort) EFS: Group2 vs1 HR=4.68 p<0.001; Group3 vs1 HR=3.22 p=0.01; Group4 vs1 HR=7.26 p<0.001
  • ADE-RS5 and pLSC6 showed consistent significant EFS/OS association in combined pediatric validation cohort ADE-RS5 EFS HR=1.38, OS HR=1.6; pLSC6 EFS HR=1.9, OS HR=2.1 (all p<0.001)
  • 5-year EFS and OS were markedly lower in high vs low pLSC6/ADE-RS5 groups in both pediatric and adult validation cohorts e.g. pediatric Group1 vs Group4 5-yr EFS 57.76% vs 29.27%, p<0.0001
Key statistics
  • fold_change 7.32-fold increase in EFS event rate per unit ADE-RS5 (single-predictor Cox regression, discovery cohort)
  • pvalue p=0.013 (ADE-RS5 high group vs MRD1 positivity, discovery cohort)
  • pvalue p<0.0001 (ADE-RS5 high group vs EFS and OS, discovery cohort)
  • fold_change HR=7.26, p<0.001 (multivariable analysis, integrated Group4 vs Group1, EFS)
  • fold_change HR=9.72, p<0.001 (multivariable analysis, integrated Group4 vs Group1, OS)
  • count n=163 (AML02 discovery/model-development cohort)
  • count ~3634 patients (1861 pediatric, 1773 adult) across 10 cohorts (combined validation cohorts)
  • fold_change ADE-RS5 EFS HR=1.38, OS HR=1.6 (p<0.001) (combined pediatric validation cohort)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The paper developed a five-gene drug resistance score (ADE-RS5) using LASSO-penalized Cox regression with extensive leave-10%-out cross-validation in a discovery cohort of 163 pediatric AML patients, then validated both ADE-RS5 and an integrated four-group score (combining ADE-RS5 with a pre-existing stemness score, pLSC6) across 10 independent cohorts totaling ~3,634 patients. Survival outcomes (EFS, OS) were compared between score groups using Cox proportional hazards models (univariate and multivariable), with MRD1 positivity assessed via logistic regression. Results were reported as hazard ratios or odds ratios with 95% confidence intervals and p-values; categorical covariate distributions across groups were summarized with p-values consistent with chi-square testing. No explicit multiplicity correction across the many comparisons, endpoints, and cohorts was described.

Replicationbiological Sample sizeDiscovery cohort N=163 stated; validation cohorts enumerated as 10 independent trials totaling ~3634 (1861 pediatric from 4 trials, 1669–1773 adult from 5–6 trials); no formal power/sample-size calculation described GroupsLow vs. high ADE-RS5; low vs. high pLSC6; four integrated groups (low/low, low/high, high/low, high/high) vs. Group-1 reference; pediatric vs. adult combined cohorts Pairingunpaired Randomization/blindingnot stated Dispersionmixed Exact p-valuesno Effect sizesyes Confidence intervalsyes Multiplicity correctionnone stated
Statistical tests used
Test Applied to n Assumptions
LASSO-penalized Cox proportional hazards regression Model development: 67 pharmacological genes vs. EFS in AML02 discovery cohort; gene stability threshold set at ≥950/1000 leave-10%-out cross-validation replications 163 not stated
Leave-10%-out cross-validation (1000 replications) Internal validation of LASSO gene selection stability in AML02 163 na
Univariate Cox proportional hazards regression ADE-RS5 (continuous) vs. EFS in AML02 discovery cohort (HR=7.32, 95% CI=3.75–14.28); ADE-RS5 group vs. EFS and OS; pLSC6 group vs. EFS and OS in each cohort 163 (discovery); 1861 pediatric and 1669 adult (validation combined) not stated
Recursive partitioning Dichotomization of continuous ADE-RS5 score into low (n=98, 60%) vs. high (n=65, 40%) groups in AML02 163 na
Logistic regression ADE-RS5 score group vs. MRD1 positivity in AML02 (OR=2.39, 95% CI=1.23–4.63, p=0.013) 163 not stated
Multivariable Cox proportional hazards regression ADE-RS5 score group and integrated pLSC6/ADE-RS5 four-group score vs. EFS and OS, adjusting for diagnostic risk group, WBC count, FLT3 status, and age in AML02 and validation cohorts 163 (discovery); combined validation cohorts (n not explicitly restated for MVA) not stated
Log-rank test (implied by Kaplan-Meier comparisons with p-values) Group-by-group survival comparisons (EFS and OS) across all discovery and validation cohorts null not stated
Chi-square or Fisher's exact test (implied) Categorical patient characteristics (gender, age group, risk group, cytogenetics, FLT3, WBC, MRD, induction response) compared across pLSC6, ADE-RS5, and integrated score groups in Table 1 1861 pediatric and 1669 adult not stated
Approaches that could also have been used
  • Continuous ADE-RS5 scores were dichotomized into two groups using recursive partitioning
    Could also: The continuous score could be retained as-is in Cox models, or patients could be divided into tertiles or quartiles — Dichotomization discards within-group variation and can inflate apparent effect sizes; using the score as a continuous predictor or in ordered quantile groups preserves the full dose-response relationship and avoids the multiple-testing concern inherent in data-driven cutpoint selection
  • Ten validation cohorts were each analyzed separately, with results summarized narratively across cohorts
    Could also: A random-effects meta-analysis pooling cohort-specific HRs could be applied across the 10 validation datasets — Meta-analytic pooling yields a single quantitative summary estimate with a formal heterogeneity statistic (I²/Q-test), making between-cohort consistency directly visible and quantified rather than assessed visually or narratively
  • Multiple endpoints (EFS, OS, MRD1 positivity), multiple score comparisons, and multiple cohorts were tested without a described multiplicity correction
    Could also: A Benjamini-Hochberg false discovery rate (FDR) correction or Bonferroni adjustment could be applied within families of related tests (e.g., all pairwise group contrasts within a cohort for a given endpoint) — When many null hypotheses are tested simultaneously, the family-wise type-I error rate rises; a pre-specified correction procedure helps readers calibrate the expected proportion of false positives among the reported findings
  • Variable selection for ADE-RS5 used LASSO-penalized Cox regression
    Could also: Elastic net (combining L1 and L2 penalties) or random survival forests could also be used for variable selection from the 67-gene panel — LASSO tends to select one variable from correlated groups arbitrarily; elastic net can retain grouped correlated predictors, and random survival forests are nonparametric and capture nonlinear or interaction effects — both are standard alternatives when gene-expression predictors are expected to be correlated
  • Five-year EFS and OS rates in Table 1 are presented with values in parentheses consistent with standard errors of Kaplan-Meier estimates
    Could also: 95% confidence intervals around the Kaplan-Meier point estimates could be reported instead of or alongside standard errors — CIs are more directly interpretable as the plausible range for the true population rate and are the conventional format recommended by reporting guidelines such as STROBE and CONSORT; SEs require an additional mental step (±1.96 × SE) to interpret as an uncertainty range
  • MRD1 positivity was analyzed with logistic regression treating it as a binary outcome
    Could also: Competing-risks regression (e.g., Fine-Gray subdistribution hazard model) or time-to-event analysis of MRD could also be used if time-to-MRD-assessment varied across patients — If time to MRD measurement differed across patients or if early death precluded MRD assessment, a competing-risks or survival framework for MRD would account for informative censoring that binary logistic regression does not address
Software: Not stated

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
7
Impact: low
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

GSE17855 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE68833 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
NCT00136084 NCT in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
NCT00703820 NCT in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Figures / tables: Fig.4Fig.5
C1
Reported
ADE-RS5 = DCTD*0.128 + TOP2A*(-0.0993) + ABCC1*0.212 + MPO*(-0.113) + CBR1*(-0.126) (Eq.1)
Reproduced
Equation applied verbatim to GSE71014; ADE-RS5 computed on all 104 samples (range -0.916..0.342, median -0.442). Coefficients reused (published), not re-derived.
partial
C2
Reported
5 of 67 candidate genes selected in >=950/1000 leave-10%-out LASSO-Cox CV on AML02 (163 pts)
Reproduced
NOT ATTEMPTED — AML02 discovery expression+EFS not public/restricted
partial
C3
Reported
High ADE-RS5 -> significantly worse OS in validation cohorts incl. CN-AML GSE71014 (meta-analysis across 10 cohorts significant; GSE71014 HR in forest plot, Fig.4/Fig.5)
Reproduced
GSE71014 (n=104, 36 events): Cox continuous HR/unit=1.752 (95% CI 0.562-5.460, p=0.334); high-vs-low 40/60 HR=1.613 (0.838-3.104, p=0.152, logrank 0.148); median split HR=1.827 (p=0.079, logrank 0.074). Direction worse-OS (HR>1) reproduced in all 3 cutpoints; single-cohort significance not reached, consistent with the pooled meta-analytic claim.
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 50/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟡2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +8

The published ADE-RS5 equation was applied verbatim to the fully public CN-AML cohort GSE71014 (n=104, 36 events) and reproduces the paper's worse-OS direction (HR>1 in all three cutpoints: continuous 1.752, 40/60 1.613, median 1.827). The single deviation — non-significance in this cohort — sits on our methodology / data-availability side: the discovery rpart cutpoints are restricted (forcing self-chosen percentile splits) and the cohort-specific HR exists only inside a forest-plot figure, so an exact numeric match is impossible. This is not an authors' defect and shows no fabrication signal; a non-significant single cohort within a significant pooled 10-cohort meta-analysis is normal. Overall a solid partial reproduction with explainable deviations.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

205.7 k
tokens (I/O) · 16.6 M incl. cache
63 min
runtime · 0.01 CPU-h
1.9 GB
peak RAM
3 (2 failed)
HPC jobs
hummel
machine