Machine Learning-Based Integrated Analysis of PANoptosis Patterns in Acute Myeloid Leukemia Reveals a Signature Predicting Survival and Immunotherapy.
The main results reproduced, with only marginal, non-material deviations.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Any deviation was negligible
- 🟡Could not use the authors’ exact input data
- 🟡Reported values were only indirectly comparable
- 🟡A deviation arose in the data or preprocessing
- 🟡A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
DESCRIBED WELL ENOUGH? Yes for the shipped core. The repo (github.com/zwxiangya/PANScore @ c9159a2, authors' own R package -- author Wei Zhang is a paper co-author) bundles the FINAL 19-gene signature + Cox coefficients (df.rda) and a 100-sample HOVON demo expression matrix (exprset.hovon.rda), with one function PANScore.calculate() implementing eq.(1). 1:1 RESULT: the deterministic risk-score computation reproduces EXACTLY -- the authors' actual function output equals an independent manual Sum(Exp*Coe) to floating-point precision (3.55e-15), 19/19 signature genes match the HOVON matrix, scores produced for all 100 demo samples. The signature is internally consistent with the paper: exactly 19 genes, and every gene named by string in the main text (CALCRL, DOCK1, CLCN5, LSP1, NRIP1, DNMT3B, ADRM1, ALDH2, NYNRIN) is present. DIFFERENT/NOT ATTEMPTED: (a) the upstream model selection (101-combo LOOCV -> RSF -> 19 genes) is not shipped, so the CHOICE of genes is not independently reproducible; (b) downstream survival statistics -- HOVON AUC 0.747/0.772/0.761, C-index, KM -- were NOT attempted because the bundled data carry no survival/clinical columns and are a 100-sample demo, not the full 618-patient HOVON cohort (full clinical data = ArrayExpress E-MTAB-3444); (c) external-cohort validation (TCGA-LAML, BeatAML, GSE12417, GSE106291), immune-infiltration, nomogram, immunotherapy/TIDE, somatic mutation, and qRT-PCR (wet-lab) are out of scope. REVIEWER FLAGS: the coefficient column in df.rda is labelled 'tcga' (duplicated as tcga.1), although the paper attributes the coefficients to Cox in HOVON -- a provenance inconsistency worth checking against Suppl Table 9. Overall: the published tool + signature reproduce faithfully on bundled data; the validation metrics are the deliberate unreproduced ~20% (data not shipped).
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 71assessed: 2026-06-15 ⛓ a6e208f7998d
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-15
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusThe study tests whether PANoptosis-related genes (PANRGs) can stratify acute myeloid leukemia (AML) into molecular subtypes and whether a PANoptosis-based prognostic signature can predict survival, the immune landscape, and immunotherapy/drug response in AML.
- ★ AML patients can be categorized into two distinct PANRG-based clusters with differing prognosis and immune characteristics finding
- ★ A 19-gene PANoptosis-related prognostic signature (PAN2RS) stratifies AML patients into high- and low-risk groups, with high-risk predicting worse survival resource
- ★ The risk score is significantly correlated with infiltration of most immune cells in the tumor microenvironment finding
- ★ A machine-learning framework of 101 model combinations across 10 algorithms (with Random Survival Forest for feature selection) was used to build and select the optimal signature method
- ★ A nomogram integrating the risk score with clinical features (age, gender, cytogenetic risk, NPM1 and FLT3-ITD mutation) predicts patient survival resource
- ★ LGR5 and VSIG4 show significant differential expression between normal and human leukemia cell lines (HL-60 and MV-4-11) finding
- The PANoptosis signature can predict immunotherapy response and identify candidate drug targets/agents for high-risk AML patients finding
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| bulk RNA-seq / microarray transcriptomic profiling (bioinformatics) | AML patient cohorts (HOVON, TCGA-LAML, BeatAML, GSE37642, GSE12417, GSE106291) | none | gene expression (log2(TPM+1)), survival, prognostic risk score | R packages limma, affy |
| consensus clustering analysis | AML patient transcriptomes (HOVON) | none | PANoptosis molecular cluster assignment | R package ConsensusClusterPlus (pam algorithm) |
| immune cell infiltration / immune microenvironment analysis | AML patient cohorts | none | immune cell infiltration, immune checkpoint gene expression, ORR | ESTIMATE, xCell, quanTIseq, ssGSEA, CIBERSORT, EPIC, MCP-counter |
| pathway enrichment analysis (GO, KEGG, GSEA, GSVA) | AML patient DEGs | none | enriched biological processes and pathways | R packages GSVA, clusterProfiler |
| machine learning survival modeling | AML cohorts (HOVON training; TCGA-LAML, GSE12417, GSE37642 testing) | none | C-index, risk score, prognostic gene selection | LOOCV; lasso, Enet, Ridge, SuperPC, StepCox, GBM, CoxBoost, plsRcox, RSF, SurvivalSVM |
| drug target / drug sensitivity prediction | leukemia cancer cell lines (CCLE/DepMap) and 337-patient ex vivo cohort | small-molecule inhibitors (122 / 106) | CERES score, IC50, AUC vs PAN2RS correlation | DepMap, CTRP, PRISM, Connectivity Map; R packages pRRophetic, impute |
| qRT-PCR | human AML cell lines HL-60 (promyelocytic) and MV-4-11 (FLT3-ITD+ myelomonocytic) vs normal | none | mRNA expression of signature genes (LGR5, VSIG4) | RPMI 1640 + 10% FBS culture; qRT-PCR (kit not stated) |
- ▼ A 19-gene PANoptosis signature (PAN2RS) was constructed; high-risk group exhibited notably worse prognosis
- – Risk score significantly correlated with infiltration of most immune cells
- – LGR5 and VSIG4 showed significant differential expression in normal vs leukemia cell lines (HL-60, MV-4-11)
- – AML cases categorized into two PANRG clusters associated with survival, immune system, and cancer pathways
- count 226 PANoptosis-related genes (PANRGs) (final merged PANoptosis gene list from MSigDB)
- count 19 genes in the PANoptosis signature (PAN2RS) (prognostic risk score signature)
- count HOVON n=618 (training) (HOVON cohort from ArrayExpress E-MTAB-3444)
- count six AML cohorts: HOVON 618, TCGA-LAML 147, BeatAML 143, GSE37642 421, GSE12417 162, GSE106291 250 (patient cohorts used)
- count 101 predictive models from 10 machine learning algorithms (LOOCV model fitting framework)
- count 89 previously published AML signatures compared (benchmarking PAN2RS predictive performance)
- other DepMap CERES correlation r < −0.45, p < 0.05 (threshold for identifying potential drug targets)
- count 337 patients and 106 inhibitors retained for drug sensitivity analysis (ex vivo drug sensitivity dataset)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This retrospective bioinformatics study trained a 19-gene PANoptosis-related risk signature (PAN2RS) in the HOVON AML cohort (n=618) using a 101-model machine learning LOOCV framework spanning 10 algorithms, then validated it across five independent cohorts. Survival discrimination was assessed via Kaplan–Meier/log-rank analysis, time-dependent ROC curves (1-, 3-, 5-year AUC), C-index, and multivariate Cox regression; drug sensitivity associations were explored via Wilcoxon rank sum tests and Spearman correlations; and a meta-analysis of hazard ratios was conducted across all six cohorts. An in vitro qRT-PCR experiment assessed expression of two signature genes in AML cell lines.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| univariate Cox regression | screening prognostic DEGs in HOVON cohort (p < 0.01 threshold); pan-cancer prognosis prediction across 33 TCGA cancer types | 618 (HOVON training set); 10295 (pan-cancer) | not stated |
| log-rank test | comparing overall survival between high/low gene expression subgroups during DEG filtering (p < 0.05); Kaplan–Meier validation of PAN2RS in all cohorts | 618 (HOVON); varied per testing cohort | not stated |
| multivariate Cox regression | independence and robustness of PAN2RS vs. clinical variables (age, gender, cytogenetic risk, NPM1, FLT3-ITD) in HOVON and TCGA-LAML; comparison of 90 signatures | 618 (HOVON); 147 (TCGA-LAML) | not stated |
| Wilcoxon rank sum test | differential drug sensitivity AUC between high and low PAN2RS groups (top vs. tail decile) in 337-patient ex vivo cohort; repeated with CTRP/PRISM-predicted AUC | 337 patients total; top and tail decile subsets (~33–34 each) | not stated |
| Spearman rank-order correlation | CERES gene dependency scores vs. PAN2RS in leukemia CCLs (threshold r < −0.45, p < 0.05); drug AUC vs. PAN2RS (threshold r < 0 or r < −0.3/−0.4); TMB vs. PAN2RS across 33 cancer types | 739 CCLs (CERES); 337 patients (drug sensitivity); 10295 patients (TMB) | not stated |
| Kolmogorov–Smirnov (KS) scoring and eXtreme Sum (XSum) scoring | CMap perturbational drug reversal scoring based on 300 DEGs between top and tail PAN2RS deciles | 2424 perturbational CMap signatures | na |
| meta-analysis (HR pooling, R package meta) | pooled hazard ratio of PAN2RS across all six AML cohorts | 6 cohorts; individual n: 618, 147, 143, 421, 162, 250 | not stated |
| time-dependent ROC analysis (survivalROC) | 1-, 3-, and 5-year AUC for PAN2RS in all six cohorts; comparison of 90 published signatures | varied per cohort | not stated |
| consensus clustering (PAM algorithm, Euclidean distance, 1000 bootstrap resamples at 80%) | unsupervised classification of HOVON AML patients into PANoptosis clusters | 618 (HOVON) | not stated |
| qRT-PCR differential expression comparison (statistical test not specified in available text) | LGR5 and VSIG4 expression in normal cells vs. HL-60 and MV-4-11 AML cell lines | null | not stated |
-
Patients were dichotomized into high- and low-risk groups using a data-derived optimal cutoff from the training cohort via survminer, then applied to testing cohorts↳ Could also: The continuous risk score could be retained as a continuous predictor in Cox models, or cutoff selection could be performed within a cross-validation loop to account for selection uncertainty — Treating the score as continuous avoids information loss from dichotomization; cross-validated cutoff selection would provide a less optimistic estimate of the cutoff's generalizability to new data
-
101 machine learning models were trained and the one with the highest mean C-index across training and testing cohorts was selected as the final model↳ Could also: A nested cross-validation or bootstrap framework could also be used to estimate the performance of the model-selection process itself, separating the selection step from the evaluation step — Choosing the best model from a large candidate pool using the same cohorts for selection and evaluation can yield an optimistic C-index estimate; nested validation quantifies how much performance is attributable to selection rather than true signal
-
FDR correction was applied to the initial DEG analysis, but subsequent survival screening (univariate Cox, log-rank), pan-cancer analyses across 33 cancer types, and multiple drug-sensitivity tests used unadjusted p-value thresholds↳ Could also: Applying FDR or Bonferroni correction to the family of log-rank tests, pan-cancer Cox analyses, or drug-sensitivity comparisons would also control the experiment-wide false-discovery rate across those test families — When many hypotheses are tested simultaneously (e.g., 33 cancer types, 106 inhibitors, multiple immune cell types), a multiplicity adjustment quantifies the expected proportion of false positives among the reported significant findings
-
Drug sensitivity differences between risk groups were assessed by comparing only the top and tail deciles (~10% each) rather than all patients↳ Could also: Comparing the full high- vs. low-risk groups (median split) or modeling PAN2RS as a continuous predictor against AUC across all patients would also characterize the dose-response relationship — Extreme-decile comparisons maximize statistical contrast but exclude the majority of patients; full-range analyses use all available data and may yield more stable and generalizable estimates of the association
-
Seven immune deconvolution algorithms (ESTIMATE, xCell, quanTIseq, ssGSEA, CIBERSORT, EPIC, MCP-counter) were applied in parallel and results reported separately per algorithm↳ Could also: A consensus or ensemble score across algorithms, or a formal inter-method agreement analysis (e.g., pairwise Spearman correlations between method outputs), could also summarize immune infiltration while characterizing method-specific variance — Different deconvolution methods can give discordant estimates for the same cell type; reporting concordance across methods highlights which immune infiltration findings are robust to algorithmic choice and which are method-dependent
-
The meta-analysis pooled hazard ratios across six cohorts using the R package meta, but the fixed-effects vs. random-effects model choice and heterogeneity assessment (e.g., I²) are not described in the available text↳ Could also: Explicitly reporting the heterogeneity statistic (I²/Q-test), the meta-analytic model used (fixed vs. random effects), and a forest plot with per-cohort HRs and confidence intervals would also characterize between-cohort variability — The six cohorts differ in platform (RNA-seq vs. microarray), patient selection, and sample size; quantifying heterogeneity informs whether a pooled HR is a meaningful summary or whether cohort-specific estimates are more informative
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-38322112
Paper: Tang et al. 2024, Int J Clin Pract. "Machine Learning-Based Integrated Analysis of PANoptosis Patterns in Acute Myeloid Leukemia..." PMID 38322112, PMCID PMC10846924, DOI 10.1155/2024/5113990.
Code: https://github.com/zwxiangya/PANScore (R package PANScore, author Wei Zhang,
one of the paper's co-authors → this is the authors' own code, P16 not invoked).
Repo = a tiny R package that bundles the final signature and computes the risk score.
Data shipped in the repo (on «infra», not «host»):
data/df.rda(614 B) — the 19 signature genes + their Cox coefficients.data/exprset.hovon.rda(9 MB) — the HOVON AML training-cohort expression matrix (demo).PANScore_0.1.0.tar.gz— the built package;R/PANScore.calculate.Ris the only function.
The pipeline (Methods §2.5–2.6, eq. 1)
- Univariate Cox in HOVON → 28 prognostic PANoptosis DEGs.
- LOOCV over 101 combinations of 10 ML algorithms (lasso, Enet, Ridge, SuperPC, StepCox, GBM, CoxBoost, plsRcox, RSF, SurvivalSVM); pick highest mean C-index → RSF.
- RSF
var.select(randomForestSRC) on 262 candidates → 19 signature genes. - Multivariate Cox in HOVON → coefficients of the 19 genes.
- Risk Score = Σ Expᵢ · Coeᵢ (eq. 1) — the per-patient PAN2RS. Deterministic.
- Optimal cutoff (survminer) → high/low groups; KM + log-rank; time-dependent AUC (survivalROC) at 1/3/5 yr; C-index (survcomp) in all 6 cohorts.
In scope (pipeline-derived, attempted)
- C1 — 19 signature genes: the gene membership shipped in
df.rdavs the genes named in the paper (Fig 3b/c, Suppl Table 9). Top-5 by RSF importance reported: CALCRL, DOCK1, CLCN5, LSP1, NRIP1. - C2 — coefficients: the per-gene Cox coefficients in
df.rda(Fig 3d / Suppl Table 9). - C3 — PAN2RS risk-score computation: run the authors'
PANScore.calculate()on the bundled HOVON expression set → reproduce the per-patient risk score exactly (the package's own deterministic demo; fully self-contained).
In scope but harder (the ~20% — attempted only if data is self-contained)
- C4 — HOVON time-dependent AUC (reported 1/3/5-yr = 0.747 / 0.772 / 0.761). Needs HOVON survival/clinical data (ArrayExpress E-MTAB-3444), which is not bundled in the repo (df.rda = coeffs only; exprset.hovon = expression only). If the bundled object carries no survival, this is recorded as not-attempted and why.
Out of scope (not pipeline / not feasible from shipped artifacts)
- The full ML model-selection (101 combos LOOCV) and RSF gene selection — upstream derivation; the repo ships only the final signature, not the selection code.
- Validation AUC/C-index in TCGA-LAML, BeatAML, GSE12417, GSE106291 (external data, not bundled).
- Consensus clustering, immune-infiltration (CIBERSORT/EPIC/xCell), nomogram, TIDE/immunotherapy, somatic-mutation, qRT-PCR of LGR5/VSIG4 in HL-60/MV-4-11 (wet-lab) — manual/external/wet-lab, not reproduced.
Honest framing
The reproducible core is the deterministic risk-score computation and the shipped signature (genes + coefficients). These let a human check, by hand, whether the published signature is internally consistent and whether the score the tool emits matches the formula in eq. 1. The upstream model selection and the downstream survival statistics on external cohorts are the unreproduced remainder.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
The deterministic published core reproduces exactly — the authors' own PANScore.calculate() recomputes the eq.(1) risk score to floating-point precision (3.55e-15) on the bundled HOVON demo, with all 19 signature genes and every main-text-named gene present. The unreproduced parts are mostly data-availability / our-scope matters: the repo ships only a 100-sample demo (no survival columns) instead of the 618-patient cohort, so the headline AUC 0.747/0.772/0.761 (Fig 4a) and the central prognostic claim could not be tested, and the gene-selection pipeline is absent. One authors-side flag remains: the coefficient column is labelled tcga despite the paper attributing coefficients to HOVON. Net: a solid, internally-consistent partial reproduction with explainable, non-fabrication deviations — hence yellow overall rather than red.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.