Machine learning-based identification of an immunotherapy-related signature to enhance outcomes and immunotherapy responses in melanoma.
The main results reproduced, with only marginal, non-material deviations.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Reported values are derivable from the shared data
- ✓The central claim held under reproduction
- 🟡Could not use the authors’ exact input data
- 🟡Reported values were only indirectly comparable
- 🟡A deviation arose in the data or preprocessing
- 🟡A deviation was attributed to the published material
- 🟡The deviation was non-trivial in magnitude
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Described well enough to reproduce the HEADLINE prognostic result, but only because the 7 signature genes are published — the cited repo ships no paper-specific code/data (just the generic Liu 2022 101-ML template; main branch is a 5-byte 'asss' file). Applying the published 7 genes (GBP5, HLA-DPB1, XBP1, CD40, GBP1, CXCL10, TNFSF13B) to the public cohorts, every gene is independently protective in TCGA-SKCM (HR 0.67-0.73, p<1e-6), and a multivariate-Cox risk score gives C-index 0.63-0.65 with 1-year time-AUC ~0.70 across TCGA-SKCM, GSE54467, GSE22153, GSE65904, plus significant KM stratification in 3/4 cohorts. Reproduced C-indices run ~0.04-0.09 below the figure-reported ~0.69-0.72 (mostly within ~2 SE; 1-year AUCs match), explainable by model choice (multiCox vs Lasso+plsRcox on the 44-gene panel) and survival-endpoint definitions (GSE65904 DSS, GSE54467 death coding). NOT attempted: the upstream gene-selection chain, the exact plsRcox model, and all downstream immune/drug/scRNA analyses. No fabrication indicators — the central signature is a genuine cross-cohort prognostic marker; the gap is exact magnitude, not credibility. Verdict: PARTIAL, provisional, human review required.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 69assessed: 2026-06-20 ⛓ b24b3fcdfc6f
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-20
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-20no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetThe paper tests whether a machine-learning-derived multi-gene signature (ITRGM) built from consensus immunotherapy prognostic genes can robustly predict immunotherapy response and prognosis in melanoma, addressing the lack of reliable biomarkers caused by tumor heterogeneity.
- ★ 66 consensus immunotherapy prognostic genes (CITPGs) were identified from the intersection of WGCNA modules, immunotherapy responder-vs-non-responder DEGs, and tumor-vs-normal DEGs finding
- ★ CITPG-high tumors show better prognosis and enriched immune activity compared to CITPG-low tumors finding
- ★ A seven-gene ITRGM signature (GBP5, HLA-DPB1, XBP1, CD40, GBP1, CXCL10, TNFSF13B) was constructed using a Lasso+plsRcox model selected from 101 machine learning algorithm combinations method
- ★ ITRGM outperforms 37 previously published signatures in predicting immunotherapy prognosis across training, testing, and meta-cohorts finding
- ★ Low-risk ITRGM patients show higher tumor mutation burden and immune cell infiltration, indicating immune-hot tumors with better prognosis finding
- ★ GBP5 expression correlates with CD8+ T cell infiltration, validated by IHC/multiplex immunofluorescence in melanoma tissue finding
- ★ ITRGM predicts immunotherapy response in 8 additional independent cohorts, including urothelial carcinoma and stomach adenocarcinoma finding
- All seven ITRGM model genes are up-regulated in immunotherapy responders finding
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| WGCNA (weighted gene co-expression network analysis) | Melanoma immunotherapy cohort PRJEB23709 | none (responder vs non-responder comparison) | gene modules correlated with immunotherapy response | R package WGCNA |
| Differential gene expression analysis | TCGA-SKCM tumor tissue vs GTEx normal tissue | none | DEGs (|log2FC|>1, FDR<0.05) | GEPIA2 |
| Differential gene expression analysis | PRJEB23709 melanoma immunotherapy responders vs non-responders | none | DEGs (|log2FC|>1, FDR<0.05) | limma R package |
| Consensus clustering | TCGA-SKCM melanoma patients | none | CITPG-high vs CITPG-low patient clusters | ConsensusClusterPlus R package |
| Machine learning model construction (10 algorithms, 101 combinations) | TCGA-SKCM (training), GSE22153/GSE54467/GSE69504 (validation) | none | C-index, risk score for prognosis/immunotherapy response | R (Enet, ridge, plsRcox, Lasso, RSF, SuperPC, CoxBoost, GBM, survival-SVM, stepwise Cox) |
| Multiplex immunofluorescence / IHC | Tissue microarray, 17 melanoma and 18 normal skin cases | none | GBP5 and CD8 co-expression/localization | GBP5 (Abcam AB313390), CD8 (Servicebio GB12068) antibodies |
| Bulk RNA sequencing | Mouse melanoma immunotherapy cohorts (GSE109485, GSE149825) | immunotherapy | ITRGM model gene expression | TISMO database |
| Single-cell RNA sequencing | SKCM_GSE115978_aPD1 melanoma dataset | anti-PD1 immunotherapy | immune landscape correlation with ITRGM genes | TISCH2 website |
- – 66 CITPGs identified from intersection of WGCNA module, responder/non-responder DEGs, and tumor/normal DEGs
- ▲ All 66 CITPGs associated with better overall survival by univariate Cox regression in TCGA-SKCM
- – Lasso regression narrowed 44 consensus genes to 7 genes with non-zero coefficients: GBP5, HLA-DPB1, XBP1, CD40, GBP1, CXCL10, TNFSF13B
- – Lasso+plsRcox model had the highest average C-index among 101 algorithm combinations tested
- – ITRGM outperformed 37 published prognostic signatures across training and validation cohorts
- ▲ Low-risk ITRGM group showed increased tumor mutation burden and immune cell infiltration
- ▲ GBP5 expression correlated with CD8+ T cell infiltration across cancer types and validated in melanoma tissue
- ▲ All seven ITRGM genes were up-regulated in immunotherapy responders
- count 66 (consensus immunotherapy prognostic genes (CITPGs) identified)
- count 44 (consensus prognostic DEGs validated across four melanoma cohorts)
- count 7 (final model genes in ITRGM signature)
- count 101 (machine learning algorithm combinations evaluated)
- fold_change >4-fold upregulation (DEGs highlighted in volcano plot between immunotherapy responders and non-responders)
- count 1808 (total cancer patients across 16 independent public cohorts used in study)
- count 459 (TCGA-SKCM patients retained after excluding incomplete clinical data)
- other 52% (5-year overall survival rate reported for combinatorial immunotherapy in melanoma (background citation))
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This is a bioinformatics/multi-omics study that used weighted gene co-expression network analysis (WGCNA), differential expression analysis, and consensus clustering to derive a set of candidate genes, then applied univariate/multivariate Cox regression and an ensemble of 10 machine-learning algorithms (101 combinations, evaluated by C-index via 10-fold cross-validation) to build a 7-gene prognostic signature (ITRGM). Performance and biological associations were assessed across multiple public transcriptomic, single-cell, and immunotherapy cohorts using survival analysis, correlation analyses, and immune-infiltration deconvolution tools, with additional IHC/immunofluorescence validation in patient and mouse tissue.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Differential expression analysis (limma package, |log2FC|>1 and FDR<0.05) | responders vs non-responders in PRJEB23709 immunotherapy cohort | — | not stated |
| Weighted gene co-expression network analysis (WGCNA) with Pearson correlation for module membership vs gene significance | gene module–immunotherapy response correlation, PRJEB23709 cohort | — | not stated |
| Kaplan-Meier survival analysis (with implied log-rank comparison) | immunotherapy responders vs non-responders (PRJEB23709) and CITPG-high vs CITPG-low clusters (TCGA-SKCM) | TCGA-SKCM n=459 (stated elsewhere in Methods) | not stated |
| Univariate Cox regression | 66 CITPGs tested individually against overall survival, TCGA-SKCM | 459 | not stated |
| Ensemble of 10 machine-learning survival algorithms across 101 combinations (final model: Lasso + plsRcox), evaluated by C-index with 10-fold cross-validation | ITRGM construction; training in TCGA-SKCM, validation in GSE22153, GSE54467, GSE69504 | TCGA-SKCM n=459; GSE54467 n=79; GSE69504 n=214; GSE22153 n=57 | not stated |
| Multivariate Cox regression and calibration curve analysis | ITRGM risk score across training/validation/meta-cohorts | — | not stated |
-
DEGs between responders and non-responders were selected using a fixed fold-change and FDR cutoff (|log2FC|>1, FDR<0.05) with limma.↳ Could also: A model-based count method such as DESeq2 or edgeR (if raw counts are available) with shrinkage estimation of fold-changes — These approaches can improve effect-size stability for genes with low counts or high variance, which may complement fold-change/FDR filtering, especially in modest-sized cohorts.
-
66 candidate genes (CITPGs) were each tested individually via univariate Cox regression against overall survival.↳ Could also: A penalized multivariate Cox model (e.g., LASSO or elastic net) fit jointly on all candidate genes as an additional screening step — Joint modeling can account for correlation among genes and reduce the number of individually tested hypotheses, which is a consideration whenever many single-gene tests are performed in the same dataset.
-
Model selection across 101 machine-learning combinations relied primarily on the concordance index (C-index) for ranking performance.↳ Could also: Complementary metrics such as time-dependent AUC, integrated Brier score, or calibration plots at multiple time points — These add information about calibration and time-varying discrimination that C-index alone does not fully capture, useful when comparing many candidate models.
-
Correlations between risk scores/module features and other variables were assessed with Pearson (module membership vs gene significance) or Spearman (risk score vs immune infiltration) coefficients.↳ Could also: Partial correlation or regression adjusting for potential confounders such as tumor purity or sequencing depth — Adjusting for known confounders can help isolate the association of interest when comparing immune infiltration estimates across samples of varying purity.
-
Survival comparisons between groups (e.g., CITPG-high vs CITPG-low, ITRGM risk groups) were visualized with Kaplan-Meier curves.↳ Could also: Reporting hazard ratios with 95% confidence intervals alongside the Kaplan-Meier plots, and checking the proportional hazards assumption — This would provide an effect-size estimate with a measure of precision and confirm that the Cox/log-rank framework's assumptions hold across the compared groups.
-
Multiple testing correction (FDR) was applied specifically to the DEG-calling step in limma.↳ Could also: Extending an explicit multiple-testing correction (e.g., Benjamini-Hochberg) to other repeated-test families in the workflow, such as the batch of univariate Cox tests across 66 genes — Applying a family-wise or FDR correction across all genes tested in a given analysis step is a standard way to control the overall false-positive rate when many hypotheses are evaluated together.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — PMID 39355255 (ITRGM melanoma immunotherapy signature)
Paper: Machine learning-based identification of an immunotherapy-related signature
to enhance outcomes and immunotherapy responses in melanoma. Front Immunol 2024;
PMCID PMC11442245; DOI 10.3389/fimmu.2024.1451103.
Authors' code: https://github.com/YuBestLab/YuBestLab.github.io
Branch with code: 101-machine-learning-algorithm
What the repo actually contains
main/masterbranches: a single 5-byteIndex.html(asss\n) — no code.- Branch
101-machine-learning-algorithm: the generic "101 machine-learning combination" template from Liu et al. 2022, Nat Commun (file41467_2022_28421_MOESM4_ESM.xlsx= the 101-model list;ML.R= helper functionsRunML/RunEval/CalRiskScore/ExtractVar/scaleData/SimpleHeatmap;scripts.R= driver). This template is NOT customised to this paper and ships no input data and no hard-coded results. The authors applied this third-party framework to their data. - Per BRIEF rule 2 (P16), applying the described third-party framework to the paper's own data is a valid reproduction route. We therefore reproduce by re-running the prognostic signature on the public cohorts using the published 7-gene list.
Reported pipeline (from Methods/Results)
- Training: TCGA-SKCM (N=459 after QC).
- Validation (untreated melanoma): GSE22153 (57), GSE54467 (79), GSE65904 (214; paper text also writes "GSE69504" — a typo; Fig. 6 uses GSE65904).
- Signature build: CITPGs (66) → DEGs between CITPG-high/low (566) → univariate Cox across 4 cohorts (44 consensus genes) → 101 ML-combinations, Lasso+plsRcox wins by mean C-index → final 7 genes: GBP5, HLA-DPB1, XBP1, CD40, GBP1, CXCL10, TNFSF13B.
- Reported performance: C-index ≈ 0.70 (TCGA), 0.71 (GSE54467), 0.72 (GSE22153), 0.69 (GSE65904) (Fig 5B, read approximately); time-AUC 1y 0.70–0.78.
IN SCOPE (pipeline-derived, attempted)
| # | Result | Pipeline | Status |
|---|---|---|---|
| C1 | 7 model genes are individually prognostic in TCGA-SKCM | univariate Cox (survival) | reproduced |
| C2 | 7-gene signature risk score → C-index per cohort | multivariate Cox (= template FinalModel='multiCox') + Lasso-Cox, evaluated by Harrell C |
reproduced (partial: ~0.04–0.09 lower) |
| C3 | Time-dependent AUC (1/3/5y) per cohort | timeROC | reproduced (1y AUC matches) |
| C4 | Signature significantly stratifies survival (KM high vs low) | median split + log-rank | partial (3/4 cohorts p<0.05) |
OUT OF SCOPE / NOT ATTEMPTED (documented, not graded)
- Upstream signature selection (WGCNA grey60 on PRJEB23709; GEPIA2 tumor-vs-normal DEGs; CITPG 3-way intersection; 566 consensus DEGs; 44-gene panel; the full 101-model tournament selecting Lasso+plsRcox). Requires several web-tool steps (GEPIA2, TIMER2, TIP) and intermediate gene lists not shipped in the repo → the 44-gene input panel is not derivable from the deposited materials without substantial extra reconstruction. We therefore take the published 7 genes as given and verify their performance.
- Exact Lasso+plsRcox model:
plsRcoxcould not be installed (CRAN download + heavy Bioc deps failed in env). Substituted the template's default final scoring (multivariate Cox on the final genes) —FinalModel<-c("panML","multiCox")[2]in the authors' ownscripts.R. - Immunotherapy-response cohorts (GSE91061, PRJEB23709, IMvigor210, …), immune infiltration (CIBERSORT/ssGSEA/ESTIMATE/TIMER2), TIDE, TMB, IPS, oncoPredict drug sensitivity, GSEA, scRNA (GSE115978), CD8 IHC (GSE243238), mouse cohorts — descriptive / web-tool / wet-lab-adjacent downstream analyses; not pipeline-core; not attempted.
Reproduction route (what was run)
«our HPC» «infra» …/reproductions/pmid-39355255/. Cohorts rebuilt to the 7 genes from:
TCGA-SKCM HiSeqV2 + curated OS (UCSC Xena); GSE22153/GSE54467/GSE65904 series matrices
- GPL61
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
The headline reproduces: applying the published 7-gene ITRGM signature to the public cohorts confirms every gene as protective (HR 0.67-0.73, p<1e-6) with C-index 0.63-0.65, 1-yr AUC ~0.70, and significant KM stratification in 3/4 cohorts — values derivable from shared data, no fabrication indicators (q5/q7 green). The deviations are explainable and sit on the input/method side: our forced multiCox-for-plsRcox substitution and self-defined cohorts/endpoints (the deposited repo is an empty generic template), giving C-indices ~0.04-0.09 below the figure-read ~0.69-0.72. Severity is moderate — magnitude and direction hold (q6 yellow). Overall solid but not 1:1: the upstream gene-selection chain and exact final model could not be run because authors deposited no paper-specific code/data, so this is yellow and warrants human review.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.