Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Machine learning-based identification of an immunotherapy-related signature to enhance outcomes and immunotherapy responses in melanoma.

Front Immunol · 2024
L1 69/100 3/4
Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +4
✓ What held up
  • Reported values are derivable from the shared data
  • The central claim held under reproduction
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🟡Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡The deviation was non-trivial in magnitude
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
69/100
Reproducibility score
0.3 SD below mean
vs. all fields · 1173 studies
🎯 Scores higher than 35% of all assessed papers rank 745 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough to reproduce the HEADLINE prognostic result, but only because the 7 signature genes are published — the cited repo ships no paper-specific code/data (just the generic Liu 2022 101-ML template; main branch is a 5-byte 'asss' file). Applying the published 7 genes (GBP5, HLA-DPB1, XBP1, CD40, GBP1, CXCL10, TNFSF13B) to the public cohorts, every gene is independently protective in TCGA-SKCM (HR 0.67-0.73, p<1e-6), and a multivariate-Cox risk score gives C-index 0.63-0.65 with 1-year time-AUC ~0.70 across TCGA-SKCM, GSE54467, GSE22153, GSE65904, plus significant KM stratification in 3/4 cohorts. Reproduced C-indices run ~0.04-0.09 below the figure-reported ~0.69-0.72 (mostly within ~2 SE; 1-year AUCs match), explainable by model choice (multiCox vs Lasso+plsRcox on the 44-gene panel) and survival-endpoint definitions (GSE65904 DSS, GSE54467 death coding). NOT attempted: the upstream gene-selection chain, the exact plsRcox model, and all downstream immune/drug/scRNA analyses. No fabrication indicators — the central signature is a genuine cross-cohort prognostic marker; the gap is exact magnitude, not credibility. Verdict: PARTIAL, provisional, human review required.

💻 Code ↗ 🗄 Data: GSE54467

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 69
    assessed: 2026-06-20 ⛓ b24b3fcdfc6f
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-20
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-20
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The paper tests whether a machine-learning-derived multi-gene signature (ITRGM) built from consensus immunotherapy prognostic genes can robustly predict immunotherapy response and prognosis in melanoma, addressing the lack of reliable biomarkers caused by tumor heterogeneity.

Core claims
  • 66 consensus immunotherapy prognostic genes (CITPGs) were identified from the intersection of WGCNA modules, immunotherapy responder-vs-non-responder DEGs, and tumor-vs-normal DEGs finding
  • CITPG-high tumors show better prognosis and enriched immune activity compared to CITPG-low tumors finding
  • A seven-gene ITRGM signature (GBP5, HLA-DPB1, XBP1, CD40, GBP1, CXCL10, TNFSF13B) was constructed using a Lasso+plsRcox model selected from 101 machine learning algorithm combinations method
  • ITRGM outperforms 37 previously published signatures in predicting immunotherapy prognosis across training, testing, and meta-cohorts finding
  • Low-risk ITRGM patients show higher tumor mutation burden and immune cell infiltration, indicating immune-hot tumors with better prognosis finding
  • GBP5 expression correlates with CD8+ T cell infiltration, validated by IHC/multiplex immunofluorescence in melanoma tissue finding
  • ITRGM predicts immunotherapy response in 8 additional independent cohorts, including urothelial carcinoma and stomach adenocarcinoma finding
  • All seven ITRGM model genes are up-regulated in immunotherapy responders finding
Experimental setups
Assay System Perturbation Readout Platform
WGCNA (weighted gene co-expression network analysis) Melanoma immunotherapy cohort PRJEB23709 none (responder vs non-responder comparison) gene modules correlated with immunotherapy response R package WGCNA
Differential gene expression analysis TCGA-SKCM tumor tissue vs GTEx normal tissue none DEGs (|log2FC|>1, FDR<0.05) GEPIA2
Differential gene expression analysis PRJEB23709 melanoma immunotherapy responders vs non-responders none DEGs (|log2FC|>1, FDR<0.05) limma R package
Consensus clustering TCGA-SKCM melanoma patients none CITPG-high vs CITPG-low patient clusters ConsensusClusterPlus R package
Machine learning model construction (10 algorithms, 101 combinations) TCGA-SKCM (training), GSE22153/GSE54467/GSE69504 (validation) none C-index, risk score for prognosis/immunotherapy response R (Enet, ridge, plsRcox, Lasso, RSF, SuperPC, CoxBoost, GBM, survival-SVM, stepwise Cox)
Multiplex immunofluorescence / IHC Tissue microarray, 17 melanoma and 18 normal skin cases none GBP5 and CD8 co-expression/localization GBP5 (Abcam AB313390), CD8 (Servicebio GB12068) antibodies
Bulk RNA sequencing Mouse melanoma immunotherapy cohorts (GSE109485, GSE149825) immunotherapy ITRGM model gene expression TISMO database
Single-cell RNA sequencing SKCM_GSE115978_aPD1 melanoma dataset anti-PD1 immunotherapy immune landscape correlation with ITRGM genes TISCH2 website
Key results
  • 66 CITPGs identified from intersection of WGCNA module, responder/non-responder DEGs, and tumor/normal DEGs
  • All 66 CITPGs associated with better overall survival by univariate Cox regression in TCGA-SKCM
  • Lasso regression narrowed 44 consensus genes to 7 genes with non-zero coefficients: GBP5, HLA-DPB1, XBP1, CD40, GBP1, CXCL10, TNFSF13B
  • Lasso+plsRcox model had the highest average C-index among 101 algorithm combinations tested
  • ITRGM outperformed 37 published prognostic signatures across training and validation cohorts
  • Low-risk ITRGM group showed increased tumor mutation burden and immune cell infiltration
  • GBP5 expression correlated with CD8+ T cell infiltration across cancer types and validated in melanoma tissue
  • All seven ITRGM genes were up-regulated in immunotherapy responders
Key statistics
  • count 66 (consensus immunotherapy prognostic genes (CITPGs) identified)
  • count 44 (consensus prognostic DEGs validated across four melanoma cohorts)
  • count 7 (final model genes in ITRGM signature)
  • count 101 (machine learning algorithm combinations evaluated)
  • fold_change >4-fold upregulation (DEGs highlighted in volcano plot between immunotherapy responders and non-responders)
  • count 1808 (total cancer patients across 16 independent public cohorts used in study)
  • count 459 (TCGA-SKCM patients retained after excluding incomplete clinical data)
  • other 52% (5-year overall survival rate reported for combinatorial immunotherapy in melanoma (background citation))

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This is a bioinformatics/multi-omics study that used weighted gene co-expression network analysis (WGCNA), differential expression analysis, and consensus clustering to derive a set of candidate genes, then applied univariate/multivariate Cox regression and an ensemble of 10 machine-learning algorithms (101 combinations, evaluated by C-index via 10-fold cross-validation) to build a 7-gene prognostic signature (ITRGM). Performance and biological associations were assessed across multiple public transcriptomic, single-cell, and immunotherapy cohorts using survival analysis, correlation analyses, and immune-infiltration deconvolution tools, with additional IHC/immunofluorescence validation in patient and mouse tissue.

Replicationunclear Sample sizeCohort sizes are stated per dataset (e.g., TCGA-SKCM n=459, GSE54467 n=79, GSE69504 n=214, GSE22153 n=57, GSE243238 n=17), but no formal power/sample-size calculation is described Groupsimmunotherapy responders vs non-responders; CITPG-high vs CITPG-low; ITRGM high-risk vs low-risk Pairingunpaired Randomization/blindingna Dispersionunclear Effect sizesyes Multiplicity correctionFDR (false discovery rate) threshold applied within limma for DEG calling
Statistical tests used
Test Applied to n Assumptions
Differential expression analysis (limma package, |log2FC|>1 and FDR<0.05) responders vs non-responders in PRJEB23709 immunotherapy cohort not stated
Weighted gene co-expression network analysis (WGCNA) with Pearson correlation for module membership vs gene significance gene module–immunotherapy response correlation, PRJEB23709 cohort not stated
Kaplan-Meier survival analysis (with implied log-rank comparison) immunotherapy responders vs non-responders (PRJEB23709) and CITPG-high vs CITPG-low clusters (TCGA-SKCM) TCGA-SKCM n=459 (stated elsewhere in Methods) not stated
Univariate Cox regression 66 CITPGs tested individually against overall survival, TCGA-SKCM 459 not stated
Ensemble of 10 machine-learning survival algorithms across 101 combinations (final model: Lasso + plsRcox), evaluated by C-index with 10-fold cross-validation ITRGM construction; training in TCGA-SKCM, validation in GSE22153, GSE54467, GSE69504 TCGA-SKCM n=459; GSE54467 n=79; GSE69504 n=214; GSE22153 n=57 not stated
Multivariate Cox regression and calibration curve analysis ITRGM risk score across training/validation/meta-cohorts not stated
Approaches that could also have been used
  • DEGs between responders and non-responders were selected using a fixed fold-change and FDR cutoff (|log2FC|>1, FDR<0.05) with limma.
    Could also: A model-based count method such as DESeq2 or edgeR (if raw counts are available) with shrinkage estimation of fold-changes — These approaches can improve effect-size stability for genes with low counts or high variance, which may complement fold-change/FDR filtering, especially in modest-sized cohorts.
  • 66 candidate genes (CITPGs) were each tested individually via univariate Cox regression against overall survival.
    Could also: A penalized multivariate Cox model (e.g., LASSO or elastic net) fit jointly on all candidate genes as an additional screening step — Joint modeling can account for correlation among genes and reduce the number of individually tested hypotheses, which is a consideration whenever many single-gene tests are performed in the same dataset.
  • Model selection across 101 machine-learning combinations relied primarily on the concordance index (C-index) for ranking performance.
    Could also: Complementary metrics such as time-dependent AUC, integrated Brier score, or calibration plots at multiple time points — These add information about calibration and time-varying discrimination that C-index alone does not fully capture, useful when comparing many candidate models.
  • Correlations between risk scores/module features and other variables were assessed with Pearson (module membership vs gene significance) or Spearman (risk score vs immune infiltration) coefficients.
    Could also: Partial correlation or regression adjusting for potential confounders such as tumor purity or sequencing depth — Adjusting for known confounders can help isolate the association of interest when comparing immune infiltration estimates across samples of varying purity.
  • Survival comparisons between groups (e.g., CITPG-high vs CITPG-low, ITRGM risk groups) were visualized with Kaplan-Meier curves.
    Could also: Reporting hazard ratios with 95% confidence intervals alongside the Kaplan-Meier plots, and checking the proportional hazards assumption — This would provide an effect-size estimate with a measure of precision and confirm that the Cox/log-rank framework's assumptions hold across the compared groups.
  • Multiple testing correction (FDR) was applied specifically to the DEG-calling step in limma.
    Could also: Extending an explicit multiple-testing correction (e.g., Benjamini-Hochberg) to other repeated-test families in the workflow, such as the batch of univariate Cox tests across 66 genes — Applying a family-wise or FDR correction across all genes tested in a given analysis step is a standard way to control the overall false-positive rate when many hypotheses are evaluated together.
Software: R/WGCNA · R/limma · R/ConsensusClusterPlus · R/survival · R/sva · R/clusterProfiler

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — PMID 39355255 (ITRGM melanoma immunotherapy signature)

Paper: Machine learning-based identification of an immunotherapy-related signature to enhance outcomes and immunotherapy responses in melanoma. Front Immunol 2024; PMCID PMC11442245; DOI 10.3389/fimmu.2024.1451103. Authors' code: https://github.com/YuBestLab/YuBestLab.github.io Branch with code: 101-machine-learning-algorithm

What the repo actually contains

  • main/master branches: a single 5-byte Index.html (asss\n) — no code.
  • Branch 101-machine-learning-algorithm: the generic "101 machine-learning combination" template from Liu et al. 2022, Nat Commun (file 41467_2022_28421_MOESM4_ESM.xlsx = the 101-model list; ML.R = helper functions RunML/RunEval/CalRiskScore/ExtractVar/scaleData/SimpleHeatmap; scripts.R = driver). This template is NOT customised to this paper and ships no input data and no hard-coded results. The authors applied this third-party framework to their data.
  • Per BRIEF rule 2 (P16), applying the described third-party framework to the paper's own data is a valid reproduction route. We therefore reproduce by re-running the prognostic signature on the public cohorts using the published 7-gene list.

Reported pipeline (from Methods/Results)

  • Training: TCGA-SKCM (N=459 after QC).
  • Validation (untreated melanoma): GSE22153 (57), GSE54467 (79), GSE65904 (214; paper text also writes "GSE69504" — a typo; Fig. 6 uses GSE65904).
  • Signature build: CITPGs (66) → DEGs between CITPG-high/low (566) → univariate Cox across 4 cohorts (44 consensus genes) → 101 ML-combinations, Lasso+plsRcox wins by mean C-index → final 7 genes: GBP5, HLA-DPB1, XBP1, CD40, GBP1, CXCL10, TNFSF13B.
  • Reported performance: C-index ≈ 0.70 (TCGA), 0.71 (GSE54467), 0.72 (GSE22153), 0.69 (GSE65904) (Fig 5B, read approximately); time-AUC 1y 0.70–0.78.

IN SCOPE (pipeline-derived, attempted)

# Result Pipeline Status
C1 7 model genes are individually prognostic in TCGA-SKCM univariate Cox (survival) reproduced
C2 7-gene signature risk score → C-index per cohort multivariate Cox (= template FinalModel='multiCox') + Lasso-Cox, evaluated by Harrell C reproduced (partial: ~0.04–0.09 lower)
C3 Time-dependent AUC (1/3/5y) per cohort timeROC reproduced (1y AUC matches)
C4 Signature significantly stratifies survival (KM high vs low) median split + log-rank partial (3/4 cohorts p<0.05)

OUT OF SCOPE / NOT ATTEMPTED (documented, not graded)

  • Upstream signature selection (WGCNA grey60 on PRJEB23709; GEPIA2 tumor-vs-normal DEGs; CITPG 3-way intersection; 566 consensus DEGs; 44-gene panel; the full 101-model tournament selecting Lasso+plsRcox). Requires several web-tool steps (GEPIA2, TIMER2, TIP) and intermediate gene lists not shipped in the repo → the 44-gene input panel is not derivable from the deposited materials without substantial extra reconstruction. We therefore take the published 7 genes as given and verify their performance.
  • Exact Lasso+plsRcox model: plsRcox could not be installed (CRAN download + heavy Bioc deps failed in env). Substituted the template's default final scoring (multivariate Cox on the final genes) — FinalModel<-c("panML","multiCox")[2] in the authors' own scripts.R.
  • Immunotherapy-response cohorts (GSE91061, PRJEB23709, IMvigor210, …), immune infiltration (CIBERSORT/ssGSEA/ESTIMATE/TIMER2), TIDE, TMB, IPS, oncoPredict drug sensitivity, GSEA, scRNA (GSE115978), CD8 IHC (GSE243238), mouse cohorts — descriptive / web-tool / wet-lab-adjacent downstream analyses; not pipeline-core; not attempted.

Reproduction route (what was run)

«our HPC» «infra» …/reproductions/pmid-39355255/. Cohorts rebuilt to the 7 genes from: TCGA-SKCM HiSeqV2 + curated OS (UCSC Xena); GSE22153/GSE54467/GSE65904 series matrices

  • GPL61
Figures / tables: Fig 4Fig 5BFig 5GFig 5CFig 5A
C1a_genes_prognostic
Reported
7 ITRGM genes protective/prognostic in TCGA-SKCM
Reproduced
all 7 HR 0.67-0.73 per SD, p<1e-6
exact
C2_TCGA
Reported
~0.70
Reproduced
0.651 (SE 0.021)
within tolerance
C2_GSE54467
Reported
~0.71
Reproduced
0.649 (SE 0.043)
partial
C2_GSE22153
Reported
~0.72
Reproduced
0.634 (SE 0.040)
partial
C2_GSE65904
Reported
~0.69
Reproduced
0.645 (SE 0.027)
within tolerance
C3_AUC1y
Reported
0.70-0.78
Reproduced
TCGA 0.701 / GSE54467 0.716 / GSE22153 0.704 / GSE65904 0.642
within tolerance
C4_KM_stratification
Reported
significant high vs low risk
Reproduced
TCGA 5.3e-10, GSE65904 1e-4, GSE22153 0.046 sig; GSE54467 0.063 borderline
partial
S1_Lasso_plsRcox_selection
Reported
Lasso+plsRcox wins 101-model tournament
Reproduced
not attempted (44-gene input panel not reconstructable from deposited materials)
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 69/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟡2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟢5. Derivability / plausibility
🟡6. Severity of the deviation
🟢7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +4

The headline reproduces: applying the published 7-gene ITRGM signature to the public cohorts confirms every gene as protective (HR 0.67-0.73, p<1e-6) with C-index 0.63-0.65, 1-yr AUC ~0.70, and significant KM stratification in 3/4 cohorts — values derivable from shared data, no fabrication indicators (q5/q7 green). The deviations are explainable and sit on the input/method side: our forced multiCox-for-plsRcox substitution and self-defined cohorts/endpoints (the deposited repo is an empty generic template), giving C-indices ~0.04-0.09 below the figure-read ~0.69-0.72. Severity is moderate — magnitude and direction hold (q6 yellow). Overall solid but not 1:1: the upstream gene-selection chain and exact final model could not be run because authors deposited no paper-specific code/data, so this is yellow and warrants human review.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

228.4 k
tokens (I/O) · 10.9 M incl. cache
51 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.