Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

A gene signature related to programmed cell death to predict immunotherapy response and prognosis in colon adenocarcinoma.

PeerJ · 2025
73/100 3/4
Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
How its reproducibility compares
73/100
Reproducibility score
at the mean
vs. all fields · 1173 studies
🎯 Scores higher than 41% of all assessed papers rank 664 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough to reproduce, and reproduces ~1:1 on its headline computational claims from FULLY PUBLIC data (clean re-run, «our HPC» «job», R 4.5.3, 16 min, ExitCode 0). The authors' repo (yangzhandr/Raw-data; zenodo 14015463 mirror) ships their analysis script + clinical/gene-set/result tables but NOT the two expression matrices and NOT the private helper lib it sources (z:/.../mg_base.R), so it is not runnable as-shipped; we re-implemented the well-specified steps (per-gene coxph, limma, ConsensusClusterPlus, timeROC, glmnet) and re-obtained expression from the cited public sources: TCGA-COAD TPM via GDC (STAR-Counts) and AC-ICAM via cBioPortal coad_silu_2022 (git-LFS). EXACT/within-tol: PCD gene count 1254, common PCD genes 1150, the HEADLINE 21 shared prognostic PCD genes (identical list), TCGA samples 428 vs 427. Applying the paper's PUBLISHED six-gene RiskScore coefficients reproduces the reported time-dependent AUCs to two decimals in BOTH cohorts (TCGA 0.688/0.696/0.679/0.718/0.699 vs 0.69/0.70/0.69/0.72/0.69; AC-ICAM 0.720/0.614/0.602/0.585/0.598 vs 0.72/0.61/0.60/0.58/0.60). The univariate-Cox screens correlate with the shipped tables at HR r=1.000 (AC-ICAM) and r=0.933 (TCGA). Count drifts (C3a 101 vs 116, C4 652 vs 484, C5 103 vs 84) are explained by re-obtaining the un-shipped TCGA TPM matrix from GDC: borderline genes cross thresholds differently but direction/effect agree (logFC r=0.970, HR r=0.998; 85-87% of shipped genes recovered, all same sign). The one genuine non-reproduction is C9: cv.glmnet(set.seed 2024) yields a null model on the public matrix, so the exact 6-gene re-selection is not reproducible (inherently seed/data-sensitive) — its performance is independently confirmed via C6/C7. NO fabrication indicators. NOT ATTEMPTED (external/un-shipped): TIDE/aneuploidy immunotherapy prediction, CIBERSORT/MCP-counter/TIMER2.0 TME, hallmark GSEA + mutation/MAF, nomogram calibration/DCA, all wet-lab assays.

💻 Code ↗ 🗄 Data: 10.5281/zenodo.14015463

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 73
    assessed: 2026-06-16 ⛓ 3a5925bf5fc8
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-22
Rubric version
not recorded
Assessed by
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

Programmed cell death (PCD)-related genes are correlated with colon adenocarcinoma (COAD) prognosis and immunotherapy response, and can be used to build a gene signature/RiskScore model for prognostic and immunotherapy prediction.

Core claims
  • COAD patients can be divided into two molecular subtypes (S1, S2) based on 21 prognostic PCD-related genes, with S1 showing worse prognosis and immunosuppressive microenvironment, S2 showing better prognosis and stronger anti-tumor immunity finding
  • A six-gene RiskScore model (protective: ATOH1, ZG16; risk: HSPA1A, SEMA4C, CDKN2A, ARHGAP4) stratifies COAD patients into high- and low-risk groups with distinct survival outcomes method
  • Silencing of CDKN2A inhibits migration and invasion and promotes apoptosis in COAD tumor cells finding
  • High-risk group patients show higher expression of immune checkpoint genes and higher TIDE scores, indicating stronger immune escape ability and less active immunotherapy response than low-risk group finding
  • A nomogram integrating Age, pathologic_M, pathologic_stage, and RiskScore shows strong and accurate prognostic prediction performance for COAD method
  • TP53 exhibits a higher somatic mutation rate in the high-risk group compared to the low-risk group finding
  • S1 subtype is enriched in tumor invasion/progression pathways (myogenesis, angiogenesis, EMT), while S2 subtype is enriched in cell proliferation pathways (G2M checkpoint, MYC targets V1, E2F targets) finding
  • 1,254 PCD-associated genes curated from a prior study were used as the candidate gene set for analysis resource
Experimental setups
Assay System Perturbation Readout Platform
Consensus clustering / bioinformatic subtyping (RNA-seq) COAD patient tumor samples (TCGA, n=427; AC-ICAM validation, n=341) none PCD-related prognostic gene expression, molecular subtype assignment (S1/S2)
CIBERSORT immune deconvolution COAD tumor samples (TCGA) none relative abundance of 22 immune cell types CIBERSORT
MCP-counter immune scoring COAD tumor samples (TCGA) none abundance scores of 3 stromal and 8 immune cell types MCP-counter
GSEA pathway enrichment COAD tumor samples (TCGA), S1 vs S2 subtypes none differentially enriched HALLMARK pathways GSEA / MSigDB HALLMARK gene set
Cox regression / LASSO Cox / stepwise regression (RiskScore modeling) COAD tumor samples (TCGA DEGs; validated in AC-ICAM) none prognostic gene selection, RiskScore coefficients, survival prediction (ROC/AUC, OS, DSS, PFI) R package glmnet
Somatic mutation analysis (SNV) COAD tumor samples (TCGA, high- vs low-risk groups) none somatic gene mutation rate differences (e.g., TP53) maftools, Mutect2
Wound healing and transwell invasion/migration assay human COAD cell line SW1116 si-CDKN2A knockdown wound closure width, number of invading/migrating cells inverted microscope
Flow cytometry apoptosis assay human COAD cell line SW1116 (and NCM460 normal colonic epithelial cells for culture) si-CDKN2A knockdown apoptotic cell percentage (Annexin V-FITC/PI staining) BD FACSAria III
Key results
  • 427 TCGA COAD samples divided into S1 and S2 subtypes at consensus matrix k=2; S1 had worse prognosis, S2 had more favorable prognosis
  • Six-gene RiskScore model formula: RiskScore=(0.208*HSPA1A)+(0.273*SEMA4C)+(−0.087*ATOH1)+(0.154*CDKN2A)+(0.163*ARHGAP4)+(−0.073*ZG16); low-risk group had better OS, DSS, and PFI than high-risk group
  • RiskScore model ROC AUC for 1-5 year survival in TCGA cohort: 0.69, 0.70, 0.69, 0.72, 0.69; in AC-ICAM validation cohort: 0.72, 0.61, 0.60, 0.58, 0.60 AUC ~0.6-0.72
  • si-CDKN2A silencing in SW1116 cells inhibited migration and invasion (wound healing, transwell) and promoted apoptosis (flow cytometry)
  • Four immune checkpoint genes (PDCD1, HAVCR2, TIGIT, LAG3) had higher expression in high-risk group; high-risk group also had higher TIDE scores
  • Nomogram combining Age, pathologic_M, pathologic_stage, and RiskScore produced calibration curves close to standard curves for 1-, 3-, and 5-year survival, and DCA showed higher net benefit than extreme curves
  • TP53 showed a higher rate of somatic mutation in the high-risk group compared to the low-risk group
  • S1 subtype enriched in myogenesis, angiogenesis, and EMT pathways; S2 subtype enriched in G2M checkpoint, MYC targets V1, and E2F targets pathways
Key statistics
  • count 427 TCGA COAD samples; 341 AC-ICAM samples (cohort sizes with survival time >30 days)
  • count 1,254 PCD-associated genes (candidate PCD gene set used for analysis)
  • count 21 PCD-related prognostic genes (genes used for consensus clustering into S1/S2 subtypes)
  • count 90 downregulated and 394 upregulated DEGs (484 total) (DEGs between S1 and S2 subtypes in TCGA, |log2FC|>log2(1.5), Padj<0.05)
  • other AUC 0.69, 0.70, 0.69, 0.72, 0.69 (years 1-5) (RiskScore model ROC performance in TCGA cohort)
  • other AUC 0.72, 0.61, 0.60, 0.58, 0.60 (years 1-5) (RiskScore model ROC performance in AC-ICAM validation cohort)
  • other coefficients: 0.208 (HSPA1A), 0.273 (SEMA4C), −0.087 (ATOH1), 0.154 (CDKN2A), 0.163 (ARHGAP4), −0.073 (ZG16) (RiskScore model gene coefficients)
  • pvalue p < 0.05 (significance threshold for Cox regression and survival analyses)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The study used consensus clustering (ConsensusClusterPlus, PAM/Pearson, 500 iterations) on 427 TCGA COAD samples to define two molecular subtypes from 21 PCD-related prognostic genes, then characterized subtypes via limma-based DEG analysis, GSEA, and CIBERSORT/MCP-counter immune deconvolution. A penalized Cox regression pipeline (univariate filter → LASSO → stepwise) derived a six-gene RiskScore model, validated in an external AC-ICAM cohort (n=341) by time-dependent ROC and Kaplan-Meier log-rank tests, and combined with clinical features in a nomogram. In vitro siRNA knockdown experiments were analyzed with Student's t-test, and bioinformatics group comparisons used the Wilcoxon rank-sum test; results were reported primarily as p-value thresholds and AUC/hazard ratio point estimates.

Replicationmixed Sample size427 TCGA COAD samples with survival > 30 days (discovery); 341 AC-ICAM samples with survival > 30 days (external validation); in vitro experiments used SW1116 cell line with si-CDKN2A knockdown; number of biological replicates for in vitro assays not stated GroupsTwo molecular subtypes (S1 vs S2); high-risk vs low-risk (median RiskScore split); in vitro si-CDKN2A knockdown vs control in SW1116 cells Pairingunpaired Randomization/blindingnot stated Dispersionnone Exact p-valuesno Effect sizesyes Confidence intervalsyes Multiplicity correctionAdjusted p-values (P_adj) applied for DEG analysis (limma), GSEA, and KEGG enrichment; specific correction method (e.g., Benjamini-Hochberg) not named; no multiplicity correction stated for the series of Cox regression screens, the seven immune checkpoint gene comparisons, or in vitro t-tests
Statistical tests used
Test Applied to n Assumptions
Univariate Cox proportional hazards regression Screening 1,254 PCD-related genes for prognostic relevance; screening 484 DEGs before LASSO; identifying independent prognostic factors for nomogram 427 (TCGA cohort) not stated
LASSO Cox regression (glmnet) Feature selection to reduce prognostically significant DEGs for RiskScore model 427 (TCGA cohort) not stated
Stepwise Cox regression Final selection of 6 RiskScore genes and their coefficients from LASSO candidates 427 (TCGA cohort) not stated
Multivariate Cox proportional hazards regression Confirmation of independent prognostic factors (Age, pathologic_M, pathologic_stage, RiskScore) for nomogram construction 427 (TCGA cohort) not stated
Log-rank test (Kaplan-Meier) Survival differences between S1/S2 subtypes (TCGA and AC-ICAM); OS, DSS, and PFI differences between high- and low-risk groups 427 (TCGA), 341 (AC-ICAM) not stated
limma moderated t-statistic (DEG analysis) Differential expression between S1 and S2 subtypes; threshold |log2FC| > log2(1.5), P_adj < 0.05 427 (TCGA cohort) not stated
GSEA (Gene Set Enrichment Analysis) Pathway enrichment comparison between S1 and S2 subtypes using HALLMARK gene sets; P_adj < 0.05 427 (TCGA cohort) not stated
KEGG enrichment analysis (clusterProfiler) Functional annotation of upregulated and downregulated DEGs between subtypes; P_adj < 0.05 null not stated
Wilcoxon rank-sum test Differences between two-group continuous variables in bioinformatics analyses (e.g., immune infiltration scores, RiskScore across clinical stages) null not stated
Student's t-test (two-sample) In vitro experimental comparisons (wound healing, transwell invasion, flow cytometry apoptosis assays with si-CDKN2A) null not stated
Spearman correlation Correlation between RiskScore and MCP-counter immune cell scores null not stated
Time-dependent ROC curve (timeROC package) 1- to 5-year predictive performance (AUC) of RiskScore model in TCGA and AC-ICAM cohorts 427 (TCGA), 341 (AC-ICAM) na
mafCompare (Fisher's exact test on somatic mutation rates) Comparison of somatic gene mutation rates between high- and low-risk groups null not stated
Approaches that could also have been used
  • The high/low-risk group boundary was set at the median RiskScore of the training (TCGA) cohort
    Could also: Optimal cutpoint selection methods (e.g., maximally selected rank statistics via R package survminer::surv_cutpoint, or pre-specified clinical quantiles) could also be used — Median splitting is cohort-dependent and sensitive to the score distribution in the training set; data-driven optimal thresholds or clinically pre-defined quantiles can improve reproducibility and reduce circularity when the same data inform both the model and the threshold
  • Feature selection combined LASSO Cox regression with subsequent stepwise (forward/backward) Cox regression to arrive at the six-gene model
    Could also: Retaining the LASSO-selected gene set directly, or applying elastic-net penalized Cox regression, could also be used — Stepwise selection can introduce model instability and selection-induced optimism; elastic-net balances LASSO sparsity with ridge-type shrinkage and is often recommended for genomic data where predictors are correlated
  • In vitro experimental comparisons used Student's t-test with the number of biological replicates unstated
    Could also: A non-parametric Wilcoxon/Mann-Whitney U test, or reporting results from a defined number of independent biological replicates with SD as the dispersion measure, could also be used — When replicate count is small (common in cell-line work), normality assumptions underlying the t-test are difficult to verify; a non-parametric alternative requires no distributional assumption, and explicit reporting of biological replicate number and dispersion allows readers to assess variability
  • Survival differences between risk groups were assessed with the log-rank test
    Could also: Restricted mean survival time (RMST) analysis could also quantify survival differences — The log-rank test assumes proportional hazards; RMST provides an absolute, time-horizon-specific summary measure of survival benefit that does not require this assumption and is directly interpretable in units of time
  • Seven immune checkpoint genes were compared between high- and low-risk groups without a stated multiplicity correction
    Could also: Applying a Benjamini-Hochberg FDR correction across the seven simultaneous comparisons could also be used — Testing multiple genes simultaneously increases the probability of at least one false positive; controlling the false discovery rate limits the expected fraction of spurious findings among those declared significant
  • Immune cell infiltration was estimated using two algorithms (CIBERSORT and MCP-counter) independently
    Could also: Additional deconvolution methods such as TIMER2.0, xCell, or EPIC could also be incorporated, and inter-algorithm concordance could be reported — Different deconvolution algorithms use distinct reference profiles and mathematical assumptions; reporting concordance across multiple methods strengthens confidence in immune microenvironment conclusions and highlights cell types where estimates agree or diverge
Software: R 4.2.0 · GraphPad Prism 8.0 · ConsensusClusterPlus (R package) · glmnet (R package) · timeROC (R package) · rms (R package) · maftools (R package) · limma (R package) · clusterProfiler (R package) · CIBERSORT · MCP-counter · TIDE

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Authors · 4
Citations
2
Impact: low
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-39950044

Paper: Zheng L, Lu J, Kong D, Zhan Y. A gene signature related to programmed cell death to predict immunotherapy response and prognosis in colon adenocarcinoma. PeerJ 2025;13:e18895. DOI 10.7717/peerj.18895.

Code: github.com/yangzhandr/Raw-data (authors' own; scripts/COAD_PCD.R, 1395 lines). Data mirror: zenodo 10.5281/zenodo.14015463 (= a zip snapshot of the same repo).

What the repo actually ships

  • scripts/COAD_PCD.R — the full analysis pipeline (authors' own code).
  • 00.origin_datas/PCD.geneSets.PMID36341760.xlsx — the 1,254 PCD gene sets (Zou 2022).
  • 00.origin_datas/TCGA/ — TCGA-COADREAD clinical + survival tables (NOT expression).
  • 00.origin_datas/coad_silu_2022/ — AC-ICAM clinical patient+sample (NOT expression).
  • 0*/*.csv|*.xlsx|*.pdf — the RESULT files (cox tables, DEGs, TIDE, figures).
  • APOPTOSIS/, raw experimental data/ — wet-lab .fcs / .tif (out of scope).

Runnability / gaps (honesty)

  • Script is NOT runnable as-shipped: it source('z:/projects/codes/mg_base.R') (a private Windows-path helper lib providing cox_batch, mg_limma_DEG, ggplotTimeROC, crbind2DataFrame, etc.) and reads two expression matrices that are not in the repo:
    • origin_datas/TCGA/Merge_TCGA-COAD_TPM.txt → re-obtained from GDC (TPM, the source the Methods cite).
    • origin_datas/coad_silu_2022/data_mrna_seq_expression.txt → re-obtained from cBioPortal datahub (study coad_silu_2022, git-LFS).
  • Therefore we re-implement the well-specified steps in standalone R (the helper functions are simple: per-gene coxph, limma, timeROC), using the shipped clinical/gene-set data + the public expression matrices.

In scope (pipeline-derived, attempted)

# result pipeline data needed
C1 427 TCGA COAD samples (OS>30d) clinical filter TCGA clin (shipped) + expr (GDC)
C2 1,254 PCD genes read gene-set xlsx shipped xlsx
C3a TCGA univariate PCD Cox: 116 sig (p<0.05) per-gene coxph TCGA expr + clin
C3b AC-ICAM univariate PCD Cox: 145 sig per-gene coxph AC-ICAM expr + clin
C3c 21 prognostic PCD genes (C3a ∩ C3b) intersection both
C4 subtype DEGs: 484 (394 up / 90 down) consensus k=2 + limma TCGA expr
C5 84 sig genes from 484 DEGs univariate Cox (coxResult.csv) per-gene coxph TCGA expr
C6 TCGA ROC AUC 1–5y: 0.69/0.70/0.69/0.72/0.69 published 6-gene RiskScore → timeROC TCGA expr
C7 AC-ICAM ROC AUC 1–5y: 0.72/0.61/0.60/0.58/0.60 published 6-gene RiskScore → timeROC AC-ICAM expr

C6/C7 use the paper-published coefficients RiskScore = 0.208·HSPA1A + 0.273·SEMA4C − 0.087·ATOH1 + 0.154·CDKN2A + 0.163·ARHGAP4 − 0.073·ZG16 so they test the reported predictive performance directly (no LASSO randomness).

Out of scope / NOT attempted (the hard ~20%) — and why

  • Exact LASSO 6-gene selection from scratch: glmnet LASSO (set.seed(2024)) + step() is version-sensitive; selecting the identical 6 genes is not robustly reproducible. We instead verify the published signature's performance (C6/C7) and report the coefficients as given.
  • Nomogram calibration / DCA / C-index figures: cosmetic, depend on rms + the private helper; low information vs effort.
  • TIDE / aneuploidy immunotherapy prediction: TIDE requires uploading the expression to the external tide.dfci.harvard.edu server (interactive); the shipped tcga.tide.res.csv is the server output, not pipeline-derivable here.
  • CIBERSORT / MCP-counter TME: needs the external infiltration_estimation_for_tcga.csv (TIMER2.0), not shipped.
  • GSEA (hallmark), mutation/maf comparison: need external gmt + a *_SNP.Rdata not in the repo; low priority.
  • Wet-lab (si-CDKN2A wound/transwell/qPCR, FCS apoptosis): manual, out of scope.

Pointers

  • repo: github.com/yangzhandr/Raw-data @ default branch main (last push
C1
Reported
427 TCGA-COAD samples (OS>30d)
Reproduced
428
within tolerance
C2
Reported
1254 PCD genes
Reproduced
1254
exact
C3common
Reported
1150 common PCD genes
Reproduced
1150
exact
C3a
Reported
116 TCGA univariate-Cox sig PCD genes
Reproduced
101 (HR r=0.933 vs shipped; 93/116 shipped recovered)
partial
C3b
Reported
145 AC-ICAM univariate-Cox sig PCD genes
Reproduced
144 (HR r=1.000 vs shipped; 144/145 shipped recovered)
within tolerance
C3c
Reported
21 shared prognostic PCD genes (HEADLINE)
Reproduced
21 (identical gene list)
exact
C4
Reported
484 subtype DEGs (394 up/90 down)
Reproduced
652 (586/66); 410/484 shipped recovered (85%), logFC r=0.970, 410/410 same sign
partial
C5
Reported
84 DEG univariate-Cox sig genes
Reproduced
103; 73/84 shipped recovered (87%), HR r=0.998
partial
C6
Reported
TCGA ROC AUC 1-5y 0.69/0.70/0.69/0.72/0.69
Reproduced
0.688/0.696/0.679/0.718/0.699
within tolerance
C7
Reported
AC-ICAM ROC AUC 1-5y 0.72/0.61/0.60/0.58/0.60
Reproduced
0.720/0.614/0.602/0.585/0.598
within tolerance
C9
Reported
6-gene LASSO(seed 2024)+step: HSPA1A,SEMA4C,ATOH1,CDKN2A,ARHGAP4,ZG16
Reproduced
cv.glmnet lambda.min=lambda.1se=0.0788 -> null (0 genes); nearest non-null step() -> HSPA1A,ATOH1,NUMBL (2/6 published recovered)
did not match

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

No assessment has been recorded yet.
🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

637.7 k
tokens (I/O) · 65.8 M incl. cache
156 min
runtime · 0.06 CPU-h
10.4 GB
peak RAM
1
HPC jobs
hummel
machine