A gene signature related to programmed cell death to predict immunotherapy response and prognosis in colon adenocarcinoma.
The main results reproduced, with only marginal, non-material deviations.
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Described well enough to reproduce, and reproduces ~1:1 on its headline computational claims from FULLY PUBLIC data (clean re-run, «our HPC» «job», R 4.5.3, 16 min, ExitCode 0). The authors' repo (yangzhandr/Raw-data; zenodo 14015463 mirror) ships their analysis script + clinical/gene-set/result tables but NOT the two expression matrices and NOT the private helper lib it sources (z:/.../mg_base.R), so it is not runnable as-shipped; we re-implemented the well-specified steps (per-gene coxph, limma, ConsensusClusterPlus, timeROC, glmnet) and re-obtained expression from the cited public sources: TCGA-COAD TPM via GDC (STAR-Counts) and AC-ICAM via cBioPortal coad_silu_2022 (git-LFS). EXACT/within-tol: PCD gene count 1254, common PCD genes 1150, the HEADLINE 21 shared prognostic PCD genes (identical list), TCGA samples 428 vs 427. Applying the paper's PUBLISHED six-gene RiskScore coefficients reproduces the reported time-dependent AUCs to two decimals in BOTH cohorts (TCGA 0.688/0.696/0.679/0.718/0.699 vs 0.69/0.70/0.69/0.72/0.69; AC-ICAM 0.720/0.614/0.602/0.585/0.598 vs 0.72/0.61/0.60/0.58/0.60). The univariate-Cox screens correlate with the shipped tables at HR r=1.000 (AC-ICAM) and r=0.933 (TCGA). Count drifts (C3a 101 vs 116, C4 652 vs 484, C5 103 vs 84) are explained by re-obtaining the un-shipped TCGA TPM matrix from GDC: borderline genes cross thresholds differently but direction/effect agree (logFC r=0.970, HR r=0.998; 85-87% of shipped genes recovered, all same sign). The one genuine non-reproduction is C9: cv.glmnet(set.seed 2024) yields a null model on the public matrix, so the exact 6-gene re-selection is not reproducible (inherently seed/data-sensitive) — its performance is independently confirmed via C6/C7. NO fabrication indicators. NOT ATTEMPTED (external/un-shipped): TIDE/aneuploidy immunotherapy prediction, CIBERSORT/MCP-counter/TIMER2.0 TME, hallmark GSEA + mutation/MAF, nomogram calibration/DCA, all wet-lab assays.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 73assessed: 2026-06-16 ⛓ 3a5925bf5fc8
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-22
- Rubric version
- not recorded
- Assessed by
- —
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetProgrammed cell death (PCD)-related genes are correlated with colon adenocarcinoma (COAD) prognosis and immunotherapy response, and can be used to build a gene signature/RiskScore model for prognostic and immunotherapy prediction.
- ★ COAD patients can be divided into two molecular subtypes (S1, S2) based on 21 prognostic PCD-related genes, with S1 showing worse prognosis and immunosuppressive microenvironment, S2 showing better prognosis and stronger anti-tumor immunity finding
- ★ A six-gene RiskScore model (protective: ATOH1, ZG16; risk: HSPA1A, SEMA4C, CDKN2A, ARHGAP4) stratifies COAD patients into high- and low-risk groups with distinct survival outcomes method
- ★ Silencing of CDKN2A inhibits migration and invasion and promotes apoptosis in COAD tumor cells finding
- ★ High-risk group patients show higher expression of immune checkpoint genes and higher TIDE scores, indicating stronger immune escape ability and less active immunotherapy response than low-risk group finding
- ★ A nomogram integrating Age, pathologic_M, pathologic_stage, and RiskScore shows strong and accurate prognostic prediction performance for COAD method
- ★ TP53 exhibits a higher somatic mutation rate in the high-risk group compared to the low-risk group finding
- S1 subtype is enriched in tumor invasion/progression pathways (myogenesis, angiogenesis, EMT), while S2 subtype is enriched in cell proliferation pathways (G2M checkpoint, MYC targets V1, E2F targets) finding
- 1,254 PCD-associated genes curated from a prior study were used as the candidate gene set for analysis resource
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Consensus clustering / bioinformatic subtyping (RNA-seq) | COAD patient tumor samples (TCGA, n=427; AC-ICAM validation, n=341) | none | PCD-related prognostic gene expression, molecular subtype assignment (S1/S2) | — |
| CIBERSORT immune deconvolution | COAD tumor samples (TCGA) | none | relative abundance of 22 immune cell types | CIBERSORT |
| MCP-counter immune scoring | COAD tumor samples (TCGA) | none | abundance scores of 3 stromal and 8 immune cell types | MCP-counter |
| GSEA pathway enrichment | COAD tumor samples (TCGA), S1 vs S2 subtypes | none | differentially enriched HALLMARK pathways | GSEA / MSigDB HALLMARK gene set |
| Cox regression / LASSO Cox / stepwise regression (RiskScore modeling) | COAD tumor samples (TCGA DEGs; validated in AC-ICAM) | none | prognostic gene selection, RiskScore coefficients, survival prediction (ROC/AUC, OS, DSS, PFI) | R package glmnet |
| Somatic mutation analysis (SNV) | COAD tumor samples (TCGA, high- vs low-risk groups) | none | somatic gene mutation rate differences (e.g., TP53) | maftools, Mutect2 |
| Wound healing and transwell invasion/migration assay | human COAD cell line SW1116 | si-CDKN2A knockdown | wound closure width, number of invading/migrating cells | inverted microscope |
| Flow cytometry apoptosis assay | human COAD cell line SW1116 (and NCM460 normal colonic epithelial cells for culture) | si-CDKN2A knockdown | apoptotic cell percentage (Annexin V-FITC/PI staining) | BD FACSAria III |
- – 427 TCGA COAD samples divided into S1 and S2 subtypes at consensus matrix k=2; S1 had worse prognosis, S2 had more favorable prognosis
- – Six-gene RiskScore model formula: RiskScore=(0.208*HSPA1A)+(0.273*SEMA4C)+(−0.087*ATOH1)+(0.154*CDKN2A)+(0.163*ARHGAP4)+(−0.073*ZG16); low-risk group had better OS, DSS, and PFI than high-risk group
- – RiskScore model ROC AUC for 1-5 year survival in TCGA cohort: 0.69, 0.70, 0.69, 0.72, 0.69; in AC-ICAM validation cohort: 0.72, 0.61, 0.60, 0.58, 0.60 AUC ~0.6-0.72
- ▼ si-CDKN2A silencing in SW1116 cells inhibited migration and invasion (wound healing, transwell) and promoted apoptosis (flow cytometry)
- ▲ Four immune checkpoint genes (PDCD1, HAVCR2, TIGIT, LAG3) had higher expression in high-risk group; high-risk group also had higher TIDE scores
- – Nomogram combining Age, pathologic_M, pathologic_stage, and RiskScore produced calibration curves close to standard curves for 1-, 3-, and 5-year survival, and DCA showed higher net benefit than extreme curves
- ▲ TP53 showed a higher rate of somatic mutation in the high-risk group compared to the low-risk group
- – S1 subtype enriched in myogenesis, angiogenesis, and EMT pathways; S2 subtype enriched in G2M checkpoint, MYC targets V1, and E2F targets pathways
- count 427 TCGA COAD samples; 341 AC-ICAM samples (cohort sizes with survival time >30 days)
- count 1,254 PCD-associated genes (candidate PCD gene set used for analysis)
- count 21 PCD-related prognostic genes (genes used for consensus clustering into S1/S2 subtypes)
- count 90 downregulated and 394 upregulated DEGs (484 total) (DEGs between S1 and S2 subtypes in TCGA, |log2FC|>log2(1.5), Padj<0.05)
- other AUC 0.69, 0.70, 0.69, 0.72, 0.69 (years 1-5) (RiskScore model ROC performance in TCGA cohort)
- other AUC 0.72, 0.61, 0.60, 0.58, 0.60 (years 1-5) (RiskScore model ROC performance in AC-ICAM validation cohort)
- other coefficients: 0.208 (HSPA1A), 0.273 (SEMA4C), −0.087 (ATOH1), 0.154 (CDKN2A), 0.163 (ARHGAP4), −0.073 (ZG16) (RiskScore model gene coefficients)
- pvalue p < 0.05 (significance threshold for Cox regression and survival analyses)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
The study used consensus clustering (ConsensusClusterPlus, PAM/Pearson, 500 iterations) on 427 TCGA COAD samples to define two molecular subtypes from 21 PCD-related prognostic genes, then characterized subtypes via limma-based DEG analysis, GSEA, and CIBERSORT/MCP-counter immune deconvolution. A penalized Cox regression pipeline (univariate filter → LASSO → stepwise) derived a six-gene RiskScore model, validated in an external AC-ICAM cohort (n=341) by time-dependent ROC and Kaplan-Meier log-rank tests, and combined with clinical features in a nomogram. In vitro siRNA knockdown experiments were analyzed with Student's t-test, and bioinformatics group comparisons used the Wilcoxon rank-sum test; results were reported primarily as p-value thresholds and AUC/hazard ratio point estimates.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Univariate Cox proportional hazards regression | Screening 1,254 PCD-related genes for prognostic relevance; screening 484 DEGs before LASSO; identifying independent prognostic factors for nomogram | 427 (TCGA cohort) | not stated |
| LASSO Cox regression (glmnet) | Feature selection to reduce prognostically significant DEGs for RiskScore model | 427 (TCGA cohort) | not stated |
| Stepwise Cox regression | Final selection of 6 RiskScore genes and their coefficients from LASSO candidates | 427 (TCGA cohort) | not stated |
| Multivariate Cox proportional hazards regression | Confirmation of independent prognostic factors (Age, pathologic_M, pathologic_stage, RiskScore) for nomogram construction | 427 (TCGA cohort) | not stated |
| Log-rank test (Kaplan-Meier) | Survival differences between S1/S2 subtypes (TCGA and AC-ICAM); OS, DSS, and PFI differences between high- and low-risk groups | 427 (TCGA), 341 (AC-ICAM) | not stated |
| limma moderated t-statistic (DEG analysis) | Differential expression between S1 and S2 subtypes; threshold |log2FC| > log2(1.5), P_adj < 0.05 | 427 (TCGA cohort) | not stated |
| GSEA (Gene Set Enrichment Analysis) | Pathway enrichment comparison between S1 and S2 subtypes using HALLMARK gene sets; P_adj < 0.05 | 427 (TCGA cohort) | not stated |
| KEGG enrichment analysis (clusterProfiler) | Functional annotation of upregulated and downregulated DEGs between subtypes; P_adj < 0.05 | null | not stated |
| Wilcoxon rank-sum test | Differences between two-group continuous variables in bioinformatics analyses (e.g., immune infiltration scores, RiskScore across clinical stages) | null | not stated |
| Student's t-test (two-sample) | In vitro experimental comparisons (wound healing, transwell invasion, flow cytometry apoptosis assays with si-CDKN2A) | null | not stated |
| Spearman correlation | Correlation between RiskScore and MCP-counter immune cell scores | null | not stated |
| Time-dependent ROC curve (timeROC package) | 1- to 5-year predictive performance (AUC) of RiskScore model in TCGA and AC-ICAM cohorts | 427 (TCGA), 341 (AC-ICAM) | na |
| mafCompare (Fisher's exact test on somatic mutation rates) | Comparison of somatic gene mutation rates between high- and low-risk groups | null | not stated |
-
The high/low-risk group boundary was set at the median RiskScore of the training (TCGA) cohort↳ Could also: Optimal cutpoint selection methods (e.g., maximally selected rank statistics via R package survminer::surv_cutpoint, or pre-specified clinical quantiles) could also be used — Median splitting is cohort-dependent and sensitive to the score distribution in the training set; data-driven optimal thresholds or clinically pre-defined quantiles can improve reproducibility and reduce circularity when the same data inform both the model and the threshold
-
Feature selection combined LASSO Cox regression with subsequent stepwise (forward/backward) Cox regression to arrive at the six-gene model↳ Could also: Retaining the LASSO-selected gene set directly, or applying elastic-net penalized Cox regression, could also be used — Stepwise selection can introduce model instability and selection-induced optimism; elastic-net balances LASSO sparsity with ridge-type shrinkage and is often recommended for genomic data where predictors are correlated
-
In vitro experimental comparisons used Student's t-test with the number of biological replicates unstated↳ Could also: A non-parametric Wilcoxon/Mann-Whitney U test, or reporting results from a defined number of independent biological replicates with SD as the dispersion measure, could also be used — When replicate count is small (common in cell-line work), normality assumptions underlying the t-test are difficult to verify; a non-parametric alternative requires no distributional assumption, and explicit reporting of biological replicate number and dispersion allows readers to assess variability
-
Survival differences between risk groups were assessed with the log-rank test↳ Could also: Restricted mean survival time (RMST) analysis could also quantify survival differences — The log-rank test assumes proportional hazards; RMST provides an absolute, time-horizon-specific summary measure of survival benefit that does not require this assumption and is directly interpretable in units of time
-
Seven immune checkpoint genes were compared between high- and low-risk groups without a stated multiplicity correction↳ Could also: Applying a Benjamini-Hochberg FDR correction across the seven simultaneous comparisons could also be used — Testing multiple genes simultaneously increases the probability of at least one false positive; controlling the false discovery rate limits the expected fraction of spurious findings among those declared significant
-
Immune cell infiltration was estimated using two algorithms (CIBERSORT and MCP-counter) independently↳ Could also: Additional deconvolution methods such as TIMER2.0, xCell, or EPIC could also be incorporated, and inter-algorithm concordance could be reported — Different deconvolution algorithms use distinct reference profiles and mathematical assumptions; reporting concordance across multiple methods strengthens confidence in immune microenvironment conclusions and highlights cell types where estimates agree or diverge
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-39950044
Paper: Zheng L, Lu J, Kong D, Zhan Y. A gene signature related to programmed cell death to predict immunotherapy response and prognosis in colon adenocarcinoma. PeerJ 2025;13:e18895. DOI 10.7717/peerj.18895.
Code: github.com/yangzhandr/Raw-data (authors' own; scripts/COAD_PCD.R, 1395 lines).
Data mirror: zenodo 10.5281/zenodo.14015463 (= a zip snapshot of the same repo).
What the repo actually ships
scripts/COAD_PCD.R— the full analysis pipeline (authors' own code).00.origin_datas/PCD.geneSets.PMID36341760.xlsx— the 1,254 PCD gene sets (Zou 2022).00.origin_datas/TCGA/— TCGA-COADREAD clinical + survival tables (NOT expression).00.origin_datas/coad_silu_2022/— AC-ICAM clinical patient+sample (NOT expression).0*/*.csv|*.xlsx|*.pdf— the RESULT files (cox tables, DEGs, TIDE, figures).APOPTOSIS/,raw experimental data/— wet-lab .fcs / .tif (out of scope).
Runnability / gaps (honesty)
- Script is NOT runnable as-shipped: it
source('z:/projects/codes/mg_base.R')(a private Windows-path helper lib providingcox_batch,mg_limma_DEG,ggplotTimeROC,crbind2DataFrame, etc.) and reads two expression matrices that are not in the repo:origin_datas/TCGA/Merge_TCGA-COAD_TPM.txt→ re-obtained from GDC (TPM, the source the Methods cite).origin_datas/coad_silu_2022/data_mrna_seq_expression.txt→ re-obtained from cBioPortal datahub (studycoad_silu_2022, git-LFS).
- Therefore we re-implement the well-specified steps in standalone R (the
helper functions are simple: per-gene
coxph,limma,timeROC), using the shipped clinical/gene-set data + the public expression matrices.
In scope (pipeline-derived, attempted)
| # | result | pipeline | data needed |
|---|---|---|---|
| C1 | 427 TCGA COAD samples (OS>30d) | clinical filter | TCGA clin (shipped) + expr (GDC) |
| C2 | 1,254 PCD genes | read gene-set xlsx | shipped xlsx |
| C3a | TCGA univariate PCD Cox: 116 sig (p<0.05) | per-gene coxph | TCGA expr + clin |
| C3b | AC-ICAM univariate PCD Cox: 145 sig | per-gene coxph | AC-ICAM expr + clin |
| C3c | 21 prognostic PCD genes (C3a ∩ C3b) | intersection | both |
| C4 | subtype DEGs: 484 (394 up / 90 down) | consensus k=2 + limma | TCGA expr |
| C5 | 84 sig genes from 484 DEGs univariate Cox (coxResult.csv) |
per-gene coxph | TCGA expr |
| C6 | TCGA ROC AUC 1–5y: 0.69/0.70/0.69/0.72/0.69 | published 6-gene RiskScore → timeROC | TCGA expr |
| C7 | AC-ICAM ROC AUC 1–5y: 0.72/0.61/0.60/0.58/0.60 | published 6-gene RiskScore → timeROC | AC-ICAM expr |
C6/C7 use the paper-published coefficients
RiskScore = 0.208·HSPA1A + 0.273·SEMA4C − 0.087·ATOH1 + 0.154·CDKN2A + 0.163·ARHGAP4 − 0.073·ZG16
so they test the reported predictive performance directly (no LASSO randomness).
Out of scope / NOT attempted (the hard ~20%) — and why
- Exact LASSO 6-gene selection from scratch: glmnet LASSO (set.seed(2024)) +
step()is version-sensitive; selecting the identical 6 genes is not robustly reproducible. We instead verify the published signature's performance (C6/C7) and report the coefficients as given. - Nomogram calibration / DCA / C-index figures: cosmetic, depend on
rms+ the private helper; low information vs effort. - TIDE / aneuploidy immunotherapy prediction: TIDE requires uploading the
expression to the external tide.dfci.harvard.edu server (interactive); the
shipped
tcga.tide.res.csvis the server output, not pipeline-derivable here. - CIBERSORT / MCP-counter TME: needs the external
infiltration_estimation_for_tcga.csv(TIMER2.0), not shipped. - GSEA (hallmark), mutation/maf comparison: need external gmt + a
*_SNP.Rdatanot in the repo; low priority. - Wet-lab (si-CDKN2A wound/transwell/qPCR, FCS apoptosis): manual, out of scope.
Pointers
- repo: github.com/yangzhandr/Raw-data @ default branch
main(last push
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.