Unveiling the immunometabolic landscape of colorectal cancer through PANoptosis-related gene expression.
Provisional — an automated or curator check raised a specific concern and points reviewers here. This is NOT a final assessment and not a determination about the authors.
The main result did not reproduce in this reproduction attempt. Where our recomputation produced values that differ from the published ones, those discrepancies are listed below. This is a single automated attempt — not peer review and not a finding of error or misconduct — and differences can also arise from data access, undocumented parameters or the computing environment. The verdict can be contested via “report an error”.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- Nothing in this column.
- 🟡Could not use the authors’ exact input data
- 🟡Reported values were only indirectly comparable
- 🟡A deviation arose in the data or preprocessing
- 🔴A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The deviation was non-trivial in magnitude
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Described well enough to reproduce the HEADLINE DEG result 1:1-class. The paper's 'code' (IOBR) is a generic third-party toolkit, not authors' scripts, so per P16 we ran the named methods (limma + sva::ComBat on the 3 GPL570 GEO discovery sets GSE41328/GSE22598/GSE23878; limma-voom on TCGA COAD+READ via recount3) on the paper's named data with the stated thresholds (|logFC|>1 & p<0.05). C1: reported 908 common DEGs vs reproduced 926 (+2.0%) despite fully independent tooling (recount3 read counts vs authors' GDC; symbol-alias matching) -> within-tol; 13/15 canonical CRC markers tested are in the intersection, confirming biological correctness. NOT attempted (the hard ~20%): consensus clustering C1=418/C2=128, LASSO 11-gene CPAN-index + survival AUC 0.62/0.60/0.62, CIBERSORT/ESTIMATE/TIDE immune scores, GSEA, and all single-cell analyses (multi-dataset, many free parameters). FLAG for human review: the stated PANoptosis-list composition 87.4/9.7/2.9% equals exactly 90/10/3 of a 103-gene list and is NOT derivable from the cited ref-23 set of 277 non-redundant genes (259/27/15) -> possible mis-citation; C2's '15-gene overlap' rests on this ambiguous list and was not reproduced rather than fabricated.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 41assessed: 2026-06-14 ⛓ 20c784954228
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-14
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusDo PANoptosis-related genes drive colorectal cancer (CRC) progression by reshaping the immune microenvironment and metabolic reprogramming, and can their differential expression be used to build a prognostic risk-stratification tool (CPAN-index) for CRC patients?
- ★ A CPAN-index prognostic model built from 11 PANoptosis-related differentially expressed genes (CPAN_DEGs) stratifies CRC patients into high-risk and low-risk groups with distinct survival and immunophenotypes. resource
- ★ The high-risk CRC group exhibits an 'invasion-metabolism-immunosuppressive' phenotype with immune tolerance and non-classical immune escape. finding
- ★ CDKN2A is upregulated in CRC cell lines and its knockdown inhibits proliferation and promotes apoptosis in vitro. finding
- ★ CDKN2A-mediated PANoptosis signaling drives CRC progression by reshaping the immune microenvironment and metabolic reprogramming. mechanism
- ★ PANoptosis-related DEGs between CRC and normal tissue are significantly enriched in metabolism, apoptosis, and immune regulation pathways. finding
- ★ scRNA-Seq shows a higher proportion of CPAN-index-positive immune cells and a lower proportion of CPAN-index-positive tumor cells, indicating a role in the tumor immune microenvironment. finding
- Apoptosis-related genes constitute the majority (87.4%) of the PANoptosis gene list. finding
- Unsupervised consensus clustering of CPAN_DEGs divides CRC into two clusters (C1, C2) with distinct immune scores and immune checkpoint expression. method
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Bulk RNA-seq differential expression and prognostic modeling (limma, LASSO, Cox) | CRC tissue (TCGA COAD/READ + GEO cohorts: GSE41328, GSE22598, GSE23878, GSE39582, GSE72970, GSE17536) | none | differentially expressed PANoptosis genes, CPAN-index risk score, survival/ROC | — |
| Single-cell RNA-seq (scRNA-Seq) cell-type analysis and module scoring | CRC samples (TISCH datasets GSE146771, EMTAB8107, GSE166555) | none | PANoptosis gene expression per cell type, proportion of CPAN-index-positive cells | TISCH database; Seurat/Harmony/SingleR/UMAP |
| Immune microenvironment deconvolution | merged CRC bulk dataset | none | immune/stromal/ESTIMATE scores, immune cell fractions, TIDE/dysfunction/CD274/IFNG/Merck18/CD8 scores | CIBERSORT, ESTIMATE, TIDE |
| RT-qPCR | CRC cell lines (LoVo, HT-29, HCT116, SW480) and normal colon epithelium HIEC-6 | CDKN2A siRNA knockdown vs negative control | CDKN2A mRNA expression | PrimeScript RT kit, SYBR GreenER Supermix, Roche480 PCR system |
| CCK-8 proliferation assay | HCT116 and SW480 CRC cells | CDKN2A siRNA knockdown | cell proliferation (absorbance at 450nm) | CCK8 (KeyGEN), microplate reader |
| Flow cytometry apoptosis assay | HCT116 and SW480 CRC cells | CDKN2A siRNA knockdown | apoptosis (Annexin/FITC and PI staining) | FITC/PI (Biosharp) |
| Transwell migration assay | HCT116 and SW480 CRC cells | CDKN2A siRNA knockdown | migrated/invaded cells (crystal violet) | — |
| Wound healing (scratch) assay | HCT116 and SW480 CRC cells (5×10^5 cells/well) | CDKN2A siRNA knockdown | wound closure at 0h and 48h | — |
- – Apoptosis-related genes comprised 87.4% of the PANoptosis gene list (pyroptosis 9.7%, necroptosis 2.9%) 87.4%
- – Of 908 DEGs common to CRC and GSE datasets, 15 genes overlapped with the PANoptosis gene list (CPAN_DEGs) 15 of 908
- ▲ CDKN2A was upregulated in CRC cell lines; knockdown inhibited proliferation and promoted apoptosis in vitro
- – CPAN-index built from 11 CPAN_DEGs distinguished high-risk vs low-risk CRC patients with the high-risk group showing immunosuppressive phenotype 11 genes
- ▲ Immune score and microenvironment (ESTIMATE) score were significantly higher in C2 cluster than C1
- ▲ Immune checkpoints (TIGIT, PDCD1LG2, PDCD1, LAG3, HAVCR2, CTLA4, CD274) significantly higher in C2 than C1, indicating immunosuppressive state
- – scRNA-Seq showed higher proportion of CPAN-index-positive immune cells and lower proportion of CPAN-index-positive tumor cells
- – Consensus clustering identified k=2 as optimal, dividing the CRC cohort into C1 (n=418) and C2 (n=128) with no significant survival difference p=0.28
- count 87.4% apoptosis, 9.7% pyroptosis, 2.9% necroptosis (composition of PANoptosis gene list)
- count 908 common DEGs; 15 in PANoptosis list (Venn overlap of CRC and GSE DEGs)
- count 11 CPAN_DEGs (genes in the CPAN-index prognostic model)
- count C1 n=418, C2 n=128 (consensus cluster sizes of CRC cohort)
- pvalue p=0.28 (p>0.05) (Kaplan-Meier survival difference between C1 and C2 clusters)
- pvalue p<0.05 (immune score and microenvironment score higher in C2 vs C1)
- other |logFC|>1 & p.Val<0.05 (DEG selection threshold (limma))
- count 4000 cells per well; 40,000 per well (transwell); 5×10^5 per well (wound healing) (cell seeding densities for functional assays)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
The study integrates public bulk RNA-seq data (TCGA and GEO cohorts) with in vitro cell-line experiments and public scRNA-seq data to characterize PANoptosis-related gene expression in colorectal cancer. Differentially expressed genes were identified with the limma R package, consensus clustering was used to stratify patients, and a LASSO-penalized Cox proportional hazards model generated a prognostic index (CPAN-index) that was externally validated in three independent GEO cohorts via ROC/AUC analysis. In vitro functional experiments (CCK-8, flow cytometry, wound healing, transwell) used at least three biological replicates per group, with comparisons by two-sided t-test and results expressed as mean ± SD.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| limma moderated t-test (via 'limma' R package) | Differential expression between CRC and normal tissues in TCGA and combined GEO datasets (|logFC|>1, p.Val<0.05); also between C1 and C2 consensus clusters | C1 n=418, C2 n=128 for cluster DEGs; TCGA+GEO merged dataset size not stated | not stated |
| ConsensusClusterPlus unsupervised clustering evaluated by PAC and CDF curves | Determination of optimal cluster number (k=2 selected) and patient stratification into C1 and C2 subtypes | C1 n=418, C2 n=128 | not stated |
| Kaplan-Meier survival analysis (log-rank test implied but not named) | Overall survival comparison of C1 vs C2 clusters (p>0.05); high-risk vs low-risk CPAN-index groups in training and validation datasets | C1 n=418, C2 n=128; validation cohort sizes not stated | not stated |
| Univariate Cox proportional hazards regression | Screening CPAN_DEGs for prognostic value prior to LASSO selection | null | not stated |
| LASSO (L1-penalized) Cox regression | Dimensionality reduction to select candidate genes for the CPAN-index; optimal lambda chosen by minimizing binomial deviance | null | not stated |
| Multivariate Cox proportional hazards regression (maximum partial likelihood estimation) | Construction of CPAN-index from 15 LASSO-selected candidate genes | null | not stated |
| ROC curve analysis (time-dependent AUC) | Validation of CPAN-index predictive performance in GSE39582, GSE17536, and GSE72970 | null | na |
| GO/KEGG overrepresentation enrichment analysis (Fisher's exact or hypergeometric, adjusted p-values reported as p.adj) | Functional annotation of 15 CPAN_DEGs: GO-BP, GO-MF, KEGG pathways (Figure 3) | 15 CPAN_DEGs | not stated |
| GSEA (Gene Set Enrichment Analysis, via 'GseaVis' R package) | Pathway enrichment comparison between high-risk and low-risk CPAN-index groups | null | not stated |
| CIBERSORT algorithm (immune deconvolution) | Immune cell proportion estimation in C1 vs C2 clusters and high-risk vs low-risk groups | null | na |
| ESTIMATE algorithm (immune/stromal scoring) | Immune score, matrix score, and microenvironment score in C1 vs C2 clusters (Figure 5) | null | na |
| Two-sided t-test | All between-group comparisons stated in Section 2.13; covers in vitro assays (CCK-8, flow cytometry, wound healing, transwell) and bulk-data group comparisons (immune/checkpoint scores between risk groups) | ≥3 biological replicates per group (stated) | not stated |
| Seurat FindMarkers (underlying statistical test not specified; Wilcoxon rank-sum by default) | Identification of genes with significant expression variation across cell types in integrated scRNA-seq data | null | not stated |
| Seurat AddModuleScore | Quantification of CPAN-index gene module activity per cell in scRNA-seq data; proportion of CPAN-index-positive cells per subpopulation calculated | null | na |
-
Group comparisons in cell-line experiments used two-sided t-tests with n≥3 biological replicates per group↳ Could also: A non-parametric Wilcoxon rank-sum (Mann-Whitney U) test could also be used — With only three replicates per group, normality cannot be formally assessed; non-parametric alternatives make fewer distributional assumptions and are commonly applied in small-sample experimental biology
-
Feature selection for the prognostic model used LASSO-penalized Cox regression (L1 penalty only)↳ Could also: Elastic net Cox regression (combining L1 and L2 penalties, alpha between 0 and 1) could also be applied — When predictors are correlated — as co-expressed genes frequently are — elastic net can improve coefficient stability by grouping correlated features rather than arbitrarily selecting one; it also tends to produce more reproducible gene sets across validation cohorts
-
Dispersion for in vitro assay data is reported as mean ± SD↳ Could also: Individual data points overlaid on bar or box plots, or 95% confidence intervals, could also convey spread — With n=3 replicates, showing individual points alongside the mean makes the raw data transparent; 95% CIs communicate both variability and estimation precision, and are increasingly recommended by journals for small-n experimental data
-
Immune cell deconvolution relied on the CIBERSORT algorithm alone↳ Could also: A second deconvolution algorithm such as TIMER2, xCell, or EPIC could also be applied in parallel — Different algorithms use different reference gene signatures and statistical assumptions; cross-validating immune composition estimates across two algorithms is a common approach for assessing robustness of findings
-
Many pairwise t-tests were conducted across immune scores, checkpoint molecules, and functional assay endpoints without a stated multiple-comparison correction↳ Could also: A Benjamini-Hochberg FDR or Holm correction could also be applied across the family of simultaneous comparisons — When many endpoints are evaluated concurrently — as in the immune checkpoint panel — a family-wise correction controls the expected proportion of false discoveries across all tests in that family
-
The log-rank test underlying the Kaplan-Meier curve comparisons is implied by the reported p-values but is not named in the text↳ Could also: Explicitly naming the log-rank test, or supplementing it with a Cox log-rank score test or restricted mean survival time (RMST) analysis, could also be reported — Naming the test aids reproducibility; RMST provides an effect-size estimate (difference in mean survival time up to a horizon) that complements the p-value and does not assume proportional hazards
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-41601652
Paper: He X, Wang W, Li L, Yin Y, Ding S. Unveiling the immunometabolic landscape of colorectal cancer through PANoptosis-related gene expression. Front Immunol 2025/2026. DOI 10.3389/fimmu.2025.1615022. PMCID PMC12832467.
Nature of the "code" link. The brief lists the code as https://github.com/IOBR/IOBR — IOBR is a generic third-party immuno-oncology R toolkit (TME deconvolution, signature scoring, batch removal), not the authors' own analysis pipeline. Per brief rule P16, applying the described third-party tools to the paper's own data is an equally valid reproduction. There is no authors' repo with parameter-pinned scripts; we reproduce by running the named tools (limma, sva/ComBat, recount3-sourced TCGA) on the paper's named data with the paper's stated thresholds.
Datasets named in the paper
- Bulk discovery (GEO, all GPL570 / Affy HG-U133 Plus 2.0): GSE41328 (n=20), GSE22598 (n=38), GSE23878 (n=59) — CRC tumour vs normal.
- TCGA COAD + READ bulk RNA-seq (tumour vs normal).
- Validation cohorts: GSE39582, GSE72970, GSE17536 (survival).
- scRNA-seq: GSE146771, GSE166555, E-MTAB-8107.
- PANoptosis gene list: from ref 23 = Song et al. 2023, Front Immunol, PMC10311484 (HCC HPAN-index), reported there as 277 non-redundant genes (apoptosis 259 / pyroptosis 27 / necroptosis 15 before de-duplication).
In scope (pipeline-derived, attempted)
| # | Reported result | Location | Pipeline | Status |
|---|---|---|---|---|
| C1 | 908 DEGs common to TCGA and GEO (|logFC|>1, p<0.05, limma) | Results/DEG | limma on GEO 3-set (ComBat-merged) ∩ limma-voom on TCGA COAD+READ | attempted (primary) |
| C2 | 15 of the 908 DEGs are PANoptosis genes | Results/DEG | C1 ∩ Song-2023 277-gene list | secondary (list provenance risk) |
| C3 | PANoptosis list composition 87.4% apop / 9.7% pyro / 2.9% necro | Results/Fig | count categories of the cited list | flagged — does not match the cited 277-gene set; 87.4/9.7/2.9 = 90/10/3 of 103 genes, implying a different list than ref 23 (see AUDIT) |
Out of scope (not attempted — why)
- Consensus clustering C1=418 / C2=128 (k=2): depends on exact TCGA sample filtering + PANoptosis-DEG feature set; high specification ambiguity.
- LASSO 11-gene CPAN-index + survival AUC 0.62/0.60/0.62: needs TCGA survival + exact LASSO seed/lambda; not pinnable (the hard last 20%).
- CIBERSORT / ESTIMATE / TIDE immune scores, GSEA, single-cell (Seurat/Harmony/ SingleR) AddModuleScore proportions (NK 70.3% etc.): multi-tool, many free parameters, several datasets — out of scope for a clear 1:1.
- CDKN2A knockdown (apoptosis/proliferation/migration p<0.0001): wet-lab, not computational — out of scope.
Strategy (80/20)
Reproduce the headline DEG intersection (C1) faithfully with the named tools and thresholds; report the GEO-side and TCGA-side DEG counts separately too (the paper prints only the intersection). Treat C2 as best-effort and C3 as an auditable inconsistency rather than a clean target. Do not chase the clustering/LASSO/immune/scRNA tail.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
The headline DEG result reproduced well — 926 common DEGs vs the reported 908 (+2.0%) despite fully independent tooling (recount3/limma-voom vs the authors' GDC/IOBR pipeline), with 13/15 canonical CRC markers present, so the DEG backbone is credible. The genuine concern is on the authors' side: the reported PANoptosis composition 87.4/9.7/2.9% equals exactly 90/10/3 of a 103-gene list and is not derivable from the cited 277-gene ref-23 set (86.0/9.0/5.0%), making the 15-gene overlap (C2) unverifiable — a possible mis-citation/too-perfect flag rather than confirmed fabrication. Severity is moderate: the central DEG finding holds while the PANoptosis-specific framing and all downstream analyses (clustering, LASSO, immune, single-cell) remain unverified, so this is a yellow case flagged for human review.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.