Integrative analysis of transcriptomic data reveals a predictive gene signature for chemoradiotherapy response in rectal cancer.
The main results reproduced, with only marginal, non-material deviations.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- Nothing in this column.
- 🟡Could not use the authors’ exact input data
- 🟡Reported values were only indirectly comparable
- 🟡A deviation arose in the data or preprocessing
- 🟡A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The deviation was non-trivial in magnitude
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Described well enough to reproduce the STEP, not the headline. The paper's headline (186-gene signature, AUC 0.80 via GLMnet/RF across 6 integrated datasets) is the hard last-20% and was deliberately not attempted. We reproduced the clearly-specified third-party pipeline step the paper names verbatim: ran Sage-Bionetworks/CMSclassifier (method='SSP', SSP.predictedCMS) on the named paper dataset GEO GSE150082 (39-sample Agilent two-color microarray) on «our HPC». The tool installs and runs cleanly and yields valid per-sample CMS calls (SSP: CMS2=9, CMS3=23, 7 unclassified). This is a different vs 1:1 outcome: there is NO single reported number tied to GSE150082 alone, because the paper's only CMS numbers are a POOLED cross-dataset statistic (chi-square N=47) computed on the authors' curated/harmonized cohort. A within-GSE150082 CMS3-vs-response cross-tab is weakly consistent with the paper's direction (CMS3 13 non-responder vs 10 responder) but the CMS4-responder claim is not evaluable here (SSP called no CMS4). KEY AUDIT FLAG: the CMS/iCMS classification described in Methods (Fig S2) is ABSENT from the authors' shipped Zenodo code (paoloAngelino/RCproject_iScience, DOI 10.5281/zenodo.17207884; grep 'CMS' = 0 hits), so the reported CMS associations are not independently reproducible from the shipped artifacts. NOT attempted: C1/C2/C3 (186-gene signature + AUC 0.80/0.62) and the exact pooled CMS/iCMS chi-square statistics (need the authors' unpublished CMS scripts + curated harmonized expression).
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 50assessed: 2026-06-14 ⛓ 4e4b0f00d319
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-14
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusCan an integrative analysis of publicly available transcriptomic datasets, combined with machine learning, identify a robust gene signature in pre-treatment biopsies that predicts response to neoadjuvant chemoradiotherapy (nCRT) in locally advanced rectal cancer (LARC)?
- ★ A 186-gene signature derived from six GEO transcriptomic datasets predicts nCRT response in LARC with AUC 0.80 in cross-validation finding
- ★ The signature is associated with CMS4 and iCMS3 molecular subtypes enriched in responders, and CMS3/iCMS2 in non-responders finding
- ★ Spatial transcriptomic (GeoMx DSP) profiling identifies compartment-specific candidate biomarkers, with tumor-associated genes holding greater predictive value than stromal genes finding
- ★ A machine learning approach integrating differential expression across multiple datasets without batch correction yields a robust, generalizable cross-validated signature method
- GSEA implicates VEGFR/EGFR signaling, cell junctions, Toll-like receptor signaling, tyrosine kinase and E2F6 pathways in treatment response mechanism
- The 186-gene signature has minimal overlap and low concordance with previously published predictive signatures finding
- The signature's predictive value is independent of patient age and sex finding
- Immune-related and epigenetic (DNA methylation, histone modification) mechanisms are central to response, with distinct contributions from tumor and stroma compartments mechanism
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Bulk transcriptomic differential expression analysis + machine learning | Pre-treatment biopsies of LARC patients (six public GEO datasets, plus four microarray subset) | none (responder vs non-responder to nCRT) | Gene expression associated with nCRT response; 186-gene predictive signature | GEO datasets; GENCODE28 annotation; GLMnet and Random Forest classifiers |
| External validation of predictive model | Domingo et al. microarray dataset (LARC) | none (responder vs non-responder) | AUC for response prediction | microarray |
| Spatial transcriptomic profiling (Digital Spatial Profiling) | 53 pre-treatment LARC biopsies (in-house HUG cohort), tumor (PanCK+) and stroma (PanCK-) compartments | none (responder vs non-responder) | Compartment-specific differentially expressed genes and pathways; pseudo-bulk signature performance | NanoString GeoMx Digital Spatial Profiling (DSP) |
| Molecular subtype classification | GEO dataset samples (LARC) | none | CMS, iCMS, and RSS subtype assignment vs response | — |
| Gene set enrichment analysis (GSEA) | Bulk signature genes; tumor and stroma compartments (GeoMx) | none | Enriched biological pathways predictive of response | — |
- – 186-gene signature achieved AUC 0.80 with both GLMnet and Random Forest in cross-validation AUC 0.80
- – CMS4 more frequent among responders, CMS3 more prevalent among non-responders X2(1,N=47)=8.1384, p=0.0043
- – iCMS3 predominant among responders, iCMS2 more prevalent among non-responders X2(1,N=226)=11.647, p=0.0006
- – Model performance on external Domingo et al. dataset was modest AUC 0.62
- – Pseudo-bulk GeoMx signature (Random Forest) showed limited discrimination of responders vs non-responders AUC 0.6071
- – GeoMx QC identified DEGs between tumor and stroma compartments reproduced across five bootstrap runs 5,056 DEGs, 406 pathways
- – Compartment-specific analysis identified genes significantly associated with response; tumor compartment held greater predictive value 13 tumor, 7 stroma, 8 opposing (of 28 interaction markers); 8 genes significant in both
- – Signature negatively correlated with Gim 2016 and Domingo 2024 signatures, weakly positive with others rho -0.27 (Gim), -0.37 (Domingo), 0.01-0.21 (others)
- correlation AUC 0.80 (Cross-validation performance of 186-gene signature (GLMnet and Random Forest))
- pvalue X2(1,N=47)=8.1384, p=0.0043 (Association between CMS subtype and treatment response)
- pvalue X2(1,N=226)=11.647, p=0.0006 (Association between iCMS subtype and treatment response)
- correlation AUC 0.62 (Model performance on external Domingo et al. microarray dataset)
- correlation AUC 0.6071 (Random Forest on pseudo-bulk GeoMx spatial data (53 biopsies))
- count 5,056 DEGs and 406 pathways (Tumor vs stroma compartment DEGs across five bootstrap runs)
- correlation rho -0.27 and -0.37 (Negative correlation of signature with Gim 2016 and Domingo 2024 signatures)
- pvalue p-values 0.21, 0.66, 0.70 (KRT17, CD55, PPBP not significantly different between responders/non-responders in our cohort)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This exploratory multi-cohort study performed differential gene expression analysis (DEA) across six publicly available GEO transcriptomic datasets (bulk RNA-seq and microarray) to identify genes associated with neoadjuvant chemoradiotherapy (nCRT) response in locally advanced rectal cancer, then applied GLMnet and Random Forest machine learning classifiers with cross-validation to derive a 186-gene predictive signature (AUC 0.80). Subtype associations were evaluated with chi-square tests, and spatial transcriptomics (GeoMx DSP, n=53 in-house biopsies) was used to identify compartment-specific candidate genes. Results were reported primarily as AUC, exact p-values for categorical associations, and mean ± SEM for spatial data, without formal multiple-testing correction for the genome-wide DEA.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Differential expression analysis (DEA; specific algorithm not named — e.g., limma, DESeq2, or edgeR) | Responders vs. non-responders across six GEO datasets (bulk RNA-seq and microarray); genes with p<0.05 carried forward | — | not stated |
| Chi-square test | Association between consensus molecular subtype (CMS) and treatment response | N=47 | not stated |
| Chi-square test | Association between immune molecular subtype (iCMS) and treatment response | N=226 | not stated |
| ANOVA | Association between 186-gene signature scores and clinical variables (age, sex) in three GEO datasets with available annotations (GSE150082, GSE133057, GSE145037) | — | not stated |
| GLMnet classifier evaluated by AUC (cross-validated) | Predictive performance of the 186-gene signature; cross-validation with random splits repeated three times across the six datasets | — | na |
| Random Forest classifier evaluated by AUC (cross-validated) | Predictive performance of the 186-gene signature; same cross-validation scheme as GLMnet | — | na |
| Spearman rank correlation (rho) | Pairwise correlation of signature scores between the 186-gene signature and ten previously published nCRT-response signatures | — | not stated |
| Differential expression analysis with compartment-by-response interaction term (specific method not named) | GeoMx spatial data: 28 genes identified with significant interaction between tissue compartment (tumor vs. stroma) and treatment response; five bootstrap runs used for QC reproducibility | n=53 | not stated |
-
DEA across thousands of genes used a nominal p<0.05 threshold to select the gene set without a described multiple-testing correction↳ Could also: Applying a false discovery rate (FDR) adjustment — such as Benjamini-Hochberg with a q-value threshold (e.g., q<0.05 or q<0.10) — is the standard approach for genome-wide DEA — FDR control quantifies the expected proportion of false positives among the declared significant genes in high-dimensional settings, making the resulting gene list more interpretable and comparable across studies
-
Cross-validation used random splits of samples pooled across all six datasets, repeated three times↳ Could also: Leave-one-dataset-out (LODO) cross-validation — training on five datasets and testing on the held-out sixth, cycling through all six — would also estimate generalization performance — LODO directly tests whether a signature trained on some cohorts generalizes to an entirely independent cohort; it respects the natural batch structure of multi-dataset studies and mirrors the prospective validation scenario the authors target
-
A chi-square test was used for the CMS subtype-by-response association with N=47↳ Could also: Fisher's exact test would also be applicable to this 2×2 contingency table, especially with the smaller N — Fisher's exact test does not rely on the large-sample chi-square approximation and provides an exact p-value when any expected cell count is small, which is a common concern in contingency tables with total N below 100
-
Spatial transcriptomic data (Figure 3D) are summarized as mean ± SEM↳ Could also: Mean ± SD, median with IQR, or a 95% confidence interval around the mean would also describe spread — SEM reflects precision of the mean estimate and decreases with larger n, which can make the spread of the data appear narrower than it is; SD or 95% CI more directly communicates biological variability and is often preferred for small group sizes
-
GLMnet and Random Forest classifiers were compared informally by inspecting that both achieved AUC=0.80↳ Could also: A paired statistical comparison of ROC curves — such as the DeLong test or bootstrap-based confidence interval for the difference in AUC — would also quantify whether the two models differ in discriminative performance — When two models yield identical point-estimate AUCs, a formal test or CI for the AUC difference indicates whether this reflects statistical equivalence or coincidental agreement, which informs which model to carry forward
-
ANOVA was performed separately in each of three datasets to test independence of the signature score from age and sex, with results described only as 'all p>0.05'↳ Could also: A linear mixed-effects model with dataset as a random effect would also test these associations while pooling evidence across all three datasets in a single analysis — Running three separate ANOVAs and combining results informally does not account for between-dataset variability; a mixed-effects model provides a single pooled estimate and a more powerful test while explicitly modeling the multi-cohort structure
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-41550766
Paper: Corrò C, et al. Integrative analysis of transcriptomic data reveals a predictive gene signature for chemoradiotherapy response in rectal cancer. iScience 2025. PMID 41550766 · PMCID PMC12803930 · DOI 10.1016/j.isci.2025.114455.
What the paper does (pipeline-derived results)
The authors integrate six public transcriptomic datasets (GSE94104, GSE93375, GSE150082, GSE133057, GSE209746, GSE145037; n=266 after filtering) plus their own spatial transcriptomics (GSE279942) to build a 186-gene signature predictive of neoadjuvant chemoradiotherapy (nCRT) response in rectal cancer.
Reported pipeline results (candidate claims)
| # | Result | Reported value | Location | Pipeline | In scope? |
|---|---|---|---|---|---|
| C1 | Signature predictive performance | AUC = 0.80 (GLMnet and Random Forest), 5-fold CV | Results; Fig 1C/1D | Custom ML across 6 integrated datasets (authors' Zenodo code) | Out — hard last-20%; needs full multi-dataset harmonization + ML training |
| C2 | External validation | AUC = 0.62 (Domingo dataset) | Fig S3 | same custom ML | Out — depends on C1 |
| C3 | Signature size | 186 genes | Highlights / Results | feature selection step of C1 | Out — derived inside C1 |
| C4 | CMS subtype ↔ response | CMS3 enriched in non-responders, CMS4 in responders; χ²(1,N=47)=8.1384, p=0.0043 | Fig S2A | CMSclassifier (SSP) on the integrated cohort | Partial — see below |
| C5 | iCMS subtype ↔ response | iCMS2 in non-responders, iCMS3 in responders; χ²(1,N=226)=11.647, p=0.0006 | Fig S2B | iCMS classifier (SSC.DQ) on integrated cohort | Out — needs the same pooled harmonization as C4/C5 |
What we reproduce (the clearly-specified, low-hanging pipeline step)
The paper states verbatim: "the reference CMS classifier … was used to predict the CMS subtypes in the samples (method='SSP'; function SSP.predictedCMS)." That tool is the third-party R package Sage-Bionetworks/CMSclassifier (the repo listed for this RU). Per study rule P16, running an existing third-party tool on the paper's own data is an equally valid reproduction.
Reproduction target (R1): Apply CMSclassifier (classifyCMS, method SSP — the
exact function named — and RF for cross-check) to one of the paper's named datasets,
GSE150082 (39-sample Agilent two-color microarray, LARC pre-treatment), and report
the produced CMS subtype distribution as a concrete, auditable data point that the
named tool runs on the named data and yields valid CMS calls.
What we deliberately do NOT attempt (and why)
- C1/C2/C3 (the 186-gene signature + AUC 0.80/0.62): requires re-harmonizing all six datasets and re-training the authors' GLMnet/RF pipeline from the Zenodo code (DOI 10.5281/zenodo.17207884). This is the hard last-20%; out of the 80/20 budget.
- Exact χ²/p-values in C4/C5: these are computed on a pooled cross-dataset subset (N=47 for CMS, N=226 for iCMS) whose sample selection + harmonization come from the authors' Zenodo code, not reconstructable 1:1 without re-running their full pipeline. We therefore reproduce the CMS-classification step itself (R1), not the pooled statistic, and (if the GSE150082 series matrix carries a response/TRG field) report the directional CMS-vs-response cross-tab within GSE150082 as a qualitative check.
Artifacts
- Code (third-party tool): https://github.com/Sage-Bionetworks/CMSclassifier (master)
- Authors' own code (not the listed repo): Zenodo 10.5281/zenodo.17207884
(
paoloAngelino/RCproject_iScience, 130 kB) - Data: GEO GSE150082 (public, downloadable series matrix + GPL annotation)
- Compute: «our HPC» SLURM (partition std), conda env on «infra»
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
We reproduced the STEP, not the headline: Sage-Bionetworks/CMSclassifier (SSP) installs and runs cleanly on the named dataset GSE150082 (CMS2=9, CMS3=23, 7 NA over 39 samples), and the CMS3-in-non-responders direction is weakly consistent (13 non-responder vs 10 responder). The paper's only reported CMS numbers are a pooled χ²(1,N=47)=8.1384, p=0.0043 (Fig S2A) over a curated, harmonized six-dataset cohort, which is not 1:1 reconstructable from public GEO data — so the deviation sits on the input/cohort-definition side, partly our self-chosen scope. The genuine authors-side flag is that the CMS/iCMS classification code is absent from the shipped Zenodo repo (grep 'CMS'=0), so Fig S2 is not independently reproducible — a missing-code gap, not proof of fabrication. Nothing in the reproduction contradicts the paper; the headline AUC=0.80 signature was deliberately not attempted, leaving the core claim untested.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.