Multimodal data integration for biologically-relevant artificial intelligence to guide adjuvant chemotherapy in stage II colorectal cancer.
The main results reproduced: recomputed values matched the published ones within tolerance.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Same input data as the authors
- 🟡Reported values were only indirectly comparable
- 🟡A deviation arose in the data or preprocessing
- 🟡A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The deviation was non-trivial in magnitude
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Described well enough to reproduce ONE clean public sub-result, NOT the headline imaging/AI pipeline. This is a large multimodal stage-II CRC paper; its core results (DeepCRC/nnUNet CT segmentation, PyRadiomics 851-feature radiomics, iHR survival-by-treatment models iHR 5.35/2.88, multiplex IF, mouse model) all depend on a private institutional CT cohort + wet-lab data (data_restricted) and were not attempted. The one fully-public, low-compute, pinnable pipeline output is the xCell cellular deconvolution on GEO GSE28702 (FOLFOX responders vs non-responders). Using the SAME tool the authors' AIRCSA repo uses (xCell 1.1.0) on the SAME public data, run on «our HPC» («job»): microvascular-EC infiltration is significantly higher in responders (p=0.0012 / 0.010 under two calibrations) = REPRODUCED 1:1; overall endothelial-cell infiltration is directionally higher in responders but does not reach significance (p=0.11-0.19) = PARTIAL. No exact p-value is printed in the paper for GSE28702 alone, so only the direction+significance qualifier is checkable. No fabrication concern: the reported direction is correct and derivable from the shipped public data; the EC significance gap is plausibly explained by CEL-level RMA vs series-matrix, a different probe-collapse, one-sided testing, or pooling the three GEO sets. Did NOT attempt: anything requiring the private imaging cohort, OMIX002301, wet-lab, or GPU segmentation.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 83assessed: 2026-06-14 ⛓ 44dee5853d81
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-14
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusCan an explainable, multimodal AI-powered radiological analyser identify stage II colorectal cancer (CRC) phenotypes that derive an overall survival benefit from adjuvant chemotherapy, and what biological (genetic/histopathological/vascular) basis underlies these imaging subtypes?
- ★ An AI-powered radiological clustering of CT images stratifies stage II CRC patients into adjuvant chemotherapy (AC)-preferable and observation-only (OO)-preferable clusters whose survival benefit from chemotherapy differs significantly (iHR=5.35). finding
- ★ The two radiological clusters differ in biological pathways related to immune and stromal cell abundance in the tumour microenvironment. finding
- ★ OO-preferable tumours show higher necrosis, haemorrhage, and tortuous vessels, whereas AC-preferable tumours show vessels with greater pericyte coverage and richer infiltration of B, CD4+-T, and CD8+-T cells into the tumour core. mechanism
- ★ Preclinical intervention on vessel morphology (anlotinib) alters predictive CT imaging/textural features, demonstrating a causal link between vasculature and imaging biomarkers. finding
- ★ An interaction-hazard-ratio (iHR) based feature-selection method using a Cox interaction term selects treatment-benefit-predictive radiomic features for unsupervised hierarchical clustering. method
- DeepCRC topology-aware deep-learning segmentation plus PyRadiomics feature extraction provides an automated, explainable imaging pipeline for stage II CRC risk stratification. resource
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Contrast-enhanced CT radiomics (DeepCRC segmentation + PyRadiomics) | Human stage II CRC patients (6 cohorts; GDPH n=405 development, YNCC n=153 validation, TCIA/COAD, GEO) | adjuvant chemotherapy vs observation (clinical, none) | radiomic features for treatment-benefit clustering / survival | DeepCRC, PyRadiomics, ITK-SNAP |
| Bulk RNA sequencing with differential expression, GSVA, and cellular deconvolution | Human CRC tumour tissue (60 patients from training set) | none (radiological cluster comparison) | differentially expressed genes, hallmark/KEGG pathway enrichment, immune/stromal TME composition | MSigDB hallmark v7.5.1, KEGG |
| Transcriptomic drug-response analysis | Public GEO CRC datasets | fluorouracil-based chemotherapy | chemotherapy benefit and underlying TME | — |
| H&E histopathology pattern evaluation | Human CRC whole-tumour sections | none | haemorrhage, necrosis, TLS, GC+TLS, tumour budding, desmoplastic reaction | — |
| Double immunohistochemistry / vessel quantification | Human CRC tumour sections | none | CD31/PanCK staining, tumour-stroma ratio, vessel junctions, mesh size, segment length | QuPath, ImageJ (Color Deconvolution, Angiogenesis Analyser) |
| Multiplex immunohistochemistry (mIHC) with spatial analysis | Human stage II CRC FFPE specimens (10 patients, GDPH) | none | vessel phenotypes and immune cell infiltration across 100 μm interface tiles | HALO image analysis software v3.2 (Indica Labs) |
| Micro-CT imaging radiomics + IHC vascular analysis | CT26 cell-line-derived xenograft mouse model (12 female BALB/c mice) | anlotinib (antiangiogenic TKI) vs saline | radiomic/textural CT features and vascular patterns (microvessel pericyte coverage index, CD31/α-SMA) | micro-CT; PASS v15.0.5 for sample size |
- – Survival benefit of chemotherapy varied significantly between AI-powered radiological clusters iHR=5.35 (95% CI 1.98–14.41), adjusted P_interaction=0.012
- ▲ AC-preferable cluster showed vessels with greater pericyte coverage and enriched B, CD4+-T, CD8+-T cell infiltration into tumour core
- ▲ OO-preferable cluster exhibited higher necrosis, haemorrhage, and tortuous vessels
- – Anlotinib-induced changes in vessel morphology produced alterations in predictive imaging features
- – Microvessel pericyte coverage index (MPI) differed between experimental and control groups in preliminary mouse data experimental 1.2±0.92 vs control 3.9±1.2
- other iHR=5.35 (95% CI 1.98, 14.41), adjusted P_interaction=0.012 (interaction hazard ratio for chemotherapy benefit between radiological clusters)
- other <5% (5-year survival benefit of adjuvant chemotherapy in stage II CRC)
- mean MPI 1.2±0.92 (experimental) vs 3.9±1.2 (control) (microvessel pericyte coverage index in mouse anlotinib vs saline groups)
- count 405 development, 153 validation (GDPH development and YNCC validation cohort sizes)
- count 60 (patients with RNA sequencing data from training set)
- other ICC threshold 0.80 (radiomic feature robustness test-retest in 30 patients)
- pvalue z-statistic <0.05 (interaction feature selection threshold for radiomic features)
- count 10 FFPE specimens; 12 BALB/c mice (6 anlotinib, 6 saline) (mIHC specimen count and mouse model group allocation)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
The study used a retrospective multicentre design to develop and validate an AI-powered radiomic risk stratification model for stage II CRC across six cohorts. Radiomic features predictive of treatment-specific survival benefit were selected via Cox proportional hazards models incorporating a treatment-by-feature interaction term (interaction hazard ratio, iHR), and unsupervised hierarchical clustering then grouped patients into two treatment-predictive subtypes. Primary efficacy of stratification was reported as an iHR with a 95% CI and adjusted P value; Benjamini-Hochberg FDR correction was applied for post-hoc multiple-comparison analyses. Biological explainability was pursued through GSVA pathway enrichment, cellular deconvolution of RNA-seq data, quantitative histopathology, and a randomised preclinical mouse experiment.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Cox proportional hazards model with treatment-by-feature interaction term (iHR, Wald test) | Per-feature selection of treatment-predictive radiomic biomarkers; reported overall interaction HR = 5.35 (95% CI 1.98–14.41) | Development cohort n = 405 (GDPH); RNA-seq subset n = 60 | not stated |
| Hierarchical clustering (unsupervised) with dendrogram and silhouette methods for optimal cluster number | Discovery of AC-preferable vs OO-preferable radiological subtypes | Development cohort n = 405 | not stated |
| Intraclass correlation coefficient (ICC, threshold 0.80) | Feature robustness test-retest study comparing original and automated segmentations | n = 30 patients | not stated |
| Gene set variation analysis (GSVA) | Hallmark (MSigDB v7.5.1) and KEGG pathway enrichment comparison between radiological clusters | n = 60 (RNA-seq patients in training set) | not stated |
| Cellular deconvolution algorithm (transcriptome-based) | Tumour microenvironment immune and stromal component estimation across radiological clusters and GEO cohorts | Not explicitly stated for each GEO cohort | not stated |
| Two-sample t-test (power/sample-size calculation formula via PASS v15.0.5) | Sample size determination for preclinical mouse model (microvessel pericyte coverage index, MPI) | n = 6 per group (total n = 12 mice) | stated |
-
Unsupervised hierarchical clustering was used to discover the two radiological subtypes, with the number of clusters chosen by dendrogram inspection and silhouette score.↳ Could also: Consensus clustering (e.g., via the R package ConsensusClusterPlus) or non-negative matrix factorisation (NMF) could also define robust subtypes. — Consensus clustering quantifies subtype stability across bootstrap resamples and provides a formal instability score, which can make the choice of cluster number more reproducible and transparent across datasets of different sizes.
-
Radiomic features were selected by running a separate Cox interaction model for each feature individually, then applying a z-statistic threshold.↳ Could also: A penalised Cox model with an interaction term (e.g., lasso or elastic-net penalisation via glmnet) could also perform simultaneous feature selection across all features. — Penalised regression handles correlated radiomic features jointly and implicitly regularises against overfitting, whereas sequential univariate screening can miss higher-order covariate structure and may be sensitive to multicollinearity among the large radiomic feature set.
-
Tumour microenvironment cellular composition was estimated from bulk RNA-seq using a deconvolution algorithm.↳ Could also: Single-cell RNA sequencing or spatially resolved transcriptomics could also characterise cell-type abundances and their spatial distributions. — Bulk deconvolution infers cell proportions from aggregate signals and relies on reference signatures; single-cell or spatial approaches resolve individual cell states and neighbourhoods directly, which can complement and validate deconvolution estimates.
-
Cluster-level pathway enrichment was assessed with GSVA using MSigDB hallmark and KEGG gene sets.↳ Could also: Gene set enrichment analysis (GSEA) with permutation-based statistics, or over-representation analysis (ORA) on differentially expressed genes, could also quantify pathway-level differences between clusters. — GSVA produces per-sample enrichment scores suitable for continuous comparisons, while GSEA uses ranked gene lists and is better suited to detecting coordinated directional shifts; ORA provides straightforward interpretability. Reporting results from more than one method can strengthen convergent conclusions.
-
Preliminary mouse model data (MPI) were reported as mean ± SD and used to power a two-sample t-test.↳ Could also: A non-parametric Mann-Whitney U test could also compare MPI between groups, particularly given the small per-group n (n = 6). — With n = 6 per group, normality is difficult to verify, and the assumption underlying the t-test is hard to assess; the Mann-Whitney U test makes no distributional assumption about the outcome and is therefore a common alternative for small-n preclinical experiments.
-
The overall survival benefit of cluster assignment was summarised with a single interaction hazard ratio derived from a Cox model.↳ Could also: Restricted mean survival time (RMST) differences between treatment arms within each cluster could also quantify the survival benefit on an absolute time scale. — The iHR summarises a relative, multiplicative treatment–subgroup interaction under the proportional hazards assumption; RMST is assumption-free in that respect and expresses the benefit as an absolute number of life-years gained over a defined horizon, which can be more directly interpretable for clinical decision-making.
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-40472802
Paper: Xie et al., Multimodal data integration for biologically-relevant AI to guide adjuvant chemotherapy in stage II colorectal cancer. EBioMedicine 2025. PMID 40472802 · PMCID PMC12171563 · DOI 10.1016/j.ebiom.2025.105789. Authors' repo: https://github.com/cx601/AIRCSA · data: GEO GSE28702 (+GSE62080, GSE69657, TCGA, TCIA, OMIX002301).
This is a large multimodal study. Most results are NOT cheaply reproducible. Triage below.
OUT OF SCOPE (not attempted — reason)
- CT segmentation (DeepCRC / nnUNet) + radiomics (PyRadiomics, 851 features) —
requires the institutional CT cohort (405 primary + 153 validation patients);
imaging data is controlled/on-request (TCIA + internal), heavy GPU compute.
data_restricted. - iHR survival-by-treatment interaction (iHR 5.35; pooled 2.88), OS benefit curves (Fig.3) — depend on the radiomic clusters from the private imaging cohort. Not reproducible without that data.
- Multiplex IF / HALO, mouse model, ImageJ angiogenesis (Fig.5–6) — wet-lab
- proprietary software (HALO v3.2), no public inputs.
non_pipeline.
- proprietary software (HALO v3.2), no public inputs.
- OMIX002301 transcriptomics — author-deposited, used for the internal radiogenomic cluster contrast (limma in RadiogenomicAnalysis.R); access TBD, not the cleanest public target.
IN SCOPE (attempted) — one clear, public, low-compute pipeline output
- xCell cellular deconvolution on GSE28702 (public, 83 samples: 42 responders /
41 non-responders to FOLFOX), endothelial-cell infiltration responder vs
non-responder.
- Pipeline: GEOquery fetch -> probe→gene collapse -> xCell (Aran 2017, the exact deconvolution tool used in the authors' AIRCSA repo) -> Wilcoxon test.
- Reported claim (Results, GSE28702 validation): "The infiltration of EC and microvascular EC was significantly higher in the responder group than that in the non-responder group in GSE28702."
- This is a P16 "third-party tool on the paper's own public data" reproduction: equally valid. Light compute, fully public inputs, a checkable directional + significance claim.
Why this target
It is the only result whose inputs are fully public (GEO), whose tool is named and runnable (xCell), and whose claim is pinnable (EC + mv EC higher in responders, significant). Everything else is gated behind private imaging / wet-lab data. Honest 80/20: we reproduce this one cleanly and do not chase the imaging-dependent 80%.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
Within the only reproducible, fully-public slice of this large multimodal CRC paper (xCell on GEO GSE28702, 42 resp/41 non-resp matched exactly), mv-EC infiltration reproduces cleanly and significantly (p=0.0012/0.010) while overall-EC reproduces in direction but loses the 'significant' qualifier (p=0.113/0.188). The deviation sits on the input/method side — series-matrix vs CEL-level RMA, probe-collapse, one- vs two-sided testing, or pooling of the three GEO sets — not in any private-data defect, and there is no fabrication concern since the reported direction is derivable from the shipped public data. The paper prints no exact GSE28702 p-value, so only a directional+significance qualifier is checkable; overall this is a solid partial reproduction of a peripheral validation claim, with the headline AI/imaging pipeline untestable (private cohort).
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.