Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Multimodal data integration for biologically-relevant artificial intelligence to guide adjuvant chemotherapy in stage II colorectal cancer.

EBioMedicine · 2025
L1 83/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +8
✓ What held up
  • Same input data as the authors
What did not (or only partly)
  • 🟡Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
83/100
Reproducibility score
0.5 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 61% of all assessed papers rank 430 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough to reproduce ONE clean public sub-result, NOT the headline imaging/AI pipeline. This is a large multimodal stage-II CRC paper; its core results (DeepCRC/nnUNet CT segmentation, PyRadiomics 851-feature radiomics, iHR survival-by-treatment models iHR 5.35/2.88, multiplex IF, mouse model) all depend on a private institutional CT cohort + wet-lab data (data_restricted) and were not attempted. The one fully-public, low-compute, pinnable pipeline output is the xCell cellular deconvolution on GEO GSE28702 (FOLFOX responders vs non-responders). Using the SAME tool the authors' AIRCSA repo uses (xCell 1.1.0) on the SAME public data, run on «our HPC» («job»): microvascular-EC infiltration is significantly higher in responders (p=0.0012 / 0.010 under two calibrations) = REPRODUCED 1:1; overall endothelial-cell infiltration is directionally higher in responders but does not reach significance (p=0.11-0.19) = PARTIAL. No exact p-value is printed in the paper for GSE28702 alone, so only the direction+significance qualifier is checkable. No fabrication concern: the reported direction is correct and derivable from the shipped public data; the EC significance gap is plausibly explained by CEL-level RMA vs series-matrix, a different probe-collapse, one-sided testing, or pooling the three GEO sets. Did NOT attempt: anything requiring the private imaging cohort, OMIX002301, wet-lab, or GPU segmentation.

💻 Code ↗ 🗄 Data: GSE28702

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 83
    assessed: 2026-06-14 ⛓ 44dee5853d81
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-14
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

Can an explainable, multimodal AI-powered radiological analyser identify stage II colorectal cancer (CRC) phenotypes that derive an overall survival benefit from adjuvant chemotherapy, and what biological (genetic/histopathological/vascular) basis underlies these imaging subtypes?

Core claims
  • An AI-powered radiological clustering of CT images stratifies stage II CRC patients into adjuvant chemotherapy (AC)-preferable and observation-only (OO)-preferable clusters whose survival benefit from chemotherapy differs significantly (iHR=5.35). finding
  • The two radiological clusters differ in biological pathways related to immune and stromal cell abundance in the tumour microenvironment. finding
  • OO-preferable tumours show higher necrosis, haemorrhage, and tortuous vessels, whereas AC-preferable tumours show vessels with greater pericyte coverage and richer infiltration of B, CD4+-T, and CD8+-T cells into the tumour core. mechanism
  • Preclinical intervention on vessel morphology (anlotinib) alters predictive CT imaging/textural features, demonstrating a causal link between vasculature and imaging biomarkers. finding
  • An interaction-hazard-ratio (iHR) based feature-selection method using a Cox interaction term selects treatment-benefit-predictive radiomic features for unsupervised hierarchical clustering. method
  • DeepCRC topology-aware deep-learning segmentation plus PyRadiomics feature extraction provides an automated, explainable imaging pipeline for stage II CRC risk stratification. resource
Experimental setups
Assay System Perturbation Readout Platform
Contrast-enhanced CT radiomics (DeepCRC segmentation + PyRadiomics) Human stage II CRC patients (6 cohorts; GDPH n=405 development, YNCC n=153 validation, TCIA/COAD, GEO) adjuvant chemotherapy vs observation (clinical, none) radiomic features for treatment-benefit clustering / survival DeepCRC, PyRadiomics, ITK-SNAP
Bulk RNA sequencing with differential expression, GSVA, and cellular deconvolution Human CRC tumour tissue (60 patients from training set) none (radiological cluster comparison) differentially expressed genes, hallmark/KEGG pathway enrichment, immune/stromal TME composition MSigDB hallmark v7.5.1, KEGG
Transcriptomic drug-response analysis Public GEO CRC datasets fluorouracil-based chemotherapy chemotherapy benefit and underlying TME
H&E histopathology pattern evaluation Human CRC whole-tumour sections none haemorrhage, necrosis, TLS, GC+TLS, tumour budding, desmoplastic reaction
Double immunohistochemistry / vessel quantification Human CRC tumour sections none CD31/PanCK staining, tumour-stroma ratio, vessel junctions, mesh size, segment length QuPath, ImageJ (Color Deconvolution, Angiogenesis Analyser)
Multiplex immunohistochemistry (mIHC) with spatial analysis Human stage II CRC FFPE specimens (10 patients, GDPH) none vessel phenotypes and immune cell infiltration across 100 μm interface tiles HALO image analysis software v3.2 (Indica Labs)
Micro-CT imaging radiomics + IHC vascular analysis CT26 cell-line-derived xenograft mouse model (12 female BALB/c mice) anlotinib (antiangiogenic TKI) vs saline radiomic/textural CT features and vascular patterns (microvessel pericyte coverage index, CD31/α-SMA) micro-CT; PASS v15.0.5 for sample size
Key results
  • Survival benefit of chemotherapy varied significantly between AI-powered radiological clusters iHR=5.35 (95% CI 1.98–14.41), adjusted P_interaction=0.012
  • AC-preferable cluster showed vessels with greater pericyte coverage and enriched B, CD4+-T, CD8+-T cell infiltration into tumour core
  • OO-preferable cluster exhibited higher necrosis, haemorrhage, and tortuous vessels
  • Anlotinib-induced changes in vessel morphology produced alterations in predictive imaging features
  • Microvessel pericyte coverage index (MPI) differed between experimental and control groups in preliminary mouse data experimental 1.2±0.92 vs control 3.9±1.2
Key statistics
  • other iHR=5.35 (95% CI 1.98, 14.41), adjusted P_interaction=0.012 (interaction hazard ratio for chemotherapy benefit between radiological clusters)
  • other <5% (5-year survival benefit of adjuvant chemotherapy in stage II CRC)
  • mean MPI 1.2±0.92 (experimental) vs 3.9±1.2 (control) (microvessel pericyte coverage index in mouse anlotinib vs saline groups)
  • count 405 development, 153 validation (GDPH development and YNCC validation cohort sizes)
  • count 60 (patients with RNA sequencing data from training set)
  • other ICC threshold 0.80 (radiomic feature robustness test-retest in 30 patients)
  • pvalue z-statistic <0.05 (interaction feature selection threshold for radiomic features)
  • count 10 FFPE specimens; 12 BALB/c mice (6 anlotinib, 6 saline) (mIHC specimen count and mouse model group allocation)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The study used a retrospective multicentre design to develop and validate an AI-powered radiomic risk stratification model for stage II CRC across six cohorts. Radiomic features predictive of treatment-specific survival benefit were selected via Cox proportional hazards models incorporating a treatment-by-feature interaction term (interaction hazard ratio, iHR), and unsupervised hierarchical clustering then grouped patients into two treatment-predictive subtypes. Primary efficacy of stratification was reported as an iHR with a 95% CI and adjusted P value; Benjamini-Hochberg FDR correction was applied for post-hoc multiple-comparison analyses. Biological explainability was pursued through GSVA pathway enrichment, cellular deconvolution of RNA-seq data, quantitative histopathology, and a randomised preclinical mouse experiment.

Replicationmixed Sample sizeClinical cohorts: development n = 405 (GDPH), external validation n = 153 (YNCC), plus four public cohorts (TCIA/COAD and three GEO datasets); mouse experiment n = 12 (6 per group), with sample size calculated via PASS software (α = 0.05 two-sided, power > 90%) GroupsAC-preferable vs OO-preferable radiological cluster; anlotinib-treated vs saline control (mouse model) Pairingunpaired Randomization/blindingstated Dispersionmixed Exact p-valuesyes Effect sizesyes Confidence intervalsyes Multiplicity correctionBenjamini-Hochberg false discovery rate (FDR)
Statistical tests used
Test Applied to n Assumptions
Cox proportional hazards model with treatment-by-feature interaction term (iHR, Wald test) Per-feature selection of treatment-predictive radiomic biomarkers; reported overall interaction HR = 5.35 (95% CI 1.98–14.41) Development cohort n = 405 (GDPH); RNA-seq subset n = 60 not stated
Hierarchical clustering (unsupervised) with dendrogram and silhouette methods for optimal cluster number Discovery of AC-preferable vs OO-preferable radiological subtypes Development cohort n = 405 not stated
Intraclass correlation coefficient (ICC, threshold 0.80) Feature robustness test-retest study comparing original and automated segmentations n = 30 patients not stated
Gene set variation analysis (GSVA) Hallmark (MSigDB v7.5.1) and KEGG pathway enrichment comparison between radiological clusters n = 60 (RNA-seq patients in training set) not stated
Cellular deconvolution algorithm (transcriptome-based) Tumour microenvironment immune and stromal component estimation across radiological clusters and GEO cohorts Not explicitly stated for each GEO cohort not stated
Two-sample t-test (power/sample-size calculation formula via PASS v15.0.5) Sample size determination for preclinical mouse model (microvessel pericyte coverage index, MPI) n = 6 per group (total n = 12 mice) stated
Approaches that could also have been used
  • Unsupervised hierarchical clustering was used to discover the two radiological subtypes, with the number of clusters chosen by dendrogram inspection and silhouette score.
    Could also: Consensus clustering (e.g., via the R package ConsensusClusterPlus) or non-negative matrix factorisation (NMF) could also define robust subtypes. — Consensus clustering quantifies subtype stability across bootstrap resamples and provides a formal instability score, which can make the choice of cluster number more reproducible and transparent across datasets of different sizes.
  • Radiomic features were selected by running a separate Cox interaction model for each feature individually, then applying a z-statistic threshold.
    Could also: A penalised Cox model with an interaction term (e.g., lasso or elastic-net penalisation via glmnet) could also perform simultaneous feature selection across all features. — Penalised regression handles correlated radiomic features jointly and implicitly regularises against overfitting, whereas sequential univariate screening can miss higher-order covariate structure and may be sensitive to multicollinearity among the large radiomic feature set.
  • Tumour microenvironment cellular composition was estimated from bulk RNA-seq using a deconvolution algorithm.
    Could also: Single-cell RNA sequencing or spatially resolved transcriptomics could also characterise cell-type abundances and their spatial distributions. — Bulk deconvolution infers cell proportions from aggregate signals and relies on reference signatures; single-cell or spatial approaches resolve individual cell states and neighbourhoods directly, which can complement and validate deconvolution estimates.
  • Cluster-level pathway enrichment was assessed with GSVA using MSigDB hallmark and KEGG gene sets.
    Could also: Gene set enrichment analysis (GSEA) with permutation-based statistics, or over-representation analysis (ORA) on differentially expressed genes, could also quantify pathway-level differences between clusters. — GSVA produces per-sample enrichment scores suitable for continuous comparisons, while GSEA uses ranked gene lists and is better suited to detecting coordinated directional shifts; ORA provides straightforward interpretability. Reporting results from more than one method can strengthen convergent conclusions.
  • Preliminary mouse model data (MPI) were reported as mean ± SD and used to power a two-sample t-test.
    Could also: A non-parametric Mann-Whitney U test could also compare MPI between groups, particularly given the small per-group n (n = 6). — With n = 6 per group, normality is difficult to verify, and the assumption underlying the t-test is hard to assess; the Mann-Whitney U test makes no distributional assumption about the outcome and is therefore a common alternative for small-n preclinical experiments.
  • The overall survival benefit of cluster assignment was summarised with a single interaction hazard ratio derived from a Cox model.
    Could also: Restricted mean survival time (RMST) differences between treatment arms within each cluster could also quantify the survival benefit on an absolute time scale. — The iHR summarises a relative, multiplicative treatment–subgroup interaction under the proportional hazards assumption; RMST is assumption-free in that respect and expresses the benefit as an absolute number of life-years gained over a defined horizon, which can be more directly interpretable for clinical decision-making.
Software: DeepCRC (topology-aware deep learning segmentation) · ITK-SNAP · PyRadiomics · GSVA (R package) · MSigDB hallmark gene sets v7.5.1 · QuPath · ImageJ with Color Deconvolution and Angiogenesis Analyser plugins · HALO image analysis software (Indica Labs) v3.2 · PASS (Power Analysis and Sample Size, NCSS) v15.0.5

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
13
Impact: medium
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-40472802

Paper: Xie et al., Multimodal data integration for biologically-relevant AI to guide adjuvant chemotherapy in stage II colorectal cancer. EBioMedicine 2025. PMID 40472802 · PMCID PMC12171563 · DOI 10.1016/j.ebiom.2025.105789. Authors' repo: https://github.com/cx601/AIRCSA · data: GEO GSE28702 (+GSE62080, GSE69657, TCGA, TCIA, OMIX002301).

This is a large multimodal study. Most results are NOT cheaply reproducible. Triage below.

OUT OF SCOPE (not attempted — reason)

  • CT segmentation (DeepCRC / nnUNet) + radiomics (PyRadiomics, 851 features) — requires the institutional CT cohort (405 primary + 153 validation patients); imaging data is controlled/on-request (TCIA + internal), heavy GPU compute. data_restricted.
  • iHR survival-by-treatment interaction (iHR 5.35; pooled 2.88), OS benefit curves (Fig.3) — depend on the radiomic clusters from the private imaging cohort. Not reproducible without that data.
  • Multiplex IF / HALO, mouse model, ImageJ angiogenesis (Fig.5–6) — wet-lab
    • proprietary software (HALO v3.2), no public inputs. non_pipeline.
  • OMIX002301 transcriptomics — author-deposited, used for the internal radiogenomic cluster contrast (limma in RadiogenomicAnalysis.R); access TBD, not the cleanest public target.

IN SCOPE (attempted) — one clear, public, low-compute pipeline output

  • xCell cellular deconvolution on GSE28702 (public, 83 samples: 42 responders / 41 non-responders to FOLFOX), endothelial-cell infiltration responder vs non-responder.
    • Pipeline: GEOquery fetch -> probe→gene collapse -> xCell (Aran 2017, the exact deconvolution tool used in the authors' AIRCSA repo) -> Wilcoxon test.
    • Reported claim (Results, GSE28702 validation): "The infiltration of EC and microvascular EC was significantly higher in the responder group than that in the non-responder group in GSE28702."
    • This is a P16 "third-party tool on the paper's own public data" reproduction: equally valid. Light compute, fully public inputs, a checkable directional + significance claim.

Why this target

It is the only result whose inputs are fully public (GEO), whose tool is named and runnable (xCell), and whose claim is pinnable (EC + mv EC higher in responders, significant). Everything else is gated behind private imaging / wet-lab data. Honest 80/20: we reproduce this one cleanly and do not chase the imaging-dependent 80%.

Figures / tables: Fig.4
mvEC_resp
Reported
microvascular EC infiltration significantly higher in FOLFOX responders vs non-responders (GSE28702)
Reproduced
higher in responders; Wilcoxon p=0.0012 (microarray cal.) / 0.010 (rnaseq=TRUE cal.) — significant
exact
EC_resp
Reported
EC infiltration significantly higher in FOLFOX responders vs non-responders (GSE28702)
Reproduced
higher in responders (direction matches) but p=0.113 (microarray) / 0.188 (rnaseq=TRUE) — not significant at 0.05
partial
cohort_n
Reported
GSE28702 = 42 responders / 41 non-responders (n=83)
Reproduced
42 / 41, n=83 (parsed from GEO sample titles)
exact

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 83/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟡2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +8

Within the only reproducible, fully-public slice of this large multimodal CRC paper (xCell on GEO GSE28702, 42 resp/41 non-resp matched exactly), mv-EC infiltration reproduces cleanly and significantly (p=0.0012/0.010) while overall-EC reproduces in direction but loses the 'significant' qualifier (p=0.113/0.188). The deviation sits on the input/method side — series-matrix vs CEL-level RMA, probe-collapse, one- vs two-sided testing, or pooling of the three GEO sets — not in any private-data defect, and there is no fabrication concern since the reported direction is derivable from the shipped public data. The paper prints no exact GSE28702 p-value, so only a directional+significance qualifier is checkable; overall this is a solid partial reproduction of a peripheral validation claim, with the headline AI/imaging pipeline untestable (private cohort).

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

93.2 k
tokens (I/O) · 4.6 M incl. cache
12 min
runtime · 0.03 CPU-h
1.9 GB
peak RAM
1
HPC jobs
hummel
machine