Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Molecular Biomarker of Drug Resistance Developed From Patient-Derived Organoids Predicts Survival of Colorectal Cancer Patients.

Front Oncol · 2022
L1 53/100 3/4
Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +6
✓ What held up
  • Reported values were directly comparable
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🟡A deviation arose in the data or preprocessing
  • 🔴A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
53/100
Reproducibility score
1.2 SD below mean
vs. all fields · 1173 studies
🎯 Scores higher than 14% of all assessed papers rank 1005 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough to attempt a clean 1:1: the paper publishes the exact 5-gene Drug-Resistant Score (DRS) formula and validates it on the public GSE40967 cohort (the brief's data accession), reporting univariate Cox P=8e-04 for OS. We applied the published model to GSE40967 (GPL570) via GEOquery on «our HPC». RESULT = PARTIAL: the biomarker's DIRECTION reproduces in every configuration (high DRS -> worse OS, HR>1) and the association is significant under the paper's described z-scored + maxstat-dichotomized pipeline (log-rank P=0.0023-0.021), but the EXACT reported P=8e-04 is not matched and the verbatim continuous-DRS Cox is non-significant (P=0.44). The gap is well-explained by two undocumented choices: (1) the paper used n=233 whereas the public cohort has 573 OS-evaluable samples and the subset basis is not stated, and (2) 'gene expression level' normalization + multi-probe collapse are unspecified (coefficients fit on TCGA FPKM, applied across 4 array platforms). No fabrication indicated; this is a reproducibility/documentation gap. NOT attempted (out of scope / hard 20%): LASSO re-derivation on TCGA-CRC (coefficients already published), the other 3 GEO validation cohorts, all wet-lab organoid work, and the Fig-8 downstream analyses (GSEA/enrichplot, TMB, ESTIMATE immune/stromal, CIBERSORT). The brief's 'code' link (GuangchuangYu/enrichplot) is a third-party GSEA visualization tool, not an analysis repo; reproduction applied the paper's published model to public data per its prose Methods (P16).

💻 Code ↗ 🗄 Data: GSE40967

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 53
    assessed: 2026-06-14 ⛓ 230618ac6839
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-14
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

Can a gene signature and drug-resistant score model derived from 5-fluorouracil (5-Fu) drug-sensitivity data of patient-derived colorectal cancer organoids serve as a molecular biomarker to predict the survival of colorectal cancer patients?

Core claims
  • A drug-resistant score model (DRSM) of five genes (CACNA1D, CIITA, PFN2, SEZ6L2, WDR78) derived from colorectal cancer organoid 5-Fu sensitivity predicts overall survival of CRC patients. resource
  • Drug-resistant score (DRS) is an independent prognostic factor for overall survival in CRC patients in the TCGA-CRC cohort. finding
  • Colorectal cancer organoids show great diversity in 5-Fu drug sensitivity, with a subset resistant and a subset sensitive. finding
  • Differentially expressed genes associated with 5-Fu resistance were identified by transcriptome sequencing of organoids before/after treatment and sensitive vs resistant organoids. method
  • The DRSM was validated across four GEO cohorts and predicts survival within different patient subgroups. finding
  • Organoid size change (day24/day0) is an effective, economical measure of organoid survival/drug sensitivity comparable to CellTiter-Glo 3D viability assay. method
  • DRS-high and DRS-low patients differ in molecular pathways, tumor mutational burden, immune response pathways, immune/stromal scores, and immune cell proportions. finding
  • Patient-derived colorectal cancer organoids can be successfully established from a majority of surgical CRC specimens. resource
Experimental setups
Assay System Perturbation Readout Platform
Organoid drug sensitivity test (organoid size change, day24/day0) 41 patient-derived colorectal cancer organoid (CRCO) lines in 3D Matrigel 10 μM 5-fluorouracil treatment Organoid size change ratio (day24/day0) as measure of survival/sensitivity ZEISS microscope (Vert.A1); Image-Pro Plus 6.0 software
Bulk RNA-seq (transcriptome sequencing) Patient-derived colorectal cancer organoids (sensitive vs resistant; before vs surviving after 5-Fu) 5-Fu treatment vs untreated comparisons Differentially expressed genes (FPKM expression levels) Illumina Novaseq, 150-bp paired-end; NEBNext Ultra RNA Library Prep Kit
Organoid generation/culture 50 surgically resected CRC tumor tissues from untreated CRC patients none Organoid culture success rate
Computational survival/biomarker modeling (LASSO regression, Cox regression, Kaplan-Meier) TCGA-CRC (TCGA-COAD/READ) and GEO cohorts (stage II-IV CRC patients) none Drug-resistant score, overall survival prediction, hazard ratios R packages: glmnet, limma, clusterProfiler, coxph, MaxStat
Key results
  • 41 organoid cultures successfully generated from 50 CRC tumor tissues (82% success rate) 82% (41/50)
  • 14 cases (34.1%) were 5-Fu sensitive and 27 (65.9%) were resistant 34.1% sensitive / 65.9% resistant
  • Drug-resistant score was an independent prognostic factor for overall survival in TCGA-CRC cohort P < 0.001
  • Five-gene DRSM developed from organoids predicts survival of CRC patients across four GEO validation cohorts
  • Organoid drug sensitivity (CRCO size day24/day0) showed great diversity across the 41 organoid lines under 10 μM 5-Fu
Key statistics
  • count 41 organoid cultures from 50 tissues (82%) (CRCO establishment success rate)
  • count 14 (34.1%) sensitive, 27 (65.9%) resistant (5-Fu sensitivity classification of organoids)
  • pvalue P < 0.001 (DRS as independent prognostic factor for OS in TCGA-CRC (multivariate analysis))
  • other 36.42% (validated cutoff value of organoid size change for sensitivity judgment)
  • count 26 genes with non-zero LASSO coefficients; 5 genes significant (P < 0.05) (gene filtration for DRSM)
  • pvalue P < 0.05 (significance threshold for DEGs and prognostic genes)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The study classified 41 patient-derived colorectal cancer organoids as 5-Fu sensitive (n=14) or resistant (n=27) using organoid size change, then identified DEGs by transcriptome sequencing analyzed with the R package limma. DEGs associated with both 5-Fu resistance and patient survival in the TCGA-CRC cohort were progressively filtered by univariate and multivariate Cox regression, then by LASSO regression with 5-fold cross-validation, yielding a 5-gene Drug-Resistant Score Model (DRSM). The model was evaluated using Kaplan-Meier curves with log-rank tests and validated across four independent GEO cohorts, with multivariate Cox regression confirming DRS as an independent prognostic factor for overall survival.

Replicationbiological Sample size50 surgically resected tumor tissues yielding 41 organoid lines (82% success rate); 14 classified sensitive, 27 resistant; TCGA-CRC cohort used for model development; four GEO cohorts used for external validation; individual cohort sample sizes not stated in available text Groups5-Fu sensitive vs resistant CRCOs (organoid level); pre-treatment vs post-treatment surviving CRCOs; DRS-high vs DRS-low CRC patients in TCGA and GEO cohorts Pairingmixed Randomization/blindingnot stated DispersionSEM Exact p-valuesno Effect sizesyes Multiplicity correctionBenjamini-Hochberg FDR
Statistical tests used
Test Applied to n Assumptions
limma empirical Bayes moderated t-statistic DEG identification: (1) 5-Fu sensitive vs resistant untreated CRCOs; (2) pre-treatment CRCOs vs post-treatment surviving CRCOs 41 organoid lines total (14 sensitive, 27 resistant); n per group for pre/post comparison not stated not stated
Gene Set Enrichment Analysis (GSEA) with Benjamini-Hochberg FDR correction Functional pathway enrichment of transcriptomic data; 1,000 permutations not stated
Univariate Cox proportional hazards regression Screening each DEG for association with overall survival in TCGA-CRC cohort; P < 0.05 threshold for retention not stated
Multivariate Cox proportional hazards regression Identification of independent prognostic factors from LASSO-retained genes; confirmation of DRS as independent prognostic factor in TCGA-CRC and GEO cohorts not stated
LASSO regression (glmnet v4.0-2) with 5-fold cross-validation Feature selection from univariate Cox-filtered genes to construct DRSM; lambda selected by minimizing cross-validated error not stated
Log-rank test Kaplan-Meier survival curve comparisons between DRS-high and DRS-low groups in TCGA and GEO cohorts not stated
MaxStat maximum rank statistic Optimal cutpoint selection to dichotomize patients into DRS-high and DRS-low groups not stated
Wilcoxon rank-sum test Comparison of two groups (specific comparisons not fully specified in available text) not stated
Two-sided Fisher exact test Analysis of contingency tables (specific comparisons not fully specified in available text) not stated
Spearman correlation and distance correlation Correlation coefficient analyses (specific variable pairs not fully specified in available text) not stated
Approaches that could also have been used
  • limma (an empirical Bayes method originally designed for microarray data) was applied to RNA-seq FPKM values for DEG identification
    Could also: DESeq2 or edgeR, which model raw RNA-seq read counts with a negative binomial distribution, could also be used for DEG analysis — DESeq2 and edgeR are purpose-built for count-based RNA-seq data and explicitly model count overdispersion; they are among the most extensively benchmarked tools for this data type and are frequently recommended when count-level data are available as the analysis input
  • An optimal data-driven cutpoint (MaxStat maximum rank statistic) was used on the TCGA training cohort to dichotomize patients into DRS-high and DRS-low groups
    Could also: The median DRS, pre-specified quartiles, or retaining DRS as a continuous predictor in Cox regression could also be used — Data-driven cutpoint optimization on the same dataset used for evaluation can inflate apparent group separation and reduce external generalizability; continuous modeling or a pre-specified split preserves statistical power and avoids this dependency
  • Univariate Cox regression (P < 0.05) was used as a pre-screening step before LASSO feature selection, running a separate test for each DEG without a multiplicity correction at that stage
    Could also: Applying LASSO directly to all DEGs without a prior univariate screen, or applying Benjamini-Hochberg FDR to the univariate p-values before retaining genes, could also be used — Running many uncorrected univariate tests before penalized regression can admit marginally significant or correlated genes; FDR-controlled pre-screening or direct penalized regression makes the selection process more self-contained and reproducible
  • Organoid size change data were summarized with SEM (standard error of the mean) from 8 replicates
    Could also: Standard deviation (SD) or a 95% confidence interval could also be used to describe the spread of replicate measurements — For a small number of replicates, SD conveys the observed biological or technical variability in the measurements directly, whereas SEM reflects precision of the mean estimate; both are informative, and the choice affects how readers interpret the magnitude of variability
  • 5-fold cross-validation was used within the TCGA training cohort to select the optimal LASSO penalty (lambda)
    Could also: Bootstrap resampling or leave-one-out cross-validation (LOOCV) could also estimate prediction error and select lambda — Bootstrap resampling provides lower-variance optimism-corrected estimates of predictive performance, particularly in moderate-sized cohorts; LOOCV is another established alternative; the choice of internal validation method can affect the stability of the selected gene set and reported performance metrics
  • Survival differences between DRS groups were evaluated with the log-rank test, which implicitly assumes proportional hazards
    Could also: A Cox model with DRS as a continuous predictor, or a restricted mean survival time (RMST) analysis, could also quantify survival differences — Treating DRS continuously in Cox regression avoids information loss from dichotomization; RMST does not require the proportional hazards assumption and provides a directly interpretable difference in survival time, which may be informative when proportional hazards has not been formally verified
Software: R/limma · R/clusterProfiler · R/glmnet 4.0-2 · R/survival (coxph) · R/MaxStat · R/pheatmap · R/enrichplot · R/TCGAbiolinks · R/GEOquery · R/forestplot · Image-Pro Plus 6.0

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
8
Impact: low
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

GSE17538 GEO in Results (http://purl.org/orb/Results)
also used by 2 papers:
GSE14333 GEO in Results (http://purl.org/orb/Results)
also used by 1 paper:
GSE87211 GEO in Results (http://purl.org/orb/Results)
also used by 1 paper:
10.17632/rnrmjkvjjc.2 DOI in Data Availability (http://purl.obolibrary.org/obo/IAO_0000611)
no other assessed paper uses this yet
GSE12945 GEO in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GSE24551 GEO in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GSE29623 GEO in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GSE33113 GEO in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GSE38832 GEO in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GSE39084 GEO in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GSE40967 GEO in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GSE71187 GEO in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
PRJNA813221 BioProject in Data Availability (http://purl.obolibrary.org/obo/IAO_0000611)
no other assessed paper uses this yet

Downstream reach in the literature

669 downstream papers · 12 datasets

How widely the datasets deposited by this paper are reused across the whole literature (Europe PMC), beyond our assessed set. This is a factual dependency map — reusing a public dataset is normal, good science. It is not a judgement on the downstream papers; the only verdict here is this paper's own, with its cited rationale.

PRJNA813221 BioProject reused by 1 papers in the literature

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — PMID 35425715

Paper: Molecular Biomarker of Drug Resistance Developed From Patient-Derived Organoids Predicts Survival of Colorectal Cancer Patients. Front Oncol 2022; PMCID PMC9004628; DOI 10.3389/fonc.2022.855674.

"Code" link in brief: https://github.com/GuangchuangYu/enrichplot (now YuLab-SMU/enrichplot) — this is a third-party visualization tool (GSEA result plots), cited as a dependency, NOT the authors' own analysis repo. Per brief rule P16, applying a third-party tool/pipeline to the paper's own data is an equally valid reproduction. The authors ship no dedicated analysis repository; the pipeline is described prose-only in Methods (limma, clusterProfiler/GSEA, glmnet/LASSO, coxph, maxstat).

Data link in brief: geo:GSE40967 — public (Marisa et al. 2013 colon-cancer cohort, GPL570 Affymetrix HG-U133 Plus 2.0; frma-normalized log2). This is one of the paper's 4 GEO validation cohorts and the one named in the brief. Public, resolves (GEO id 200040967), clinical incl. os.event + os.delay (months).

Results classified

IN SCOPE (pipeline-derived, reproducible from public data)

Result Pipeline Reproducible?
DRSM validation on GSE40967: 5-gene Drug-Resistant Score predicts OS; univariate Cox P = 8e-04 (Results "Validation of the DRSM", Fig 6D) Apply published DRS formula to GSE40967 expression → univariate Cox vs OS YES — primary target. Formula + coefficients fully published; data public.
(stretch) Same on GSE17538 P=0.0016, GSE87211 P=0.018, GSE38832 P=0.0044 same Possible but other accessions, fiddly per-cohort clinical parsing = the hard 20%

The published model (verbatim, Methods "Development of the DRSM"):

DRS = GEL(CACNA1D)*-0.0563 + GEL(CIITA)*-0.0356 + GEL(PFN2)*0.0332
    + GEL(SEZ6L2)*0.0378 + GEL(WDR78)*-0.0386

maxstat used to split DRS-high/low; coxph for univariate/multivariate.

OUT OF SCOPE (wet-lab / own non-public data / manual)

  • Organoid culture, 41 CRCO lines, 5-Fu drug-sensitivity assay (wet-lab).
  • Organoid transcriptome sequencing + DEGs (limma) — raw organoid RNA-seq not in the brief's data link; the 5 genes are the published output, so we validate the output model, not re-derive it.
  • LASSO derivation on TCGA-CRC (could be attempted via TCGAbiolinks but TCGA download + LASSO refit = heavy + the 26→5 gene selection has manual coxph filtering steps; the coefficients are already published, so re-deriving them is the hard 20% and not required to validate the model).
  • GSEA/enrichplot pathway figures, TMB, ESTIMATE immune/stromal score, CIBERSORT immune cell proportions (Fig 8) — depend on TCGA DRS-high/low grouping; downstream.

Plan

Reproduce the GSE40967 univariate-Cox OS validation of the published 5-gene DRS (the brief's named data point). Honest 1:1: does applying the paper's exact formula to the public GSE40967 expression reproduce a significant OS association (reported P=8e-04)? Compute on «our HPC» (GEOquery + survival + maxstat conda env).

DRS_GSE40967_uniCox_P
Reported
univariate Cox P = 8e-04 (5-gene DRS vs OS, GSE40967)
Reproduced
verbatim continuous-Cox P=0.44 (non-sig); under paper's z-scored + maxstat-dichotomized pipeline log-rank P=0.0023-0.021 (significant, same direction); exact 8e-04 not matched
partial
DRS_GSE40967_n
Reported
n = 233
Reproduced
573 OS-evaluable of 585 public GSE40967-GPL570 samples
did not match
DRS_direction
Reported
high DRS = worse OS
Reproduced
HR_high-vs-low = 1.48-2.13 (>1) in all variants
exact

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 53/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟢2. Endpoint comparability
🟡3. Location of the main deviation
🔴4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +6

The published 5-gene DRS model reproduces in direction robustly (HR_high/low 1.48–2.13, high DRS = worse OS) and recovers significance under the paper's described z-scored + maxstat-dichotomized pipeline (log-rank P=0.0023–0.021), but the exact reported univariate Cox P=8e-04 is not matched and the verbatim continuous-Cox is non-significant (P=0.44). The deviation lies on the input/authors' side: an undocumented n=233 subset (vs 573 OS-evaluable public samples) and unspecified normalization/probe-collapse, not the published coefficients. Severity is moderate — magnitude and direction hold and the association is recoverable — with no fabrication indicator; this is a reproducibility and documentation gap.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

123.6 k
tokens (I/O) · 7.3 M incl. cache
14 min
runtime · 0.02 CPU-h
1.9 GB
peak RAM
2
HPC jobs
hummel
machine