Corpus 1,286 assessed · 1,187 scored · 648 reproduced ≥75 · 174 flagged ·∅ 73.9/100
← New search

Molecular Biomarker of Drug Resistance Developed From Patient-Derived Organoids Predicts Survival of Colorectal Cancer Patients.

Front Oncol · 2022
L1 53/100 3/4
Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +6
✓ What held up
  • Reported values were directly comparable
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🟡A deviation arose in the data or preprocessing
  • 🔴A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
53/100
Reproducibility score
1.2 SD below mean
vs. all fields · 1187 studies
🎯 Scores higher than 14% of all assessed papers rank 1014 of 1187 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough to attempt a clean 1:1: the paper publishes the exact 5-gene Drug-Resistant Score (DRS) formula and validates it on the public GSE40967 cohort (the brief's data accession), reporting univariate Cox P=8e-04 for OS. We applied the published model to GSE40967 (GPL570) via GEOquery on «our HPC». RESULT = PARTIAL: the biomarker's DIRECTION reproduces in every configuration (high DRS -> worse OS, HR>1) and the association is significant under the paper's described z-scored + maxstat-dichotomized pipeline (log-rank P=0.0023-0.021), but the EXACT reported P=8e-04 is not matched and the verbatim continuous-DRS Cox is non-significant (P=0.44). The gap is well-explained by two undocumented choices: (1) the paper used n=233 whereas the public cohort has 573 OS-evaluable samples and the subset basis is not stated, and (2) 'gene expression level' normalization + multi-probe collapse are unspecified (coefficients fit on TCGA FPKM, applied across 4 array platforms). No fabrication indicated; this is a reproducibility/documentation gap. NOT attempted (out of scope / hard 20%): LASSO re-derivation on TCGA-CRC (coefficients already published), the other 3 GEO validation cohorts, all wet-lab organoid work, and the Fig-8 downstream analyses (GSEA/enrichplot, TMB, ESTIMATE immune/stromal, CIBERSORT). The brief's 'code' link (GuangchuangYu/enrichplot) is a third-party GSEA visualization tool, not an analysis repo; reproduction applied the paper's published model to public data per its prose Methods (P16).

💻 Code ↗ 🗄 Data: GSE40967

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 53
    assessed: 2026-06-14 ⛓ 230618ac6839
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-14
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-09-19

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

A gene signature and drug-resistant score model built from 5-fluorouracil (5-Fu) sensitivity data of patient-derived colorectal cancer organoids can serve as a molecular biomarker to predict the survival of colorectal cancer patients.

Core claims
  • Patient-derived colorectal cancer organoids (CRCOs) can be established from surgical CRC tissue with an 82% success rate (41/50) resource
  • Organoid size change assay effectively differentiates 5-Fu sensitive (34.1%) versus resistant (65.9%) CRCOs, reflecting clinically observed heterogeneity in 5-Fu response finding
  • Transcriptome sequencing of CRCOs identified DEGs associated with 5-Fu resistance, including molecules and pathways previously reported in 5-Fu resistance finding
  • A five-gene drug-resistant score model (DRSM: CACNA1D, CIITA, PFN2, SEZ6L2, WDR78) was constructed via LASSO regression from organoid-derived DEGs in the TCGA-CRC cohort method
  • The drug-resistant score (DRS) is an independent prognostic factor for overall survival in CRC patients in the TCGA-CRC cohort (P < 0.001) finding
  • The DRSM was validated in four independent GEO cohorts and predicts survival across different patient subgroups finding
  • DRS-high and DRS-low patients differ in molecular pathways, tumor mutational burden, immune-related pathways, immune/stromal scores, and immune cell composition finding
  • Organoid size change is as effective as the CellTiter-Glo 3D cell viability assay for judging drug sensitivity, while being more economical and easier to use method
Experimental setups
Assay System Perturbation Readout Platform
organoid culture establishment primary CRC tumor tissue (human surgical specimens) none organoid formation success rate
drug sensitivity assay (organoid size change) colorectal cancer organoids (CRCOs) 10 μM 5-Fu treatment organoid size change ratio (day24/day0) ZEISS Vert.A1 microscope; Image-Pro Plus 6.0
bulk transcriptome sequencing (RNA-seq) CRCOs (5-Fu sensitive vs resistant; pre- vs post-5-Fu treatment) 5-Fu treatment vs untreated differentially expressed genes (DEGs) Illumina Novaseq; NEBNext Ultra RNA Library Prep Kit
functional enrichment analysis (GSEA) CRCO transcriptome data 5-Fu resistance status enriched molecular pathways clusterProfiler / enrichplot R packages
gene expression and survival cohort analysis TCGA-COAD/READ CRC patient cohort none (observational) DRS, overall survival, TMB, immune score, stromal score, immune cell proportion TCGAbiolinks
gene expression and survival validation 4 GEO CRC patient cohorts none (observational) survival prediction by DRSM GEOquery
LASSO and Cox regression modeling TCGA-CRC gene expression and survival data none gene coefficients / drug-resistant score model glmnet (v4.0-2) R package; coxph
Key results
  • 41 of 50 (82%) CRC tissues yielded viable organoid cultures 82%
  • 14/41 (34.1%) CRCOs were 5-Fu sensitive and 27/41 (65.9%) were resistant
  • DEGs identified in sensitive vs resistant untreated CRCOs and in pre- vs post-5-Fu-treatment CRCOs overlapped known 5-Fu resistance genes/pathways
  • Five-gene DRSM (CACNA1D, CIITA, PFN2, SEZ6L2, WDR78) constructed from 26 candidate genes via LASSO with 5-fold cross-validation
  • DRS is an independent predictor of overall survival in multivariate Cox analysis in the TCGA-CRC cohort P < 0.001
  • DRSM predicted survival across four GEO validation cohorts and across different patient subgroups
  • DRS-high vs DRS-low groups differed in TMB, immune-response pathways, immune score, stromal score, and immune cell proportions
  • A 36.42% cutoff of organoid size change (day24/d0) was used to classify 5-Fu sensitivity, consistent with prior validated methodology 36.42%
Key statistics
  • pvalue P < 0.001 (DRS as independent prognostic factor for OS, multivariate Cox analysis, TCGA-CRC cohort)
  • count 41/50 (82%) (CRCO organoid culture establishment success rate)
  • count 14/41 (34.1%) sensitive, 27/41 (65.9%) resistant (5-Fu drug sensitivity classification of CRCO lines)
  • other 36.42% cutoff (size day24/d0) (cutoff for classifying 5-Fu sensitivity of CRCOs)
  • count 26 genes with non-zero coefficients (candidate genes retained after LASSO regression filtration)
  • count 5 genes with P < 0.05 (genes retained as independent prognostic factors for the DRSM)
  • pvalue P < 0.05 (significance threshold for DEG identification via limma)
  • other DRS = GEL(CACNA1D)×-0.0563 + GEL(CIITA)×-0.0356 + GEL(PFN2)×0.0332 + GEL(SEZ6L2)×0.0378 + GEL(WDR78)×-0.0386 (drug-resistant score (DRS) model formula)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The study classified 41 patient-derived colorectal cancer organoids as 5-Fu sensitive (n=14) or resistant (n=27) using organoid size change, then identified DEGs by transcriptome sequencing analyzed with the R package limma. DEGs associated with both 5-Fu resistance and patient survival in the TCGA-CRC cohort were progressively filtered by univariate and multivariate Cox regression, then by LASSO regression with 5-fold cross-validation, yielding a 5-gene Drug-Resistant Score Model (DRSM). The model was evaluated using Kaplan-Meier curves with log-rank tests and validated across four independent GEO cohorts, with multivariate Cox regression confirming DRS as an independent prognostic factor for overall survival.

Replicationbiological Sample size50 surgically resected tumor tissues yielding 41 organoid lines (82% success rate); 14 classified sensitive, 27 resistant; TCGA-CRC cohort used for model development; four GEO cohorts used for external validation; individual cohort sample sizes not stated in available text Groups5-Fu sensitive vs resistant CRCOs (organoid level); pre-treatment vs post-treatment surviving CRCOs; DRS-high vs DRS-low CRC patients in TCGA and GEO cohorts Pairingmixed Randomization/blindingnot stated DispersionSEM Exact p-valuesno Effect sizesyes Multiplicity correctionBenjamini-Hochberg FDR
Statistical tests used
Test Applied to n Assumptions
limma empirical Bayes moderated t-statistic DEG identification: (1) 5-Fu sensitive vs resistant untreated CRCOs; (2) pre-treatment CRCOs vs post-treatment surviving CRCOs 41 organoid lines total (14 sensitive, 27 resistant); n per group for pre/post comparison not stated not stated
Gene Set Enrichment Analysis (GSEA) with Benjamini-Hochberg FDR correction Functional pathway enrichment of transcriptomic data; 1,000 permutations not stated
Univariate Cox proportional hazards regression Screening each DEG for association with overall survival in TCGA-CRC cohort; P < 0.05 threshold for retention not stated
Multivariate Cox proportional hazards regression Identification of independent prognostic factors from LASSO-retained genes; confirmation of DRS as independent prognostic factor in TCGA-CRC and GEO cohorts not stated
LASSO regression (glmnet v4.0-2) with 5-fold cross-validation Feature selection from univariate Cox-filtered genes to construct DRSM; lambda selected by minimizing cross-validated error not stated
Log-rank test Kaplan-Meier survival curve comparisons between DRS-high and DRS-low groups in TCGA and GEO cohorts not stated
MaxStat maximum rank statistic Optimal cutpoint selection to dichotomize patients into DRS-high and DRS-low groups not stated
Wilcoxon rank-sum test Comparison of two groups (specific comparisons not fully specified in available text) not stated
Two-sided Fisher exact test Analysis of contingency tables (specific comparisons not fully specified in available text) not stated
Spearman correlation and distance correlation Correlation coefficient analyses (specific variable pairs not fully specified in available text) not stated
Approaches that could also have been used
  • limma (an empirical Bayes method originally designed for microarray data) was applied to RNA-seq FPKM values for DEG identification
    Could also: DESeq2 or edgeR, which model raw RNA-seq read counts with a negative binomial distribution, could also be used for DEG analysis — DESeq2 and edgeR are purpose-built for count-based RNA-seq data and explicitly model count overdispersion; they are among the most extensively benchmarked tools for this data type and are frequently recommended when count-level data are available as the analysis input
  • An optimal data-driven cutpoint (MaxStat maximum rank statistic) was used on the TCGA training cohort to dichotomize patients into DRS-high and DRS-low groups
    Could also: The median DRS, pre-specified quartiles, or retaining DRS as a continuous predictor in Cox regression could also be used — Data-driven cutpoint optimization on the same dataset used for evaluation can inflate apparent group separation and reduce external generalizability; continuous modeling or a pre-specified split preserves statistical power and avoids this dependency
  • Univariate Cox regression (P < 0.05) was used as a pre-screening step before LASSO feature selection, running a separate test for each DEG without a multiplicity correction at that stage
    Could also: Applying LASSO directly to all DEGs without a prior univariate screen, or applying Benjamini-Hochberg FDR to the univariate p-values before retaining genes, could also be used — Running many uncorrected univariate tests before penalized regression can admit marginally significant or correlated genes; FDR-controlled pre-screening or direct penalized regression makes the selection process more self-contained and reproducible
  • Organoid size change data were summarized with SEM (standard error of the mean) from 8 replicates
    Could also: Standard deviation (SD) or a 95% confidence interval could also be used to describe the spread of replicate measurements — For a small number of replicates, SD conveys the observed biological or technical variability in the measurements directly, whereas SEM reflects precision of the mean estimate; both are informative, and the choice affects how readers interpret the magnitude of variability
  • 5-fold cross-validation was used within the TCGA training cohort to select the optimal LASSO penalty (lambda)
    Could also: Bootstrap resampling or leave-one-out cross-validation (LOOCV) could also estimate prediction error and select lambda — Bootstrap resampling provides lower-variance optimism-corrected estimates of predictive performance, particularly in moderate-sized cohorts; LOOCV is another established alternative; the choice of internal validation method can affect the stability of the selected gene set and reported performance metrics
  • Survival differences between DRS groups were evaluated with the log-rank test, which implicitly assumes proportional hazards
    Could also: A Cox model with DRS as a continuous predictor, or a restricted mean survival time (RMST) analysis, could also quantify survival differences — Treating DRS continuously in Cox regression avoids information loss from dichotomization; RMST does not require the proportional hazards assumption and provides a directly interpretable difference in survival time, which may be informative when proportional hazards has not been formally verified
Software: R/limma · R/clusterProfiler · R/glmnet 4.0-2 · R/survival (coxph) · R/MaxStat · R/pheatmap · R/enrichplot · R/TCGAbiolinks · R/GEOquery · R/forestplot · Image-Pro Plus 6.0

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
8
Impact: low
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

GSE17538 GEO in Results (http://purl.org/orb/Results)
also used by 2 papers:
GSE14333 GEO in Results (http://purl.org/orb/Results)
also used by 1 paper:
GSE87211 GEO in Results (http://purl.org/orb/Results)
also used by 1 paper:
10.17632/rnrmjkvjjc.2 DOI in Data Availability (http://purl.obolibrary.org/obo/IAO_0000611)
no other assessed paper uses this yet
GSE12945 GEO in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GSE24551 GEO in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GSE29623 GEO in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GSE33113 GEO in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GSE38832 GEO in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GSE39084 GEO in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GSE40967 GEO in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GSE71187 GEO in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
PRJNA813221 BioProject in Data Availability (http://purl.obolibrary.org/obo/IAO_0000611)
no other assessed paper uses this yet

Downstream reach in the literature

669 downstream papers · 12 datasets

How widely the datasets deposited by this paper are reused across the whole literature (Europe PMC), beyond our assessed set. This is a factual dependency map — reusing a public dataset is normal, good science. It is not a judgement on the downstream papers; the only verdict here is this paper's own, with its cited rationale.

PRJNA813221 BioProject reused by 1 papers in the literature

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — PMID 35425715

Paper: Molecular Biomarker of Drug Resistance Developed From Patient-Derived Organoids Predicts Survival of Colorectal Cancer Patients. Front Oncol 2022; PMCID PMC9004628; DOI 10.3389/fonc.2022.855674.

"Code" link in brief: https://github.com/GuangchuangYu/enrichplot (now YuLab-SMU/enrichplot) — this is a third-party visualization tool (GSEA result plots), cited as a dependency, NOT the authors' own analysis repo. Per brief rule P16, applying a third-party tool/pipeline to the paper's own data is an equally valid reproduction. The authors ship no dedicated analysis repository; the pipeline is described prose-only in Methods (limma, clusterProfiler/GSEA, glmnet/LASSO, coxph, maxstat).

Data link in brief: geo:GSE40967 — public (Marisa et al. 2013 colon-cancer cohort, GPL570 Affymetrix HG-U133 Plus 2.0; frma-normalized log2). This is one of the paper's 4 GEO validation cohorts and the one named in the brief. Public, resolves (GEO id 200040967), clinical incl. os.event + os.delay (months).

Results classified

IN SCOPE (pipeline-derived, reproducible from public data)

Result Pipeline Reproducible?
DRSM validation on GSE40967: 5-gene Drug-Resistant Score predicts OS; univariate Cox P = 8e-04 (Results "Validation of the DRSM", Fig 6D) Apply published DRS formula to GSE40967 expression → univariate Cox vs OS YES — primary target. Formula + coefficients fully published; data public.
(stretch) Same on GSE17538 P=0.0016, GSE87211 P=0.018, GSE38832 P=0.0044 same Possible but other accessions, fiddly per-cohort clinical parsing = the hard 20%

The published model (verbatim, Methods "Development of the DRSM"):

DRS = GEL(CACNA1D)*-0.0563 + GEL(CIITA)*-0.0356 + GEL(PFN2)*0.0332
    + GEL(SEZ6L2)*0.0378 + GEL(WDR78)*-0.0386

maxstat used to split DRS-high/low; coxph for univariate/multivariate.

OUT OF SCOPE (wet-lab / own non-public data / manual)

  • Organoid culture, 41 CRCO lines, 5-Fu drug-sensitivity assay (wet-lab).
  • Organoid transcriptome sequencing + DEGs (limma) — raw organoid RNA-seq not in the brief's data link; the 5 genes are the published output, so we validate the output model, not re-derive it.
  • LASSO derivation on TCGA-CRC (could be attempted via TCGAbiolinks but TCGA download + LASSO refit = heavy + the 26→5 gene selection has manual coxph filtering steps; the coefficients are already published, so re-deriving them is the hard 20% and not required to validate the model).
  • GSEA/enrichplot pathway figures, TMB, ESTIMATE immune/stromal score, CIBERSORT immune cell proportions (Fig 8) — depend on TCGA DRS-high/low grouping; downstream.

Plan

Reproduce the GSE40967 univariate-Cox OS validation of the published 5-gene DRS (the brief's named data point). Honest 1:1: does applying the paper's exact formula to the public GSE40967 expression reproduce a significant OS association (reported P=8e-04)? Compute on «our HPC» (GEOquery + survival + maxstat conda env).

DRS_GSE40967_uniCox_P
Reported
univariate Cox P = 8e-04 (5-gene DRS vs OS, GSE40967)
Reproduced
verbatim continuous-Cox P=0.44 (non-sig); under paper's z-scored + maxstat-dichotomized pipeline log-rank P=0.0023-0.021 (significant, same direction); exact 8e-04 not matched
partial
DRS_GSE40967_n
Reported
n = 233
Reproduced
573 OS-evaluable of 585 public GSE40967-GPL570 samples
did not match
DRS_direction
Reported
high DRS = worse OS
Reproduced
HR_high-vs-low = 1.48-2.13 (>1) in all variants
exact

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 53/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟢2. Endpoint comparability
🟡3. Location of the main deviation
🔴4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +6

The published 5-gene DRS model reproduces in direction robustly (HR_high/low 1.48–2.13, high DRS = worse OS) and recovers significance under the paper's described z-scored + maxstat-dichotomized pipeline (log-rank P=0.0023–0.021), but the exact reported univariate Cox P=8e-04 is not matched and the verbatim continuous-Cox is non-significant (P=0.44). The deviation lies on the input/authors' side: an undocumented n=233 subset (vs 573 OS-evaluable public samples) and unspecified normalization/probe-collapse, not the published coefficients. Severity is moderate — magnitude and direction hold and the association is recoverable — with no fabrication indicator; this is a reproducibility and documentation gap.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at [email protected].

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

123.6 k
tokens (I/O) · 7.3 M incl. cache
14 min
runtime · 0.02 CPU-h
1.9 GB
peak RAM
2
HPC jobs
hummel
machine