Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Exploiting convergent phenotypes to derive a pan-cancer cisplatin response gene expression signature.

NPJ Precis Oncol · 2023
L1 50/100 PQI 83
Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
✓ What held up
  • Reported values were directly comparable
  • Any deviation was negligible
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
50/100
Reproducibility score
1.4 SD below mean
vs. all fields · 1173 studies
🎯 Scores higher than 8% of all assessed papers rank 1026 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough to attempt a faithful 1:1, but NOT completed. Target: the central cissig 19-gene signature (Table 1) and the epithelial GDSC cell-line count (429/430), via the authors' own deterministic pipeline (download_data -> clean_data -> extract_cissig; fold seed=0, SAM random.seed=1, MTP seed=1, Spearman co-expression on TCGA). On «our HPC» («job») the conda R/Bioconductor env built and loaded cleanly (samr 3.0, limma 3.58.1, multtest 2.58.0, RTCGAToolbox 2.32.1) and a faithful one-script port ran, but it aborted at the very FIRST data fetch: the authors' hardcoded GDSC RMA basal-expression URL (cancerrxgene.org/gdsc1000/.../Cell_line_RMA_proc_basalExp.txt.zip) returns HTTP 410 Gone (the GDSC site reorganised). No value was independently regenerated. What IS established and auditable: the paper's 19-gene Table 1 is exactly the gene list printed by the authors' committed knitted HTML (internal consistency, no fabrication signal); and a paper-vs-code discrepancy of 429 (Methods text) vs 430 (code output) cell lines. NOT attempted: the predictive-modelling AUCs (Table 2) and GSE48276/GSE70691 clinical KM/Cox analyses (Fig 6) - out of scope per 80/20 (heavy ML / plot-and-p-value based, downstream of the signature). Path to finish: swap get_exp() to a live GDSC RMA mirror; the rest of the pipeline was specified and ready to run within the 12h job.

💻 Code ↗ 🗄 Data: GSE48276

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 50
    assessed: 2026-06-15 ⛓ bc0e15bb832b
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-15
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

Because tumors with disparate genetic/mutational backgrounds can independently evolve convergent phenotypes (e.g., drug resistance), a consensus gene expression signature derived across many genomically diverse cancer cell lines can predict chemotherapy response (specifically cisplatin) better than mutation-based markers, without relying on a single genomic marker.

Core claims
  • A convergent-phenotype-based seed gene/co-expression method can extract consensus gene expression signatures predictive of response to chemotherapeutic drugs in the GDSC database method
  • The derived Cisplatin Response Signature (CisSig, 19 genes) predicts cisplatin response (IC50) within GDSC epithelial cancer cell lines finding
  • CisSig expression trends align with clinical use of cisplatin across disease sites in independent TCGA and Total Cancer Care (TCC) datasets finding
  • CisSig shows preliminary validation predicting overall survival in a small cohort of muscle-invasive bladder cancer patients treated with cisplatin-containing chemotherapy finding
  • CisSig's predictive performance (hazard ratio) exceeds the top 95% of a null distribution built from 1000 random gene signatures of the same length finding
  • CisSig score is largely independent of mutation status in common DNA damage response genes (BRCA1/2, PTEN, RAD51C/D, ATM, etc.), with only PTEN showing a significant association finding
  • CisSig gene list: ADAT2, ATP1B3, CDIN1, C1QBP, CDC7, CDCA7, FKBP14, KRT5, LRRC8C, LY6K, MMP10, NPM3, PSAT1, RIOK1, SLFN11, STOML2, USP31, WDR3, ZNF750 resource
  • The seed gene extraction approach is inspired by Buffa et al.'s hypoxia metagene method, using differential expression to define seeds and co-expression network trimming to refine the signature method
Experimental setups
Assay System Perturbation Readout Platform
Differential gene expression analysis (limma, SAM, multtest) 429 epithelial-based cancer cell lines, GDSC database none (comparison of top vs bottom 20% cisplatin IC50 responders) genes over-expressed in cisplatin-sensitive state (seed genes) RMA-normalized microarray
Co-expression network analysis TCGA RNA-seq data, epithelial-based normal and tumor tissue samples none genes highly co-expressed with seed genes (connectivity seeds) RNA-seq
Signature quality control analysis (sigQC package) TCGA epithelial-origin clinical samples none intra-signature correlation, mean/median correlation, standard deviation, coefficient of variance, composite QC score
Gene expression heatmap and CisSig score (median normalized expression) GDSC epithelial cancer cell lines (top/bottom IC50 quintiles) none CisSig gene expression Z-scores; CisSig score microarray
Cell Line Persistence Curve (Kaplan-Meier-like survival analysis) and null distribution comparison GDSC epithelial cancer cell lines none IC50 distribution/hazard ratio between high and low CisSig score cohorts vs 1000 random signature nulls
Mutation status analysis (chi-square test) GDSC epithelial cancer cell lines none presence of mutation in 16 DNA damage response genes vs CisSig score (top/bottom half)
Predictive modeling (simple/elastic net/L1/L2 linear regression, simple/elastic net/L1 logistic regression, random forest) GDSC epithelial cancer cell lines none IC50 predicted as continuous or binary outcome from CisSig score or all gene expression
Clinical outcome validation Muscle-invasive bladder cancer (MIBC) patient cohort cisplatin-containing chemotherapy overall survival
Key results
  • Spearman correlation between IC50 and AUC as cisplatin response metrics in GDSC cell lines r=0.84, p<0.001
  • CisSig score significantly different between cisplatin-sensitive and -resistant GDSC cell lines (Wilcoxon rank-sum test) p<0.0001
  • IC50 distribution significantly different between high and low CisSig score cohorts (log-rank test on Cell Line Persistence Curve) p<0.0001; median IC50 3.98 vs 7.93 log2(µM)
  • CisSig's hazard ratio performance exceeds the top 95% of a null distribution of 1000 random gene signatures of equal length
  • Of 16 DNA damage response genes tested, only PTEN showed a statistically significant relationship between mutation status and CisSig score after multiple testing correction p=0.042
  • Elastic net logistic regression using all gene expression to predict binary IC50 in quintile-selected cell lines AUC=0.94
  • L2 linear regression using all gene expression to predict continuous IC50 in quintile-selected cell lines corr. coef.=0.81
  • Simple linear regression using CisSig score alone to predict continuous IC50 corr. coef.=0.51 (all cell lines), 0.74 (quintiles)
Key statistics
  • correlation r=0.84, p<0.001 (Spearman correlation between IC50 and AUC drug response metrics)
  • pvalue p<0.0001 (Wilcoxon rank-sum test, CisSig score in sensitive vs resistant cell lines)
  • pvalue p<0.0001 (log-rank test, IC50 distribution between high and low CisSig score cohorts)
  • pvalue p=0.042 (chi-square test, PTEN mutation status vs CisSig score)
  • count 429 (epithelial-based cancer cell lines used from GDSC database)
  • count 19 (number of genes included in final CisSig signature)
  • other AUC=0.94 (elastic net logistic regression, quintile-selected cell lines, binary IC50 prediction)
  • correlation corr. coef.=0.81 (L2 linear regression, quintile-selected cell lines, continuous IC50 prediction)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

CisSig was derived from 429 epithelial cancer cell lines (GDSC) via a three-stage pipeline: differential gene expression analysis (limma, SAM, multtest) comparing top and bottom 20% cisplatin IC50 responders across five cross-validation folds, followed by co-expression network filtering in TCGA epithelial samples to select the final 19-gene signature. Validation employed Wilcoxon rank-sum tests, a novel log-rank-based 'Cell Line Persistence Curve' benchmarked against a 1000-signature permutation null distribution, chi-square tests for genomic associations, and a suite of regression and classification models evaluated by correlation coefficient and AUC. The paper also reports preliminary survival analysis in a small MIBC patient cohort, though that section is not fully reproduced in the provided text.

Replicationunclear Sample size429 epithelial-based GDSC cell lines stated; TCGA clinical samples used for co-expression (count not specified in provided text); MIBC patient cohort described only as 'small cohort' in the abstract; no formal power calculation described GroupsCisplatin-sensitive vs. resistant cell lines (top/bottom 20% IC50 for DE analysis; top/bottom quintile for validation); high vs. low CisSig score quintiles; TCGA disease sites grouped by standard-of-care cisplatin use; MIBC patients on cisplatin-containing chemotherapy Pairingunpaired Randomization/blindingnot stated Dispersionnone Exact p-valuesno Effect sizesyes Confidence intervalsno Multiplicity correctionMethod not named; paper states 'correcting for multiple hypothesis testing' for 16 chi-square tests (only PTEN reached significance at p = 0.042 after correction); for DE analysis, the intersection of all three methods (limma, SAM, multtest) is described as a stringency strategy to reduce false discovery rate rather than a named correction procedure
Statistical tests used
Test Applied to n Assumptions
Spearman correlation Concordance between IC50 and AUC drug response metrics across GDSC epithelial cell lines (Supplementary Fig. 1) 429 epithelial-based cell lines not stated
limma linear model (differential gene expression) Top vs. bottom 20% cisplatin IC50 responders, each of 5 cross-validation folds (Fig. 2b–c) 343–344 cell lines per fold not stated
SAM (Significance Analysis of Microarrays) Top vs. bottom 20% cisplatin IC50 responders, each of 5 cross-validation folds (Fig. 2b–c) 343–344 cell lines per fold not stated
multtest (multiple testing / permutation-based differential expression) Top vs. bottom 20% cisplatin IC50 responders, each of 5 cross-validation folds (Fig. 2b–c) 343–344 cell lines per fold not stated
Wilcoxon rank-sum test CisSig scores between top and bottom IC50 quintile cell lines (Fig. 3b) Top and bottom 20% of 429 cell lines (~86 per group, implied by quintile split) not stated
Log-rank test Cell Line Persistence Curves for high vs. low CisSig score quintile cohorts (Fig. 3c) Top and bottom 20% of 429 cell lines (~86 per group, implied by quintile split) not stated
Permutation null distribution (hazard ratio from 1000 random same-length gene signatures) CisSig Cell Line Persistence Curve performance vs. empirical null (Fig. 3d) 429 epithelial-based cell lines; 1000 random signatures not stated
Chi-square test (×16 genes) Mutation status of 16 DNA damage response genes vs. top/bottom half of CisSig score (Supplementary Fig. 4) 429 epithelial-based cell lines not stated
Simple linear regression, elastic net (L1+L2), L1 (Lasso), and L2 (Ridge) penalized linear regression; simple logistic regression and penalized logistic regression variants; random forest (implied from Table 2 structure) Prediction of cisplatin IC50 as continuous or binary outcome in GDSC epithelial cell lines (Table 2) 429 (all cell lines) or quintile subset (~172, implied) per model variant not stated
Approaches that could also have been used
  • Seed gene selection required agreement across all three DE methods (limma, SAM, multtest intersection) as an informal strategy to reduce false discovery
    Could also: A single DE method (e.g., limma or DESeq2) with a formal Benjamini-Hochberg FDR threshold (e.g., FDR < 0.05) combined with a minimum fold-change filter could also select high-confidence DE genes — A named FDR procedure makes the type-I error rate explicit, reproducible, and interpretable; pairing it with an effect-size filter guards against statistically significant but biologically negligible differences that can arise with n ~ 344
  • DE analysis used only the top and bottom 20% of IC50 responders, discarding the middle 60% of cell lines
    Could also: DE analysis with IC50 as a continuous covariate across all 429 cell lines (e.g., limma linear trend, or per-gene Spearman correlation with IC50) could also identify genes whose expression tracks cisplatin response — Using all available data retains the full response spectrum and avoids sensitivity to the choice of the 20% threshold, which can materially change which genes are selected as seeds
  • CisSig score was defined as the unweighted median normalized expression of the 19 signature genes
    Could also: A weighted sum using trained regression coefficients (e.g., from the elastic-net model in Table 2) or the first principal component of the 19-gene expression matrix could also serve as a single summary score — Weighted or PC-based scores account for differential gene contributions and inter-gene correlations; the median is straightforward, robust to individual outlier genes, and easy to compute in clinical settings
  • Signature performance was assessed by dichotomizing cell lines at quintile boundaries (top/bottom 20%) for the Wilcoxon test and Cell Line Persistence Curve
    Could also: A Spearman correlation between continuous CisSig score and continuous IC50 across all 429 cell lines, or a Cox model with CisSig score as a continuous predictor in the persistence-curve framework, could also characterize the association — Continuous analyses use all available data, avoid threshold-dependence, and yield a single interpretable effect size (e.g., correlation coefficient or hazard ratio per unit score increase) over the full cell line spectrum
  • Multiple prediction models (Table 2) appear to be evaluated on cell lines that overlap with the derivation set, with model performance reported without a fully held-out external test partition described in the provided text
    Could also: Nested cross-validation (inner loop for hyperparameter tuning, outer loop for unbiased performance estimation) or a pre-specified held-out test set would also estimate generalization performance — Evaluating models on data used during signature derivation or on non-independent partitions can yield optimistic performance estimates; nested CV or a pre-specified holdout partition provides estimates that better approximate expected performance on new, unseen data
  • Sixteen chi-square tests were conducted to assess mutation-status associations with CisSig score, with multiple-testing correction applied but the method not named
    Could also: Logistic regression with continuous CisSig score as the predictor (rather than dichotomized) and mutation status as the binary outcome, corrected via Benjamini-Hochberg FDR, could also test these associations; Fisher's exact test is preferable when expected cell counts are small — Logistic regression preserves the full information in the continuous score; naming the correction method (e.g., Benjamini-Hochberg) makes the analysis reproducible; Fisher's exact test avoids the chi-square approximation when any expected count falls below 5
Software: R/limma · SAM (samr R package or standalone SAM software) · R/multtest · R/sigQC

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
15
Impact: medium
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

E-MTAB-3610 ArrayExpress in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-37076665 (cissig)

Paper: Scarborough et al., Exploiting convergent phenotypes to derive a pan-cancer cisplatin response gene expression signature. NPJ Precis Oncol 2023. DOI 10.1038/s41698-023-00375-y · PMC10115855. Repo: https://github.com/jessicascarborough/cissig (authors' own code; commit c8317dc0a6ab6b8ee66dc98f2e817014f1a3c2a5, 2023-03-14). This IS the authors' pipeline, so P16 (third-party tool) does not apply — we run their code on their data.

The pipeline (what produces the central result)

Signature extraction = download_data.Rmdclean_data.Rmdextract_cissig.Rmd. Deterministic by design:

  • 5-fold partition create_test_groups(..., seed = 0)
  • SAM (samr::SAM, random.seed = 1, nperms = 10000)
  • limma eBayes (deterministic)
  • multtest MTP (seed = 1, B = 1000 bootstrap)
  • per-fold seed genes = intersection(SAM ∩ limma ∩ multtest), up-regulated in sensitive
  • co-expression on TCGA: Spearman affinity → top-5% membership → connectivity → top-20% connected seeds per fold
  • final cissig = genes appearing in ≥3 of the 5 folds Params (fixed in code & paper): deCompPerc=0.20, rmExtremePerc=0, samPerm=10000, multPerm=1000, nFolds=5.

Inputs (all fetched on «our HPC» compute node, kept on «infra»):

  • GDSC RMA basal expression (Cell_line_RMA_proc_basalExp.txt, cancerrxgene.org)
  • GDSC2 fitted dose-response xlsx + Cell_Lines_Details.xlsx (Sanger FTP)
  • TCGA RSEM-normalised RNA-seq for 22 epithelial sites via RTCGAToolbox/Firehose (runDate 20160128)

IN SCOPE (pipeline-derived; attempted)

  • C1 — cissig 19-gene signature (Table 1; shipped extract_cissig.html). CORE result.
  • C2 — number of epithelial GDSC cisplatin cell lines (Methods: 429; code plot: 430).

OUT OF SCOPE / not attempted (the hard ~20%)

  • Predictive-modelling AUCs (Table 2, model_sig_gdsc.Rmd) — many ML methods, resampling; depends on C1 + heavy tuning. Optional, not attempted first.
  • GSE48276 / GSE70691 clinical KM curves & Cox models (Fig 6) — plot/p-value heavy, depends on the extracted signature. Not attempted (lower-value, plot-based).
  • TCC analysis (Data/TCC/...), gene-set/enrichment, wet-lab — out of scope.

Reproduction strategy

Re-run the authors' signature-extraction pipeline on «our HPC» (one SLURM job, conda R/Bioconductor env built in-job). Compare regenerated gene list to Table 1 / the shipped html, and the epithelial cell-line count to the Methods/code value. Note alias: paper Table 1 CDIN1 = code output C15orf41 (same gene, renamed).

Figures / tables: Table
C1
Reported
cissig = 19 genes (Table 1): ADAT2, ATP1B3, CDIN1, C1QBP, CDC7, CDCA7, FKBP14, KRT5, LRRC8C, LY6K, MMP10, NPM3, PSAT1, RIOK1, SLFN11, STOML2, USP31, WDR3, ZNF750
Reproduced
NOT regenerated (pipeline blocked at GDSC RMA expression download, HTTP 410 Gone). Reported value verified internally: paper Table 1 == authors' shipped extract_cissig.html, 19/19 (CDIN1 = renamed C15orf41).
partial
C2
Reported
429 epithelial GDSC cisplatin cell lines (Methods); 430 in authors' clean_data.html plot title
Reproduced
NOT regenerated (same 410 Gone blocker)
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 50/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟢2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟡5. Derivability / plausibility
🟢6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

177.4 k
tokens (I/O) · 14.2 M incl. cache
51 min
runtime · 0.02 CPU-h
2.5 GB
peak RAM
1 (1 failed)
HPC jobs
hummel
machine