Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Machine learning developed an intratumor heterogeneity signature for predicting clinical outcome and immunotherapy benefit in bladder cancer.

Transl Androl Urol · 2024
L1 40/100 3/4
⚑ Flagged for review — a reproduced result did not match the reported value

Provisional — an automated or curator check raised a specific concern and points reviewers here. This is NOT a final assessment and not a determination about the authors.

Why this verdict

The main result did not reproduce in this reproduction attempt. Where our recomputation produced values that differ from the published ones, those discrepancies are listed below. This is a single automated attempt — not peer review and not a finding of error or misconduct — and differences can also arise from data access, undocumented parameters or the computing environment. The verdict can be contested via “report an error”.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Main result did not reproduce
Decisive
From: Q5 · Derivability / plausibility 🔴
Main result did not reproduce
Decisive
From: Q7 · Core claim 🔴
Main result did not reproduce
Decisive
From: Q8 · Severity of the miss (overall human judgment) 🔴
✓ What held up
  • Reported values were directly comparable
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🔴A deviation arose in the data or preprocessing
  • 🔴A deviation was attributed to the published material
  • 🔴Reported values were not (fully) derivable from the shared data
  • 🔴The deviation was non-trivial in magnitude
  • 🔴The central claim did not (fully) hold under reproduction
  • 🔴Overall, the reproduction showed a material discrepancy
How its reproducibility compares
40/100
Reproducibility score
1.9 SD below mean
vs. all fields · 1173 studies
🎯 Scores higher than 4% of all assessed papers rank 1126 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Pipeline-derived results reproduced via the paper's PUBLISHED locked 17-gene Enet(a=0.2) IRS (formula verified VERBATIM, grade exact) applied to independently-obtained public data on «our HPC» (R 4.3.3; «job»). All 5 in-scope cohorts attempted (TCGA training + 4 GEO validation), 17/17 genes mapped in each. RESULT = MISMATCH overall but nuanced: the locked model reproduces the reported discrimination in only 1/5 cohorts (GSE32894 C 0.713 ~ reported 0.77-0.78), partially in GSE13507 (1-yr only), and FAILS in the TRAINING cohort TCGA-BLCA (locked C 0.547, AUC 0.56/0.53/0.51 vs reported 0.744/0.791/0.816) plus GSE31684 and GSE48276. STRONGEST ANOMALY: reported TCGA TRAINING AUCs exceed the in-sample refit ceiling of these exact 17 genes (C 0.625) -- a model cannot beat its own apparent ceiling on its training data. The reported 0.69 avg C-index matches the REFIT/apparent ceiling (mean 0.729), not locked external validation (mean 0.575). The genes DO carry real prognostic signal (refit C 0.62-0.88; GSE32894 reproduces cleanly), so this is NOT wholesale fabrication -- most likely undisclosed preprocessing / a different undeposited TCGA matrix and/or reporting apparent vs locked performance. Fabrication flag MODERATE, PROVISIONAL: needs human audit + the authors' (undeposited) exact cohort matrices. Datasets profiled: 5 (4 GEO open arrays + TCGA-BLCA RNA-seq); GSE48276's GPL14951 annotation is missing on GEO (mapped via shared ILMN ids). Large intermediates kept on «infra» (~3.4 GB); only small results copied here.

💻 Code ↗ 🗄 Data: GSE31684

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 43
    assessed: 2026-06-19 ⛓ 779f00d001c8
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-22
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

Intratumor heterogeneity (ITH) drives bladder cancer progression, metastasis, and treatment resistance, so genes correlated with ITH can be used to build a prognostic signature that predicts clinical outcome and immunotherapy benefit in bladder cancer.

Core claims
  • An integrative machine learning procedure (10 methods, 101 algorithm combinations) identified an Enet (alpha=0.2)-based intratumor heterogeneity-related signature (IRS) with the highest average C-index (0.69) across TCGA and GEO cohorts method
  • IRS is an independent risk factor for overall survival in bladder cancer by univariate and multivariate Cox analysis in TCGA and all GEO datasets finding
  • IRS outperforms clinical parameters (age, gender, grade, stage) and 52 previously published prognostic signatures in predicting bladder cancer outcome (higher C-index) finding
  • Low IRS score is associated with higher immune-activated cell infiltration, cytolytic activity, and T cell co-stimulation finding
  • Low IRS score correlates with lower TIDE score, lower immune escape score, higher PD-1/CTLA4 immunophenoscore, higher TMB, and better immunotherapy response across IMvigor210, GSE91061, and GSE78220 cohorts finding
  • High IRS score is associated with higher IC50 values (lower drug sensitivity) for common chemotherapy and targeted therapy agents finding
  • 20 candidate ITH-related genes (17 retained in final IRS formula) were identified as potential prognostic biomarkers via univariate Cox analysis of DEGs between high/low ITH score groups resource
  • DEPTH2 algorithm was used to calculate ITH scores from TCGA bladder cancer transcriptomic data method
Experimental setups
Assay System Perturbation Readout Platform
bulk mRNA transcriptomic profiling + DEPTH2 ITH scoring bladder cancer, TCGA cohort (n=396) none ITH score, differentially expressed genes DEPTH2 algorithm; limma package
integrative machine learning signature construction (10 methods, 101 combinations, LOOCV) bladder cancer, TCGA + GEO (GSE13507, GSE31684, GSE32984, GSE48276) none risk score (IRS), C-index R (Enet, survminer, survivalROC)
immune infiltration deconvolution (TIMER, xCell, MCP-counter, CIBERSORT, CIBERSORT-ABS, EPIC, quanTIseq, ssGSEA, ESTIMATE) bladder cancer, TCGA cohort none immune cell abundance, immune/stromal/ESTIMATE scores R packages (GSVA, ESTIMATE, ggpubr)
immunotherapy response scoring (TIDE, immunophenoscore, TMB) bladder cancer, TCGA cohort none TIDE score, immunophenoscore, TMB score TIDE website; TCIA website
immunotherapy clinical cohort survival/response analysis bladder cancer, IMvigor210 (anti-PD-L1, n=298) anti-PD-1/PD-L1 therapy response rate, overall survival
immunotherapy clinical cohort survival/response analysis skin cutaneous melanoma, GSE91061 (n=98) anti-PD-1 therapy response rate, overall survival
immunotherapy clinical cohort survival/response analysis skin cutaneous melanoma, GSE78220 (n=28) anti-PD-1/CTLA4 therapy response rate, overall survival
drug sensitivity prediction (IC50 estimation) bladder cancer, TCGA cohort in silico chemotherapy/targeted therapy drug panel predicted IC50 values oncoPredict R package; Genomics of Drug Sensitivity in Cancer (GDSC)
Key results
  • Enet (alpha=0.2)-based IRS had the highest average C-index among 101 model combinations C-index=0.69
  • IRS predicted OS in TCGA with high AUC at 1, 3, 5 years AUC=0.744/0.791/0.816
  • High ITH score associated with worse overall survival P<0.001
  • Low IRS score correlated with higher immune-activated cell infiltration (CD8+ T cells, NK cells, macrophages M1) and immune function scores
  • Low IRS score associated with better predicted immunotherapy benefit (lower TIDE/immune escape, higher immunophenoscore/TMB, higher response rate)
  • High IRS score associated with higher IC50 (lower sensitivity) to chemotherapy and targeted therapy drugs
  • IRS had higher C-index than clinical parameters and 52 published prognostic signatures
  • In IMvigor210, responders had lower IRS score than non-responders; high IRS score had worse OS and lower response rate P=0.006 (OS); P<0.01 (response rate)
Key statistics
  • count 396 (TCGA bladder cancer cohort size)
  • fold_change |Log2 FC| >1.5 (cutoff for DEGs between high/low ITH score groups)
  • count 625 (number of DEGs identified between high and low ITH score groups)
  • other C-index=0.69 (average C-index of optimal Enet (alpha=0.2) IRS model)
  • other AUC=0.744, 0.791, 0.816 (1-, 3-, 5-year ROC AUC for OS prediction by IRS in TCGA)
  • pvalue P<0.001 (OS difference between high and low ITH score groups)
  • pvalue P=0.006 (OS difference between high and low IRS score in IMvigor210 cohort)
  • count 17 (number of genes in final Enet-based IRS formula)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This is a retrospective bioinformatics/machine-learning study using public bulk transcriptomic cohorts (TCGA and several GEO series) to derive an intratumor-heterogeneity-related signature (IRS) in bladder cancer. Differentially expressed genes were identified with limma, prognostic genes were selected via univariate Cox regression, and an integrative pipeline of 10 machine-learning algorithms (101 combinations) selected an optimal model by average C-index across cohorts. Downstream comparisons between IRS-defined groups (survival, immune infiltration, immunotherapy response, drug sensitivity) were tested with Student's t-test, ANOVA, chi-square/Fisher's exact test, and Cox regression, with results reported largely via C-index, AUC, hazard ratios, and significance thresholds.

Replicationbiological Sample sizePer-cohort sample sizes are stated (TCGA n=396; GSE13507 n=165; GSE31684 n=90; GSE32984 n=223; GSE48276 n=73; IMvigor210 n=298; GSE91061 n=98; GSE78220 n=28); no a priori power calculation described GroupsLow vs high ITH score; low vs high IRS (risk score) group; immunotherapy responders vs non-responders Pairingunpaired Randomization/blindingna Dispersionunclear Exact p-valuesno Effect sizesyes Confidence intervalsyes Multiplicity correctionnone stated
Statistical tests used
Test Applied to n Assumptions
Unpaired Student's t-test Comparisons of continuous scores (e.g., immune cell scores, checkpoint expression, IC50 values) between high vs low IRS/ITH score groups Per-cohort n as listed (e.g., TCGA n=396; GSE13507 n=165; GSE31684 n=90; GSE32984 n=223; GSE48276 n=73; IMvigor210 n=298; GSE91061 n=98; GSE78220 n=28) not stated
One-way ANOVA Comparisons across more than two groups (e.g., clinical stage categories) 'as appropriate' Per-cohort n as listed above not stated
Chi-square test / Fisher's exact test Categorical clinical characteristic comparisons 'as appropriate' not stated
Univariate and multivariate Cox proportional hazards regression Identifying independent risk factors and IRS's independent prognostic value (Figure 3C,3D) TCGA and all GEO cohorts (per-cohort n as listed) not stated
Kaplan-Meier survival analysis with reported P values (e.g., P<0.001, P=0.006) OS comparison between high/low ITH or IRS score groups (Figure 1B, Figure 2B-2F, Figure 5G-5I) Per-cohort n as listed above not stated
Gene set enrichment analysis (GSEA) / ssGSEA Functional pathway and immune-function enrichment differences between IRS groups (Figures 4, 7) TCGA cohort na
Approaches that could also have been used
  • Many pairwise comparisons between high vs low IRS groups (immune cells, checkpoints, functional scores) across multiple figure panels were each tested with Student's t-test without a stated multiplicity adjustment.
    Could also: Applying a false discovery rate procedure (e.g., Benjamini-Hochberg) across the family of comparisons within a panel or analysis. — This would explicitly control the expected proportion of false positives when many related endpoints are examined together.
  • Continuous score comparisons between two groups used the unpaired Student's t-test.
    Could also: A non-parametric alternative such as the Mann-Whitney U test. — This does not require a normality assumption and can be a useful complement for immune/expression scores that may be skewed, without needing to verify distributional assumptions.
  • Differentially expressed genes were defined using a fold-change threshold together with a raw P value cutoff (P<0.05).
    Could also: Using limma's adjusted (FDR/Benjamini-Hochberg) P values as the significance criterion. — This accounts for the large number of genes tested simultaneously in genome-wide expression comparisons, which is a standard practice for this type of analysis.
  • The optimal high/low IRS cutoff was determined using a data-driven cutpoint procedure (surv_cutpoint).
    Could also: A pre-specified split such as the median value, or modeling IRS as a continuous variable in Cox regression. — A pre-specified or continuous approach avoids the potential for cutpoint optimization to inflate apparent group separation and can ease comparison with other studies using different cutoffs.
  • Survival differences between groups were summarized with a Kaplan-Meier plot and a P value, without the underlying test being explicitly named.
    Could also: Explicitly reporting the log-rank test (or a Cox-model likelihood-ratio test) alongside the hazard ratio and its confidence interval. — Naming the specific test and pairing it with an effect size (HR with CI) gives readers both the significance and the magnitude/precision of the survival difference.
  • Model discriminative performance was primarily summarized as a single C-index value per cohort.
    Could also: Reporting the C-index together with a confidence interval (e.g., via bootstrap resampling) and calibration statistics. — This conveys the precision of the discrimination estimate rather than a point estimate alone, which can be useful when comparing performance across many cohorts and competing signatures.
Software: R 3.5.0 · limma (R package) · survminer (surv_cutpoint) · survivalROC (R package) · GSVA (R package, ssGSEA) · oncoPredict (R package)

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Figures / tables: Figure 2AFigure 2BFigure 2DFigure 2CFigure 2EFigure 2F
IRS_formula
Reported
17-gene Enet(a=0.2) risk-score formula + coefficients
Reproduced
verified verbatim vs PMC BioC full text; 17/17 genes map in every cohort
exact
AUC_TCGA
Reported
AUC 0.744/0.791/0.816 (TRAINING cohort)
Reproduced
locked model AUC 0.556/0.525/0.511, C-index 0.547 (refit ceiling 0.625)
did not match
AUC_GSE31684
Reported
AUC 0.677/0.690/0.697
Reproduced
locked model AUC 0.540/0.569/0.576, C-index 0.529 (refit ceiling 0.686)
did not match
AUC_GSE13507
Reported
AUC 0.679/0.680/0.705
Reproduced
locked model AUC 0.672/0.574/0.616, C-index 0.559 (refit 0.725); 1-yr matches only
partial
AUC_GSE32894
Reported
AUC 0.772/0.789/0.783
Reproduced
locked model AUC 0.694/0.766/0.759, C-index 0.713 (refit 0.880); largely reproduces
partial
AUC_GSE48276
Reported
AUC NA/0.655/0.720
Reproduced
locked model AUC 0.485/0.522/0.542, C-index 0.497 (refit 0.625)
did not match
Cindex_avg
Reported
average C-index 0.69 (Enet a=0.2, optimal)
Reproduced
mean locked C 0.575 vs mean refit-ceiling C 0.729 across 4 GEO cohorts; 0.69 reproducible only as apparent/refit
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 40/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟢2. Endpoint comparability
🔴3. Location of the main deviation
🔴4. Cause of the deviation
🔴5. Derivability / plausibility
🔴6. Severity of the deviation
🔴7. Core claim
🔴8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Main result did not reproduce
Decisive
From: Q5 · Derivability / plausibility 🔴
Main result did not reproduce
Decisive
From: Q7 · Core claim 🔴
Main result did not reproduce
Decisive
From: Q8 · Severity of the miss (overall human judgment) 🔴

The paper publishes its model in full and the 17-gene Enet(a=0.2) formula was verified verbatim, yet applying that locked model to standard public TCGA-BLCA and GSE31684 data reproduces none of the reported discrimination — training-cohort AUC 0.744/0.791/0.816 collapses to 0.556/0.525/0.511 (C 0.547) and validation C-index 0.69 to ~0.53, robust across six probe-collapse/scaling variants. Decisively, the reported AUCs exceed even the in-sample refit ceiling on the authors' own training cohort (TCGA C 0.625), so the values are not derivable from the shared model + public data and cannot be blamed on our preprocessing. The genes do carry ~0.69 signal when refit, but the locked coefficients do not transfer; this is an authors'-side over-optimism/fabrication concern, severity severe, central claim not confirmed.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

527.8 k
tokens (I/O) · 37 M incl. cache
107 min
runtime · 0.03 CPU-h
1.7 GB
peak RAM
3
HPC jobs
hummel
machine