Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Stage-stratified molecular profiling of non-muscle-invasive bladder cancer enhances biological, clinical, and therapeutic insight.

Cell Rep Med · 2021
L1 75/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
  • Reported values are derivable from the shared data
What did not (or only partly)
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
75/100
Reproducibility score
at the mean
vs. all fields · 1173 studies
🎯 Scores higher than 45% of all assessed papers rank 612 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

DESCRIBED WELL ENOUGH for the expression arm, partial 1:1 on public data. The harvested code github.com/seandavi/ngCGH is a harvesting FALSE POSITIVE: the paper explicitly reports no original code, and ngCGH (pseudo-CGH from NGS tumour/normal BAMs) is incompatible with the only public data (GSE163209 = Affymetrix HTA 2.0 expression CELs); the NGS/copy-number data ngCGH would need is at EGA controlled-access (EGAS00001005765/766/767, not obtainable). Instead we reproduced the paper's clearly-specified expression NORMALIZATION on the 217 public CELs with a standard third-party tool (Bioconductor oligo + pd.hta.2.0 RMA) and compared value-for-value to the authors' deposited GSE163209 matrix. Result: same feature namespace (67528 transcript clusters) and full sample set (217) recovered, values positively correlated (per-sample Pearson median 0.84, Spearman 0.79) but NOT value-identical (median|diff| ~1.0 log2). The gap is a summarization-method difference (oligo 'core' meta-probesets vs authors' APT HTA-2_0.r3 Psrs.mps; feature counts 70523 vs 67528), not a data discrepancy; no fabrication signal (deposited values are plain, internally-consistent log2 RMA, trend-reproducible from shipped CELs). NOT ATTEMPTED (hard 20% / out of scope): exact APT value match (needs Thermo HTA-2_0.r3 library files + Affymetrix Power Tools); NMF expression subtypes E1-E4 / TaE1-3 / T1 (need supp sample-class tables + subjective rank choice); UROMOL2021 & LundTax assignment, RTN regulons; copy-number subtypes CN1-CN4 and mutation profiles (derive from EGA controlled-access sequencing). Engineering notes: fixed bioconda pd.hta.2.0 R4.5 Rd-install failure (source install --no-help) and preprocessCore 'pthread_create code 22' on cgroup-limited SLURM (rebuilt with --disable-threading).

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 75
    assessed: 2026-06-15 ⛓ dde751ddb1bf
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-15
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

Because stage Ta and T1 non-muscle-invasive bladder cancers (NMIBCs) show distinct clinical behavior, the authors hypothesized that analyzing each tumor stage separately, rather than all NMIBCs combined, would yield deeper biological understanding and more clinically meaningful (particularly prognostic) molecular subclassification.

Core claims
  • Stage-stratified molecular subclassification of Ta and T1 tumors provides greater biological understanding and more clinically meaningful information than subtypes derived from all NMIBCs combined. finding
  • Expression and CN subtypes derived from all NMIBCs are largely driven by tumor stage and provide no additional prognostic information beyond stage for T1 tumors. finding
  • Both Ta and T1 tumors contain immune-infiltrated subtypes, with the infiltrated Ta subtype (TaE3) showing improved recurrence-free survival linked to an anti-tumor immune response. finding
  • T1 subtypes show differential prevalence of cisplatin sensitivity-related (DDR/ERCC2) mutations, suggesting subtypes that may benefit from chemo- or immunotherapy. finding
  • Mutations affecting enhancer activation (COMPASS-like complex components, CREBBP/EP300) are highly prevalent in Ta tumors. mechanism
  • This study provides the largest whole-exome sequence dataset for T1 bladder tumors. resource
  • Two-stage non-negative matrix factorization (NMF) and regulon analysis were used to discover and characterize transcriptional subtypes. method
  • APOBEC mutational signatures (SBS2/SBS13) are a dominant mutational process in NMIBC, more pronounced in T1 than Ta tumors. mechanism
Experimental setups
Assay System Perturbation Readout Platform
DNA copy number analysis (genome-wide) 113 stage Ta and 104 high-grade stage T1 primary bladder tumors (human) none copy number gains/losses, fraction of genome altered (FGA), CN clusters
Genome-wide mRNA expression profiling 113 stage Ta and 104 stage T1 bladder tumors (human) none gene expression subtypes, signature scores, regulon activity
Whole-exome sequencing 58 stage T1 bladder tumors with paired blood (human) none somatic SNVs, mutational signatures, driver genes, TMB mean 87× coverage, 89% of bases >30×
Targeted sequencing Ta tumors and T1 tumors not analyzed by WES, with paired blood (human) none mutation frequencies, TMB, SBS signatures
Computational immune deconvolution (ESTIMATE) stage Ta bladder tumors (human) none immune score, immune cell type infiltration, PD-L1/CD274 expression
Key results
  • Combined NMIBC expression and CN subtypes were strongly associated with tumor stage but provided no additional prognostic information; no PFS differences between subtypes within T1 tumors alone.
  • LundTax and UROMOL2021 classifications showed no significant relationship to PFS in stage T1 samples only. p = 0.46 and p = 0.67
  • The infiltrated Ta subtype TaE3 had improved RFS, with elevated ESTIMATE immune scores and PD-L1 expression and higher cytolytic (GZMA/PRF1) activity.
  • FGFR3 mutations were the most common alteration in Ta tumors. 62%
  • Mutations predicted to affect enhancer activation were present in 73% of Ta samples (65% COMPASS-like component, 34% CREBBP/EP300). 73%
  • FGFR3 and KMT2D mutations were associated with higher tumor mutational burden in Ta tumors.
  • In T1 tumors, more SNVs were present in the chromosomally stable subtype T1CN1, and medium/high TMB associated with better PFS.
  • APOBEC signatures (SBS2/SBS13) dominated 79% of T1 samples; APOBEC3A/3B expression higher in T1 than Ta. 79%
Key statistics
  • count 113 stage Ta and 104 high-grade stage T1 tumors (fresh-frozen tumors with paired blood analyzed)
  • count 49,477 somatic SNVs (mean 868, median 450 per sample) (whole-exome sequencing of T1 samples)
  • other 7.9 mutations per megabase (median 7.55) (mean non-synonymous mutation frequency in Ta tumors by targeted sequencing)
  • pvalue p = 0.0016 (TSC1 mutations more common in GS2 Ta tumors)
  • pvalue p < 0.0001 (interferon signaling highest in TaE3; T cell and macrophage infiltration in TaE3)
  • other 35% of SNVs (29% of tumors ≥50%, 61% ≥25%) (APOBEC SBS2/SBS13 contribution in Ta tumors)
  • count 23 genes with predicted driver function (dNdScv analysis of T1 whole-exome data)
  • other 55 months median follow-up; recurrence in 45% Ta and 47% T1; 1 Ta and 16 T1 progressed to MIBC/metastatic (clinical follow-up (107 Ta, 88 T1 individuals))

Statistical methods review

Model: opus

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This is an observational multi-omics cohort study of 113 stage Ta and 104 stage T1 primary bladder tumors, integrating copy-number, mutation, and transcriptome profiling. Molecular subtypes were derived by unsupervised clustering (copy number) and non-negative matrix factorization (expression), and associations were tested with categorical tests (chi-square, Fisher's exact), nonparametric group comparisons (Mann-Whitney, Kruskal-Wallis with Dunn's), and survival analysis (log-rank for RFS/PFS). Results are presented largely as subtype-stratified comparisons with significance thresholds and box-style summaries (median, quartiles, min/max).

Replicationbiological Sample size113 Ta and 104 T1 tumors with paired blood; whole-exome for 58 T1 (mean 87x coverage), targeted sequencing for others; no formal power/sample-size calculation stated Groupsmolecular subtypes (CN and expression) within all NMIBC and stratified by stage Ta vs T1; mutation-status and immune-score groups Pairingunpaired Randomization/blindingna Dispersionmixed Exact p-valuesyes Effect sizesno Confidence intervalsno Multiplicity correctionBonferroni correction for chi-square tests; Dunn's multiple-comparison correction following Kruskal-Wallis
Statistical tests used
Test Applied to n Assumptions
Chi-square test (with Bonferroni correction) relationships among stage, CN subtype, expression subtype, and DDR-gene mutation distribution (Figures 1B–1E, 3E, 3H) not stated
Fisher's exact test distribution of ERCC2 mutations across T1 CN subtypes (Figure 3F) na
Mann-Whitney U test TMB by FGFR3/KMT2D mutation in Ta (Figure 2D); APOBEC3A/3B expression Ta vs T1 (Figure 3C); TMB by DDR-gene mutation in T1 (Figure 3I) na
Kruskal-Wallis test with Dunn's multiple-comparison correction 12-gene progression risk score and CIS score across subtypes (Figures 1I, 1J); immune score and CD274/PD-L1 across Ta subtypes (Figure 2H); TMB across T1 CN subtypes (Figure 3B) na
Log-rank test PFS and RFS by CN subtype and expression subtype, by PIK3CA and TP53 mutation, and by immune score (Figures 1G, 1H, 2E, 2G, 2I, 3G, S2E, S2F) 107 Ta and 88 T1 with follow-up (median 55 months) na
Differential expression across subtypes (threshold p < 0.0001 noted) gene expression signatures/heatmaps (Figure 1F) and pairwise subtype comparisons not stated
Approaches that could also have been used
  • Group spread in box-style plots is reported as median with 25th/75th percentiles and min/max, while some figures show mean and SD.
    Could also: A consistent dispersion convention across all figures (e.g., median with IQR for the nonparametric comparisons, plus a 95% confidence interval) could also be used. — A single, uniform convention matched to the test family helps readers compare panels directly and conveys uncertainty around the central estimate alongside raw spread.
  • Subtype-versus-survival relationships are assessed with the log-rank test.
    Could also: A Cox proportional-hazards model could also be applied. — A Cox model would additionally provide hazard ratios with confidence intervals and allow adjustment for covariates such as stage, grade, and treatment, quantifying effect magnitude beyond a single p-value.
  • Several pairwise group comparisons of continuous measures (e.g., TMB, immune scores) use Mann-Whitney or Kruskal-Wallis with Dunn's correction and are reported with p-values.
    Could also: Reporting an accompanying effect size (e.g., rank-biserial correlation, Cliff's delta, or median difference with a 95% CI) could also be included. — Effect sizes communicate the practical magnitude of differences, which is especially informative when subgroup sample sizes vary.
  • Categorical associations among subtypes and mutations use chi-square tests with Bonferroni correction.
    Could also: Fisher's exact test (already used elsewhere in the paper) or a Benjamini-Hochberg FDR could also be applied across this family of contingency tests. — Exact tests are well suited to sparse contingency tables with low expected counts, and an FDR approach can offer greater power than Bonferroni when many associations are screened.
  • Multiplicity is handled per analysis with Bonferroni and Dunn's corrections.
    Could also: A single pre-specified study-wide multiplicity framework (e.g., Benjamini-Hochberg FDR across all reported subtype comparisons) could also be reported. — A unified correction scope makes the overall false-discovery profile of the many comparisons explicit in one place.
  • Immune-score survival analysis dichotomizes the cohort into top and bottom quartiles.
    Could also: Modeling the immune score as a continuous variable could also be used. — Retaining the full continuous range preserves statistical power and avoids dependence on a specific cut-point, while still allowing visualization of thresholds.
Software: dNdScv (driver-gene identification) · NMF (non-negative matrix factorization) for transcriptional subtyping · ESTIMATE (immune/stromal scoring) · COSMIC SBS / mutational signature analysis

Result convergence & founder nodes

Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
59
Impact: high
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

1x10 in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
also used by 2 papers:
RRID:SCR_001876 RRID in Article (http://semanticscience.org/resource/SIO_001029)
also used by 1 paper:
RRID:SCR_002105 in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
also used by 1 paper:
RRID:SCR_010910 RRID in Article (http://semanticscience.org/resource/SIO_001029)
also used by 1 paper:
RRID:SCR_014583 RRID in Article (http://semanticscience.org/resource/SIO_001029)
also used by 1 paper:
5IVW PDBe in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
A21777 in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
C57387 in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
E-MTAB-4321 in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
EGAS00001005765 EGA in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
EGAS00001005766 EGA in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
EGAS00001005767 EGA in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
RRID:SCR_000559 RRID in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet
RRID:SCR_000595 RRID in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet
RRID:SCR_002260 RRID in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet
RRID:SCR_003199 RRID in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet
RRID:SCR_003201 RRID in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet
RRID:SCR_005109 in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
RRID:SCR_006525 RRID in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet
RRID:SCR_006791 in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
RRID:SCR_006849 RRID in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet
RRID:SCR_011841 RRID in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet
RRID:SCR_017093 RRID in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet
RRID:SCR_018718 RRID in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-35028613

Paper: Hurst et al., Stage-stratified molecular profiling of non-muscle-invasive bladder cancer enhances biological, clinical, and therapeutic insight. Cell Rep Med 2021;2(12):100472. PMID 35028613 · PMCID PMC8714941 · DOI 10.1016/j.xcrm.2021.100472.

Data & code availability (verbatim from the paper)

  • Microarray (expression): GEO GSE163209public. Affymetrix HTA 2.0, 217 NMIBC (113 stage Ta + 104 stage T1). Deposited processed matrix is transcript-cluster-level log2 RMA (IDs TC*.hg.1; libs HTA-2_0.r3.pgf + HTA-2_0.r3.Psrs.mps).
  • Raw sequencing (copy number + mutations): EGA EGAS00001005765 / EGAS00001005766 / EGAS00001005767controlled access (not obtainable).
  • Code: "This paper does not report original code." (explicit statement.)

On the harvested code link github.com/seandavi/ngCGH

  • ngCGH = "tools for producing pseudo-CGH of next-generation sequencing data"; it consumes tumour/normal NGS BAM pairs to emit copy-number/CGH log2 ratios (last commit 2016, Python 2, no license).
  • It is NOT the authors' code (paper reports none) and is incompatible with the only public data (GSE163209 is expression CEL arrays, not NGS BAMs).
  • The NGS data ngCGH would need is the EGA copy-number/exome data, which is controlled access. → The harvested code↔data pair is a text-mining false positive; running ngCGH as-assigned is not feasible on public data.

In scope (pipeline-derived, reproducible from PUBLIC data)

The paper's expression pipeline is precisely specified and uses a named third-party tool (Affymetrix Power Tools apt-probeset-summarize, rma). Per brief P16, applying an equivalent third-party tool to the paper's own data is equally valid.

  • R1 (PRIMARY): Expression normalization. Re-run RMA on the 217 public CELs with Bioconductor oligo + pd.hta.2.0 (target='core' → same TC*.hg.1 IDs) and compare value-for-value to the authors' deposited GSE163209 normalized matrix. → quantitative 1:1 concordance (per-sample correlation, |diff|).
  • R2 (structural): cohort = 217 NMIBC = 113 Ta + 104 T1 on HTA 2.0 (GPL17586).

Out of scope (not attempted, with reason)

  • Copy-number subtypes (CN1–CN4), mutational profiles — derive from EGA controlled-access sequencing data → data_restricted, not obtainable.
  • NMF expression subtypes (E1–E4, Ta:3, T1:4), UROMOL2021 / LundTax class assignment, RTN regulons — downstream of R1; rank/label assignment needs the paper's supplementary sample↔class tables and subjective NMF rank selection (the hard ~20%). Noted, not attempted; R1 reproduces the upstream matrix they all build on.
Figures / tables: Fig 1Fig S3EFig 1A
R1
Reported
GSE163209 deposited transcript-cluster log2 RMA matrix (Affymetrix HTA 2.0, 217 NMIBC; APT apt-probeset-summarize rma, HTA-2_0.r3.Psrs.mps)
Reproduced
Independent oligo RMA (target=core) on same 217 public CELs: 67528 common transcript clusters, all 217 samples; per-sample Pearson median 0.840 (min 0.637), Spearman 0.786, overall r 0.836, median|diff| 1.00 log2
partial
R2
Reported
217 NMIBC = 113 stage Ta + 104 stage T1 on Affymetrix HTA 2.0 (GPL17586)
Reproduced
217 = 113 Ta + 104 T1 confirmed from GEO sample characteristics + processed 217 CELs
exact

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 75/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟢5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

202.2 k
tokens (I/O) · 17 M incl. cache
47 min
runtime · 0.32 CPU-h
33.5 GB
peak RAM
3 (2 failed)
HPC jobs
hummel
machine