Bacillus Calmette-Guérin Treatment Changes the Tumor Microenvironment of Non-Muscle-Invasive Bladder Cancer.
Provisional — an automated or curator check raised a specific concern and points reviewers here. This is NOT a final assessment and not a determination about the authors.
The main result did not reproduce in this reproduction attempt. Where our recomputation produced values that differ from the published ones, those discrepancies are listed below. This is a single automated attempt — not peer review and not a finding of error or misconduct — and differences can also arise from data access, undocumented parameters or the computing environment. The verdict can be contested via “report an error”.
- Nothing in this column.
- 🔴Could not use the authors’ exact input data
- 🔴Reported values were only indirectly comparable
- 🟡A deviation arose in the data or preprocessing
- 🔴A deviation was attributed to the published material
- 🔴Reported values were not (fully) derivable from the shared data
- 🟡The deviation was non-trivial in magnitude
- 🟡The central claim did not (fully) hold under reproduction
- 🔴Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
PARTIAL reproduction. Only the somatic-WES modality has deposited raw data (PRJNA574742 = 42 WXS runs; no RNA-seq, no TCR-seq), so the RNA-seq DEG (C4), CIBERSORT (C6) and TCR/migmap (C5) claims are data_unavailable=mismatch. A full GATK4 best-practices somatic pipeline (BWA-MEM b37 -> MarkDup -> BQSR -> Mutect2 matched T/N -> contamination -> FilterMutectCalls, Ensembl GRCh37.75 annotation, ~70x downsampled) ran on the 3 matched-normal patients P4/P10/P17 = 7 tumour timepoints. C1 (TMB): pipeline reproduced and per-tumour somatic burden obtained (exon PASS 671-6896; coding TMB mean 36/median 23 per Mb) but ~3-5x HIGHER than the reported 6.8/Mb and 10-1466 range -> partial (reproduced-but-higher; single-caller Mutect2 vs the paper's Mutect2-intersect-Strelka2, no PoN, dropped gnomAD). C2 (TP53 33%): TP53 somatic present in 6/7 tumours (driver confirmed) but cohort % not reproducible -> partial. C3: KRT4/TTN/AHNAK2 present, canonical urothelial drivers strongly recovered, but FDFT1 (paper's top gene, 53%) detected in 0/7 -> flagged as possible original artifact; partial. Contamination <1% for 6/7 (P17_s2 hypermutator 7.1%). Grades provisional; human reviewer decides. Not a clean 1:1 on absolute TMB; methodologically-explained over-call. Honest partial.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 26assessed: 2026-06-15 ⛓ c9e673627821
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-24
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetHow does BCG (Bacillus Calmette–Guérin) treatment affect the tumor microenvironment, genomic evolution, and immune landscape between primary and relapsed non-muscle-invasive bladder cancer (NMIBC) tumors?
- ★ BCG therapy has bidirectional effects on tumor evolution and immune checkpoint landscape, with a significant reduction in neoantigen burden percentage in relapsed tumors finding
- ★ A remarkable proportion of subclonal mutations are unique to matched pre- or post-treatment tumors, suggesting BCG-induced and/or spatial heterogeneity finding
- ★ Relapsed tumors show a shift in mutational signatures with enrichment of aristolochic acid (AA)-associated mutations, implying AA exposure may be associated with tumor recurrence finding
- ★ Immune checkpoint regulation gene expression is enhanced in relapsed tumors, suggesting combination of immune checkpoint inhibitors with BCG may be an effective treatment strategy finding
- ★ TCR sequencing reveals treatment-associated changes in the T-cell repertoire between primary and relapsed tumors finding
- AKAP13 mutations are significantly associated with increased risk of recurrence after BCG treatment finding
- Relapsed tumors have a significantly higher chromosome instability (CIN) score than primary tumors finding
- Whole-exome sequencing, RNA sequencing, and immune repertoire sequencing of time-serial matched tumors provide a systematic discovery approach for tracking genomic and immune dynamics under BCG treatment method
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Whole-exome sequencing (WES) | 36 TUR tumor samples from 24 NMIBC patients plus matched normal tissue | BCG treatment (pre- vs post-/relapsed comparison) | somatic mutations, tumor mutation burden, copy number variation, mutational signatures | — |
| RNA sequencing | 19 matched primary and relapsed NMIBC tumor samples | BCG treatment | differentially expressed genes, neoantigen expression validation, MHC-I score | — |
| Immune repertoire sequencing (TCR-seq) | NMIBC primary and relapsed tumor samples | BCG treatment | T-cell receptor repertoire changes | — |
| Sanger sequencing | NMIBC tumor samples | none (validation of WES calls) | TP53 mutation validation | — |
| PyClone phylogenetic/cellular prevalence analysis | Matched time-serial tumor samples from 4 patients with ≥3 tumor samples | BCG treatment over time | subclonal composition, intratumor heterogeneity, cellular prevalence of variant clusters | — |
| OncoNEM clonal phylogeny analysis | NMIBC tumor samples with ≥2 time points | BCG treatment over time | tumor evolution model, clonal origin/phylogeny | — |
| GISTIC2.0 copy number analysis | Primary vs relapsed NMIBC tumor samples | BCG treatment | chromosome arm/gene-level amplifications and deletions (e.g., CDKN2A, E2F3, PPARG, FGFR3) | — |
| MSIsensor microsatellite instability analysis | NMIBC tumor samples | none | microsatellite instability percentage/status | — |
- ▼ Percentage of neoantigen burden significantly decreased in relapsed tumors compared with primary tumors p=0.0078
- ▲ Relapsed tumors had a significantly higher chromosome instability (CIN) score than primary tumors p=0.0078
- ▲ Aristolochic acid (AA)-associated mutational signature was highly enriched in relapsed tumors 16.2% contribution, cosine-similarity 0.955
- ▲ APOBEC-associated mutation signature 13 (C>G via APOBEC1 in TpC context) was more prominent in relapsed tumors than primary 72.4%
- ▲ AKAP13 mutations associated with increased risk of recurrence after BCG treatment hazard ratio=7.6 (95% CI 1.9-29), p=0.0035
- – 159 differentially expressed genes identified between primary and relapsed tumors, predominantly upregulated and enriched for immune/infection-related pathways 156 up, 4 down (FDR<0.05, fold change>2)
- – No statistically significant change in overall tumor mutation burden between primary and relapsed tumors p=0.76
- ▲ Primary tumors that later recurred had higher tumor mutation burden than primary tumors without recurrence p=0.04
- pvalue p=0.0078 (Wilcoxon signed-rank test, neoantigen percentage decrease in relapsed vs primary tumors)
- pvalue p=0.0078 (Wilcoxon signed-rank test, CIN score higher in relapsed vs primary tumors)
- other hazard ratio=7.6, 95% CI 1.9-29, p=0.0035, adjusted p=0.059 (AKAP13 mutation association with recurrence risk)
- pvalue p=0.76 (Wilcoxon signed-rank test, TMB primary vs relapsed tumors, no significant difference)
- pvalue p=0.04 (Wilcoxon signed-rank test, TMB in primary tumors with vs without recurrence)
- count 159 DEGs (156 upregulated, 4 downregulated) (differential expression analysis between primary and relapsed tumors, FDR<0.05 and fold change>2)
- other cosine-similarity=0.955 (match of relapsed-tumor mutational signature to COSMIC aristolochic acid signature)
- mean 6.8 mutations per Mb (range 10-1,466 mutations) (average tumor mutation burden across the patient cohort)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This prospective longitudinal cohort study enrolled 24 NMIBC patients and analyzed 36 matched whole-exome sequencing samples, 19 RNA-seq samples, and paired immune repertoire sequences from primary and post-BCG relapsed tumors. Genomic and transcriptomic comparisons between primary and relapsed tumors were primarily conducted with nonparametric Wilcoxon tests and Fisher's exact tests, while differential gene expression was filtered by FDR < 0.05 and fold change > 2. Survival-associated mutations were evaluated via hazard ratio, and mutational signatures, copy number aberrations, and clonal architecture were characterized with specialized bioinformatic tools (Bayesian NMF, GISTIC2.0, PyClone, OncoNEM).
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Wilcoxon signed-rank test | CIN score comparison between relapsed and primary tumors (Figure S5) | 36 WES tumor samples from 24 patients (matched pairs) | not stated |
| Wilcoxon signed-rank test (labeled 'Wilcoxon test' in figure) | TMB comparison between primary (n=15) and relapsed (n=21) tumors (Figure 3A) | primary n=15, relapsed n=21 | not stated |
| Wilcoxon signed-rank test | TMB comparison between study NMIBC cohort and TCGA MIBC cohort (Figure 3A) | TCGA n=413 | not stated |
| Wilcoxon signed-rank test | TMB comparison between study NMIBC cohort and BGI MIBC cohort (Figure 3A) | BGI n=99 | not stated |
| Wilcoxon signed-rank test | TMB comparison between primary tumors with vs. without subsequent recurrence | not stated separately | not stated |
| Wilcoxon signed-rank test | Percentage of neoantigen burden comparison between relapsed and primary tumors (Figure 3B) | primary n=15, relapsed n=21 | not stated |
| Wilcoxon signed-rank test | MHC-I score expression difference between primary and relapsed tumors (Figure 6C) | n=19 RNA-seq samples | not stated |
| Fisher's exact test | APOBEC-mediated mutation enrichment in relapsed vs. primary tumors (Figure S6) | not stated | not stated |
| Cox proportional hazards (hazard ratio reported) | AKAP13 mutation association with recurrence risk after BCG (Table S6, Figure S3B) | n=24 patients | not stated |
| FDR-controlled differential expression analysis (specific tool not named); FDR < 0.05 and fold change > 2 | DEG analysis between primary and relapsed tumors (Figure 6A, Table S8) | n=19 RNA-seq samples | not stated |
| MutSigCV (internally corrected mutation significance scoring) | Identification of significantly mutated genes across all tumor samples (Table S5) | 36 WES samples | not stated |
| GISTIC2.0 (q-value-based copy number significance) | Chromosomal aberrations in relapsed vs. primary tumors (Figure S4) | 36 WES samples | not stated |
| Bayesian non-negative matrix factorization (NMF) with COSMIC cosine similarity matching | Mutational signature decomposition in primary and relapsed tumor sets separately (Figure 2) | primary n=15 (2 signatures), relapsed n=21 (3 signatures) | not stated |
| Gene set enrichment analysis (GSEA) with KEGG pathway database | Pathway enrichment of DEGs between primary and relapsed tumors (Figure 6B, Table S9) | n=19 RNA-seq samples | not stated |
-
Multiple pairwise Wilcoxon tests were applied across several genomic outcomes (TMB, neoantigen percentage, CIN score, MHC-I score) comparing primary vs. relapsed tumors without a stated family-wise correction across that set of tests.↳ Could also: A Bonferroni, Holm, or Benjamini-Hochberg correction applied across the suite of Wilcoxon pairwise comparisons could also be used to bound the probability of any false positive across outcomes. — When several related outcomes are each tested independently, a family-wise or FDR correction over the set makes the overall error rate explicit; this is especially informative for readers interpreting which individual comparisons to weight most.
-
Longitudinal (pre/post) comparisons were handled primarily with paired Wilcoxon tests, treating each time-point pair independently.↳ Could also: A linear mixed-effects model (or generalized mixed model for count outcomes) with patient as a random effect and time point as a fixed effect could also be used to jointly model the repeated-measures structure. — Mixed-effects models explicitly account for within-patient correlation across multiple time points and can jointly estimate the effect of BCG treatment while adjusting for covariates such as grade and stage, providing a unified framework for the longitudinal design.
-
Differential gene expression between primary and relapsed tumors was reported with FDR < 0.05 and fold change > 2 thresholds; the specific DE tool was not named.↳ Could also: DESeq2 or edgeR, standard tools for count-based RNA-seq with small and unbalanced sample sizes, could also be applied with the paired patient structure modeled as a blocking factor (e.g., ~patient + condition in DESeq2). — Explicitly modeling the matched patient structure in a count-based framework can reduce residual variance from inter-patient heterogeneity and increase sensitivity; naming the tool and version supports computational reproducibility.
-
Mutational signatures were decomposed by Bayesian NMF with the number of signatures (K) chosen separately as 2 for primary and 3 for relapsed sets; the selection criterion for K is not described.↳ Could also: SigProfilerExtractor or the R package MutationalPatterns, which evaluate a range of K values using cross-validation, reconstruction error, and silhouette stability metrics, could also be used. — Reporting a stability criterion for K selection makes the signature number choice transparent and allows readers to assess how robustly the decomposition generalizes; these tools also facilitate systematic comparison with the full COSMIC v3 reference catalogue.
-
APOBEC enrichment in relapsed vs. primary tumors was assessed with a single unpaired Fisher's exact test (p=0.29).↳ Could also: A paired McNemar's test using matched pre/post tumor pairs from the same patient, or a mixed-effects logistic regression with patient as a random effect, could also be used. — Because pre- and post-treatment samples from the same patient share genetic background, a paired or patient-stratified test accounts for within-patient correlation and may be more sensitive to a treatment-related shift in APOBEC activity.
-
Mutation counts were summarized with mean ± SD (e.g., 37.26 ± 15.7 unique mutations in primary tumors; 34.50 ± 34.46 in relapsed tumors).↳ Could also: Median and IQR or median and range could also summarize mutation count distributions. — Mutation count distributions are typically right-skewed; the relapsed group SD (34.46) nearly equals the mean (34.50), suggesting pronounced skew or outliers where the median and IQR would be more robust summary statistics.
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
Downstream reach in the literature
0 downstream papers · 1 datasetsHow widely the datasets deposited by this paper are reused across the whole literature (Europe PMC), beyond our assessed set. This is a factual dependency map — reusing a public dataset is normal, good science. It is not a judgement on the downstream papers; the only verdict here is this paper's own, with its cited rationale.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-35311085
Paper: Strandgaard T. et al. (mislabelled in registry; actual: Liu D., Abbosh P., et al.) "Bacillus Calmette-Guérin Treatment Changes the Tumor Microenvironment of Non-Muscle-Invasive Bladder Cancer." Front Oncol 2022;12:842182. PMID 35311085, PMCID PMC8930202, DOI 10.3389/fonc.2022.842182.
Registry code link: https://github.com/mikessh/migmap (MIGMAP — an IgBLAST wrapper for Rep-Seq/immune-repertoire reads, by mikessh). The paper itself does not cite migmap; it used IgBlast + MiRMAP for TCRβ CDR3 repertoire / Shannon diversity. migmap is a drop-in equivalent of that step (P16: third-party tool on the paper's data is valid) — but only if the TCR-seq raw reads are obtainable.
Registry data link: SRA BioProject PRJNA574742.
What the deposited data actually contains (verified via ENA, 2026-06-15)
PRJNA574742 = 42 runs, ALL WXS / GENOMIC (whole-exome), HiSeq X Ten.
There is NO RNA-seq and NO TCR-seq/AMPLICON data in the BioProject.
Sample naming: Pn_c = matched normal/control, Pn_sX = tumour timepoints.
Only 6 patients have a matched normal deposited (P4, P10, P11, P14, P15, P17);
the other ~18 patients are tumour-only (no paired normal).
In scope (pipeline-derived, data present)
| Result | Pipeline (paper) | Reproducible? |
|---|---|---|
| Somatic mutation burden / TMB (avg 6.8 mut/Mb, range 10–1466 mut) | Trimmomatic → BWA-MEM → Picard MarkDup → GATK BQSR → MuTect2 ∩ Strelka2 → ANNOVAR | YES, partial — only for the 6 matched-pair patients deposited |
| Presence of recurrent driver mutations (TP53, FGFR3, KRT4, FDFT1, TTN, AHNAK2 …) | same somatic pipeline | YES (presence), partial |
| Cohort-level gene frequencies ("FDFT1 53%", "TP53 33%" …) | same | NO — needs all 24 patients; only 6 have matched normals |
| Mutation signatures (APOBEC, AA), CNV (CNVkit/GISTIC), neoantigens (pVACseq) | downstream of somatic calls | out of 80/20 scope; not attempted |
Out of scope — data_unavailable (raw data NOT deposited)
- RNA-seq DEGs (159 DEGs) — STAR/StringTie2/edgeR — no RNA-seq in PRJNA574742.
- CIBERSORT / TIMER / quanTIseq immune deconvolution — needs RNA-seq.
- TCR repertoire / Shannon Diversity Index (p=0.013), CDR3 length peaks — the
IgBlast/MiRMAP (≈ migmap) result — no TCR-seq/AMPLICON reads deposited.
→ The registry's own code artefact (migmap) is therefore not runnable on any
deposited data; recorded as
data_unavailablefor the repertoire claim.
Out of scope — non-pipeline / wet-lab
Sanger validation, MiSeq TCR amplification protocol, clinical survival (KM/Cox on clinical metadata), follow-up months.
Reproduction target (80/20)
Run the paper's somatic-WES best-practice pipeline (BWA-MEM b37 → MarkDuplicates →
BQSR → Mutect2 matched tumour-normal → FilterMutectCalls → snpEff) on the
matched-pair patients P4, P10, P17 (3 normals + 7 tumours = 10 exomes), compute
per-tumour somatic mutation counts and TMB, and compare against the reported
average 6.8 mut/Mb and the 10–1466 absolute-count range, plus check for the
named driver genes. Documented substitutions: b37 (Homo_sapiens_assembly19) for
"UCSC hg19"; GATK4 Mutect2 for MuTect2 (GATK3.8); single-caller (Mutect2) instead
of Mutect2∩Strelka2 intersection (will over-count vs the paper's stricter
intersection — noted in grading); snpEff for ANNOVAR. P14/P11/P15 not run (compute
budget; same pipeline would apply).
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.