Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Identification of succinylation-related genes in bladder cancer: integration of single-cell and transcriptomic data.

Front Immunol · 2026
L1 78/100 PQI 93
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +7
✓ What held up
  • Any deviation was negligible
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🟡Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
78/100
Reproducibility score
at the mean
vs. all fields · 1173 studies
🎯 Scores higher than 51% of all assessed papers rank 533 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough to reproduce the low-hanging, public-data parts 1:1, but NOT the headline model. CLEAR 1:1 (C2): the 4 succinylation-related genes the paper reports as OS-associated on TCGA-BLCA reproduce within-tol on UCSC-Xena TCGA-BLCA via univariate Cox -- SIRT6 0.578->0.587, SIRT7 0.696->0.726, OXCT1 1.163->1.140, SUCLA2 1.369->1.449; all same direction, all p<0.01, CIs overlap (our n=424 vs paper's 404 explains ~2-6% HR deltas). PINNED TOOL (C1, P16): TIDEpy installed and run on the pinned GSE13507 data -- 256 samples cleanly reduced to exactly 165 primary tumors (matches paper), TIDE mean 0.071, responder 50.9%; graded partial because the paper prints no GSE13507 TIDE number to match (it shows TIDE only on TCGA risk groups). NOT ATTEMPTED: the 3-gene risk model (KCTD16/GSDMB/CD3D) and its GSE13507 split HRG=73/LRG=92 -- the LASSO/Cox beta coefficients are absent from text+supplements and the cutoff's applicability to GSE13507 is unstated, so the split is not deterministically reproducible (docs_insufficient, the under-specified ~20%, not chased). Also out of 80/20 scope: scRNA-seq GSE135337, CIBERSORT/ESTIMATE/GSEA/maftools/pRRophetic, nomogram/time-ROC, in-house Soochow DEGs, and all wet-lab. No fabrication concern: the checked values are derivable from the shipped public data.

💻 Code ↗ 🗄 Data: GSE13507

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 78
    assessed: 2026-06-16 ⛓ 57ca35be82f0
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-16
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-16
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

Although succinylation is linked to tumor progression, its role in bladder cancer (BLCA) remains understudied; this study aims to identify and validate prognostic succinylation-related genes (SRGs) in BLCA and elucidate their impact on the tumor microenvironment by integrating single-cell and transcriptomic data.

Core claims
  • KCTD16, CD3D and GSDMB are succinylation-related prognostic genes in BLCA finding
  • A risk model combining the risk score of these genes with age and N stage robustly predicts BLCA outcomes (AUC > 0.7) resource
  • The high-risk group displays enhanced immune evasion with higher TIDE score, reflecting an 'inflamed yet dysfunctional' TME state finding
  • Single-cell analysis identifies epithelial cells as key subpopulations, with additional involvement of T cells and fibroblasts finding
  • Scissor+ cells correlate with the high-risk phenotype and exhibit pseudotime-dependent expression patterns finding
  • Knockdown of GSDMB and KCTD16 significantly promotes T24 cell proliferation, supporting their tumor-suppressive roles mechanism
  • Integration of bulk transcriptomic and scRNA-seq data via the Scissor algorithm maps bulk risk signatures to single-cell resolution method
  • GSDMB is upregulated while KCTD16 and CD3D are downregulated in BLCA finding
Experimental setups
Assay System Perturbation Readout Platform
bulk RNA-seq / transcriptomic sequencing (differential expression) 15 paired BLCA and adjacent normal tissues (Soochow-BLCA cohort, The Fourth Affiliated Hospital of Soochow University) none differentially expressed genes (DESeq2)
bulk RNA-seq (prognostic modeling / survival analysis) TCGA-BLCA, 404 of 416 patient tissues none risk score, survival, prognostic gene expression
bulk RNA-seq (external validation) GEO GSE13507, 165 primary BLCA samples none risk stratification / survival validation
single-cell RNA-seq (scRNA-seq) GSE135337, 7 primary BLCA tissues and 1 adjacent normal tissue none cell-type clustering, prognostic gene expression, Scissor/pseudotime/cell communication Seurat v5.1.0
RT-qPCR BLCA tissue / cells (in vitro validation) none mRNA expression of GSDMB, KCTD16, CD3D
CCK-8 proliferation assay T24 bladder cancer cell line siRNA knockdown of GSDMB and KCTD16 cell proliferation
TIDE immune escape analysis TCGA-BLCA samples none TIDE score by risk group TIDEpy
in silico drug sensitivity analysis TCGA-BLCA samples none IC50 values of drugs between risk groups pRRophetic v0.5 / GDSC database
Key results
  • Risk model incorporating risk score, age and N stage showed robust predictive accuracy AUC > 0.7
  • High-risk group displayed enhanced immune evasion with higher TIDE score
  • RT-qPCR showed significant upregulation of GSDMB in BLCA
  • RT-qPCR showed significant downregulation of KCTD16 and CD3D in BLCA
  • Knockdown of GSDMB and KCTD16 significantly promoted T24 cell proliferation
  • Prognostic gene expression differed significantly between NMIBC and MIBC subtypes
  • Scissor+ cells correlated with high-risk phenotype and showed pseudotime-dependent expression
Key statistics
  • other AUC > 0.7 (risk model predictive accuracy (time-dependent ROC))
  • pvalue p < 0.001 (higher TIDE score in high-risk group (immune evasion))
  • pvalue p < 0.05 (RT-qPCR differential expression of GSDMB, KCTD16, CD3D in BLCA)
  • count 404 (TCGA-BLCA samples retained for prognostic analysis (from 416))
  • count 15 paired (Soochow-BLCA paired tumor/normal tissues for DEG screening)
  • count 165 (GSE13507 primary BLCA samples for external validation)
  • count 20 SRGs (succinylation-related genes compiled from literature)
  • other ~75% NMIBC / 25% MIBC (diagnostic distribution of BLCA subtypes)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This study used a multi-stage computational pipeline to identify succinylation-related prognostic genes in bladder cancer. DESeq2 differential expression and ssGSEA pathway scoring were used to nominate candidate genes, which were narrowed by univariable Cox regression, LASSO, and multivariable Cox regression (with explicit proportional-hazards testing) to yield a weighted risk score. Performance was assessed with Kaplan-Meier/log-rank tests and time-dependent ROC curves in a training cohort (TCGA-BLCA, n=404) and an external validation cohort (GSE13507, n=165). Functional, immune, mutational, and drug-sensitivity characterisation used Spearman correlation, Wilcoxon tests, GSEA, ESTIMATE, and TIDE, while single-cell analyses applied UMAP clustering, SingleR annotation, Scissor bulk-to-single-cell integration, CellChat, and pseudotime analysis.

Replicationmixed Sample sizeTCGA-BLCA training n=404 (after excluding 12 samples with incomplete data from 416); external validation GSE13507 n=165; Soochow-BLCA n=15 paired samples; scRNA-seq GSE135337 n=8 samples (7 BLCA + 1 normal); no formal power calculation described for any cohort GroupsBLCA vs adjacent normal (Soochow); high vs low ssGSEA score; high-risk vs low-risk group; NMIBC vs MIBC; high vs low TMB; multiple clinical-variable strata Pairingmixed Randomization/blindingnot stated Dispersionunclear Exact p-valuesno Effect sizesyes Confidence intervalsyes Multiplicity correctionBenjamini-Hochberg FDR applied in DESeq2 (adjusted p<0.05 threshold) and in clusterProfiler GSEA (adjusted p<0.05); no correction stated for multiple Wilcoxon tests, multiple Spearman correlations, or Cox screening across many genes
Statistical tests used
Test Applied to n Assumptions
DESeq2 Wald test (negative-binomial model) Differential expression between BLCA and adjacent normal tissue in the Soochow-BLCA cohort 15 BLCA + 15 adjacent normal (paired samples) not stated
Log-rank test (Kaplan-Meier) Overall survival comparison between high- and low-ssGSEA-score groups; high- vs low-risk groups in TCGA-BLCA training and GSE13507 validation 404 (TCGA training); 165 (GSE13507 validation) not stated
Univariable Cox proportional hazards regression Screening of candidate prognostic genes (p<0.01) and clinical variables (p<0.05) for OS association; PH assumption tested (p>0.05 required for retention) 404 (TCGA-BLCA) stated
LASSO regression (L1-penalised Cox model, 10-fold cross-validation) Dimensionality reduction and multicollinearity elimination among univariably significant prognostic candidates 404 (TCGA-BLCA) not stated
Multivariable Cox proportional hazards regression with stepwise selection Identification of independent prognostic genes and clinical factors; PH assumption tested for all retained variables 404 (TCGA-BLCA) stated
Wilcoxon rank-sum test Risk scores vs clinical variables (age, gender, TNM stage); TME/ESTIMATE/TIDE scores across risk groups; TMB between risk groups; IC50 values between risk groups; differential expression of prognostic genes across cell types in scRNA-seq 404 (TCGA bulk); cell counts not stated (scRNA-seq) not stated
Spearman rank correlation Prognostic genes vs 20 known SRGs (linkET); prognostic genes vs all TCGA genes (pre-ranking for GSEA); TMB vs risk score 404 (TCGA-BLCA) not stated
Gene Set Enrichment Analysis (GSEA, pre-ranked by Spearman correlation) Biological pathway enrichment per prognostic gene using KEGG gene sets (MSigDB c2.cp.kegg.v7.4) 404 (TCGA-BLCA) not stated
Time-dependent ROC curve / AUC at 1, 2, and 3 years Predictive accuracy of risk score model and nomogram in TCGA-BLCA training set and GSE13507 validation set 404 (training); 165 (validation) not stated
Approaches that could also have been used
  • The optimal cutpoint for stratifying patients into high- and low-risk (and high- and low-ssGSEA-score) groups was determined data-adaptively using the survminer package in the training cohort
    Could also: A pre-specified cutpoint (e.g., median, clinical quartiles) or retaining the risk score as a continuous variable in a Cox model could also be used — Data-adaptive optimal cutpoints maximise the apparent survival difference in the training data, which can inflate reported separation; a pre-specified or continuous approach avoids this additional analytical degree of freedom and is more directly transferable to independent cohorts
  • Multiple Wilcoxon rank-sum tests were performed across several related families of comparisons (clinical variables, TME scores, TIDE scores, TMB, drug IC50 values, scRNA-seq cell types) without stated correction for multiplicity
    Could also: Benjamini-Hochberg FDR correction applied within each family of related Wilcoxon tests could also be reported — Explicitly controlling the expected false-discovery rate within each comparison family is a standard practice when many simultaneous tests are conducted; it makes the nominal significance level more interpretable alongside the uncorrected results
  • The 15 paired BLCA and adjacent normal samples from the Soochow-BLCA cohort were analysed with DESeq2, but the available text does not describe whether the paired (within-patient) structure was modelled in the DESeq2 design formula
    Could also: Including patient ID as a blocking factor in the DESeq2 design formula (~patient + condition) explicitly models within-patient correlation — Accounting for the paired structure reduces residual variance, can increase statistical power to detect true differential expression, and more accurately reflects the study design
  • LASSO (L1 penalty only) was applied for variable selection among prognostic candidate genes
    Could also: Elastic net regularisation (mixing L1 and L2 penalties, alpha between 0 and 1) could also be applied — Elastic net can handle groups of correlated predictors more stably than LASSO, which tends to select one variable arbitrarily from a correlated cluster; this may be relevant when co-expressed candidate genes are considered together
  • Predictive discrimination was assessed using time-dependent AUC at three fixed time points (1, 2, 3 years)
    Could also: Harrell's concordance index (C-index) integrated over the full follow-up period could also be reported alongside the time-point AUCs — The C-index provides a global summary of discrimination across all observed event times rather than at pre-specified landmarks, and is widely reported in prognostic model papers as a complementary measure
  • Internal model performance was evaluated on the same TCGA-BLCA dataset used for model building, with external validation provided by GSE13507
    Could also: Bootstrap-based internal validation (e.g., 500–1,000 resamples) within TCGA-BLCA to estimate an optimism-corrected C-index or AUC could also be performed — Bootstrap optimism correction quantifies overfitting within the training data independently of external validation, providing a more conservative estimate of expected model performance in new samples and complementing the external validation result
Software: R/DESeq2 1.42.0 · R/ggplot2 3.4.1 · R/ComplexHeatmap 2.14.0 · R/GSVA 1.42.0 · R/survminer 0.4.9 · R/ggvenn 0.1.9 · R/clusterProfiler 4.2.2 · Cytoscape 3.10.3 · R/survival 3.7.0 · R/glmnet 4.1.8 · R/timeROC 0.4 · R/rms 6.9.0 · R/psych 2.1.6 · R/GseaVis 0.1.0 · R/maftools 2.18.0 · R/pRRophetic 0.5 · R/Seurat 5.1.0 · R/SingleR 2.4.0 · R/Scissor 2.0.0 · R/CellChat 1.6.1 · R/linkET 0.0.7.4 · R/SCP 0.5.6 · TIDE platform (TIDEpy) · ESTIMATE algorithm · STRING database · CellMarker database

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
0
Impact: low
Foundation confidence
Built on 1 assessed reference(s) · mean reproducibility 38/100
built largely on non-reproducible work
Topics

Assessed papers, coloured by verdict. Click a node to open it.

Built on (assessed references) (1)
Cited by (assessed papers) (0)
  • No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

GSE13507 GEO in Results (http://purl.org/orb/Results)
also used by 1 paper:
GSE135337 GEO in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet

Downstream reach in the literature

100 downstream papers · 1 datasets

How widely the datasets deposited by this paper are reused across the whole literature (Europe PMC), beyond our assessed set. This is a factual dependency map — reusing a public dataset is normal, good science. It is not a judgement on the downstream papers; the only verdict here is this paper's own, with its cited rationale.

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-42220482

Paper: Identification of succinylation-related genes in bladder cancer: integration of single-cell and transcriptomic data. Front Immunol 2026; DOI 10.3389/fimmu.2026.1797389. Pinned by brief: Code = TIDEpy (github.com/jingxinfu/TIDEpy) · Data = GEO GSE13507.

The paper is a multi-tool integrative bioinformatics study (TCGA-BLCA training, GSE13507 external validation, GSE135337 scRNA-seq, plus an in-house Soochow cohort and wet-lab RT-qPCR/CCK-8). Below: which reported results are pipeline-derived and in scope vs out.

In scope (attempted)

id result pipeline data feasibility
C1 TIDE immune-escape scoring (Fig 5) — the pinned tool TIDEpy (TIDE) GSE13507 (pinned data) RUN. P16: third-party tool on the paper's own data. Paper applies TIDE to TCGA-BLCA risk groups and prints no GSE13507 TIDE number, so this is a tool-runs-clean demonstration → expected grade partial.
C2 4 SRGs significantly associated with OS (Results / Suppl. Cox) — SIRT6 HR 0.578, SIRT7 0.696, OXCT1 1.163, SUCLA2 1.369 univariate Cox (R survival) TCGA-BLCA (public, UCSC Xena) RUN. Deterministic given expression+OS; directly checkable 1:1 (HR + 95% CI).

Out of scope / not attempted (with reason)

  • 3-gene risk model (KCTD16/GSDMB/CD3D) + GSE13507 split HRG=73/LRG=92 (Fig 3). docs_insufficient for a 1:1: the LASSO/Cox β coefficients are NOT reported in text or supplementary tables, and the paper does not state whether the TCGA-derived cutoff −2.002354 was re-applied to GSE13507 or re-fit. Without coefficients the split is not deterministically reproducible. This is the under-specified ~20% — not chased.
  • Full LASSO selection (24→3 genes), nomogram, time-ROC AUC>0.7 — non-deterministic LASSO seed + unreported coefficients. Out.
  • DEGs (3,769) in Soochow-BLCA cohort — in-house data not deposited. Out.
  • scRNA-seq (GSE135337): 39,899 cells, 13 clusters, Scissor 222+/267− — heavy, many under-specified params (resolution, marker sets). Not low-hanging. Out (80/20).
  • CIBERSORT/ESTIMATE/GSEA/maftools/pRRophetic panels — secondary, multi-step, parameter-sensitive. Out (80/20).
  • RT-qPCR, CCK-8 proliferation (Fig 9) — wet-lab. Out of scope by definition.

Honest framing

The brief pins a third-party tool (TIDEpy) — P16 makes running it on the paper's data a valid reproduction. We do that (C1) and add one clean deterministic 1:1 check on public TCGA data (C2). We explicitly do not claim to reproduce the headline risk model, because its coefficients are not shipped.

Figures / tables: Fig 5
C1
Reported
TIDE applied to BLCA (Fig 5); no GSE13507 TIDE value printed
Reproduced
TIDEpy ran clean on 165 GSE13507 primary tumors; TIDE mean 0.071, median -0.019; responder 84/165 = 50.9%
partial
C2a
Reported
SIRT6 univariate Cox HR=0.578 (0.426-0.786)
Reproduced
HR=0.587 (0.448-0.768), p=1.0e-4, n=424 (TCGA-BLCA)
within tolerance
C2b
Reported
SIRT7 univariate Cox HR=0.696 (0.527-0.919)
Reproduced
HR=0.726 (0.589-0.895), p=2.7e-3, n=424
within tolerance
C2c
Reported
OXCT1 univariate Cox HR=1.163 (1.036-1.306)
Reproduced
HR=1.140 (1.047-1.242), p=2.5e-3, n=424
within tolerance
C2d
Reported
SUCLA2 univariate Cox HR=1.369 (1.026-1.872)
Reproduced
HR=1.449 (1.164-1.804), p=9.1e-4, n=424
within tolerance

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 78/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟡2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟡5. Derivability / plausibility
🟢6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +7

The supporting claim — four succinylation-related genes (SIRT6, SIRT7, OXCT1, SUCLA2) associated with OS on TCGA-BLCA — reproduces cleanly within tolerance (HRs match direction, all p<0.01, overlapping CIs); the small ~2-6% deltas are explained by our n=424 vs the paper's 404 sample filtering. No fabrication concern: these values are derivable from the shipped public data. However, the paper's headline 3-gene risk model (KCTD16/GSDMB/CD3D) and its GSE13507 split could not be reproduced because the LASSO/Cox β coefficients and cutoff are absent from text and supplements — an authors-side documentation gap. Net: a solid, fabrication-free partial reproduction whose central prognostic claim remains unverified.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

113.8 k
tokens (I/O) · 5.2 M incl. cache
12 min
runtime · 0.02 CPU-h
1.9 GB
peak RAM
1
HPC jobs
hummel
machine