Identification of succinylation-related genes in bladder cancer: integration of single-cell and transcriptomic data.
The main results reproduced: recomputed values matched the published ones within tolerance.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Any deviation was negligible
- 🟡Could not use the authors’ exact input data
- 🟡Reported values were only indirectly comparable
- 🟡A deviation arose in the data or preprocessing
- 🟡A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Described well enough to reproduce the low-hanging, public-data parts 1:1, but NOT the headline model. CLEAR 1:1 (C2): the 4 succinylation-related genes the paper reports as OS-associated on TCGA-BLCA reproduce within-tol on UCSC-Xena TCGA-BLCA via univariate Cox -- SIRT6 0.578->0.587, SIRT7 0.696->0.726, OXCT1 1.163->1.140, SUCLA2 1.369->1.449; all same direction, all p<0.01, CIs overlap (our n=424 vs paper's 404 explains ~2-6% HR deltas). PINNED TOOL (C1, P16): TIDEpy installed and run on the pinned GSE13507 data -- 256 samples cleanly reduced to exactly 165 primary tumors (matches paper), TIDE mean 0.071, responder 50.9%; graded partial because the paper prints no GSE13507 TIDE number to match (it shows TIDE only on TCGA risk groups). NOT ATTEMPTED: the 3-gene risk model (KCTD16/GSDMB/CD3D) and its GSE13507 split HRG=73/LRG=92 -- the LASSO/Cox beta coefficients are absent from text+supplements and the cutoff's applicability to GSE13507 is unstated, so the split is not deterministically reproducible (docs_insufficient, the under-specified ~20%, not chased). Also out of 80/20 scope: scRNA-seq GSE135337, CIBERSORT/ESTIMATE/GSEA/maftools/pRRophetic, nomogram/time-ROC, in-house Soochow DEGs, and all wet-lab. No fabrication concern: the checked values are derivable from the shipped public data.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 78assessed: 2026-06-16 ⛓ 57ca35be82f0
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-16
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-16no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusAlthough succinylation is linked to tumor progression, its role in bladder cancer (BLCA) remains understudied; this study aims to identify and validate prognostic succinylation-related genes (SRGs) in BLCA and elucidate their impact on the tumor microenvironment by integrating single-cell and transcriptomic data.
- ★ KCTD16, CD3D and GSDMB are succinylation-related prognostic genes in BLCA finding
- ★ A risk model combining the risk score of these genes with age and N stage robustly predicts BLCA outcomes (AUC > 0.7) resource
- ★ The high-risk group displays enhanced immune evasion with higher TIDE score, reflecting an 'inflamed yet dysfunctional' TME state finding
- ★ Single-cell analysis identifies epithelial cells as key subpopulations, with additional involvement of T cells and fibroblasts finding
- ★ Scissor+ cells correlate with the high-risk phenotype and exhibit pseudotime-dependent expression patterns finding
- ★ Knockdown of GSDMB and KCTD16 significantly promotes T24 cell proliferation, supporting their tumor-suppressive roles mechanism
- Integration of bulk transcriptomic and scRNA-seq data via the Scissor algorithm maps bulk risk signatures to single-cell resolution method
- ★ GSDMB is upregulated while KCTD16 and CD3D are downregulated in BLCA finding
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| bulk RNA-seq / transcriptomic sequencing (differential expression) | 15 paired BLCA and adjacent normal tissues (Soochow-BLCA cohort, The Fourth Affiliated Hospital of Soochow University) | none | differentially expressed genes (DESeq2) | — |
| bulk RNA-seq (prognostic modeling / survival analysis) | TCGA-BLCA, 404 of 416 patient tissues | none | risk score, survival, prognostic gene expression | — |
| bulk RNA-seq (external validation) | GEO GSE13507, 165 primary BLCA samples | none | risk stratification / survival validation | — |
| single-cell RNA-seq (scRNA-seq) | GSE135337, 7 primary BLCA tissues and 1 adjacent normal tissue | none | cell-type clustering, prognostic gene expression, Scissor/pseudotime/cell communication | Seurat v5.1.0 |
| RT-qPCR | BLCA tissue / cells (in vitro validation) | none | mRNA expression of GSDMB, KCTD16, CD3D | — |
| CCK-8 proliferation assay | T24 bladder cancer cell line | siRNA knockdown of GSDMB and KCTD16 | cell proliferation | — |
| TIDE immune escape analysis | TCGA-BLCA samples | none | TIDE score by risk group | TIDEpy |
| in silico drug sensitivity analysis | TCGA-BLCA samples | none | IC50 values of drugs between risk groups | pRRophetic v0.5 / GDSC database |
- – Risk model incorporating risk score, age and N stage showed robust predictive accuracy AUC > 0.7
- ▲ High-risk group displayed enhanced immune evasion with higher TIDE score
- ▲ RT-qPCR showed significant upregulation of GSDMB in BLCA
- ▼ RT-qPCR showed significant downregulation of KCTD16 and CD3D in BLCA
- ▲ Knockdown of GSDMB and KCTD16 significantly promoted T24 cell proliferation
- – Prognostic gene expression differed significantly between NMIBC and MIBC subtypes
- – Scissor+ cells correlated with high-risk phenotype and showed pseudotime-dependent expression
- other AUC > 0.7 (risk model predictive accuracy (time-dependent ROC))
- pvalue p < 0.001 (higher TIDE score in high-risk group (immune evasion))
- pvalue p < 0.05 (RT-qPCR differential expression of GSDMB, KCTD16, CD3D in BLCA)
- count 404 (TCGA-BLCA samples retained for prognostic analysis (from 416))
- count 15 paired (Soochow-BLCA paired tumor/normal tissues for DEG screening)
- count 165 (GSE13507 primary BLCA samples for external validation)
- count 20 SRGs (succinylation-related genes compiled from literature)
- other ~75% NMIBC / 25% MIBC (diagnostic distribution of BLCA subtypes)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This study used a multi-stage computational pipeline to identify succinylation-related prognostic genes in bladder cancer. DESeq2 differential expression and ssGSEA pathway scoring were used to nominate candidate genes, which were narrowed by univariable Cox regression, LASSO, and multivariable Cox regression (with explicit proportional-hazards testing) to yield a weighted risk score. Performance was assessed with Kaplan-Meier/log-rank tests and time-dependent ROC curves in a training cohort (TCGA-BLCA, n=404) and an external validation cohort (GSE13507, n=165). Functional, immune, mutational, and drug-sensitivity characterisation used Spearman correlation, Wilcoxon tests, GSEA, ESTIMATE, and TIDE, while single-cell analyses applied UMAP clustering, SingleR annotation, Scissor bulk-to-single-cell integration, CellChat, and pseudotime analysis.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| DESeq2 Wald test (negative-binomial model) | Differential expression between BLCA and adjacent normal tissue in the Soochow-BLCA cohort | 15 BLCA + 15 adjacent normal (paired samples) | not stated |
| Log-rank test (Kaplan-Meier) | Overall survival comparison between high- and low-ssGSEA-score groups; high- vs low-risk groups in TCGA-BLCA training and GSE13507 validation | 404 (TCGA training); 165 (GSE13507 validation) | not stated |
| Univariable Cox proportional hazards regression | Screening of candidate prognostic genes (p<0.01) and clinical variables (p<0.05) for OS association; PH assumption tested (p>0.05 required for retention) | 404 (TCGA-BLCA) | stated |
| LASSO regression (L1-penalised Cox model, 10-fold cross-validation) | Dimensionality reduction and multicollinearity elimination among univariably significant prognostic candidates | 404 (TCGA-BLCA) | not stated |
| Multivariable Cox proportional hazards regression with stepwise selection | Identification of independent prognostic genes and clinical factors; PH assumption tested for all retained variables | 404 (TCGA-BLCA) | stated |
| Wilcoxon rank-sum test | Risk scores vs clinical variables (age, gender, TNM stage); TME/ESTIMATE/TIDE scores across risk groups; TMB between risk groups; IC50 values between risk groups; differential expression of prognostic genes across cell types in scRNA-seq | 404 (TCGA bulk); cell counts not stated (scRNA-seq) | not stated |
| Spearman rank correlation | Prognostic genes vs 20 known SRGs (linkET); prognostic genes vs all TCGA genes (pre-ranking for GSEA); TMB vs risk score | 404 (TCGA-BLCA) | not stated |
| Gene Set Enrichment Analysis (GSEA, pre-ranked by Spearman correlation) | Biological pathway enrichment per prognostic gene using KEGG gene sets (MSigDB c2.cp.kegg.v7.4) | 404 (TCGA-BLCA) | not stated |
| Time-dependent ROC curve / AUC at 1, 2, and 3 years | Predictive accuracy of risk score model and nomogram in TCGA-BLCA training set and GSE13507 validation set | 404 (training); 165 (validation) | not stated |
-
The optimal cutpoint for stratifying patients into high- and low-risk (and high- and low-ssGSEA-score) groups was determined data-adaptively using the survminer package in the training cohort↳ Could also: A pre-specified cutpoint (e.g., median, clinical quartiles) or retaining the risk score as a continuous variable in a Cox model could also be used — Data-adaptive optimal cutpoints maximise the apparent survival difference in the training data, which can inflate reported separation; a pre-specified or continuous approach avoids this additional analytical degree of freedom and is more directly transferable to independent cohorts
-
Multiple Wilcoxon rank-sum tests were performed across several related families of comparisons (clinical variables, TME scores, TIDE scores, TMB, drug IC50 values, scRNA-seq cell types) without stated correction for multiplicity↳ Could also: Benjamini-Hochberg FDR correction applied within each family of related Wilcoxon tests could also be reported — Explicitly controlling the expected false-discovery rate within each comparison family is a standard practice when many simultaneous tests are conducted; it makes the nominal significance level more interpretable alongside the uncorrected results
-
The 15 paired BLCA and adjacent normal samples from the Soochow-BLCA cohort were analysed with DESeq2, but the available text does not describe whether the paired (within-patient) structure was modelled in the DESeq2 design formula↳ Could also: Including patient ID as a blocking factor in the DESeq2 design formula (~patient + condition) explicitly models within-patient correlation — Accounting for the paired structure reduces residual variance, can increase statistical power to detect true differential expression, and more accurately reflects the study design
-
LASSO (L1 penalty only) was applied for variable selection among prognostic candidate genes↳ Could also: Elastic net regularisation (mixing L1 and L2 penalties, alpha between 0 and 1) could also be applied — Elastic net can handle groups of correlated predictors more stably than LASSO, which tends to select one variable arbitrarily from a correlated cluster; this may be relevant when co-expressed candidate genes are considered together
-
Predictive discrimination was assessed using time-dependent AUC at three fixed time points (1, 2, 3 years)↳ Could also: Harrell's concordance index (C-index) integrated over the full follow-up period could also be reported alongside the time-point AUCs — The C-index provides a global summary of discrimination across all observed event times rather than at pre-specified landmarks, and is widely reported in prognostic model papers as a complementary measure
-
Internal model performance was evaluated on the same TCGA-BLCA dataset used for model building, with external validation provided by GSE13507↳ Could also: Bootstrap-based internal validation (e.g., 500–1,000 resamples) within TCGA-BLCA to estimate an optimism-corrected C-index or AUC could also be performed — Bootstrap optimism correction quantifies overfitting within the training data independently of external validation, providing a more conservative estimate of expected model performance in new samples and complementing the external validation result
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
Assessed papers, coloured by verdict. Click a node to open it.
- No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
Downstream reach in the literature
100 downstream papers · 1 datasetsHow widely the datasets deposited by this paper are reused across the whole literature (Europe PMC), beyond our assessed set. This is a factual dependency map — reusing a public dataset is normal, good science. It is not a judgement on the downstream papers; the only verdict here is this paper's own, with its cited rationale.
- Identification of distinct basal and luminal subtype... 2014 · 1,327 cites
- PrognoScan: a new database for meta-analysis of the... 2009 · 772 cites
- Predictive value of progression-related gene classif... 2010 · 310 cites
- Siglec15 shapes a non-inflamed tumor microenvironmen... 2021 · 277 cites
- Combination of a novel gene expression signature wit... 2012 · 174 cites
- An EMT-related gene signature for the prognosis of h... 2020 · 158 cites
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-42220482
Paper: Identification of succinylation-related genes in bladder cancer: integration of single-cell and transcriptomic data. Front Immunol 2026; DOI 10.3389/fimmu.2026.1797389. Pinned by brief: Code = TIDEpy (github.com/jingxinfu/TIDEpy) · Data = GEO GSE13507.
The paper is a multi-tool integrative bioinformatics study (TCGA-BLCA training, GSE13507 external validation, GSE135337 scRNA-seq, plus an in-house Soochow cohort and wet-lab RT-qPCR/CCK-8). Below: which reported results are pipeline-derived and in scope vs out.
In scope (attempted)
| id | result | pipeline | data | feasibility |
|---|---|---|---|---|
| C1 | TIDE immune-escape scoring (Fig 5) — the pinned tool | TIDEpy (TIDE) | GSE13507 (pinned data) | RUN. P16: third-party tool on the paper's own data. Paper applies TIDE to TCGA-BLCA risk groups and prints no GSE13507 TIDE number, so this is a tool-runs-clean demonstration → expected grade partial. |
| C2 | 4 SRGs significantly associated with OS (Results / Suppl. Cox) — SIRT6 HR 0.578, SIRT7 0.696, OXCT1 1.163, SUCLA2 1.369 | univariate Cox (R survival) |
TCGA-BLCA (public, UCSC Xena) | RUN. Deterministic given expression+OS; directly checkable 1:1 (HR + 95% CI). |
Out of scope / not attempted (with reason)
- 3-gene risk model (KCTD16/GSDMB/CD3D) + GSE13507 split HRG=73/LRG=92 (Fig 3).
docs_insufficientfor a 1:1: the LASSO/Cox β coefficients are NOT reported in text or supplementary tables, and the paper does not state whether the TCGA-derived cutoff −2.002354 was re-applied to GSE13507 or re-fit. Without coefficients the split is not deterministically reproducible. This is the under-specified ~20% — not chased. - Full LASSO selection (24→3 genes), nomogram, time-ROC AUC>0.7 — non-deterministic LASSO seed + unreported coefficients. Out.
- DEGs (3,769) in Soochow-BLCA cohort — in-house data not deposited. Out.
- scRNA-seq (GSE135337): 39,899 cells, 13 clusters, Scissor 222+/267− — heavy, many under-specified params (resolution, marker sets). Not low-hanging. Out (80/20).
- CIBERSORT/ESTIMATE/GSEA/maftools/pRRophetic panels — secondary, multi-step, parameter-sensitive. Out (80/20).
- RT-qPCR, CCK-8 proliferation (Fig 9) — wet-lab. Out of scope by definition.
Honest framing
The brief pins a third-party tool (TIDEpy) — P16 makes running it on the paper's data a valid reproduction. We do that (C1) and add one clean deterministic 1:1 check on public TCGA data (C2). We explicitly do not claim to reproduce the headline risk model, because its coefficients are not shipped.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
The supporting claim — four succinylation-related genes (SIRT6, SIRT7, OXCT1, SUCLA2) associated with OS on TCGA-BLCA — reproduces cleanly within tolerance (HRs match direction, all p<0.01, overlapping CIs); the small ~2-6% deltas are explained by our n=424 vs the paper's 404 sample filtering. No fabrication concern: these values are derivable from the shipped public data. However, the paper's headline 3-gene risk model (KCTD16/GSDMB/CD3D) and its GSE13507 split could not be reproduced because the LASSO/Cox β coefficients and cutoff are absent from text and supplements — an authors-side documentation gap. Net: a solid, fabrication-free partial reproduction whose central prognostic claim remains unverified.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.