Corpus 1,286 assessed · 1,187 scored · 648 reproduced ≥75 · 174 flagged ·∅ 73.9/100
← New search

Identification of succinylation-related genes in bladder cancer: integration of single-cell and transcriptomic data.

Front Immunol · 2026
L1 78/100 PQI 93
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +7
✓ What held up
  • Any deviation was negligible
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🟡Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
78/100
Reproducibility score
at the mean
vs. all fields · 1187 studies
🎯 Scores higher than 52% of all assessed papers rank 534 of 1187 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough to reproduce the low-hanging, public-data parts 1:1, but NOT the headline model. CLEAR 1:1 (C2): the 4 succinylation-related genes the paper reports as OS-associated on TCGA-BLCA reproduce within-tol on UCSC-Xena TCGA-BLCA via univariate Cox -- SIRT6 0.578->0.587, SIRT7 0.696->0.726, OXCT1 1.163->1.140, SUCLA2 1.369->1.449; all same direction, all p<0.01, CIs overlap (our n=424 vs paper's 404 explains ~2-6% HR deltas). PINNED TOOL (C1, P16): TIDEpy installed and run on the pinned GSE13507 data -- 256 samples cleanly reduced to exactly 165 primary tumors (matches paper), TIDE mean 0.071, responder 50.9%; graded partial because the paper prints no GSE13507 TIDE number to match (it shows TIDE only on TCGA risk groups). NOT ATTEMPTED: the 3-gene risk model (KCTD16/GSDMB/CD3D) and its GSE13507 split HRG=73/LRG=92 -- the LASSO/Cox beta coefficients are absent from text+supplements and the cutoff's applicability to GSE13507 is unstated, so the split is not deterministically reproducible (docs_insufficient, the under-specified ~20%, not chased). Also out of 80/20 scope: scRNA-seq GSE135337, CIBERSORT/ESTIMATE/GSEA/maftools/pRRophetic, nomogram/time-ROC, in-house Soochow DEGs, and all wet-lab. No fabrication concern: the checked values are derivable from the shipped public data.

💻 Code ↗ 🗄 Data: GSE13507

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 78
    assessed: 2026-06-16 ⛓ 57ca35be82f0
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-16
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-16
no human curator yet
Last updated
2026-09-19

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The study tests whether succinylation-related genes (SRGs), identified by integrating single-cell and bulk transcriptomic data, have prognostic value in bladder cancer (BLCA) and are linked to tumor microenvironment (TME) remodeling and immune evasion.

Core claims
  • KCTD16, CD3D, and GSDMB were identified as prognostic succinylation-related genes in BLCA finding
  • A risk score model combining these three genes with age and N stage shows robust predictive accuracy (AUC > 0.7) finding
  • The high-risk group displays enhanced immune evasion, indicated by higher TIDE scores, and the TME appears 'inflamed yet dysfunctional' finding
  • Single-cell analysis identifies epithelial cells as the key subpopulation associated with prognostic risk, with additional involvement of T cells and fibroblasts finding
  • Scissor+ cells (mapped from high-risk bulk signature) correlate with the high-risk phenotype and show pseudotime-dependent expression patterns finding
  • RT-qPCR validation shows GSDMB is upregulated and KCTD16/CD3D are downregulated in BLCA tissue finding
  • Knockdown of GSDMB and KCTD16 significantly promotes T24 bladder cancer cell proliferation, supporting tumor-suppressive roles finding
  • The Scissor algorithm can map bulk transcriptome-derived risk signatures onto single-cell data to identify high-risk cell subpopulations method
Experimental setups
Assay System Perturbation Readout Platform
bulk RNA-seq (transcriptomic sequencing) Soochow-BLCA cohort: 15 paired BLCA and adjacent normal tissues none (tumor vs. adjacent normal) differentially expressed genes (DEGs)
bulk RNA-seq (TCGA transcriptomic data) TCGA-BLCA, 404 tumor samples with clinical data none SRG activity score (ssGSEA), survival outcomes, risk score
bulk RNA-seq (GEO transcriptomic data) GSE13507, 165 primary BLCA samples none external validation of risk stratification and survival
scRNA-seq GSE135337: 7 primary BLCA tissues and 1 adjacent normal tissue none cell type annotation, prognostic gene expression across cell types, NMIBC vs MIBC distribution
somatic mutation profiling TCGA-BLCA samples none mutation frequency, tumor mutational burden (TMB)
in silico drug sensitivity prediction TCGA-BLCA samples none predicted IC50 values for GDSC drugs GDSC database via pRRophetic
RT-qPCR BLCA tissue (human) none mRNA expression of KCTD16, CD3D, GSDMB
CCK-8 proliferation assay T24 bladder cancer cell line knockdown of GSDMB and KCTD16 cell proliferation
Key results
  • KCTD16, CD3D, and GSDMB identified as prognostic genes in BLCA
  • Risk model incorporating the 3 genes plus age and N stage showed robust predictive accuracy AUC > 0.7
  • High-risk group had significantly higher TIDE score than low-risk group p < 0.001
  • Epithelial cells identified as key subpopulation; T cells and fibroblasts also implicated
  • GSDMB expression significantly upregulated in BLCA by RT-qPCR p < 0.05
  • KCTD16 and CD3D expression significantly downregulated in BLCA by RT-qPCR p < 0.05
  • Knockdown of GSDMB and KCTD16 significantly promoted T24 cell proliferation
  • Scissor+ cells correlated with high-risk phenotype and showed pseudotime-dependent expression
Key statistics
  • pvalue p < 0.001 (TIDE score comparison between high-risk and low-risk groups)
  • pvalue p < 0.05 (RT-qPCR differential expression of GSDMB, KCTD16, CD3D in BLCA vs normal)
  • count 404 BLCA samples (TCGA-BLCA cohort retained for prognostic analysis)
  • count 15 paired tumor/normal samples (Soochow-BLCA supplementary training cohort)
  • count 165 primary BLCA samples (GSE13507 external validation cohort)
  • count 7 primary BLCA tissues + 1 adjacent normal tissue (GSE135337 scRNA-seq dataset)
  • other AUC > 0.7 (Predictive accuracy of the risk score model)
  • fold_change |log2FC| > 0.5, adjusted p < 0.05 (Threshold used to define differentially expressed genes)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This study used a multi-stage computational pipeline to identify succinylation-related prognostic genes in bladder cancer. DESeq2 differential expression and ssGSEA pathway scoring were used to nominate candidate genes, which were narrowed by univariable Cox regression, LASSO, and multivariable Cox regression (with explicit proportional-hazards testing) to yield a weighted risk score. Performance was assessed with Kaplan-Meier/log-rank tests and time-dependent ROC curves in a training cohort (TCGA-BLCA, n=404) and an external validation cohort (GSE13507, n=165). Functional, immune, mutational, and drug-sensitivity characterisation used Spearman correlation, Wilcoxon tests, GSEA, ESTIMATE, and TIDE, while single-cell analyses applied UMAP clustering, SingleR annotation, Scissor bulk-to-single-cell integration, CellChat, and pseudotime analysis.

Replicationmixed Sample sizeTCGA-BLCA training n=404 (after excluding 12 samples with incomplete data from 416); external validation GSE13507 n=165; Soochow-BLCA n=15 paired samples; scRNA-seq GSE135337 n=8 samples (7 BLCA + 1 normal); no formal power calculation described for any cohort GroupsBLCA vs adjacent normal (Soochow); high vs low ssGSEA score; high-risk vs low-risk group; NMIBC vs MIBC; high vs low TMB; multiple clinical-variable strata Pairingmixed Randomization/blindingnot stated Dispersionunclear Exact p-valuesno Effect sizesyes Confidence intervalsyes Multiplicity correctionBenjamini-Hochberg FDR applied in DESeq2 (adjusted p<0.05 threshold) and in clusterProfiler GSEA (adjusted p<0.05); no correction stated for multiple Wilcoxon tests, multiple Spearman correlations, or Cox screening across many genes
Statistical tests used
Test Applied to n Assumptions
DESeq2 Wald test (negative-binomial model) Differential expression between BLCA and adjacent normal tissue in the Soochow-BLCA cohort 15 BLCA + 15 adjacent normal (paired samples) not stated
Log-rank test (Kaplan-Meier) Overall survival comparison between high- and low-ssGSEA-score groups; high- vs low-risk groups in TCGA-BLCA training and GSE13507 validation 404 (TCGA training); 165 (GSE13507 validation) not stated
Univariable Cox proportional hazards regression Screening of candidate prognostic genes (p<0.01) and clinical variables (p<0.05) for OS association; PH assumption tested (p>0.05 required for retention) 404 (TCGA-BLCA) stated
LASSO regression (L1-penalised Cox model, 10-fold cross-validation) Dimensionality reduction and multicollinearity elimination among univariably significant prognostic candidates 404 (TCGA-BLCA) not stated
Multivariable Cox proportional hazards regression with stepwise selection Identification of independent prognostic genes and clinical factors; PH assumption tested for all retained variables 404 (TCGA-BLCA) stated
Wilcoxon rank-sum test Risk scores vs clinical variables (age, gender, TNM stage); TME/ESTIMATE/TIDE scores across risk groups; TMB between risk groups; IC50 values between risk groups; differential expression of prognostic genes across cell types in scRNA-seq 404 (TCGA bulk); cell counts not stated (scRNA-seq) not stated
Spearman rank correlation Prognostic genes vs 20 known SRGs (linkET); prognostic genes vs all TCGA genes (pre-ranking for GSEA); TMB vs risk score 404 (TCGA-BLCA) not stated
Gene Set Enrichment Analysis (GSEA, pre-ranked by Spearman correlation) Biological pathway enrichment per prognostic gene using KEGG gene sets (MSigDB c2.cp.kegg.v7.4) 404 (TCGA-BLCA) not stated
Time-dependent ROC curve / AUC at 1, 2, and 3 years Predictive accuracy of risk score model and nomogram in TCGA-BLCA training set and GSE13507 validation set 404 (training); 165 (validation) not stated
Approaches that could also have been used
  • The optimal cutpoint for stratifying patients into high- and low-risk (and high- and low-ssGSEA-score) groups was determined data-adaptively using the survminer package in the training cohort
    Could also: A pre-specified cutpoint (e.g., median, clinical quartiles) or retaining the risk score as a continuous variable in a Cox model could also be used — Data-adaptive optimal cutpoints maximise the apparent survival difference in the training data, which can inflate reported separation; a pre-specified or continuous approach avoids this additional analytical degree of freedom and is more directly transferable to independent cohorts
  • Multiple Wilcoxon rank-sum tests were performed across several related families of comparisons (clinical variables, TME scores, TIDE scores, TMB, drug IC50 values, scRNA-seq cell types) without stated correction for multiplicity
    Could also: Benjamini-Hochberg FDR correction applied within each family of related Wilcoxon tests could also be reported — Explicitly controlling the expected false-discovery rate within each comparison family is a standard practice when many simultaneous tests are conducted; it makes the nominal significance level more interpretable alongside the uncorrected results
  • The 15 paired BLCA and adjacent normal samples from the Soochow-BLCA cohort were analysed with DESeq2, but the available text does not describe whether the paired (within-patient) structure was modelled in the DESeq2 design formula
    Could also: Including patient ID as a blocking factor in the DESeq2 design formula (~patient + condition) explicitly models within-patient correlation — Accounting for the paired structure reduces residual variance, can increase statistical power to detect true differential expression, and more accurately reflects the study design
  • LASSO (L1 penalty only) was applied for variable selection among prognostic candidate genes
    Could also: Elastic net regularisation (mixing L1 and L2 penalties, alpha between 0 and 1) could also be applied — Elastic net can handle groups of correlated predictors more stably than LASSO, which tends to select one variable arbitrarily from a correlated cluster; this may be relevant when co-expressed candidate genes are considered together
  • Predictive discrimination was assessed using time-dependent AUC at three fixed time points (1, 2, 3 years)
    Could also: Harrell's concordance index (C-index) integrated over the full follow-up period could also be reported alongside the time-point AUCs — The C-index provides a global summary of discrimination across all observed event times rather than at pre-specified landmarks, and is widely reported in prognostic model papers as a complementary measure
  • Internal model performance was evaluated on the same TCGA-BLCA dataset used for model building, with external validation provided by GSE13507
    Could also: Bootstrap-based internal validation (e.g., 500–1,000 resamples) within TCGA-BLCA to estimate an optimism-corrected C-index or AUC could also be performed — Bootstrap optimism correction quantifies overfitting within the training data independently of external validation, providing a more conservative estimate of expected model performance in new samples and complementing the external validation result
Software: R/DESeq2 1.42.0 · R/ggplot2 3.4.1 · R/ComplexHeatmap 2.14.0 · R/GSVA 1.42.0 · R/survminer 0.4.9 · R/ggvenn 0.1.9 · R/clusterProfiler 4.2.2 · Cytoscape 3.10.3 · R/survival 3.7.0 · R/glmnet 4.1.8 · R/timeROC 0.4 · R/rms 6.9.0 · R/psych 2.1.6 · R/GseaVis 0.1.0 · R/maftools 2.18.0 · R/pRRophetic 0.5 · R/Seurat 5.1.0 · R/SingleR 2.4.0 · R/Scissor 2.0.0 · R/CellChat 1.6.1 · R/linkET 0.0.7.4 · R/SCP 0.5.6 · TIDE platform (TIDEpy) · ESTIMATE algorithm · STRING database · CellMarker database

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
0
Impact: low
Foundation confidence
Built on 1 assessed reference(s) · mean reproducibility 38/100
built largely on non-reproducible work
Topics

Assessed papers, coloured by verdict. Click a node to open it.

Built on (assessed references) (1)
Cited by (assessed papers) (0)
  • No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

GSE13507 GEO in Results (http://purl.org/orb/Results)
also used by 1 paper:
GSE135337 GEO in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet

Downstream reach in the literature

100 downstream papers · 1 datasets

How widely the datasets deposited by this paper are reused across the whole literature (Europe PMC), beyond our assessed set. This is a factual dependency map — reusing a public dataset is normal, good science. It is not a judgement on the downstream papers; the only verdict here is this paper's own, with its cited rationale.

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-42220482

Paper: Identification of succinylation-related genes in bladder cancer: integration of single-cell and transcriptomic data. Front Immunol 2026; DOI 10.3389/fimmu.2026.1797389. Pinned by brief: Code = TIDEpy (github.com/jingxinfu/TIDEpy) · Data = GEO GSE13507.

The paper is a multi-tool integrative bioinformatics study (TCGA-BLCA training, GSE13507 external validation, GSE135337 scRNA-seq, plus an in-house Soochow cohort and wet-lab RT-qPCR/CCK-8). Below: which reported results are pipeline-derived and in scope vs out.

In scope (attempted)

id result pipeline data feasibility
C1 TIDE immune-escape scoring (Fig 5) — the pinned tool TIDEpy (TIDE) GSE13507 (pinned data) RUN. P16: third-party tool on the paper's own data. Paper applies TIDE to TCGA-BLCA risk groups and prints no GSE13507 TIDE number, so this is a tool-runs-clean demonstration → expected grade partial.
C2 4 SRGs significantly associated with OS (Results / Suppl. Cox) — SIRT6 HR 0.578, SIRT7 0.696, OXCT1 1.163, SUCLA2 1.369 univariate Cox (R survival) TCGA-BLCA (public, UCSC Xena) RUN. Deterministic given expression+OS; directly checkable 1:1 (HR + 95% CI).

Out of scope / not attempted (with reason)

  • 3-gene risk model (KCTD16/GSDMB/CD3D) + GSE13507 split HRG=73/LRG=92 (Fig 3). docs_insufficient for a 1:1: the LASSO/Cox β coefficients are NOT reported in text or supplementary tables, and the paper does not state whether the TCGA-derived cutoff −2.002354 was re-applied to GSE13507 or re-fit. Without coefficients the split is not deterministically reproducible. This is the under-specified ~20% — not chased.
  • Full LASSO selection (24→3 genes), nomogram, time-ROC AUC>0.7 — non-deterministic LASSO seed + unreported coefficients. Out.
  • DEGs (3,769) in Soochow-BLCA cohort — in-house data not deposited. Out.
  • scRNA-seq (GSE135337): 39,899 cells, 13 clusters, Scissor 222+/267− — heavy, many under-specified params (resolution, marker sets). Not low-hanging. Out (80/20).
  • CIBERSORT/ESTIMATE/GSEA/maftools/pRRophetic panels — secondary, multi-step, parameter-sensitive. Out (80/20).
  • RT-qPCR, CCK-8 proliferation (Fig 9) — wet-lab. Out of scope by definition.

Honest framing

The brief pins a third-party tool (TIDEpy) — P16 makes running it on the paper's data a valid reproduction. We do that (C1) and add one clean deterministic 1:1 check on public TCGA data (C2). We explicitly do not claim to reproduce the headline risk model, because its coefficients are not shipped.

Figures / tables: Fig 5
C1
Reported
TIDE applied to BLCA (Fig 5); no GSE13507 TIDE value printed
Reproduced
TIDEpy ran clean on 165 GSE13507 primary tumors; TIDE mean 0.071, median -0.019; responder 84/165 = 50.9%
partial
C2a
Reported
SIRT6 univariate Cox HR=0.578 (0.426-0.786)
Reproduced
HR=0.587 (0.448-0.768), p=1.0e-4, n=424 (TCGA-BLCA)
within tolerance
C2b
Reported
SIRT7 univariate Cox HR=0.696 (0.527-0.919)
Reproduced
HR=0.726 (0.589-0.895), p=2.7e-3, n=424
within tolerance
C2c
Reported
OXCT1 univariate Cox HR=1.163 (1.036-1.306)
Reproduced
HR=1.140 (1.047-1.242), p=2.5e-3, n=424
within tolerance
C2d
Reported
SUCLA2 univariate Cox HR=1.369 (1.026-1.872)
Reproduced
HR=1.449 (1.164-1.804), p=9.1e-4, n=424
within tolerance

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 78/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟡2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟡5. Derivability / plausibility
🟢6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +7

The supporting claim — four succinylation-related genes (SIRT6, SIRT7, OXCT1, SUCLA2) associated with OS on TCGA-BLCA — reproduces cleanly within tolerance (HRs match direction, all p<0.01, overlapping CIs); the small ~2-6% deltas are explained by our n=424 vs the paper's 404 sample filtering. No fabrication concern: these values are derivable from the shipped public data. However, the paper's headline 3-gene risk model (KCTD16/GSDMB/CD3D) and its GSE13507 split could not be reproduced because the LASSO/Cox β coefficients and cutoff are absent from text and supplements — an authors-side documentation gap. Net: a solid, fabrication-free partial reproduction whose central prognostic claim remains unverified.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at [email protected].

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

113.8 k
tokens (I/O) · 5.2 M incl. cache
12 min
runtime · 0.02 CPU-h
1.9 GB
peak RAM
1
HPC jobs
hummel
machine