STAT3-dependent analysis reveals PDK4 as independent predictor of recurrence in prostate cancer.
The main results reproduced, with only marginal, non-material deviations.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Same input data as the authors
- ✓Reported values were directly comparable
- ✓Reported values are derivable from the shared data
- ✓The central claim held under reproduction
- 🟡A deviation arose in the data or preprocessing
- 🟡A deviation was attributed to the published material
- 🟡The deviation was non-trivial in magnitude
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Described well enough to reproduce the dataset-specific result 1:1. The RU-assigned accession GSE120741 is the paper's NCI validation cohort, used (Fig EV3) for STAT3-vs-metabolic-signature ssGSEA correlations - NOT for the headline PDK4 recurrence-survival result (that uses MSKCC/GSE21032, a different accession, out of scope here). GEO ships a processed expression matrix (normalized log2 ComBat-corrected read counts, GeneSymbol x 92 tumor samples), so STAR re-alignment was intentionally skipped (80/20); this is the P16 third-party-tool case (ssGSEA via GSVA on the paper's own data). Reproduced on «our HPC» SLURM: ssGSEA (GSVA 2.4.4) of KEGG_OXIDATIVE_PHOSPHORYLATION and KEGG_RIBOSOME (msigdbr 26.1.0), then Pearson correlation of STAT3 expression vs each signature. PRIMARY claims match essentially exactly: OXPHOS rho -0.77 -> -0.749 (|d|=0.021), Ribosome rho -0.82 -> -0.802 (|d|=0.018), same sign and significance. SECONDARY STAT3-target claim: magnitude (0.365 vs 0.39) and p (3.5e-4 vs 1.5e-4) match strikingly but SIGN is opposite using MSigDB AZARE_STAT3_TARGETS; the paper cites two STAT3-target sets (Azare 2007; Carpenter & Lo 2014) and does not pin the exact gene list, so the signature identity is ambiguous - flagged for human review, not a fabrication signal. NOT attempted: PDK4 MSKCC survival (different accession), wet-lab/metabolomics/proteomics/PET (non-pipeline), and exact BH-adjusted p (paper adjusts jointly across all datasets x signatures; we report raw cor.test p, consistent in magnitude). Overall: a clean partial - the two primary GSE120741 correlations reproduce 1:1; one secondary correlation sign-mismatches on an under-specified gene set.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 66assessed: 2026-06-15 ⛓ 97aadbb6eed9
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-15
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusBy comparing low-STAT3 to high-STAT3 prostate cancer at the transcriptomic and proteomic levels, the study tests which biological processes correlate with STAT3 expression in order to identify markers associated with earlier biochemical recurrence (BCR) for risk stratification.
- ★ Low STAT3 expression in primary prostate cancer is associated with up-regulated OXPHOS and ribosomal biosynthesis at the transcriptomic level. finding
- ★ STAT3 expression is inversely correlated with TCA cycle/OXPHOS activity, corroborated at the proteomic level in laser-microdissected human and murine FFPE prostate samples. finding
- ★ PDK4 gene expression, a key regulator of the TCA cycle, is significantly down-regulated in low-STAT3 patients and low PDK4 associates with higher risk of biochemical recurrence. finding
- ★ PDK4 is an independent prognostic marker predicting biochemical recurrence independent of ISUP grading, clinical/pathological staging, and pre-surgical PSA levels. finding
- ★ A weighted gene co-expression network (WGCNA) of TCGA-PRAD RNA-Seq data yields gene modules (OXPHOS cluster 2, Ribosome cluster 3) negatively correlated with STAT3 expression. method
- STAT3 gene expression reflects its transcriptional activity, correlating positively with pY-STAT3 protein and STAT3 target gene signatures. mechanism
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| bulk RNA-seq (TCGA PRAD), differential expression + KEGG/GO/EGSEA/WGCNA analysis | human prostate adenocarcinoma (TCGA-PRAD, 498 patients; 397/382 with clinical data) | none (low vs high STAT3 patient stratification) | differentially expressed genes, enriched pathways, gene co-expression modules, eigengene correlations | — |
| Reverse Phase Protein Array (RPPA) | human prostate adenocarcinoma (TCGA PRAD) | none | tyrosine-phosphorylated (pY) STAT3 protein levels correlated with STAT3 cpm | — |
| ssGSEA gene signature scoring | human PCa validation cohorts (NCI n=91 GSE120741; VPC n=43; RAS n=33) | none | KEGG OXPHOS and Ribosome signature scores correlated with STAT3 log cpm | — |
| shotgun proteomics of laser-microdissected FFPE tissue | human and murine prostate FFPE samples | none / murine Pten Stat3 deletion model context | TCA cycle/OXPHOS protein levels relative to STAT3 | — |
- ▲ 1,194 genes significantly differentially expressed between low and high STAT3, with Ribosome and OXPHOS among top up-regulated KEGG pathways in low STAT3 1,194 genes (log-FC ≥ 1, adj.P ≤ 0.05)
- ▼ OXPHOS cluster 2 and Ribosomal cluster 3 eigengenes negatively correlated with STAT3 expression cluster 2 ρ=−0.67; cluster 3 ρ=−0.74
- ▲ Epigenetic cluster 11 eigengene positively correlated with STAT3 expression ρ=0.59 (adj.P=8e-36)
- ▼ STAT3 log cpm negatively correlated with KEGG OXPHOS signature across three independent validation cohorts NCI ρ=−0.77; VPC ρ=−0.53; RAS ρ=−0.57
- ▲ STAT3 target gene signatures positively correlated with STAT3 expression AZARE ρ=0.67; STAT3 TARGETS UP ρ=0.46; pY-STAT3 ρ=0.24
- – Overlap of differentially expressed genes with WGCNA modules: 316 with OXPHOS cluster 2 and 103 with Ribosome cluster 3 316 and 103 genes
- ▲ Clusters 10 and 12 positively correlated with Gleason score and pT risk, with no overlap with STAT3-correlated clusters cluster 12 vs GSC ρ=0.51; cluster 10 vs GSC ρ=0.44
- correlation ρ=−0.74, adj.P-value=1e-65 (Ribosomal cluster 3 eigengene vs STAT3 expression (WGCNA))
- correlation ρ=−0.67, adj.P-value=7e-50 (OXPHOS cluster 2 eigengene vs STAT3 expression (WGCNA))
- correlation ρ=−0.77, adj.P-value=4.53e-19 (STAT3 log cpm vs KEGG OXPHOS, NCI cohort)
- correlation ρ=0.67, adj.P-value=4.72e-63 (STAT3 vs AZARE STAT3 TARGETS signature (ssGSEA), TCGA PRAD)
- correlation ρ=0.24, P-value=7.8e-06 (STAT3 log cpm vs pY-STAT3 RPPA protein, TCGA PRAD)
- pvalue P-value=2.5e-05 (STAT3 target genes up-regulated in high vs low STAT3 (roast gene set test))
- count 1,194 differentially expressed genes (low STAT3 (n=100) vs high STAT3 (n=100) comparison)
- count BCR=52 of 397; 13,932 genes, 382 patients, 13 modules (TCGA PRAD clinical/WGCNA cohort composition)
Statistical methods review
Model: opusA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This is an observational, computational study combining transcriptomics (TCGA-PRAD RNA-Seq and three additional public PCa cohorts) with shotgun proteomics of laser-microdissected human and murine FFPE samples. The main analytic approach compared low-versus-high STAT3 patient groups via differential gene expression, built a weighted gene co-expression network (WGCNA), performed gene-set/pathway enrichment (KEGG/GO overexpression, EGSEA, roast, ssGSEA), and used Pearson correlation to relate STAT3 and module eigengenes to molecular signatures and clinical traits, with PDK4 then evaluated as a predictor of biochemical recurrence. Results were reported with correlation coefficients and Benjamini–Hochberg-adjusted P-values.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Differential expression test (thresholds log-FC ≥ 1, adj. P ≤ 0.05) | low STAT3 vs high STAT3 in TCGA PRAD (1,194 DE genes) | n = 100 low STAT3, n = 100 high STAT3 (498 total ranked into tertile-like groups) | not stated |
| KEGG/GO overexpression (enrichment) analysis | DE genes and WGCNA clusters (Fig 1B, Fig 2A,C,D,E) | — | na |
| EGSEA gene set testing | KEGG signaling/metabolic and Hallmark gene sets, low vs high STAT3 (Figs EV1, EV2) | — | not stated |
| roast gene set test | STAT3 TARGETS UP gene set, high vs low STAT3 (P = 2.5e-05) | — | na |
| Pearson correlation | STAT3 log cpm vs pY-STAT3 RPPA, ssGSEA signatures, and module eigengenes/clinical traits across TCGA, NCI, VPC, RAS data sets (Figs 2B, EV3) | TCGA n = 498/397/382; NCI n = 91; VPC n = 43; RAS n = 33 | not stated |
| ssGSEA single-sample enrichment scoring | STAT3 target, OXPHOS, and Ribosome signatures across cohorts | — | na |
| One-way / multi-way ANOVA with Tukey HSD post-hoc | STAT3 pathway gene expression vs clinical traits (e.g., SOCS3 across GSC groups) | — | not stated |
-
Patients were dichotomized into low vs high STAT3 by ranking into the top and bottom 20% quantiles (n = 100 each) with the middle group set aside.↳ Could also: STAT3 could also be analyzed as a continuous variable in a regression model alongside the differential-expression or correlation analyses. — A continuous treatment uses all 498 samples and retains information that grouping into quantiles necessarily collapses, which some analysts prefer for power and to avoid threshold dependence.
-
Group differences and signature relationships were summarized primarily with correlation coefficients (ρ) and adjusted P-values.↳ Could also: Reporting 95% confidence intervals for the correlation and effect estimates would also be informative. — Confidence intervals convey the precision of an estimate in addition to its point value, complementing the P-value-based reporting.
-
Pearson correlation was used to relate STAT3 expression to protein levels, signatures, and module eigengenes.↳ Could also: A rank-based Spearman correlation could also be applied to these relationships. — Spearman captures monotonic associations without assuming linearity or normality, which can be useful when distributions are skewed or relationships are non-linear.
-
PDK4 and STAT3 pathway genes were related to clinical traits using ANOVA and correlation, with recurrence framed via group comparisons.↳ Could also: Time-to-event modeling such as Kaplan–Meier with the log-rank test and Cox proportional-hazards regression could also be used for the recurrence endpoint. — Survival models explicitly account for follow-up time and censoring and yield hazard ratios with confidence intervals when assessing independent prognostic value (the visible text does not specify the exact recurrence model used).
-
Multiplicity was handled with Benjamini–Hochberg FDR across correlation and enrichment families and Tukey HSD for ANOVA post-hoc tests.↳ Could also: Alternative error-rate controls such as Bonferroni/Holm (family-wise) or storey q-value (FDR) could also be applied. — Different procedures balance sensitivity and stringency differently; the choice depends on whether family-wise error or false-discovery rate is the priority for a given family of tests.
-
Differential expression used fixed thresholds (log-FC ≥ 1, adj. P ≤ 0.05) to define significant genes.↳ Could also: Threshold-free gene-set approaches (e.g., GSEA on a ranked list) could also be used to summarize the differential signal. — Ranked, threshold-free methods avoid sensitivity to a specific fold-change cutoff and can detect coordinated shifts in gene sets that individual-gene thresholds may miss.
Result convergence & founder nodes
Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.
-
Epigenetic gene co-expression module eigengene positively correlated with STAT3 expressionRNA-seq human prostate adenocarcinoma up 2020×1papers★ This paper is the founder (earliest)
-
Gene co-expression modules positively correlated with Gleason score and pT risk, distinct from STAT3-correlated modulesRNA-seq human prostate adenocarcinoma up 2020×1papers★ This paper is the founder (earliest)
-
OXPHOS pathway up-regulated in low-STAT3 tumors, i.e. negatively associated with STAT3 expressionRNA-seq human prostate adenocarcinoma down 2020×1papers★ This paper is the founder (earliest)
-
Ribosomal gene co-expression module eigengene negatively correlated with STAT3 expressionRNA-seq human prostate adenocarcinoma down 2020×1papers★ This paper is the founder (earliest)
-
STAT3 target gene signatures positively correlated with STAT3 expressionRNA-seq human prostate adenocarcinoma up 2020×1papers★ This paper is the founder (earliest)
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-32323921
Paper: Oberhuber et al. 2020, Mol Syst Biol 16:e9247. "STAT3-dependent analysis reveals PDK4 as independent predictor of recurrence in prostate cancer." PMID 32323921 · PMCID PMC7178451 · DOI 10.15252/msb.20199247
Registry-assigned artifacts for THIS reproduction unit:
- Code: https://github.com/alexdobin/STAR (third-party aligner — P16 case; not authors' own repo)
- Data: GEO GSE120741 (= "The Netherlands Cancer Institute / NCI" cohort, Stelloo et al.; BioProject PRJNA494345). Used by Oberhuber et al. as a public validation cohort.
Key finding from reading the paper (Methods + Results + Fig EV3)
GSE120741 is NOT the dataset behind the headline PDK4-recurrence survival result. The PDK4 biochemical-recurrence survival analysis (Kaplan-Meier / Cox) is done on the MSKCC cohort (GSE21032), which is a different accession not assigned to this RU.
The role of GSE120741 (NCI, n = 91) in the paper is the STAT3-vs-metabolic-signature correlation (Fig EV3B–C and associated text). The paper reports (verbatim values):
- STAT3 log-CPM vs ssGSEA KEGG "OXPHOS" signature — NCI: ρ = −0.77, adj. P = 4.53e-19
- STAT3 log-CPM vs ssGSEA KEGG "Ribosome" signature — NCI: ρ = −0.82, adj. P = 8.33e-23
- STAT3 log-CPM vs STAT3-TARGET signature (AZARE STAT3 targets) — NCI: ρ = −0.39, adj. P = 1.5e-04
Method per paper: "significant negative Pearson correlation of STAT3 log cpm with KEGG 'OXPHOS' / 'Ribosome' signatures derived by ssGSEA (Barbie et al, 2009)"; p-values adjusted by Benjamini–Hochberg. Survival pkgs (survival v3.1-8, survminer v0.4.6) are for the MSKCC survival part — out of scope for this accession.
IN SCOPE (pipeline-derived, tied to GSE120741) — what we reproduce
| id | reported (NCI cohort, Fig EV3 / text) | pipeline |
|---|---|---|
| C1_oxphos_rho | Pearson ρ(STAT3 logCPM, ssGSEA KEGG_OXIDATIVE_PHOSPHORYLATION) = −0.77, adj.P 4.53e-19 | ssGSEA (GSVA) + cor.test on GSE120741 GE table |
| C2_ribosome_rho | Pearson ρ(STAT3 logCPM, ssGSEA KEGG_RIBOSOME) = −0.82, adj.P 8.33e-23 | same |
| C3_stat3target_rho | Pearson ρ(STAT3 logCPM, STAT3-target signature) = −0.39, adj.P 1.5e-04 | same (secondary; signature identity less certain) |
| C0_n | NCI cohort n = 91 | sample count of GE table |
Primary targets: C1, C2 (well-specified KEGG sets, strong effect → robust to method). Secondary: C3 (the exact "AZARE STAT3 targets" gene list is less precisely pinned).
OUT OF SCOPE (not attempted, with reason)
- PDK4 Kaplan-Meier / Cox biochemical-recurrence survival — uses GSE21032 (MSKCC), a different accession not assigned to this RU. Not the GSE120741 result.
- All wet-lab / mouse Stat3-KO / metabolomics / proteomics / PET imaging — non-pipeline.
- STAR read alignment from raw FASTQ — GEO ships a processed gene-expression matrix (GSE120741_Porto_ge_table.txt.gz), so re-aligning from SRA adds no value to reproducing the reported correlation; we use the shipped processed matrix (the standard 80/20 choice).
Reproduction approach
- «our HPC» job: download GSE120741_Porto_ge_table.txt.gz (GEO FTP) to «infra».
- log-CPM normalize (edgeR) if the matrix is raw counts; else use shipped values.
- ssGSEA (GSVA, method="ssgsea") of KEGG_OXIDATIVE_PHOSPHORYLATION and KEGG_RIBOSOME (gene sets from msigdbr/MSigDB) across all samples.
- Pearson correlation: STAT3 expression vs each signature score; BH-adjust.
- Compare ρ to −0.77 / −0.82 (sign + magnitude). within-tol = |Δρ| ≤ 0.10 and same sign.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
On the in-scope accession (GSE120741, the NCI validation cohort for the Fig EV3 STAT3-vs-metabolic correlations), the two primary claims reproduce essentially 1:1 from the authors' own shipped matrix: OXPHOS ρ −0.77→−0.749 and Ribosome ρ −0.82→−0.802, same sign and significance. The only conflict is the secondary C3 STAT3-target correlation, which matches in magnitude/p (0.39 vs 0.365; 1.5e-4 vs 3.5e-4) but flips sign — attributable to an under-specified gene set in the Methods (two sets cited, none pinned), i.e. partly an authors-side documentation gap and partly our signature choice, not a fabrication signal. A minor n=92-vs-91 input difference (one unspecified QC drop) does not affect the robust correlations. Overall: solid partial reproduction with explainable deviations on a non-headline claim.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.