Metabolic reprogramming and prognostic insights in molecular landscapes driven by glycolysis in ovarian cancer.
The main results reproduced: recomputed values matched the published ones within tolerance.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- Nothing in this column.
- 🟡Could not use the authors’ exact input data
- 🟡Reported values were only indirectly comparable
- 🟡A deviation arose in the data or preprocessing
- 🔴A deviation was attributed to the published material
- 🔴Reported values were not (fully) derivable from the shared data
- 🟡The deviation was non-trivial in magnitude
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Described well enough to test the central claim, but NOT to reproduce the exact reported numbers. The repo (github.com/mingwei3516/OC @ a780c8e) ships 28 R scripts with hardcoded Windows paths, no README, no data, no model coefficients, and no set.seed; the raw data bundle is only on a login-walled Chinese cloud (jianguoyun). So we could not re-run the exact 30-gene screen / 10-gene LASSO / reported AUC+p (and model.R uses an unseeded random split with a post-hoc AUC>0.68 acceptance filter => optimistic bias; paper text says LASSO lambda=15 but code uses lambda.min). Instead we ran a DIFFERENT-but-valid reproduction: applied the paper's named 10-gene signature + the repo's own modelling recipe to a fully public cohort (GSE26193, 107 OC, 76 events) on «our HPC». Outcome = PARTIAL / mostly-supports: the signature stratifies OS strongly (full-cohort logrank p<1e-3, AUC 0.74-0.81, comparable to reported >0.685), and per-gene HR directions are 7/7 coherent with the paper's qPCR (Fig.9) and subtype (Fig.2) characterisations. A small held-out half is underpowered (logrank p=0.36) as expected for n=53. NOT attempted (out of 80/20 scope): exact TCGA+GTEx 457-DEG step (needs GeneCards 4110-gene list + reassembled TCGA/GTEx), consensus clustering, CIBERSORT/ssGSEA/ESTIMATE immune analyses, scRNA (GSE154600/GSE150864), and qRT-PCR (wet-lab). The 10-gene signature's prognostic value reproduces; its exact derivation is not independently checkable from the shipped artifacts.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 77assessed: 2026-06-14 ⛓ 3331f8edd823
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-14
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetGlycolysis-related genes (GRGs), reflecting the Warburg effect, harbor prognostic biomarkers and therapeutic targets whose expression patterns can stratify ovarian cancer patients by molecular subtype and clinical outcome.
- ★ 457 differentially expressed GRGs were identified between OC and normal ovarian tissue, of which 30 were significantly associated with prognosis finding
- ★ OC can be classified into three GRG-based molecular subtypes (A, B, C), with cluster C showing the worst prognosis and activation of tumor-associated pathways finding
- ★ A ten-gene GRG prognostic signature (LMCD1, L1CAM, MYCN, GALT, IDO1, RPL18, XBP1, LPAR3, RUNX3, PLCG1) built via LASSO-Cox regression robustly predicts OC survival across multiple cohorts method
- ★ High-risk and low-risk groups defined by the GRG score show significant differences in tumor immune microenvironment composition finding
- ★ Single-cell RNA-seq analysis identifies GRG expression heterogeneity across stromal and malignant cell populations in the tumor microenvironment finding
- ★ qRT-PCR in OVCAR-3 vs IOSE-80 cells largely confirms the predicted differential expression direction of the model genes finding
- Drug sensitivity analysis identified 48 compounds with differential IC50 between risk groups, including dasatinib and foretinib as more effective in high-risk patients resource
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| bulk RNA-seq / transcriptomic differential expression analysis | TCGA-OC tumor samples vs GTEx normal ovarian tissue | none | differentially expressed glycolysis-related genes | TCGA/GTEx databases |
| univariate Cox regression | TCGA-OC-GSE26193 cohort | none | association of GRG expression with survival prognosis | — |
| copy number variation (CNV) analysis | TCGA-OC-GSE26193 cohort | none | chromosomal gains/losses in GRGs | TCGA database |
| consensus clustering | TCGA-OC-GSE26193 cohort | none | molecular subtype classification (k=3) based on 30 prognostic GRGs | — |
| KEGG pathway enrichment / GSEA | TCGA-OC-GSE26193 cohort, subtype C vs A | none | tumor-associated signaling pathway activation | — |
| immune infiltration analysis (stromal/immune scoring, cell-type deconvolution) | TCGA-OC-GSE26193 cohort, high-risk vs low-risk groups | none | immune cell type proportions, stromal and immune scores | — |
| single-cell RNA sequencing | OV_GSE154600 dataset (TISCH database) | none | GRG expression across 26 cell clusters / 11 cell types in tumor microenvironment | TISCH database |
| qRT-PCR | IOSE-80 (normal ovarian epithelial) vs OVCAR-3 (ovarian cancer) cell lines | none (cell line comparison) | mRNA expression of the 10 model genes | qRT-PCR |
- – 457 GRGs differentially expressed between OC and normal ovarian tissue; 30 significantly correlated with prognosis (18 risk-associated, 12 protective)
- ▼ Three molecular subtypes identified; cluster C had the worst prognosis, cluster A the most favorable
- – Ten-gene GRG signature predicted 1/3/5-year OS with AUC >0.685 in training set and >0.583 in testing set AUC>0.685 (train), AUC>0.583 (test)
- ▼ External validation confirmed lower survival in high-risk group in GSE53963 and GSE140082 cohorts P=0.014 (GSE53963); P=0.023 (GSE140082)
- ▼ M1 macrophage proportion decreased significantly with increasing risk score R=-0.29, P=1.7×10⁻⁶
- – High-risk group showed more resting dendritic cells and M0 macrophages but fewer activated dendritic cells, M1 macrophages, activated CD4+ memory T cells, and follicular helper T cells versus low-risk group
- – qRT-PCR showed L1CAM, LMCD1, PLCG1, RUNX3 upregulated and GALT, MYCN, XBP1 downregulated in OVCAR-3 vs IOSE-80; no significant difference for IDO1, LPAR3, RPL18
- ▼ Drug sensitivity analysis identified 48 drugs with differential sensitivity between risk groups; high-risk group more sensitive to dasatinib and foretinib (lower IC50) 48 drugs
- count 457 (differentially expressed glycolysis-related genes (TCGA-OC vs GTEx))
- count 30 (GRGs significantly correlated with OC prognosis)
- correlation R = -0.29, P = 1.7×10⁻⁶ (M1 macrophage proportion vs risk score)
- pvalue P = 0.014 (high- vs low-risk survival difference in GSE53963 cohort)
- pvalue P = 0.023 (high- vs low-risk survival difference in GSE140082 cohort)
- other AUC > 0.685 (training set) (1/3/5-year OS prediction accuracy of the ten-gene model)
- other AUC > 0.583 (testing set) (1/3/5-year OS prediction accuracy of the ten-gene model)
- count 48 (drugs with differential sensitivity between high- and low-risk groups)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This retrospective bioinformatics study integrated transcriptomic and clinical data from TCGA, GTEx, and GEO to characterize glycolysis-related gene (GRG) signatures in ovarian cancer. Differentially expressed GRGs were screened by univariate Cox regression, consensus clustering defined three molecular subtypes, and a ten-gene prognostic risk score was constructed via LASSO-penalized Cox regression with 10-fold cross-validation. The model was validated in two independent GEO cohorts and cross-checked experimentally by qRT-PCR in two cell lines; results were reported using Kaplan-Meier curves, time-dependent ROC AUC, a nomogram with calibration plot, and Spearman correlations with immune infiltration estimates.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Univariate Cox proportional-hazards regression | Screening 457 differentially expressed GRGs for association with overall survival in the TCGA-OC-GSE26193 cohort; threshold P < 0.05 yielded 30 significant GRGs | — | not stated |
| LASSO-penalized Cox regression with 10-fold cross-validation | Feature selection reducing 30 prognosis-associated GRGs to a ten-gene prognostic signature; λ minimising CV error reported as 15 | — | not stated |
| Consensus clustering (k = 3) | Classification of OC samples into three molecular subtypes based on expression profiles of 30 prognosis-related GRGs; k = 3 chosen as optimal | — | not stated |
| Kaplan-Meier survival analysis with log-rank test (implied) | Overall survival comparison among three subtypes (P < 0.001) and between high- vs low-risk GRG score groups in training, testing, and external validation cohorts (P = 0.014 for GSE53963; P = 0.023 for GSE140082) | GSE53963 N = 174; GSE140082 N = 380; main TCGA-OC-GSE26193 cohort size not stated | not stated |
| Time-dependent ROC analysis (AUC) | 1-, 3-, and 5-year overall survival prediction by the GRG model in training (AUC > 0.685) and testing (AUC > 0.583) subsets | — | na |
| Spearman rank correlation | Association between continuous GRG risk score and proportions of immune cell types in the TME (e.g., M1 macrophages: R = −0.29, P = 1.7 × 10⁻⁶) | — | not stated |
| KEGG pathway enrichment analysis and Gene Set Enrichment Analysis (GSEA) | Biological pathway characterization of molecular subtypes and the ten model genes, including Hallmark glycolysis enrichment | — | not stated |
| qRT-PCR gene expression comparison (specific statistical test not stated) | Validation of differential expression of ten model genes in normal ovarian epithelial cells (IOSE-80) vs ovarian cancer cells (OVCAR-3) | — | not stated |
-
457 univariate Cox regressions were screened at a nominal P < 0.05 threshold without a stated multiple testing correction, yielding 30 candidate GRGs↳ Could also: Apply a Benjamini-Hochberg FDR correction across all tests before selecting candidate genes — At α = 0.05 with 457 tests, roughly 23 false positives are expected by chance; an FDR correction would quantify and limit this inflation, clarifying which associations are likely genuine and making the gene list more reproducible across independent datasets
-
Consensus clustering with k = 3 was selected as the optimal subtype solution based on the consensus matrix↳ Could also: Supplement consensus clustering with quantitative internal validity indices—silhouette coefficient, gap statistic, or cophenetic correlation—reported across a range of k values (e.g., k = 2–6) — Presenting multiple stability metrics alongside the consensus matrix allows readers to independently assess how much better k = 3 separates samples than k = 2 or k = 4, strengthening the justification for the chosen subtype number
-
LASSO-Cox regression was used for feature selection, producing a ten-gene signature↳ Could also: Use elastic net-penalized Cox regression (combining L1 and L2 penalties) or bootstrap stability selection to assess which genes are consistently chosen across resamples — LASSO can be unstable when predictors are correlated, potentially selecting different gene subsets in different samples; elastic net or stability selection would quantify the reproducibility of each gene's inclusion, supporting the generalizability of the ten-gene signature
-
Model discrimination was summarized with time-dependent AUC at three fixed time points (1, 3, 5 years)↳ Could also: Also report Harrell's concordance index (C-statistic) with a 95% confidence interval as an overall discrimination summary — The C-statistic is the standard global discrimination metric for Cox survival models; reporting it alongside time-point-specific AUCs would enable direct comparison with previously published ovarian cancer prognostic models and provide a single summary measure of ranking ability
-
Internal validation used a single random split of the TCGA-OC-GSE26193 cohort into training and testing subsets↳ Could also: Use repeated k-fold cross-validation or bootstrap-based internal validation averaging over many splits — A single random split can yield optimistic or pessimistic performance estimates depending on which patients fall into each partition; repeated cross-validation or bootstrapping averages across many partitions, providing a more stable and less split-dependent internal performance estimate
-
The qRT-PCR cell-line experiment compared IOSE-80 vs OVCAR-3 without stating the statistical test, number of biological replicates, or measure of dispersion↳ Could also: Report a two-sample t-test or Mann-Whitney U test with the number of biological and technical replicates, and present means with SD or individual data points — With only two cell lines the comparison is inherently descriptive; explicitly naming the test, replicate count, and a spread measure would allow readers to judge the evidential weight of the experimental validation relative to the bioinformatic findings
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-40707588
Paper: Wang M et al. (2025) Metabolic reprogramming and prognostic insights in molecular landscapes driven by glycolysis in ovarian cancer. Sci Rep. DOI 10.1038/s41598-025-12350-7 · PMID 40707588 · PMCID PMC12290113.
Code: https://github.com/mingwei3516/OC (commit a780c8ecfa225579daeaaa81eea690ccae39c689, pushed 2025-06-30).
28 standalone R scripts, no README, no data files, no set.seed(), hardcoded
Windows paths («path»). Each script reads tab-delimited
input files that are NOT shipped (symbol.txt=TCGA, GSE26193.txt, time.txt,
diff.txt=GRG list, uniSigExpTime.txt, merge.txt, …).
Data named in paper: TCGA-OV (429), GTEx ovary (88), GSE26193 (107), GSE53963 (174), GSE140082 (380), GSE26712, GSE154600 (scRNA), GSE150864. Glycolysis gene list (4110 GRGs) from GeneCards keyword "glycolysis". Raw bundle on jianguoyun (Chinese cloud, login wall).
Pipeline-derived results (candidate in-scope)
| # | Reported result | Pipeline / pkg | Inputs needed | Reproducible from public data? |
|---|---|---|---|---|
| R1 | 457 GRGs differentially expressed (TCGA vs GTEx) | limma | TCGA-OV + GTEx + GeneCards 4110 list | Partly — needs GeneCards list (login/terms) + reassembled TCGA/GTEx |
| R2 | 30 GRGs prognostic (univ. Cox); 18 poor / 12 favorable | survival coxph | merged TCGA-OC-GSE26193 + survival + GRG list | Partly — needs merged cohort + GRG list |
| R3 | k=3 consensus clusters (subtypes A/B/C); C worst, P<0.001 | ConsensusClusterPlus | 30-gene expr on merged cohort | Depends on R2 inputs |
| R4 | 10-gene LASSO signature: LMCD1,L1CAM,MYCN,GALT,IDO1,RPL18,XBP1,LPAR3,RUNX3,PLCG1 | glmnet LASSO-Cox + step | uniSigExpTime on merged cohort | Not exactly — random split, no seed, post-hoc AUC filter |
| R5 | Train AUC >0.685; test AUC >0.583 (1/3/5y) | timeROC | risk scores | Not exactly (depends R4) |
| R6 | External validation: GSE53963 p=0.014; GSE140082 p=0.023 (low-risk better OS) | survival logrank | model coef + GEO data | Partly — needs coef (Table S2, not in repo) |
| R7 | qRT-PCR up/down of 10 genes (OVCAR-3 vs IOSE-80) | wet-lab | — | OUT OF SCOPE (wet-lab) |
| R8 | CIBERSORT/ssGSEA/ESTIMATE immune (e.g. M1 macro R=-0.29) | CIBERSORT/GSVA/estimate | risk groups | Depends R4 |
| R9 | scRNA: 26 clusters/11 types; RPL18 in all 11; L1CAM malignant | TISCH/Seurat | GSE154600 | Heavy, out of 80/20 |
DECISION — what we attempt (80/20, clean public data points)
The paper's headline numbers all hang off a merged TCGA+GTEx+GSE26193 cohort plus a
GeneCards gene list and model coefficients that are not shipped with the code, and
the LASSO model is fit on an unseeded random 50/50 split with a post-hoc AUC>0.68
acceptance filter (model.R) — so the exact 10 genes / AUC / p-values are not
deterministically reproducible from the artifacts. We do not chase those.
Instead we test the paper's central reproducible claim on a fully public dataset
(GSE26193, the RU's pinned accession; 107 OC, OS available) using the paper's own
named 10-gene set and the repo's own modelling recipe (singleGRG.Sur.R + model.R):
- C1 (per-gene prognostic signal): univariate Cox HR + KM logrank p for each of the 10 signature genes on GSE26193. Data point = how many of 10 are individually prognostic (P<0.05) and their HR direction. Tests R2/R4 plausibility.
- C2 (signature stratifies survival): fit a multivariable Cox risk score on the 10 genes (model.R recipe), median-split high/low, KM logrank p + timeROC AUC@1/3/5y on GSE26193. Tests the headline claim (R4/R5/R6): does this glycolysis signature carry prognostic signal in an independent public OC cohort?
Out of scope / not attempted: R7 (wet-lab qPCR); exact reproduction of R1/R3/R8/R9 (need GeneCards list, jianguoyun bundle, scRNA — beyond 80/20); reproducing the exact reported AUC/p numbers (impossible without shipped coefficients + seed).
**Honesty note (possible-fabric
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
The paper's central biological claim — the 10-gene glycolysis signature is prognostic in ovarian cancer with the stated risk directions — reproduces qualitatively on an independent public cohort (GSE26193: full-cohort logrank p<1e-3, AUC 0.74–0.81 vs reported >0.685, and 7/7 per-gene HR directions coherent with the paper's qPCR/subtype claims). However, the exact reported numbers are not derivable from shipped artifacts: the repo ships no data, no coefficients and no set.seed(), the raw bundle is login-walled, model.R uses a post-hoc AUC>0.68 acceptance filter (optimistic bias), and the text says LASSO λ=15 while the code uses lambda.min. So the deviation sits mainly on data availability and authors'-side reproducibility defects, compounded by our self-chosen cohort/refit — severity is moderate (direction/magnitude hold) but derivability fails outright.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.