Identification and validation of a metabolic-related gene risk model predicting the prognosis of lung, colon, and breast cancers.
The main results reproduced, with only marginal, non-material deviations.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Same input data as the authors
- ✓Reported values were directly comparable
- 🟡A deviation arose in the data or preprocessing
- 🟡A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The deviation was non-trivial in magnitude
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Described well enough for the DEG step -> reproduced 1:1 for the primary breast dataset. Reproduced the differential-expression up/down gene counts via the paper's described tool (GEO2R = GEOquery+limma) on the 3 public GEO datasets (GSE42568/GSE21510/GSE18842) on «our HPC». Tumor/normal sample splits matched the paper EXACTLY (104/17, 123/25, 46/45), confirming correct data+grouping. With gene-symbol collapse + GEO2R quantile normalization, breast GSE42568 reproduces essentially 1:1 (up 2082->2117 +1.7%, down 2119->2249 +6.1%, total 9652->9971 +3.3%); lung and colon down/total within 3-12%. One flagged anomaly: colon GSE21510 reported up=1594 (< its down=2390) is not reproducible (we get up3165>down~2285) and is the single value not derivable under the standard pipeline -> flagged for human review (possible up/down swap or undocumented filter), not asserted as fabrication. NOT attempted (hard 20% / out of scope): the repo's own ML survival-regression metrics (KNN/SVR/XGB MSE/MAE/MAPE) because the input TCGA 46-gene matrices are not shipped (docs_insufficient for inputs); GEPIA2 and STRING/Cytoscape PPI steps (external interactive web tools); qRT-PCR (wet-lab).
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 74assessed: 2026-06-14 ⛓ 61d0ae1cac2a
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-14
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusThe study investigates whether metabolic reprogramming follows shared molecular gene signatures across breast (BRC), colorectal (CRC), and lung (LUC) cancers, aiming to identify common metabolism-related genes (MRGs) usable as universal diagnostic/prognostic biomarkers and to test machine learning for predicting overall survival.
- ★ 540 DEGs are shared across BRC, CRC, and LUC, yielding 46 metabolism-related genes (MRGs) and 20 key/hub MRGs common to all three cancers. finding
- ★ Of the 20 key/hub MRGs, 11 are prognostically significant for overall survival. finding
- ★ Support vector regression (SVR) predicts overall patient survival with remarkable accuracy, supporting clinical utility of MRG-based ML models. finding
- ★ qRT-PCR validation in cancer cell lines confirmed key MRG expression profiles, with some showing cell-type-specific patterns. finding
- ★ Shared MRGs represent potential biomarkers for early detection, prognosis, and metabolic-targeted therapy across BRC, CRC, and LUC. resource
- An integrated pipeline (GEO/TCGA mining, GO/KEGG, GSEA-MSigDB, STRING/CytoHubba PPI, GEPIA2, ML, qRT-PCR) identifies and validates key/hub MRGs. method
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Bulk gene expression microarray DEG analysis (GEO2R) | Human BRC tumor vs adjacent non-tumor tissue (GSE42568) | none | Differentially expressed genes (log2FC, adjusted p-value) | GEO2R R 4.2.2; GEOquery 2.66.0; limma 3.54.0 |
| Bulk gene expression microarray DEG analysis (GEO2R) | Human CRC tumor vs adjacent non-tumor tissue (GSE21510) | none | Differentially expressed genes (log2FC, adjusted p-value) | GEO2R R 4.2.2; GEOquery; limma |
| Bulk gene expression microarray DEG analysis (GEO2R) | Human LUC tumor vs non-tumor tissue (GSE18842) | none | Differentially expressed genes (log2FC, adjusted p-value) | GEO2R R 4.2.2; GEOquery; limma |
| PPI network and hub gene analysis | Common MRGs across BRC/CRC/LUC | none | Hub/key MRGs by closeness centrality | STRING; CytoHubba in Cytoscape 3.7.2 |
| Gene expression and Kaplan-Meier survival analysis (GEPIA2) | TCGA BRC, CRC, LUC tissues vs normal | none | TPM expression; overall survival in high- vs low-risk groups | GEPIA2; TCGA |
| qRT-PCR (2^-ΔΔCT) | Cancer cell lines HCT116 (CRC), MDA-MB231 (BRC), H1299/NCI-H1299 (LUC) vs normal VH10 fibroblasts | none | Fold change of 20 key/hub MRG transcripts; β-actin control | SYBR green PCR Core reagents (Promega); cDNA Synthesis Kit (Thermo); Qubit RNA BR (Invitrogen) |
| Machine learning regression for overall survival prediction (KNN, SVR, XGBoost) | TCGA-BRCA, TCGA-LUNG, TCGA-COAD RNA-Seq + clinical OS data | none | Predicted overall survival time using 46 MRGs | scikit-learn v1.5.2; XGBoost v2.0.3; Illumina HiSeq 2000 |
- – 540 DEGs overlapped across BRC, CRC, and LUC out of 11,384 DEGs analyzed 540 of 11,384
- – 46 MRGs and 20 key/hub MRGs identified common to all three cancer types 46 MRGs; 20 hub MRGs
- – 11 key MRGs were prognostically significant 11 of 20
- – SVR achieved remarkable accuracy in predicting overall survival
- – qRT-PCR confirmed key MRG expression profiles, some cell-type-specific
- – 1028 MRGs of central metabolic pathways screened from MSigDB/GSEA before intersection 1028
- count 11,384 DEGs analyzed; 540 overlapping (Total DEGs and common DEGs across BRC, CRC, LUC)
- count BRC GSE42568: 104 tumor, 17 adjacent non-tumor (BRC microarray sample sizes)
- count CRC GSE21510: 123 tumor, 25 adjacent non-tumor (CRC microarray sample sizes)
- count LUC GSE18842: 46 tumor, 45 non-tumor (LUC microarray sample sizes)
- other adjusted p-value ≤ 0.05, log2FC ≥ 1 (up) or ≤ -1 (down) (DEG significance thresholds (Benjamini-Hochberg FDR))
- count Final ML datasets: BRC 1215, LUC 1077, CRC/COAD 430 samples (Preprocessed TCGA samples for ML; 80/20 train/test, 10 repeats)
- count 1 BRCA and 15 LUNG samples missing OS imputed via KNN (k=3) (Missing overall survival value handling)
- pvalue p ≤ 0.05 (Statistical significance cutoff for qRT-PCR (Student's t-test) and survival analysis)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
The study identified differentially expressed genes (DEGs) from three GEO microarray datasets (BRC, CRC, LUC) using limma with Benjamini-Hochberg FDR correction, then intersected overlapping DEGs with MSigDB metabolic gene sets and built PPI networks in STRING/CytoHubba to derive hub genes. Prognostic significance of hub genes was assessed via Kaplan-Meier log-rank tests and Cox proportional hazard models in GEPIA2 against TCGA survival data. Cell-line expression was validated by qRT-PCR with Student's t-test, and three machine learning regressors (KNN, SVR, XGBoost) were benchmarked for overall survival prediction using 80/20 train-test splits repeated ten times with different random seeds.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| limma moderated t-test with Benjamini-Hochberg FDR adjustment | DEG identification: tumor vs adjacent non-tumorous tissue in BRC (GSE42568), CRC (GSE21510), and LUC (GSE18842) | BRC: 104 tumor + 17 normal; CRC: 123 tumor + 25 normal; LUC: 46 tumor + 45 normal | not stated |
| Log-rank test | Kaplan-Meier overall survival comparison of high-risk vs low-risk expression groups for hub MRGs via GEPIA2/TCGA | TCGA BRCA n=1215, LUNG n=1077, COAD n=430 | not stated |
| Cox proportional hazard model | Overall survival modeling with 95% CI for key MRGs via GEPIA2/TCGA | TCGA BRCA n=1215, LUNG n=1077, COAD n=430 | not stated |
| Two-sample Student's t-test | qRT-PCR expression comparison of cancer cell lines (HCT116, MDA-MB231, H1299) vs normal cell line (VH10); analyzed in GraphPad Prism v9 | — | not stated |
| Over-representation / enrichment analysis with adjust-pValue correction (WebGestalt) | GO (BP/MF/CC) and KEGG pathway enrichment of DEGs | — | not stated |
-
Student's t-test was used to compare qRT-PCR expression between cancer and normal cell lines, with n not reported↳ Could also: A non-parametric Mann-Whitney U (Wilcoxon rank-sum) test could also be applied — When biological replicates are few, as is common in cell-line experiments, normality cannot be reliably assessed; a rank-based test makes no distributional assumption and is robust under small n
-
Patients were dichotomized into high-risk and low-risk groups by expression level for Kaplan-Meier survival analysis↳ Could also: Continuous Cox proportional hazard regression incorporating expression as a continuous covariate could also be used — Continuous modeling avoids information loss from arbitrary median dichotomization, provides a hazard ratio per unit change in expression, and retains statistical power
-
No multiplicity correction was described for the survival analyses conducted across 11 (or 20) hub genes↳ Could also: A Benjamini-Hochberg FDR or Bonferroni correction applied across the family of survival tests could also be used — Testing multiple genes for prognostic significance inflates the probability of at least one false discovery; a familywise correction makes the degree of control explicit and is consistent with the correction already applied in the DEG step
-
ML model performance was estimated using a single 80/20 train-test split repeated ten times with different random seeds↳ Could also: Stratified k-fold cross-validation (e.g., 5- or 10-fold) could also be used — k-fold CV uses all samples for both training and validation across folds, yielding a lower-variance performance estimate and a natural standard deviation across folds that quantifies uncertainty
-
Default hyperparameters were used for all three ML algorithms without tuning↳ Could also: Cross-validated hyperparameter search (grid search or random search with a held-out validation set) could also be applied — Default settings are not optimized for any particular dataset; systematic tuning can improve predictive performance and also helps distinguish the effect of gene selection from the effect of model configuration
-
Missing overall survival values were handled with single KNN imputation (k=3) prior to ML modeling↳ Could also: Multiple imputation by chained equations (MICE) could also be used to address missing survival data — Multiple imputation generates several plausible completed datasets and pools results, explicitly propagating the uncertainty introduced by imputation into downstream estimates, whereas single imputation treats imputed values as if they were observed
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-39779736
Paper: Khan et al. 2025, Sci Rep — "Identification and validation of a metabolic-related gene risk model predicting the prognosis of lung, colon, and breast cancers." DOI 10.1038/s41598-025-85366-8. PMCID PMC11711664.
Repo: https://github.com/jiyauddin0786/metabolism_related_genes
commit 8494aa8138ca86cabe61bcf45ddfd3a3e3e9f818 (branch main).
Contents: a single notebook 01_survival.ipynb (ML survival-time regression on
TCGA) + precomputed per-seed metric outputs. No DEG / GEPIA / PPI code — those
steps were run on external web tools (GEO2R, GEPIA2, STRING/Cytoscape).
Reported pipeline steps and reproducibility classification
| # | Reported result | Pipeline / tool | In scope? | Why |
|---|---|---|---|---|
| 1 | DEG up/down counts per cancer (GSE42568, GSE21510, GSE18842), thresholds adj.P≤0.05 & |log2FC|≥1, BH-FDR | GEO2R = GEOquery + limma | YES (primary) | Tool & thresholds explicit; data fully public; third-party tool on paper's data is valid (BRIEF P16). Deterministic, crisp integer claims. |
| 2 | 540 common DEGs across 3 cancers; 46 common with 1028 MSigDB MRGs; 20 PPI hub genes | set intersection + STRING/Cytoscape closeness | partial | intersection derivable once #1 reproduced; PPI hub selection = external Cytoscape (web), not attempted |
| 3 | GEPIA2 expression up/down validation on TCGA; KM overall-survival per gene | GEPIA2 web server | NO | external interactive web tool, no shipped code/params; out of scope |
| 4 | ML survival-time regression metrics (KNN/SVR/XGB) MSE/MAE/MAPE on TCGA 46-gene matrices | repo 01_survival.ipynb (scikit-learn/xgboost) |
attempted/deprioritised | authors' OWN code, BUT input matrices (brc_raw,luc_raw,crc_raw,*_survival.txt) are NOT shipped (lived on author Google Drive); reconstructing exact TCGA/Xena matrices = the hard 20%, exact MSE match improbable → docs_insufficient for inputs |
| 5 | qRT-PCR validation in cell lines | wet-lab | NO | non-pipeline |
| 6 | MRG curation (1028 genes from MSigDB) | manual GSEA/MSigDB | NO | manual/external |
Decision (80/20)
- Reproduce result #1 — DEG up/down counts for the three GEO datasets via a GEO2R-faithful limma script (GEOquery + limma), run on «our HPC». Primary target is GSE42568 (the RU's named accession, breast); colon (GSE21510) and lung (GSE18842) added because the same lightweight pipeline covers them.
- Skip the hard 20%: ML metrics (#4) require unshipped input matrices; PPI hub selection (#2) and GEPIA2 (#3) are external interactive tools. Documented, not fabricated.
Reported values to compare (from full text, Results)
- GSE42568 (breast): 2082 up / 2119 down
- GSE21510 (colon): 1594 up / 2390 down
- GSE18842 (lung): 1415 up / 1784 down
- (also "total DEGs": BRC 9652, CRC 20408→ note: paper's "total" wording is ambiguous and does not equal up+down; we compare the well-defined up/down counts which carry explicit thresholds.)
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
For the reproduced DEG step the data and grouping are verified exact (104/17, 123/25, 46/45) and the breast GSE42568 dataset reproduces essentially 1:1 (up 2082→2117, down 2119→2249, total 9652→9971), with lung/colon down & total within ~3–11% — explainable by our self-chosen GEO2R quantile-norm + gene-symbol collapse. The one substantive anomaly is colon GSE21510 up=1594, which is not derivable (we get ~3165, +98.6%) while its down/total reproduce within 4%, suggesting an authors-side up/down swap or undocumented filter — flagged, not asserted as fabrication. The paper's central claim (a metabolic-gene survival risk model, e.g. SVR MSEs 0.776/1.168/1.061) could not be tested because its TCGA input matrices were never deposited, so confirmation is limited to the upstream DEG step. Overall: solid with explainable deviations plus one flagged non-derivable value → yellow.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.