Corpus 1,286 assessed · 1,187 scored · 648 reproduced ≥75 · 174 flagged ·∅ 73.9/100
← New search

Identification and validation of a metabolic-related gene risk model predicting the prognosis of lung, colon, and breast cancers.

Sci Rep · 2025
L1 74/100 PQI 91
Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +7
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
What did not (or only partly)
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
74/100
Reproducibility score
at the mean
vs. all fields · 1187 studies
🎯 Scores higher than 44% of all assessed papers rank 649 of 1187 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough for the DEG step -> reproduced 1:1 for the primary breast dataset. Reproduced the differential-expression up/down gene counts via the paper's described tool (GEO2R = GEOquery+limma) on the 3 public GEO datasets (GSE42568/GSE21510/GSE18842) on «our HPC». Tumor/normal sample splits matched the paper EXACTLY (104/17, 123/25, 46/45), confirming correct data+grouping. With gene-symbol collapse + GEO2R quantile normalization, breast GSE42568 reproduces essentially 1:1 (up 2082->2117 +1.7%, down 2119->2249 +6.1%, total 9652->9971 +3.3%); lung and colon down/total within 3-12%. One flagged anomaly: colon GSE21510 reported up=1594 (< its down=2390) is not reproducible (we get up3165>down~2285) and is the single value not derivable under the standard pipeline -> flagged for human review (possible up/down swap or undocumented filter), not asserted as fabrication. NOT attempted (hard 20% / out of scope): the repo's own ML survival-regression metrics (KNN/SVR/XGB MSE/MAE/MAPE) because the input TCGA 46-gene matrices are not shipped (docs_insufficient for inputs); GEPIA2 and STRING/Cytoscape PPI steps (external interactive web tools); qRT-PCR (wet-lab).

💻 Code ↗ 🗄 Data: GSE42568

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 74
    assessed: 2026-06-14 ⛓ 61d0ae1cac2a
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-14
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-09-19

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The study tests whether metabolic reprogramming in breast, colorectal, and lung cancers follows a shared/common pattern, aiming to identify metabolism-related genes (MRGs) common to all three cancers that could serve as universal early diagnostic and prognostic biomarkers.

Core claims
  • 540 DEGs overlap across breast, colorectal, and lung cancers out of 11,384 DEGs analyzed finding
  • 46 metabolism-related genes (MRGs) are shared across all three studied cancer types finding
  • 20 key/hub MRGs were identified via PPI network analysis common to all three cancers finding
  • 11 key MRGs are prognostically significant for overall survival finding
  • qRT-PCR validation of key MRGs in cancer cell lines confirmed expression profiles, with some genes showing cell-type-specific patterns finding
  • Support vector regressor (SVR) exhibited remarkable accuracy in predicting overall survival using the identified MRGs finding
  • Integrated bioinformatics pipeline (GEO2R, WebGestalt, STRING, Cytoscape/CytoHubba, GEPIA2) used to identify and characterize key/hub MRGs method
  • Three machine learning regression algorithms (KNN, SVM/SVR, XGBoost) applied to predict overall survival time from MRG expression and clinical data method
Experimental setups
Assay System Perturbation Readout Platform
microarray gene expression / GEO2R DEG analysis human breast cancer tissue (GSE42568: 104 tumor, 17 adjacent non-tumor) none differentially expressed genes (adjust-pValue ≤0.05, log2FC ≥1 or ≤-1) GEO2R (R 4.2.2), GEOquery 2.66.0, limma 3.54.0
microarray gene expression / GEO2R DEG analysis human colorectal cancer tissue (GSE21510: 123 tumor, 25 adjacent non-tumor) none differentially expressed genes GEO2R
microarray gene expression / GEO2R DEG analysis human lung cancer tissue (GSE18842: 46 tumor, 45 non-tumor) none differentially expressed genes GEO2R
GO/KEGG/PANTHER functional enrichment analysis common DEGs from BRC, CRC, LUC datasets none enriched biological processes and pathways WebGestalt (W130-137)
protein-protein interaction (PPI) network analysis MRGs intersected with common DEGs across BRC, CRC, LUC none hub gene ranking via topological measures (e.g., closeness) STRING; Cytoscape 3.7.2 with CytoHubba plugin
gene expression and Kaplan-Meier survival analysis TCGA tumor vs normal tissue (BRC, CRC, LUC) none key MRG expression levels and overall survival (high-risk vs low-risk groups) GEPIA2
quantitative real-time PCR (qRT-PCR) cancer cell lines (HCT116-CRC, MDA-MB231-BRC, H1299-LUC) vs normal fibroblast cell line VH10 none relative expression (fold change) of 20 key/hub MRGs vs β-actin SYBR green PCR Core reagents (Promega); 2^-ΔΔCT method
machine learning regression (KNN, SVR, XGBoost) TCGA-BRCA (1215), TCGA-LUNG (1077), TCGA-COAD (430) patients with RNA-seq and clinical data none predicted overall survival time from 46 MRG expression profiles scikit-learn v1.5.2; XGBoost v2.0.3
Key results
  • 540 DEGs overlapped across BRC, CRC, and LUC out of 11,384 DEGs analyzed
  • 46 MRGs identified as involved in all three studied cancer types
  • 20 key/hub MRGs identified from the PPI network common to all cancers
  • 11 of the key MRGs were prognostically significant
  • qRT-PCR validation confirmed expression profiles of key MRGs, with some genes showing cell-type-specific expression patterns
  • SVR showed remarkable accuracy in predicting overall survival, outperforming other models
  • Missing OS values were imputed via KNN (k=3) for 1 BRC and 15 LUNG clinical samples k=3
  • Final ML datasets comprised 1215 BRC, 1077 LUC, and 430 CRC patient samples with matched gene expression and survival data
Key statistics
  • count 11,384 DEGs analyzed (total DEGs across all three cancer datasets)
  • count 540 common DEGs (DEGs overlapping across BRC, CRC, and LUC)
  • count 1028 MRGs screened (metabolism-related genes from MSigDB central metabolic pathways)
  • count 46 MRGs (MRGs common to all three cancer types)
  • count 20 key/hub MRGs (hub genes identified via PPI network/CytoHubba)
  • count 11 prognostically significant key MRGs (key MRGs with prognostic significance for overall survival)
  • pvalue adjust-pValue ≤ 0.05, log2FC ≥1 (up) or ≤-1 (down) (DEG significance thresholds via GEO2R/limma)
  • count BRC 104 tumor/17 normal; CRC 123 tumor/25 normal; LUC 46 tumor/45 normal (sample sizes of GEO tumor vs non-tumor tissue datasets)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The study identified differentially expressed genes (DEGs) from three GEO microarray datasets (BRC, CRC, LUC) using limma with Benjamini-Hochberg FDR correction, then intersected overlapping DEGs with MSigDB metabolic gene sets and built PPI networks in STRING/CytoHubba to derive hub genes. Prognostic significance of hub genes was assessed via Kaplan-Meier log-rank tests and Cox proportional hazard models in GEPIA2 against TCGA survival data. Cell-line expression was validated by qRT-PCR with Student's t-test, and three machine learning regressors (KNN, SVR, XGBoost) were benchmarked for overall survival prediction using 80/20 train-test splits repeated ten times with different random seeds.

Replicationmixed Sample sizeGEO and TCGA sample sizes stated per dataset and cancer type; no formal power calculation or sample size justification described Groupstumor vs adjacent non-tumorous tissue (DEG); high-risk vs low-risk expression groups (survival); cancer cell lines vs normal fibroblast cell line (qRT-PCR) Pairingunpaired Randomization/blindingnot stated Dispersionunclear Exact p-valuesno Effect sizesyes Confidence intervalsyes Multiplicity correctionBenjamini-Hochberg FDR
Statistical tests used
Test Applied to n Assumptions
limma moderated t-test with Benjamini-Hochberg FDR adjustment DEG identification: tumor vs adjacent non-tumorous tissue in BRC (GSE42568), CRC (GSE21510), and LUC (GSE18842) BRC: 104 tumor + 17 normal; CRC: 123 tumor + 25 normal; LUC: 46 tumor + 45 normal not stated
Log-rank test Kaplan-Meier overall survival comparison of high-risk vs low-risk expression groups for hub MRGs via GEPIA2/TCGA TCGA BRCA n=1215, LUNG n=1077, COAD n=430 not stated
Cox proportional hazard model Overall survival modeling with 95% CI for key MRGs via GEPIA2/TCGA TCGA BRCA n=1215, LUNG n=1077, COAD n=430 not stated
Two-sample Student's t-test qRT-PCR expression comparison of cancer cell lines (HCT116, MDA-MB231, H1299) vs normal cell line (VH10); analyzed in GraphPad Prism v9 not stated
Over-representation / enrichment analysis with adjust-pValue correction (WebGestalt) GO (BP/MF/CC) and KEGG pathway enrichment of DEGs not stated
Approaches that could also have been used
  • Student's t-test was used to compare qRT-PCR expression between cancer and normal cell lines, with n not reported
    Could also: A non-parametric Mann-Whitney U (Wilcoxon rank-sum) test could also be applied — When biological replicates are few, as is common in cell-line experiments, normality cannot be reliably assessed; a rank-based test makes no distributional assumption and is robust under small n
  • Patients were dichotomized into high-risk and low-risk groups by expression level for Kaplan-Meier survival analysis
    Could also: Continuous Cox proportional hazard regression incorporating expression as a continuous covariate could also be used — Continuous modeling avoids information loss from arbitrary median dichotomization, provides a hazard ratio per unit change in expression, and retains statistical power
  • No multiplicity correction was described for the survival analyses conducted across 11 (or 20) hub genes
    Could also: A Benjamini-Hochberg FDR or Bonferroni correction applied across the family of survival tests could also be used — Testing multiple genes for prognostic significance inflates the probability of at least one false discovery; a familywise correction makes the degree of control explicit and is consistent with the correction already applied in the DEG step
  • ML model performance was estimated using a single 80/20 train-test split repeated ten times with different random seeds
    Could also: Stratified k-fold cross-validation (e.g., 5- or 10-fold) could also be used — k-fold CV uses all samples for both training and validation across folds, yielding a lower-variance performance estimate and a natural standard deviation across folds that quantifies uncertainty
  • Default hyperparameters were used for all three ML algorithms without tuning
    Could also: Cross-validated hyperparameter search (grid search or random search with a held-out validation set) could also be applied — Default settings are not optimized for any particular dataset; systematic tuning can improve predictive performance and also helps distinguish the effect of gene selection from the effect of model configuration
  • Missing overall survival values were handled with single KNN imputation (k=3) prior to ML modeling
    Could also: Multiple imputation by chained equations (MICE) could also be used to address missing survival data — Multiple imputation generates several plausible completed datasets and pools results, explicitly propagating the uncertainty introduced by imputation into downstream estimates, whereas single imputation treats imputed values as if they were observed
Software: R / GEO2R (GEOquery + limma) R 4.2.2; GEOquery 2.66.0; limma 3.54.0 · WebGestalt W130-137 · Cytoscape / CytoHubba 3.7.2 · GraphPad Prism 9 · scikit-learn 1.5.2 · XGBoost 2.0.3 · GEPIA2 · STRING · Venny 2.1

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
2
Impact: low
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

GSE18842 GEO in Methods (http://purl.org/orb/Methods)
also used by 1 paper:
GSE21510 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE42568 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-39779736

Paper: Khan et al. 2025, Sci Rep — "Identification and validation of a metabolic-related gene risk model predicting the prognosis of lung, colon, and breast cancers." DOI 10.1038/s41598-025-85366-8. PMCID PMC11711664.

Repo: https://github.com/jiyauddin0786/metabolism_related_genes commit 8494aa8138ca86cabe61bcf45ddfd3a3e3e9f818 (branch main). Contents: a single notebook 01_survival.ipynb (ML survival-time regression on TCGA) + precomputed per-seed metric outputs. No DEG / GEPIA / PPI code — those steps were run on external web tools (GEO2R, GEPIA2, STRING/Cytoscape).

Reported pipeline steps and reproducibility classification

# Reported result Pipeline / tool In scope? Why
1 DEG up/down counts per cancer (GSE42568, GSE21510, GSE18842), thresholds adj.P≤0.05 & |log2FC|≥1, BH-FDR GEO2R = GEOquery + limma YES (primary) Tool & thresholds explicit; data fully public; third-party tool on paper's data is valid (BRIEF P16). Deterministic, crisp integer claims.
2 540 common DEGs across 3 cancers; 46 common with 1028 MSigDB MRGs; 20 PPI hub genes set intersection + STRING/Cytoscape closeness partial intersection derivable once #1 reproduced; PPI hub selection = external Cytoscape (web), not attempted
3 GEPIA2 expression up/down validation on TCGA; KM overall-survival per gene GEPIA2 web server NO external interactive web tool, no shipped code/params; out of scope
4 ML survival-time regression metrics (KNN/SVR/XGB) MSE/MAE/MAPE on TCGA 46-gene matrices repo 01_survival.ipynb (scikit-learn/xgboost) attempted/deprioritised authors' OWN code, BUT input matrices (brc_raw,luc_raw,crc_raw,*_survival.txt) are NOT shipped (lived on author Google Drive); reconstructing exact TCGA/Xena matrices = the hard 20%, exact MSE match improbable → docs_insufficient for inputs
5 qRT-PCR validation in cell lines wet-lab NO non-pipeline
6 MRG curation (1028 genes from MSigDB) manual GSEA/MSigDB NO manual/external

Decision (80/20)

  • Reproduce result #1 — DEG up/down counts for the three GEO datasets via a GEO2R-faithful limma script (GEOquery + limma), run on «our HPC». Primary target is GSE42568 (the RU's named accession, breast); colon (GSE21510) and lung (GSE18842) added because the same lightweight pipeline covers them.
  • Skip the hard 20%: ML metrics (#4) require unshipped input matrices; PPI hub selection (#2) and GEPIA2 (#3) are external interactive tools. Documented, not fabricated.

Reported values to compare (from full text, Results)

  • GSE42568 (breast): 2082 up / 2119 down
  • GSE21510 (colon): 1594 up / 2390 down
  • GSE18842 (lung): 1415 up / 1784 down
  • (also "total DEGs": BRC 9652, CRC 20408→ note: paper's "total" wording is ambiguous and does not equal up+down; we compare the well-defined up/down counts which carry explicit thresholds.)
deg_brca_up
Reported
2082
Reproduced
2117
within tolerance
deg_brca_down
Reported
2119
Reproduced
2249
within tolerance
deg_brca_total
Reported
9652
Reproduced
9971
within tolerance
deg_luc_up
Reported
1415
Reproduced
1564
partial
deg_luc_down
Reported
1784
Reproduced
1862
within tolerance
deg_crc_down
Reported
2390
Reproduced
2285
within tolerance
deg_crc_total
Reported
20408
Reproduced
19806
within tolerance
deg_crc_up
Reported
1594
Reproduced
3165
did not match
group_splits_all3
Reported
104/17,123/25,46/45
Reproduced
104/17,123/25,46/45
exact

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 74/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +7

For the reproduced DEG step the data and grouping are verified exact (104/17, 123/25, 46/45) and the breast GSE42568 dataset reproduces essentially 1:1 (up 2082→2117, down 2119→2249, total 9652→9971), with lung/colon down & total within ~3–11% — explainable by our self-chosen GEO2R quantile-norm + gene-symbol collapse. The one substantive anomaly is colon GSE21510 up=1594, which is not derivable (we get ~3165, +98.6%) while its down/total reproduce within 4%, suggesting an authors-side up/down swap or undocumented filter — flagged, not asserted as fabrication. The paper's central claim (a metabolic-gene survival risk model, e.g. SVR MSEs 0.776/1.168/1.061) could not be tested because its TCGA input matrices were never deposited, so confirmation is limited to the upstream DEG step. Overall: solid with explainable deviations plus one flagged non-derivable value → yellow.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at [email protected].

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

124.9 k
tokens (I/O) · 8.7 M incl. cache
15 min
runtime · 0.03 CPU-h
1.6 GB
peak RAM
3
HPC jobs
hummel
machine