Elucidating the Prognostic and Therapeutic Implications of Insulin Resistance Genes in Breast Cancer: A Machine Learning-Powered Analysis.
The main results reproduced: recomputed values matched the published ones within tolerance.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Same input data as the authors
- ✓No authors-side cause for any deviation
- ✓Any deviation was negligible
- ✓The central claim held under reproduction
- 🟡Reported values were only indirectly comparable
- 🟡A deviation arose in the data or preprocessing
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Described well enough for a PARTIAL 1:1 on the one clearly-specified, data-tied result. The paper ships NO authors' code (the brief's GitHub link is ggplot2) and the brief's designated dataset GSE20685 is a validation cohort, not the training set. The paper's reproducible end product is a fully-specified 7-gene IRRS risk-score formula. We applied it verbatim to GSE20685 (327 samples, GPL570) on «our HPC» and re-ran the Fig 1G survival analysis: the signature significantly stratifies OS in the reported direction (HR=1.68, high=worse; C-index 0.60), with log-rank p=0.020 vs the reported 0.012 -- same order of magnitude, both significant, conclusion intact. Graded within-tol; the exact-p gap is expected (TCGA-fit coefficients applied to microarray log2 units; unspecified probe-collapse). NOT attempted (out of scope, the hard 20%): GeneCards IRG retrieval (non-versioned -> non-reproducible, flagged), LASSO-Cox feature selection on TCGA (no code/seed), XGBoost/SVM classifiers, immune deconvolution, enrichment, and scRNA -- none have shipped code and none are tied to the brief's GSE20685. No fabrication of results to avoid a drop; the central prognostic claim genuinely reproduces.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 84assessed: 2026-06-14 ⛓ eef212420178
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-14
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15no human curator yet
- Last updated
- 2026-09-19
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetAlthough insulin resistance (IR) has been linked to tumorigenesis and cancer progression, its prognostic relevance in breast cancer has not been fully elucidated, so the study tests whether insulin resistance-related genes (IRGs) can be used to build a robust prognostic signature for breast cancer patients.
- ★ A seven-gene IRG prognostic signature (LIFR, EZR, TBC1D4, NSF, RPL5, SAA1, PGK1) predicts overall survival in breast cancer across training and four validation cohorts finding
- ★ The insulin resistance risk score (IRRS) formula combining the seven hub gene expressions weighted by LASSO coefficients method
- ★ Higher IRRS is significantly associated with adverse clinical parameters including PAM50 subtype, tumor stage, tumor size, grade, and lymph node involvement finding
- ★ A nomogram combining IRRS and PAM50 subtype provides individualized survival prediction with IRRS contributing more than PAM50 method
- ★ Tumor microenvironment composition (tumor purity, immune/stromal infiltration) differs between high- and low-IRRS groups finding
- ★ Hub genes are enriched in T cells, B cells, and epithelial cells at single-cell resolution finding
- ★ Machine learning models (XGBoost and SVM) built on the IRG feature set classify patients with external validation in METABRIC and GSE96058 method
- The IRRS-based signature outperforms several previously published breast cancer prognostic biomarkers by C-index finding
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| bulk RNA-seq/microarray differential expression (limma) | TCGA-BRCA tumor vs normal breast tissue | none | differentially expressed genes (log2FC, adjusted p) | R package limma v4.1.0 |
| univariate/multivariate Cox regression + LASSO Cox regression | TCGA-BRCA and METABRIC breast cancer patients | none | hub IRGs with non-zero coefficients, hazard ratios | R packages survival, glmnet v4.1.8 |
| Kaplan-Meier survival analysis (log-rank test) | TCGA-BRCA, METABRIC, GSE96058, GSE20685, GSE7390 breast cancer cohorts | none (stratified by IRRS) | overall survival, DSS, DFI, PFI, PFS | R packages survival, survminer |
| tumor microenvironment deconvolution (ESTIMATE, CIBERSORT) | TCGA-BRCA breast cancer samples | none (stratified by high/low IRRS) | tumor purity, immune score, stromal score, 22 immune cell type abundances | R packages ESTIMATE v1.0.13, CIBERSORT |
| GSEA / GO / KEGG functional enrichment | TCGA-BRCA high- vs low-IRRS groups | none | enriched pathways/gene sets | R package clusterProfiler v4.2.2 |
| single-cell RNA sequencing (UMAP, cell type annotation, UCell scoring) | 5 triple-negative breast cancer tumor samples (SCP1106) | none | hub gene enrichment per cell type, UCell signature score | R packages Seurat v4.1, UCell v2.0 |
| machine learning classification (XGBoost, SVM) | TCGA-BRCA train/test split; external validation in METABRIC and GSE96058 | none | patient risk group classification, ROC-AUC | R packages xgboost, e1071 |
| immunohistochemistry (IHC) | 10 FFPE breast cancer patient tissue sections (ER+, HER2+, TNBC) | none | protein expression of CD163, CD8α, PGK1 | Pannoramic 250 Flash III scanner |
- – LASSO Cox regression of 24 overlapping IRGs yielded a 7-gene prognostic signature (LIFR, EZR, TBC1D4, NSF, RPL5, SAA1, PGK1)
- ▲ Combined IRRS showed higher hazard ratio for BC risk than any individual gene HR = 4.368, 95% CI 2.810-6.792
- ▼ High-IRRS group had significantly reduced overall survival across training and all four validation cohorts
- ▼ High IRRS significantly associated with poorer DSS, DFI, PFI, and PFS in TCGA-BRCA
- – Nomogram combining IRRS and PAM50 subtype predicted 1-, 3-, 5-year OS with good calibration (except 10-year) C-index 0.64 (TCGA-BRCA), 0.59 (METABRIC)
- ▼ Tumor purity was significantly lower in the low-IRRS group compared to high-IRRS group
- – Hub genes were more enriched in T cells, B cells, and epithelial cells at single-cell level
- ▲ IRRS-based signature outperformed previously published BC biomarkers based on C-index comparison
- count 6907 BC samples across 10 cohorts (total compiled dataset after excluding patients lacking OS data)
- count 1115 significant IRGs from 2828 candidate IRGs and 6324 DEGs (differential expression analysis TCGA-BRCA tumor (1095) vs normal (113))
- count 54 and 290 significant IRGs in TCGA-BRCA and METABRIC respectively; 24 overlapping (univariate Cox regression results, common IRGs used for LASSO)
- fold_change HR = 4.368, 95% CI: 2.810-6.792, p < 0.0001 (combined IRRS hazard ratio for BC prognosis)
- pvalue TCGA p<0.0001; METABRIC p<0.048; GSE20685 p=0.012; GSE96058 p=1.8x10^-11; GSE7390 p=0.027 (log-rank test OS comparison high vs low IRRS across cohorts)
- pvalue DSS p=1.4x10^-4; DFI p=5.8x10^-3; PFI p=8.1x10^-5; PFS p<8.1x10^-0.5 (log-rank test additional survival endpoints in TCGA-BRCA)
- other C-index = 0.64 (TCGA-BRCA), 0.59 (METABRIC) (nomogram predictive performance)
- count 307 patients across 4 neoadjuvant therapy cohorts (GSE191127, GSE20181, GSE18728, GSE225078) (cohorts used to evaluate predictive capacity for therapeutic response)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
The paper developed an insulin resistance risk score (IRRS) for breast cancer prognosis by filtering differentially expressed genes via limma, then applying univariate Cox regression and LASSO Cox regression to select a seven-gene signature from TCGA-BRCA. Kaplan–Meier curves with log-rank tests were used to compare overall survival between median-split risk groups across one training and four validation cohorts; XGBoost and SVM classifiers (5-fold cross-validated grid search, ROC-AUC optimized) provided an independent machine learning validation layer. Results were reported as hazard ratios with 95% confidence intervals, exact log-rank p-values, C-index values, and ROC-AUC scores.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| limma moderated t-statistic (linear model) | Differential expression between 1095 BC and 113 normal samples in TCGA-BRCA; threshold |log2FC| > 0.5 and adjusted p < 0.01 | 1095 BC + 113 normal | not stated |
| Univariate Cox proportional hazards regression | Prognostic IRG identification separately in TCGA-BRCA (p < 0.01) and METABRIC (p < 0.01); 54 and 290 significant IRGs identified, respectively | 1095 (TCGA-BRCA); METABRIC n not stated for this step | not stated |
| LASSO Cox regression (penalized) | Feature selection from 24 IRGs overlapping between TCGA-BRCA and METABRIC univariate screens, yielding 7-gene signature | 1095 (TCGA-BRCA training cohort) | not stated |
| Log-rank test with Kaplan–Meier survival curves | OS comparison between high- and low-IRRS groups in training cohort and four validation cohorts; also DSS, DFI, PFI, PFS in TCGA-BRCA; subtype-stratified OS in supplementary | 1095 training; ~5500 validation (METABRIC + GSE96058 + GSE20685 + GSE7390 combined) | not stated |
| Univariate and multivariate Cox proportional hazards regression | Variable selection for nomogram from IRRS, T stage, M stage, overall stage, PAM50; final model retains IRRS and PAM50 (p < 0.05) | TCGA-BRCA (1095); METABRIC (n not stated for this step) | not stated |
| Student's t-test (unpaired, two-sample; paired status not stated) | IRG expression differences between normal and tumor tissues; IRRS changes before and after neoadjuvant therapy | not stated for these specific comparisons | not stated |
| Wilcoxon rank-sum test | Association between IRRS and clinicopathological features: PAM50 subtypes, AJCC T and N stages, lymph node status | not stated per comparison | not stated |
| XGBoost classifier with recursive feature elimination and 5-fold cross-validated grid search (ROC-AUC) | ML classification of patient risk groups; validated in METABRIC and GSE96058 | ~767 training (~70% of TCGA-BRCA), ~328 testing (~30%) | not stated |
| SVM with RBF kernel, 5-fold cross-validated grid search (ROC-AUC) | Parallel ML classification using same RFE-selected features as XGBoost | same 70/30 TCGA-BRCA split | not stated |
| GSEA with GO and KEGG gene sets (permutation-based) | Pathway enrichment between high- and low-IRRS groups in TCGA-BRCA; hallmark GSEA between high- and low-UCell-score cells in scRNA-seq data | not stated | na |
-
Patients were dichotomized into high- and low-IRRS groups using the cohort median as a fixed cutpoint for all Kaplan–Meier analyses↳ Could also: IRRS could be modeled as a continuous predictor in Cox regression, or a data-driven optimal cutpoint method (e.g., maximally selected rank statistics via the R package maxstat) could be applied — Median dichotomization discards within-group variation and may yield different cutpoints across cohorts; treating IRRS continuously or using an optimized cutpoint preserves statistical power and improves cross-cohort reproducibility
-
Multiple log-rank tests were performed across five cohorts and five survival endpoints without an explicit multiplicity correction for this family of comparisons↳ Could also: A Bonferroni correction or Benjamini–Hochberg FDR adjustment could also be applied across the full family of survival comparisons — Conducting many log-rank tests inflates the family-wise type I error rate; an explicit correction allows readers to assess whether individual p-values remain significant after accounting for the number of tests
-
Seven hub genes were selected via LASSO Cox regression, which applies an L1 penalty and tends to select one gene arbitrarily from a correlated group↳ Could also: Elastic net Cox regression (combining L1 and L2 penalties, also available in glmnet with alpha < 1) could also be used for feature selection — Co-regulated insulin resistance genes are likely correlated; elastic net selects groups of correlated features more stably than pure LASSO, potentially yielding a more robust and reproducible signature
-
Student's t-test was used to compare IRG expression between normal and tumor tissues and to assess IRRS change before and after neoadjuvant therapy↳ Could also: A Wilcoxon rank-sum test (for unpaired comparisons) or Wilcoxon signed-rank test (for paired before/after data) could also be used, consistent with the non-parametric approach already applied for clinical feature associations in the same paper — Log-transformed gene expression data can retain skew; non-parametric tests require fewer distributional assumptions and the paired signed-rank test would additionally exploit the within-patient pairing in the pre/post-therapy comparison
-
Machine learning model performance was evaluated primarily using ROC-AUC on a single 70/30 random split for internal testing↳ Could also: Repeated k-fold cross-validation or bootstrap-based internal validation, together with calibration metrics (e.g., Brier score or calibration plots), could also be reported alongside ROC-AUC — A single random split introduces variability in the performance estimate; repeated resampling yields a more stable estimate, and calibration metrics complement AUC by assessing whether predicted probabilities match observed event rates—relevant when the score is intended to guide clinical decisions
-
Immune cell infiltration was estimated using CIBERSORT as the sole deconvolution method with default parameters↳ Could also: Additional deconvolution tools such as TIMER2, EPIC, or quanTIseq could also be applied and compared — Different deconvolution algorithms use distinct reference gene expression matrices and statistical assumptions; reporting concordance across multiple methods or noting where estimates converge strengthens confidence in immune infiltration conclusions
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
Downstream reach in the literature
199 downstream papers · 2 datasetsHow widely the datasets deposited by this paper are reused across the whole literature (Europe PMC), beyond our assessed set. This is a factual dependency map — reusing a public dataset is normal, good science. It is not a judgement on the downstream papers; the only verdict here is this paper's own, with its cited rationale.
- PrognoScan: a new database for meta-analysis of the... 2009 · 772 cites
- Survival analysis across the entire transcriptome id... 2021 · 751 cites
- Metabolic enzyme expression highlights a key role fo... 2014 · 494 cites
- The splicing factor SRSF1 regulates apoptosis and pr... 2012 · 358 cites
- MYC-driven accumulation of 2-hydroxyglutarate is ass... 2014 · 349 cites
- GOBO: gene expression-based outcome for breast cance... 2011 · 339 cites
- Clinical Value of RNA Sequencing-Based Classifiers f... 2018 · 165 cites
- bc-GenExMiner 4.5: new mining module computes breast... 2021 · 142 cites
- Abundance of Regulatory T Cell (Treg) as a Predictiv... 2020 · 95 cites
- Bulk and single-cell transcriptome profiling reveal... 2021 · 89 cites
- A lncRNA prognostic signature associated with immune... 2020 · 86 cites
- Prognosis and Dissection of Immunosuppressive Microe... 2022 · 73 cites
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-40427728
Paper: Elucidating the Prognostic and Therapeutic Implications of Insulin Resistance Genes in Breast Cancer: A Machine Learning-Powered Analysis. Biology 2025, 14(5):539. PMID 40427728 · PMCID PMC12109394 · DOI 10.3390/biology14050539.
Brief's pointers: Code = github.com/tidyverse/ggplot2 (a generic plotting
library, NOT the authors' analysis code); Data = geo:GSE20685.
Key facts established from the paper
- No authors' code repository exists. The cited GitHub link is just ggplot2. There is no shipped pipeline, no scripts, no parameter files. The paper names R packages (limma, glmnet, xgboost, e1071, survival, rms, ESTIMATE, CIBERSORT, clusterProfiler, Seurat, UCell, ggplot2) but ships none of the glue code.
- The designated dataset GSE20685 is a validation cohort, not the training set. Training was on TCGA-BRCA (1095). Validation cohorts: METABRIC (1906), GSE96058 (3069), GSE20685 (327), GSE7390 (198).
- The final product is a fully-specified 7-gene risk score (IRRS):
IRRS = 0.040·EZR − 0.046·LIFR − 0.138·TBC1D4 − 0.0105·SAA1 + 0.0218·NSF − 0.0566·RPL5 + 0.464·PGK1(Results / risk-score section). - Reported GSE20685 result: Kaplan–Meier OS, median IRRS split into high/low risk groups → log-rank p = 0.012 (Figure 1G; "GSE20658" in the panel is a typo for GSE20685). This is the one clear, identifiable numeric result tied to the brief's designated dataset.
IN SCOPE (clearly-specified, low-hanging, pipeline-derivable — the 80%)
C1 — Prognostic validation of the published IRRS on GSE20685. Apply the published 7-gene IRRS formula verbatim to GSE20685 expression (GPL570, Affymetrix HG-U133 Plus 2.0, 327 samples), median-split into high/low risk, run KM OS + log-rank. Expected: p ≈ 0.012, high IRRS = worse OS. This is a third-party-tool-on-the-paper's-own-data reproduction (P16): we run the paper's own published model on the paper's designated dataset. The formula is the "existing tool"; GSE20685 is the data.
Secondary (same pipeline, free): Cox HR (high vs low) and Harrell's C-index on GSE20685 — direction-of-effect and discrimination sanity checks.
OUT OF SCOPE (the hard ~20%, not attempted — and why)
- IRG retrieval from GeneCards (keyword "insulin resistance", relevance > 3.0, → 2828 genes). GeneCards is a live, unversioned web DB; the exact gene set on the authors' query date is not recoverable. Non-reproducible by construction. Possible-fabrication flag: the 2828→1115 numbers are not derivable from any shipped artifact.
- LASSO-Cox feature selection that produced the 7 genes. Trained on TCGA-BRCA with no shipped seed/lambda/code; the output (7 genes + coefficients) is what we validate instead. We do not re-derive the selection.
- XGBoost / SVM classifiers (accuracy/AUC tables), immune deconvolution (ESTIMATE/CIBERSORT), enrichment (clusterProfiler), and scRNA (Seurat/UCell, SCP1106). No code, many cohorts, heavy — outside the 80% and not tied to the brief's designated GSE20685.
- Other validation cohorts (METABRIC/GSE96058/GSE7390). We validate the one the brief designates (GSE20685); the others would be the same procedure.
Pipeline named per in-scope result
C1: published linear risk-score (fixed coefficients) → survival::survdiff
(log-rank) + survival::coxph (HR, C-index), on GSEMatrix expression fetched
via GEOquery. All compute on «our HPC» (SLURM), data on «infra».
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
The one clearly-specified, data-tied claim reproduces: applying the paper's published 7-gene IRRS verbatim to GSE20685 yields a significant OS split (log-rank p=0.020 vs reported 0.012, both significant) in the correct direction (HR=1.68, high=worse), so the central prognostic conclusion holds. The exact-p gap is a technical/expected deviation on our/methodology side (TCGA-fit coefficients applied to Affymetrix log2 units, unspecified probe collapse), not an authors' defect. The genuine concern is data/code availability: no authors' code ships (cited GitHub is ggplot2) and the upstream 2828 GeneCards IRG / 1115-after-DE counts are not derivable from any shipped artifact (possible-fabrication flag), keeping q5 and overall at yellow. Severity is negligible for the reproduced endpoint; overall a solid PARTIAL.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at [email protected].
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.