Elucidating the Prognostic and Therapeutic Implications of Insulin Resistance Genes in Breast Cancer: A Machine Learning-Powered Analysis.
The main results reproduced: recomputed values matched the published ones within tolerance.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Same input data as the authors
- ✓No authors-side cause for any deviation
- ✓Any deviation was negligible
- ✓The central claim held under reproduction
- 🟡Reported values were only indirectly comparable
- 🟡A deviation arose in the data or preprocessing
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Described well enough for a PARTIAL 1:1 on the one clearly-specified, data-tied result. The paper ships NO authors' code (the brief's GitHub link is ggplot2) and the brief's designated dataset GSE20685 is a validation cohort, not the training set. The paper's reproducible end product is a fully-specified 7-gene IRRS risk-score formula. We applied it verbatim to GSE20685 (327 samples, GPL570) on «our HPC» and re-ran the Fig 1G survival analysis: the signature significantly stratifies OS in the reported direction (HR=1.68, high=worse; C-index 0.60), with log-rank p=0.020 vs the reported 0.012 -- same order of magnitude, both significant, conclusion intact. Graded within-tol; the exact-p gap is expected (TCGA-fit coefficients applied to microarray log2 units; unspecified probe-collapse). NOT attempted (out of scope, the hard 20%): GeneCards IRG retrieval (non-versioned -> non-reproducible, flagged), LASSO-Cox feature selection on TCGA (no code/seed), XGBoost/SVM classifiers, immune deconvolution, enrichment, and scRNA -- none have shipped code and none are tied to the brief's GSE20685. No fabrication of results to avoid a drop; the central prognostic claim genuinely reproduces.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 84assessed: 2026-06-14 ⛓ eef212420178
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-14
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusAlthough insulin resistance (IR) is linked to tumorigenesis and cancer progression, its prognostic relevance in breast cancer is not fully understood; the study tests whether a prognostic signature built from insulin resistance-related genes (IRGs) can robustly stratify breast cancer patients by survival risk and predict treatment response.
- ★ A seven-gene insulin resistance signature (LIFR, EZR, TBC1D4, NSF, RPL5, SAA1, PGK1) yields an insulin resistance risk score (IRRS) with high predictive power for overall survival in breast cancer. finding
- ★ Higher IRRS is significantly associated with worse overall survival and more aggressive clinicopathological features across multiple independent cohorts. finding
- ★ Lower IRRS is associated with more favorable clinical outcomes, including enhanced response to neoadjuvant therapy. finding
- ★ A clinicopathological nomogram integrating IRRS and PAM50 subtype predicts 1-, 3-, 5-, and 10-year overall survival and outperforms previously published BC biomarkers by C-index. method
- ★ Machine learning models (XGBoost and SVM) built on the seven hub genes provide a clinically applicable scoring system validated in external cohorts. method
- At single-cell resolution, the seven hub genes are more enriched in T cells, B cells, and epithelial cells. finding
- The IR-based signature is a resource for risk stratification and personalized therapeutic strategies in breast cancer. resource
- High-IRRS tumors show higher tumor purity and differing immune/stromal TME composition than low-IRRS tumors. finding
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Bulk transcriptome analysis / differential expression and Cox/LASSO prognostic modeling | TCGA-BRCA breast cancer cohort (1095 tumor, 113 normal) | none | differentially expressed IRGs, hazard ratios, IRRS, overall survival | — |
| Prognostic model external validation (Kaplan-Meier OS) | METABRIC, GSE96058, GSE20685, GSE7390 BC cohorts | none | overall survival stratified by IRRS group | — |
| Neoadjuvant therapy response analysis | GSE191127, GSE20181, GSE18728, GSE225078 (307 neoadjuvant-treated BC patients) | neoadjuvant therapy | change in IRRS before/after therapy, treatment response | — |
| Tumor microenvironment deconvolution (ESTIMATE, CIBERSORT) | TCGA-BRCA | none | tumor purity, immune/stromal scores, relative abundance of 22 immune cell types | — |
| GSEA / GO / KEGG functional enrichment | TCGA-BRCA high- vs low-IRRS groups | none | enriched pathways/gene sets | — |
| Single-cell RNA sequencing analysis (Seurat, FastMNN, UMAP, UCell) | 5 triple-negative breast cancer samples (SCP1106, Single Cell Portal) | none | hub gene enrichment per cell type, UCell signature score | — |
| Machine learning classification (XGBoost and SVM with RFE) | TCGA-BRCA (70% train/30% test), validated in METABRIC and GSE96058 | none | patient group prediction, ROC-AUC, feature importance | Intel i7, 32 GB RAM; R xgboost / e1071 |
| Immunohistochemistry (IHC) | FFPE tissue from 10 BC patients (ER+, HER2+, TNBC; stages I-III), First People's Hospital of Changzhou | none | protein staining of CD163, CD8α, PGK1 | Pannoramic 250 Flash III scanner (3DHISTECH); anti-CD163/CD8α/PGK1 antibodies (Zen-Bioscience) |
- ▲ Combined IRRS factor showed a significantly higher hazard ratio than any individual gene HR = 4.368, 95% CI: 2.810–6.792, p < 0.0001
- ▼ High-IRRS group had significantly reduced overall survival vs low-IRRS group in TCGA p < 0.0001
- ▼ High-IRRS associated with reduced OS in GSE96058 validation cohort p = 1.8 × 10^-11
- ▼ High-IRRS associated with reduced OS in METABRIC, GSE20685, and GSE7390 validation cohorts METABRIC p < 0.048; GSE20685 p = 0.012; GSE7390 p = 0.027
- ▼ High IRRS associated with poorer DSS, DFI, PFI, and PFS in TCGA DSS p = 1.4 × 10^-4; DFI p = 5.8 × 10^-3; PFI p = 8.1 × 10^-5
- ▲ EZR, NSF, and PGK1 identified as risk genes with elevated expression linked to poorer outcomes
- – Nomogram (IRRS + PAM50) showed good discrimination by concordance index, outperforming prior BC biomarkers C-index 0.64 (TCGA), 0.59 (METABRIC)
- ▼ Tumor purity of low-IRRS group significantly lower than high-IRRS group
- other HR = 4.368, 95% CI: 2.810–6.792, p < 0.0001 (IRRS combined hazard ratio for BC risk)
- pvalue 1.8 × 10^-11 (OS difference high vs low IRRS, GSE96058)
- pvalue p < 0.0001 (OS difference high vs low IRRS, TCGA)
- other C-index 0.64 (Nomogram predictive performance in TCGA-BRCA)
- other C-index 0.59 (Nomogram predictive performance in METABRIC)
- count 6324 significant DEGs; 1115 significant IRGs (DE analysis of 1095 tumor vs 113 normal in TCGA-BRCA)
- count seven hub genes from 24 common IRGs via LASSO (54 (TCGA) and 290 (METABRIC) univariate Cox significant IRGs, 24 common)
- count 6907 BC samples across 10 cohorts (total compiled dataset after excluding samples lacking OS)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
The paper developed an insulin resistance risk score (IRRS) for breast cancer prognosis by filtering differentially expressed genes via limma, then applying univariate Cox regression and LASSO Cox regression to select a seven-gene signature from TCGA-BRCA. Kaplan–Meier curves with log-rank tests were used to compare overall survival between median-split risk groups across one training and four validation cohorts; XGBoost and SVM classifiers (5-fold cross-validated grid search, ROC-AUC optimized) provided an independent machine learning validation layer. Results were reported as hazard ratios with 95% confidence intervals, exact log-rank p-values, C-index values, and ROC-AUC scores.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| limma moderated t-statistic (linear model) | Differential expression between 1095 BC and 113 normal samples in TCGA-BRCA; threshold |log2FC| > 0.5 and adjusted p < 0.01 | 1095 BC + 113 normal | not stated |
| Univariate Cox proportional hazards regression | Prognostic IRG identification separately in TCGA-BRCA (p < 0.01) and METABRIC (p < 0.01); 54 and 290 significant IRGs identified, respectively | 1095 (TCGA-BRCA); METABRIC n not stated for this step | not stated |
| LASSO Cox regression (penalized) | Feature selection from 24 IRGs overlapping between TCGA-BRCA and METABRIC univariate screens, yielding 7-gene signature | 1095 (TCGA-BRCA training cohort) | not stated |
| Log-rank test with Kaplan–Meier survival curves | OS comparison between high- and low-IRRS groups in training cohort and four validation cohorts; also DSS, DFI, PFI, PFS in TCGA-BRCA; subtype-stratified OS in supplementary | 1095 training; ~5500 validation (METABRIC + GSE96058 + GSE20685 + GSE7390 combined) | not stated |
| Univariate and multivariate Cox proportional hazards regression | Variable selection for nomogram from IRRS, T stage, M stage, overall stage, PAM50; final model retains IRRS and PAM50 (p < 0.05) | TCGA-BRCA (1095); METABRIC (n not stated for this step) | not stated |
| Student's t-test (unpaired, two-sample; paired status not stated) | IRG expression differences between normal and tumor tissues; IRRS changes before and after neoadjuvant therapy | not stated for these specific comparisons | not stated |
| Wilcoxon rank-sum test | Association between IRRS and clinicopathological features: PAM50 subtypes, AJCC T and N stages, lymph node status | not stated per comparison | not stated |
| XGBoost classifier with recursive feature elimination and 5-fold cross-validated grid search (ROC-AUC) | ML classification of patient risk groups; validated in METABRIC and GSE96058 | ~767 training (~70% of TCGA-BRCA), ~328 testing (~30%) | not stated |
| SVM with RBF kernel, 5-fold cross-validated grid search (ROC-AUC) | Parallel ML classification using same RFE-selected features as XGBoost | same 70/30 TCGA-BRCA split | not stated |
| GSEA with GO and KEGG gene sets (permutation-based) | Pathway enrichment between high- and low-IRRS groups in TCGA-BRCA; hallmark GSEA between high- and low-UCell-score cells in scRNA-seq data | not stated | na |
-
Patients were dichotomized into high- and low-IRRS groups using the cohort median as a fixed cutpoint for all Kaplan–Meier analyses↳ Could also: IRRS could be modeled as a continuous predictor in Cox regression, or a data-driven optimal cutpoint method (e.g., maximally selected rank statistics via the R package maxstat) could be applied — Median dichotomization discards within-group variation and may yield different cutpoints across cohorts; treating IRRS continuously or using an optimized cutpoint preserves statistical power and improves cross-cohort reproducibility
-
Multiple log-rank tests were performed across five cohorts and five survival endpoints without an explicit multiplicity correction for this family of comparisons↳ Could also: A Bonferroni correction or Benjamini–Hochberg FDR adjustment could also be applied across the full family of survival comparisons — Conducting many log-rank tests inflates the family-wise type I error rate; an explicit correction allows readers to assess whether individual p-values remain significant after accounting for the number of tests
-
Seven hub genes were selected via LASSO Cox regression, which applies an L1 penalty and tends to select one gene arbitrarily from a correlated group↳ Could also: Elastic net Cox regression (combining L1 and L2 penalties, also available in glmnet with alpha < 1) could also be used for feature selection — Co-regulated insulin resistance genes are likely correlated; elastic net selects groups of correlated features more stably than pure LASSO, potentially yielding a more robust and reproducible signature
-
Student's t-test was used to compare IRG expression between normal and tumor tissues and to assess IRRS change before and after neoadjuvant therapy↳ Could also: A Wilcoxon rank-sum test (for unpaired comparisons) or Wilcoxon signed-rank test (for paired before/after data) could also be used, consistent with the non-parametric approach already applied for clinical feature associations in the same paper — Log-transformed gene expression data can retain skew; non-parametric tests require fewer distributional assumptions and the paired signed-rank test would additionally exploit the within-patient pairing in the pre/post-therapy comparison
-
Machine learning model performance was evaluated primarily using ROC-AUC on a single 70/30 random split for internal testing↳ Could also: Repeated k-fold cross-validation or bootstrap-based internal validation, together with calibration metrics (e.g., Brier score or calibration plots), could also be reported alongside ROC-AUC — A single random split introduces variability in the performance estimate; repeated resampling yields a more stable estimate, and calibration metrics complement AUC by assessing whether predicted probabilities match observed event rates—relevant when the score is intended to guide clinical decisions
-
Immune cell infiltration was estimated using CIBERSORT as the sole deconvolution method with default parameters↳ Could also: Additional deconvolution tools such as TIMER2, EPIC, or quanTIseq could also be applied and compared — Different deconvolution algorithms use distinct reference gene expression matrices and statistical assumptions; reporting concordance across multiple methods or noting where estimates converge strengthens confidence in immune infiltration conclusions
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
Downstream reach in the literature
199 downstream papers · 2 datasetsHow widely the datasets deposited by this paper are reused across the whole literature (Europe PMC), beyond our assessed set. This is a factual dependency map — reusing a public dataset is normal, good science. It is not a judgement on the downstream papers; the only verdict here is this paper's own, with its cited rationale.
- PrognoScan: a new database for meta-analysis of the... 2009 · 772 cites
- Survival analysis across the entire transcriptome id... 2021 · 751 cites
- Metabolic enzyme expression highlights a key role fo... 2014 · 494 cites
- The splicing factor SRSF1 regulates apoptosis and pr... 2012 · 358 cites
- MYC-driven accumulation of 2-hydroxyglutarate is ass... 2014 · 349 cites
- GOBO: gene expression-based outcome for breast cance... 2011 · 339 cites
- Clinical Value of RNA Sequencing-Based Classifiers f... 2018 · 165 cites
- bc-GenExMiner 4.5: new mining module computes breast... 2021 · 142 cites
- Abundance of Regulatory T Cell (Treg) as a Predictiv... 2020 · 95 cites
- Bulk and single-cell transcriptome profiling reveal... 2021 · 89 cites
- A lncRNA prognostic signature associated with immune... 2020 · 86 cites
- Prognosis and Dissection of Immunosuppressive Microe... 2022 · 73 cites
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-40427728
Paper: Elucidating the Prognostic and Therapeutic Implications of Insulin Resistance Genes in Breast Cancer: A Machine Learning-Powered Analysis. Biology 2025, 14(5):539. PMID 40427728 · PMCID PMC12109394 · DOI 10.3390/biology14050539.
Brief's pointers: Code = github.com/tidyverse/ggplot2 (a generic plotting
library, NOT the authors' analysis code); Data = geo:GSE20685.
Key facts established from the paper
- No authors' code repository exists. The cited GitHub link is just ggplot2. There is no shipped pipeline, no scripts, no parameter files. The paper names R packages (limma, glmnet, xgboost, e1071, survival, rms, ESTIMATE, CIBERSORT, clusterProfiler, Seurat, UCell, ggplot2) but ships none of the glue code.
- The designated dataset GSE20685 is a validation cohort, not the training set. Training was on TCGA-BRCA (1095). Validation cohorts: METABRIC (1906), GSE96058 (3069), GSE20685 (327), GSE7390 (198).
- The final product is a fully-specified 7-gene risk score (IRRS):
IRRS = 0.040·EZR − 0.046·LIFR − 0.138·TBC1D4 − 0.0105·SAA1 + 0.0218·NSF − 0.0566·RPL5 + 0.464·PGK1(Results / risk-score section). - Reported GSE20685 result: Kaplan–Meier OS, median IRRS split into high/low risk groups → log-rank p = 0.012 (Figure 1G; "GSE20658" in the panel is a typo for GSE20685). This is the one clear, identifiable numeric result tied to the brief's designated dataset.
IN SCOPE (clearly-specified, low-hanging, pipeline-derivable — the 80%)
C1 — Prognostic validation of the published IRRS on GSE20685. Apply the published 7-gene IRRS formula verbatim to GSE20685 expression (GPL570, Affymetrix HG-U133 Plus 2.0, 327 samples), median-split into high/low risk, run KM OS + log-rank. Expected: p ≈ 0.012, high IRRS = worse OS. This is a third-party-tool-on-the-paper's-own-data reproduction (P16): we run the paper's own published model on the paper's designated dataset. The formula is the "existing tool"; GSE20685 is the data.
Secondary (same pipeline, free): Cox HR (high vs low) and Harrell's C-index on GSE20685 — direction-of-effect and discrimination sanity checks.
OUT OF SCOPE (the hard ~20%, not attempted — and why)
- IRG retrieval from GeneCards (keyword "insulin resistance", relevance > 3.0, → 2828 genes). GeneCards is a live, unversioned web DB; the exact gene set on the authors' query date is not recoverable. Non-reproducible by construction. Possible-fabrication flag: the 2828→1115 numbers are not derivable from any shipped artifact.
- LASSO-Cox feature selection that produced the 7 genes. Trained on TCGA-BRCA with no shipped seed/lambda/code; the output (7 genes + coefficients) is what we validate instead. We do not re-derive the selection.
- XGBoost / SVM classifiers (accuracy/AUC tables), immune deconvolution (ESTIMATE/CIBERSORT), enrichment (clusterProfiler), and scRNA (Seurat/UCell, SCP1106). No code, many cohorts, heavy — outside the 80% and not tied to the brief's designated GSE20685.
- Other validation cohorts (METABRIC/GSE96058/GSE7390). We validate the one the brief designates (GSE20685); the others would be the same procedure.
Pipeline named per in-scope result
C1: published linear risk-score (fixed coefficients) → survival::survdiff
(log-rank) + survival::coxph (HR, C-index), on GSEMatrix expression fetched
via GEOquery. All compute on «our HPC» (SLURM), data on «infra».
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
The one clearly-specified, data-tied claim reproduces: applying the paper's published 7-gene IRRS verbatim to GSE20685 yields a significant OS split (log-rank p=0.020 vs reported 0.012, both significant) in the correct direction (HR=1.68, high=worse), so the central prognostic conclusion holds. The exact-p gap is a technical/expected deviation on our/methodology side (TCGA-fit coefficients applied to Affymetrix log2 units, unspecified probe collapse), not an authors' defect. The genuine concern is data/code availability: no authors' code ships (cited GitHub is ggplot2) and the upstream 2828 GeneCards IRG / 1115-after-DE counts are not derivable from any shipped artifact (possible-fabrication flag), keeping q5 and overall at yellow. Severity is negligible for the reproduced endpoint; overall a solid PARTIAL.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.