Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Construction and Validation of an Immune Infiltration-Related Gene Signature for the Prediction of Prognosis and Therapeutic Response in Breast Cancer.

Front Immunol · 2021
L1 51/100 3/4
⚑ Flagged for review — a reproduced result did not match the reported value

Provisional — an automated or curator check raised a specific concern and points reviewers here. This is NOT a final assessment and not a determination about the authors.

Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Main result did not reproduce
Decisive
From: Q5 · Derivability / plausibility 🔴
Main result did not reproduce
Decisive
From: Q7 · Core claim 🔴
Main result did not reproduce
Decisive
From: Q8 · Severity of the miss (overall human judgment) 🔴
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
What did not (or only partly)
  • 🔴A deviation arose in the data or preprocessing
  • 🔴A deviation was attributed to the published material
  • 🔴Reported values were not (fully) derivable from the shared data
  • 🔴The deviation was non-trivial in magnitude
  • 🔴The central claim did not (fully) hold under reproduction
  • 🔴Overall, the reproduction showed a material discrepancy
How its reproducibility compares
51/100
Reproducibility score
1.3 SD below mean
vs. all fields · 1173 studies
🎯 Scores higher than 13% of all assessed papers rank 1018 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

DESCRIBED WELL ENOUGH (for the validation step) -> PARTIAL, with a flagged internal inconsistency. The paper ships no analysis code (the listed 'code' is pRRophetic, a third-party drug tool); but the headline external-validation result is reproducible because the 15-gene signature + LASSO coefficients are published in Suppl Table S2 and the validation cohort GSE20685 (n=327, GPL570) is fully public. On «our HPC» («job», R 4.3.3 + GEOquery 2.70) I recomputed IRS = sum(coef x z-scored expression) on GSE20685 and re-ran the Fig-5F survival test. RESULT: (R3) n=327 matches EXACTLY; (R1) the signature significantly stratifies OS -- log-rank p=0.0177 at the paper's -0.03 cutoff, 0.0133 median-split, continuous-Cox p=0.0098 -- same significance and order of magnitude as the reported 0.0091 (within-tol). BUT (R2) the effect DIRECTION is INVERTED: with the exact published (all-positive) coefficients, high-IRS = BETTER survival (HR 0.58), opposite the paper's 'high-IRS worse OS / meta HR 2.72'. (F1) Root cause: Suppl Table S2 lists all 15 coefficients as positive, contradicting the paper's own text ('192 protective + 1 risk') and Fig 3F (negative=protective); the genes are immune-activation markers (MS4A1/CD20, CCR9, KLRC3) whose high expression is biologically expected to predict BETTER BRCA prognosis -- consistent with our inversion. Flagged as a possible-fabrication/internal-inconsistency note for the human auditor. NOT ATTEMPTED (out of scope / hard 20%): signature construction (ImmuCellAI->WGCNA->LASSO on TCGA, no code), TCGA/METABRIC cohorts (access-gated / custom norm), pRRophetic drug-IC50 ranking and immunotherapy 57%/90% (keyed off TCGA IRS deciles + CTRP/PRISM/TCIA). One caveat: FAM92B (1/15) has no GPL570 probe so 14/15 genes used; does not change the aggregate sign. Exact normalization ('normalized mRNA expression') is under-specified in the paper -> per-cohort gene z-score used.

💻 Code ↗ 🗄 Data: GSE20685

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 51
    assessed: 2026-06-15 ⛓ 4f78204c3cbc
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-15
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

Can an immune infiltration-related gene signature derived from tumor-infiltrating immune cells be constructed to robustly predict prognosis and therapeutic (immunotherapy and chemotherapy) response in breast cancer patients?

Core claims
  • A 15-gene immune infiltration-related signature (IRS) predicts overall survival in breast cancer, with higher IRS indicating worse prognosis finding
  • The prognostic value of the gene signature was validated across independent cohorts (TCGA, METABRIC, GSE20685, GSE21653) and by meta-analysis finding
  • Higher IRS is associated with reduced sensitivity to immunotherapy; low-IRS patients have a higher immunotherapy response rate finding
  • The signature predicts chemotherapy drug sensitivity, identifying candidate compounds more effective in high-IRS patients finding
  • WGCNA combined with univariate Cox and LASSO Cox regression was used to build the immune-related gene signature method
  • A nomogram combining IRS with clinicopathological features was constructed to quantify individual risk and survival probability resource
  • Eight immune cell types (Tfh, CD8 T, Tcm, MAIT, CD4 T, NK, Tgd, Th2) act as protective factors associated with better overall survival in breast cancer finding
Experimental setups
Assay System Perturbation Readout Platform
Immune cell infiltration scoring (24 immune cells) TCGA breast cancer patients (n=1007) none immune cell infiltration scores ImmuCellAI
Bulk RNA-seq gene expression / signature training TCGA breast cancer cohort (n=1007) none gene expression, overall survival, IRS Illumina HiSeq RNA-Seq
Microarray gene expression / external validation METABRIC breast cancer cohort (n=1761) none gene expression, overall survival
Microarray gene expression / external validation GEO cohort GSE20685 (n=327) none gene expression, overall survival GPL570
Microarray gene expression / external validation GEO cohort GSE21653 (n=266) none gene expression, overall survival GPL570
Immunotherapy response prediction (anti-PD1/anti-CTLA4) TCGA breast cancer patients drug (immune checkpoint inhibitors, in silico) predicted response vs nonresponse ImmuCellAI
Chemotherapy drug sensitivity prediction TCGA breast cancer samples; cancer cell lines (CTRP 481 compounds/835 CCLs; PRISM 1448 compounds/482 CCLs) drug AUC of dose-response curve (drug sensitivity) pRRophetic; CTRP; PRISM Repurposing 19Q4
Proteomics breast cancer (CPTAC) none gene-level protein abundance CPTAC
Key results
  • Patients in the high-IRS group exhibited worse overall survival in the TCGA training cohort P<0.001
  • Patients with higher immune infiltration scores (Th2, Tfh, Tcm, MAIT, NK, CD4 T, Tgd, CD8 T) exhibited better overall survival
  • Immunotherapy response rate in the low-IRS group was much higher than in the high-IRS group P<0.001
  • 14 of 15 signature genes showed higher expression in the low-IRS group; only SYTL3 (risk gene) was higher in high-IRS group
  • Red WGCNA module correlated with infiltration scores of six protective immune cells and was identified as the immune-related module
  • LASSO Cox regression selected 15 genes with nonzero coefficients from 193 prognostic candidates optimal λ=0.014
  • Compounds with negative IRS-AUC correlation identified as more effective in high-IRS patients Spearman r<−0.25 (CTRP) or <−0.30 (PRISM)
Key statistics
  • count 15 immune-related genes in signature (final LASSO-selected gene signature)
  • count 193 candidates (192 protective, 1 risk) (univariate Cox p<0.05 from 802 red-module genes)
  • pvalue P<0.001 (high vs low IRS OS difference in TCGA training set)
  • pvalue P<0.001 (chi-square test immunotherapy response between IRS groups)
  • other optimal cutoff -0.03 (IRS high/low group separation from surv_cutpoint)
  • other λ=0.014 (optimal LASSO lambda via 10-fold cross-validation)
  • other AUC 0.80–0.91 (ImmuCellAI immunotherapy response prediction accuracy)
  • correlation Spearman r<−0.25 (CTRP), r<−0.30 (PRISM) (IRS vs drug AUC correlation threshold)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The paper constructs a 15-gene immune infiltration-related risk score (IRS) for breast cancer prognosis using a sequential bioinformatics pipeline: univariate Cox regression to select prognosis-related immune cells, WGCNA to identify co-expressed gene modules, a second round of univariate Cox filtering (p<0.05) to narrow 802 candidate genes to 193, and LASSO Cox regression with 10-fold cross-validation to derive the final signature. The IRS was validated in four independent cohorts (TCGA, METABRIC, GSE20685, GSE21653) via Kaplan-Meier/log-rank testing and time-dependent ROC analysis, with a fixed cutoff of −0.03 applied uniformly across all datasets. Chemotherapy response was predicted by correlating compound AUC values with IRS using Spearman correlation, and immunotherapy response was assessed by chi-square test comparing response rates across IRS groups.

Replicationbiological Sample sizeSample sizes stated by dataset (TCGA n=1007, METABRIC n=1761, GSE20685 n=327, GSE21653 n=266); no formal power calculation or sample size justification reported GroupsHigh-IRS vs low-IRS breast cancer patients; also immunotherapy responders vs non-responders; immune cell infiltration high vs low Pairingunpaired Randomization/blindingnot stated Dispersionnone Exact p-valuesno Multiplicity correctionnone stated
Statistical tests used
Test Applied to n Assumptions
Univariate Cox proportional hazards regression Immune cell infiltration scores (24 cells) versus overall survival; identification of prognosis-related cells 1007 (TCGA training set) not stated
WGCNA (Pearson correlation with soft thresholding, power=12) Whole-transcriptome gene co-expression module identification and correlation with immune cell infiltration scores 1007 (TCGA training set) not stated
Univariate Cox proportional hazards regression (802 tests) 802 red-module genes versus overall survival; p<0.05 threshold applied to select 193 candidates 1007 (TCGA training set) not stated
LASSO Cox regression with 10-fold cross-validation Selection of 15-gene signature from 193 univariate-significant candidates; optimal λ=0.014 1007 (TCGA training set) not stated
Kaplan-Meier survival analysis with log-rank test High-IRS vs low-IRS overall survival comparison in training (TCGA) and three validation cohorts (METABRIC, GSE20685, GSE21653); also per-immune-cell survival curves (8 cells) TCGA n=1007; METABRIC n=1761; GSE20685 n=327; GSE21653 n=266 not stated
Chi-square test Immunotherapy response rate (responder vs non-responder) across high-IRS and low-IRS groups TCGA subset (exact n not stated) not stated
Spearman rank correlation AUC drug sensitivity values vs IRS for compounds in CTRP (481 compounds, 835 CCLs) and PRISM (1448 compounds, 482 CCLs); filter thresholds r<−0.25 (CTRP) or r<−0.30 (PRISM) TCGA clinical samples (exact n not stated) not stated
Time-dependent ROC (tROC) Predictive accuracy of the IRS model across training and validation cohorts multiple cohorts (exact per-cohort n as above) na
Meta-analysis (pooled hazard ratio) Pooled prognostic value of the IRS across all four datasets using the 'meta' R package combined across TCGA, METABRIC, GSE20685, GSE21653 not stated
Gene Set Enrichment Analysis (GSEA) Validation of immune status differences between high- and low-IRS groups using c5.go.v7.2.entrez.gmt gene set 1007 (TCGA training set) not stated
Approaches that could also have been used
  • 802 univariate Cox regressions were run simultaneously and filtered at a nominal p<0.05 threshold (yielding 193 candidates)
    Could also: Apply Benjamini-Hochberg false discovery rate (FDR) correction across the 802 tests before passing candidates to LASSO — With 802 parallel tests at α=0.05, approximately 40 false positives are expected by chance; FDR correction reduces the rate of spurious candidates entering the LASSO step, which can affect which genes receive nonzero coefficients
  • Patients were dichotomized into high- and low-IRS groups using a single data-derived optimal cutoff (−0.03) applied uniformly across all cohorts
    Could also: Treat IRS as a continuous variable in Cox regression, or use pre-specified tertiles/quartiles, rather than a single median- or optimum-derived cutoff — Data-derived binary cutoffs can be sensitive to the training distribution; analyzing IRS continuously (hazard ratio per unit) avoids information loss from dichotomization and is less prone to optimistic bias when the same cutoff is applied to independent sets
  • Primary survival comparisons used log-rank tests on Kaplan-Meier curves stratified by IRS group
    Could also: Supplement with multivariate Cox regression adjusting for established clinicopathological covariates (age, stage, ER/PR/HER2 status, grade) — Log-rank tests are univariate; multivariate Cox models assess whether the IRS adds independent prognostic information beyond known clinical factors, which is an important question for clinical utility
  • Drug sensitivity associations were identified by Spearman correlation thresholds (r<−0.25/−0.30) applied across 481 and 1448 compounds respectively, without multiplicity correction
    Could also: Apply BH-FDR correction across all compound-IRS correlations within each dataset before applying the r-magnitude filter — With hundreds of simultaneous correlations, many nominal associations may be false positives; FDR-adjusted q-values would allow the magnitude threshold to be interpreted alongside a controlled error rate
  • LASSO Cox regression alone was used for final gene selection from the 193 candidate genes
    Could also: Use elastic net Cox regression (combining L1 and L2 penalties) as an alternative variable selection approach — When candidate predictors are correlated (as co-expression module genes tend to be), elastic net can select groups of correlated genes more stably than LASSO, which tends to pick one representative arbitrarily from a correlated cluster
  • Predictive performance was summarized using time-dependent AUC from tROC analysis
    Could also: Also report Harrell's C-index (concordance statistic) with a 95% confidence interval — The C-index is the standard summary measure of discrimination for survival models and is more directly comparable across published breast cancer prognostic signatures; reporting both tROC-AUC and C-index would facilitate cross-study comparison
Software: R 4.0.3 · R/survival 3.2-7 · R/glmnet 4.0-2 · R/survminer 0.4.3 · R/WGCNA · R/clusterProfiler 3.18.0 · R/timeROC 0.4 · R/meta 4.15-1 · R/rms · R/ggplot2 3.3.2 · R/pRRophetic 4.15-1 · R/estimate · ImmuCellAI (web tool)

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
14
Impact: medium
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-33986754

Paper: Peng Y, et al. Construction and Validation of an Immune Infiltration-Related Gene Signature for the Prediction of Prognosis and Therapeutic Response in Breast Cancer. Front Immunol 2021. PMID 33986754 / PMCID PMC8110914 / DOI 10.3389/fimmu.2021.666137.

"Code" link (brief): https://github.com/paulgeeleher/pRRophetic — this is a third-party drug-response prediction tool, NOT the authors' analysis pipeline. The authors released no code for the signature construction / survival validation. Per brief rule P16, applying a public tool/known signature to the paper's own public data is an equally valid reproduction.

Data: GEO GSE20685 (Kao et al. 2011; Taiwan breast cancer cohort, n=327, Affymetrix HG-U133 Plus 2.0 / GPL570) — the paper's external validation set #3. Public, fully obtainable via GEOquery.

Signature (fully specified, machine-readable): 15 immune-related genes + LASSO coefficients are published in Supplementary Table S2 (DataSheet_2.csv). IRS = Σ(coefficient × normalized mRNA expression); cohort cutoff -0.03 (derived by surv_cutpoint on TCGA, applied to all cohorts).

Pipeline-derived results, in scope (low-hanging, clearly specified)

  • R1 — GSE20685 Kaplan–Meier external validation (Fig 5F). Reported log-rank P = 0.0091; high-IRS group reported to have worse OS. Recompute IRS on GSE20685 from the 15 published coefficients, split at the paper's cutoff, run log-rank + Cox. This is the single cleanest, fully-public, paper-specific pipeline output. ATTEMPTED.
  • R2 — Direction of the IRS effect (meta-analysis HR = 2.72, Fig 5G; "high-IRS worse OS"). Checked as part of R1 (sign of the Cox HR). ATTEMPTED.

Out of scope (not attempted — the optional hard ~20%, with reasons)

  • Signature construction (ImmuCellAI 24-cell scores → univariate Cox → WGCNA red module → LASSO Cox on TCGA): no authors' code; needs exact TCGA-BRCA preprocessing + ImmuCellAI server output; not reproducible 1:1, and re-deriving it would not yield the same 15 genes/coefficients. The published signature is the documented entry point — we validate it, we don't re-derive it.
  • TCGA / METABRIC cohorts (Figs 4, 5D, 6): METABRIC is access-gated (cBioPortal /Synapse), TCGA needs the same custom normalization; GSE20685 is the clean public surrogate for the validation claim.
  • pRRophetic drug-IC50 ranking (CTRP/PRISM, Fig 9) and immunotherapy response 57%/90% (Fig 8G): both keyed off TCGA IRS deciles + external databases (CTRP/PRISM/TCIA), not reproducible without the TCGA IRS pipeline.
  • IRS↔immune-cell correlations (DataSheet_5/6, TIMER/TCIA): require recomputing IRS on TCGA + downloading TIMER/TCIA scores — out of the 80/20 low-hanging set.

Reproduction stance

Known published signature (15 coefficients, Suppl. Table S2) applied to the paper's own public validation cohort (GSE20685) at the documented entry point, on «our HPC». Deterministic given normalization choice → graded by tolerance on the p-value and by sign on the effect direction.

Figures / tables: Figure 5FFig 5FFig 3F
R1
Reported
0.0091 (GSE20685 KM log-rank p, Fig 5F)
Reproduced
0.0177 (fixed cutoff -0.03) / 0.0133 (median) / 0.0098 (continuous Cox)
within tolerance
R2
Reported
high-IRS WORSE OS (meta HR=2.72)
Reproduced
high-IRS BETTER OS, HR=0.58 (group) / 0.956 (continuous) -> direction INVERTED
did not match
R3
Reported
327 (validation n, OS>30d)
Reproduced
327
exact
F1
Reported
192 protective(neg)+1 risk(pos) coefficients (text/Fig 3F)
Reproduced
all 15 published coefficients POSITIVE (0.835-1.455)
did not match

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 51/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🔴3. Location of the main deviation
🔴4. Cause of the deviation
🔴5. Derivability / plausibility
🔴6. Severity of the deviation
🔴7. Core claim
🔴8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Main result did not reproduce
Decisive
From: Q5 · Derivability / plausibility 🔴
Main result did not reproduce
Decisive
From: Q7 · Core claim 🔴
Main result did not reproduce
Decisive
From: Q8 · Severity of the miss (overall human judgment) 🔴

The external-validation step reproduces well on the public side — n=327 matches exactly (R3) and the 15-gene signature still significantly stratifies OS (p=0.0177 vs reported 0.0091, within-tol R1). But using the exact published all-positive coefficients (Suppl Table S2) the prognostic direction inverts: high-IRS = BETTER OS (HR 0.58), opposite the paper's high-IRS-worse / meta HR=2.72 claim. This sits on the authors' side: the deposited coefficient signs contradict the paper's own text ('192 protective + 1 risk') and Fig 3F, an internal inconsistency that mechanically drives the inversion and is biologically consistent with our result (MS4A1/CD20, CCR9, KLRC3 immune-activation genes). Severe and fabrication-suspect on the directional claim, hence red on q5/q7/q8.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

120.2 k
tokens (I/O) · 7.9 M incl. cache
16 min
runtime · 0.02 CPU-h
2.5 GB
peak RAM
2 (1 failed)
HPC jobs
hummel
machine