Corpus 1,286 assessed · 1,187 scored · 648 reproduced ≥75 · 174 flagged ·∅ 73.9/100
← New search

Construction and Validation of an Immune Infiltration-Related Gene Signature for the Prediction of Prognosis and Therapeutic Response in Breast Cancer.

Front Immunol · 2021
L1 51/100 3/4
⚑ Flagged for review — a reproduced result did not match the reported value

Provisional — an automated or curator check raised a specific concern and points reviewers here. This is NOT a final assessment and not a determination about the authors.

Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Main result did not reproduce
Decisive
From: Q5 · Derivability / plausibility 🔴
Main result did not reproduce
Decisive
From: Q7 · Core claim 🔴
Main result did not reproduce
Decisive
From: Q8 · Severity of the miss (overall human judgment) 🔴
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
What did not (or only partly)
  • 🔴A deviation arose in the data or preprocessing
  • 🔴A deviation was attributed to the published material
  • 🔴Reported values were not (fully) derivable from the shared data
  • 🔴The deviation was non-trivial in magnitude
  • 🔴The central claim did not (fully) hold under reproduction
  • 🔴Overall, the reproduction showed a material discrepancy
How its reproducibility compares
51/100
Reproducibility score
1.3 SD below mean
vs. all fields · 1187 studies
🎯 Scores higher than 13% of all assessed papers rank 1027 of 1187 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

DESCRIBED WELL ENOUGH (for the validation step) -> PARTIAL, with a flagged internal inconsistency. The paper ships no analysis code (the listed 'code' is pRRophetic, a third-party drug tool); but the headline external-validation result is reproducible because the 15-gene signature + LASSO coefficients are published in Suppl Table S2 and the validation cohort GSE20685 (n=327, GPL570) is fully public. On «our HPC» («job», R 4.3.3 + GEOquery 2.70) I recomputed IRS = sum(coef x z-scored expression) on GSE20685 and re-ran the Fig-5F survival test. RESULT: (R3) n=327 matches EXACTLY; (R1) the signature significantly stratifies OS -- log-rank p=0.0177 at the paper's -0.03 cutoff, 0.0133 median-split, continuous-Cox p=0.0098 -- same significance and order of magnitude as the reported 0.0091 (within-tol). BUT (R2) the effect DIRECTION is INVERTED: with the exact published (all-positive) coefficients, high-IRS = BETTER survival (HR 0.58), opposite the paper's 'high-IRS worse OS / meta HR 2.72'. (F1) Root cause: Suppl Table S2 lists all 15 coefficients as positive, contradicting the paper's own text ('192 protective + 1 risk') and Fig 3F (negative=protective); the genes are immune-activation markers (MS4A1/CD20, CCR9, KLRC3) whose high expression is biologically expected to predict BETTER BRCA prognosis -- consistent with our inversion. Flagged as a possible-fabrication/internal-inconsistency note for the human auditor. NOT ATTEMPTED (out of scope / hard 20%): signature construction (ImmuCellAI->WGCNA->LASSO on TCGA, no code), TCGA/METABRIC cohorts (access-gated / custom norm), pRRophetic drug-IC50 ranking and immunotherapy 57%/90% (keyed off TCGA IRS deciles + CTRP/PRISM/TCIA). One caveat: FAM92B (1/15) has no GPL570 probe so 14/15 genes used; does not change the aggregate sign. Exact normalization ('normalized mRNA expression') is under-specified in the paper -> per-cohort gene z-score used.

💻 Code ↗ 🗄 Data: GSE20685

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 51
    assessed: 2026-06-15 ⛓ 4f78204c3cbc
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-15
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-09-19

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

Immune infiltration is associated with breast cancer progression, and a gene signature derived from immune infiltration-related genes can predict patient prognosis and response to immunotherapy/chemotherapy better than current models.

Core claims
  • A higher immune infiltration-related risk score (IRS) indicates worse prognosis and lower sensitivity to immunotherapy in breast cancer patients finding
  • A 15-gene immune infiltration-related signature (SSYTL3/SYTL3, MS4A1, FAM92B, GBP2, LGALS2, SPINK2, AMPD1, STAR, TDGF1, KLRC3, BCL2L14, CCR9, CCL1, TAPBPL, FAM159A) was constructed using WGCNA and LASSO Cox regression to predict survival method
  • The red WGCNA module (802 genes) was identified as the immune-related module, enriched for T cell activation and immune receptor activity finding
  • Tfh, CD8 T, Tcm, MAIT, CD4 T, NK, Tgd and Th2 cell infiltration act as protective factors for breast cancer survival finding
  • The IRS signature's prognostic value was validated across TCGA training set and three independent validation cohorts (METABRIC, GSE20685, GSE21653) via meta-analysis finding
  • A nomogram combining IRS with clinicopathological features improves risk stratification and individualized survival prediction method
  • Candidate chemotherapeutic compounds with lower estimated AUC (higher sensitivity) in the high-IRS group were identified via CTRP and PRISM drug sensitivity datasets resource
Experimental setups
Assay System Perturbation Readout Platform
Univariate Cox-PH regression on immune cell infiltration scores TCGA breast cancer patients (n=1007) none prognosis-related immune cell infiltration scores (ImmuCellAI) Illumina HiSeq RNA-Seq / ImmuCellAI
Weighted gene coexpression network analysis (WGCNA) TCGA breast cancer transcriptome data none gene modules correlated with protective immune cell infiltration scores WGCNA R package
LASSO Cox regression 193 candidate genes from red WGCNA module, TCGA training set none immune-related gene signature with nonzero coefficients (15 genes) glmnet R package
Kaplan-Meier survival analysis / log-rank test TCGA (training) and METABRIC, GSE20685, GSE21653 (validation) breast cancer cohorts none (stratified by IRS high/low) overall survival by IRS group survival R package
Gene set enrichment analysis (GSEA) TCGA high-IRS vs low-IRS breast cancer samples none immune-related pathway/gene set enrichment (c5.go.v7.2 gene set) clusterProfiler R package
Immunotherapy response prediction (anti-PD1/anti-CTLA4) TCGA breast cancer patients none predicted responder vs nonresponder status by IRS group ImmuCellAI
Chemotherapy drug sensitivity prediction cancer cell lines (CTRP: 835 CCLs/481 compounds; PRISM: 482 CCLs/1448 compounds) and TCGA samples none predicted AUC drug response correlated with IRS pRRophetic R package
Time-dependent ROC (tROC) analysis TCGA and validation cohorts none predictive power/accuracy of IRS model for survival timeROC R package
Key results
  • 11 prognosis-related infiltrating immune cells identified by univariate Cox regression
  • Higher infiltration scores of Th2, Tfh, Tcm, MAIT, NK, CD4 T, Tgd and CD8 T cells associated with better overall survival
  • 193 candidate genes significantly related to prognosis identified from 802 red-module genes (univariate Cox p<0.05)
  • 15 genes with nonzero LASSO coefficients selected as final gene signature at optimal lambda lambda=0.014
  • Patients in high-IRS group had significantly worse overall survival than low-IRS group in TCGA training set P<0.001
  • Proportion of immunotherapy responders was much higher in the low-IRS group than the high-IRS group P<0.001
  • NK and T cell chemotaxis/commitment genes enriched in IRS-low tumors by GSEA
  • 14 of 15 signature genes showed higher expression in low-IRS group; SYTL3 (risk gene) showed the opposite pattern
Key statistics
  • pvalue P<0.001 (Kaplan-Meier OS difference between high- and low-IRS groups in TCGA training cohort)
  • pvalue P<0.001 (chi-square test comparing immunotherapy response rates between high- and low-IRS groups)
  • pvalue P<0.05 (univariate Cox regression threshold used to select 193 candidate genes from red WGCNA module)
  • count 15 genes (final immune-related gene signature from LASSO Cox regression)
  • count 193 candidate genes (prognosis-associated genes identified from red module genes)
  • correlation Spearman r<-0.25 (CTRP) or r<-0.30 (PRISM) (threshold for selecting drug compounds negatively correlated with IRS)
  • fold_change log2FC>0.05 (threshold for lower estimated AUC (higher drug sensitivity) in high-IRS group)
  • count n=1007 (TCGA), n=1761 (METABRIC), n=327 (GSE20685), n=266 (GSE21653) (sample sizes of training and validation cohorts)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The paper constructs a 15-gene immune infiltration-related risk score (IRS) for breast cancer prognosis using a sequential bioinformatics pipeline: univariate Cox regression to select prognosis-related immune cells, WGCNA to identify co-expressed gene modules, a second round of univariate Cox filtering (p<0.05) to narrow 802 candidate genes to 193, and LASSO Cox regression with 10-fold cross-validation to derive the final signature. The IRS was validated in four independent cohorts (TCGA, METABRIC, GSE20685, GSE21653) via Kaplan-Meier/log-rank testing and time-dependent ROC analysis, with a fixed cutoff of −0.03 applied uniformly across all datasets. Chemotherapy response was predicted by correlating compound AUC values with IRS using Spearman correlation, and immunotherapy response was assessed by chi-square test comparing response rates across IRS groups.

Replicationbiological Sample sizeSample sizes stated by dataset (TCGA n=1007, METABRIC n=1761, GSE20685 n=327, GSE21653 n=266); no formal power calculation or sample size justification reported GroupsHigh-IRS vs low-IRS breast cancer patients; also immunotherapy responders vs non-responders; immune cell infiltration high vs low Pairingunpaired Randomization/blindingnot stated Dispersionnone Exact p-valuesno Multiplicity correctionnone stated
Statistical tests used
Test Applied to n Assumptions
Univariate Cox proportional hazards regression Immune cell infiltration scores (24 cells) versus overall survival; identification of prognosis-related cells 1007 (TCGA training set) not stated
WGCNA (Pearson correlation with soft thresholding, power=12) Whole-transcriptome gene co-expression module identification and correlation with immune cell infiltration scores 1007 (TCGA training set) not stated
Univariate Cox proportional hazards regression (802 tests) 802 red-module genes versus overall survival; p<0.05 threshold applied to select 193 candidates 1007 (TCGA training set) not stated
LASSO Cox regression with 10-fold cross-validation Selection of 15-gene signature from 193 univariate-significant candidates; optimal λ=0.014 1007 (TCGA training set) not stated
Kaplan-Meier survival analysis with log-rank test High-IRS vs low-IRS overall survival comparison in training (TCGA) and three validation cohorts (METABRIC, GSE20685, GSE21653); also per-immune-cell survival curves (8 cells) TCGA n=1007; METABRIC n=1761; GSE20685 n=327; GSE21653 n=266 not stated
Chi-square test Immunotherapy response rate (responder vs non-responder) across high-IRS and low-IRS groups TCGA subset (exact n not stated) not stated
Spearman rank correlation AUC drug sensitivity values vs IRS for compounds in CTRP (481 compounds, 835 CCLs) and PRISM (1448 compounds, 482 CCLs); filter thresholds r<−0.25 (CTRP) or r<−0.30 (PRISM) TCGA clinical samples (exact n not stated) not stated
Time-dependent ROC (tROC) Predictive accuracy of the IRS model across training and validation cohorts multiple cohorts (exact per-cohort n as above) na
Meta-analysis (pooled hazard ratio) Pooled prognostic value of the IRS across all four datasets using the 'meta' R package combined across TCGA, METABRIC, GSE20685, GSE21653 not stated
Gene Set Enrichment Analysis (GSEA) Validation of immune status differences between high- and low-IRS groups using c5.go.v7.2.entrez.gmt gene set 1007 (TCGA training set) not stated
Approaches that could also have been used
  • 802 univariate Cox regressions were run simultaneously and filtered at a nominal p<0.05 threshold (yielding 193 candidates)
    Could also: Apply Benjamini-Hochberg false discovery rate (FDR) correction across the 802 tests before passing candidates to LASSO — With 802 parallel tests at α=0.05, approximately 40 false positives are expected by chance; FDR correction reduces the rate of spurious candidates entering the LASSO step, which can affect which genes receive nonzero coefficients
  • Patients were dichotomized into high- and low-IRS groups using a single data-derived optimal cutoff (−0.03) applied uniformly across all cohorts
    Could also: Treat IRS as a continuous variable in Cox regression, or use pre-specified tertiles/quartiles, rather than a single median- or optimum-derived cutoff — Data-derived binary cutoffs can be sensitive to the training distribution; analyzing IRS continuously (hazard ratio per unit) avoids information loss from dichotomization and is less prone to optimistic bias when the same cutoff is applied to independent sets
  • Primary survival comparisons used log-rank tests on Kaplan-Meier curves stratified by IRS group
    Could also: Supplement with multivariate Cox regression adjusting for established clinicopathological covariates (age, stage, ER/PR/HER2 status, grade) — Log-rank tests are univariate; multivariate Cox models assess whether the IRS adds independent prognostic information beyond known clinical factors, which is an important question for clinical utility
  • Drug sensitivity associations were identified by Spearman correlation thresholds (r<−0.25/−0.30) applied across 481 and 1448 compounds respectively, without multiplicity correction
    Could also: Apply BH-FDR correction across all compound-IRS correlations within each dataset before applying the r-magnitude filter — With hundreds of simultaneous correlations, many nominal associations may be false positives; FDR-adjusted q-values would allow the magnitude threshold to be interpreted alongside a controlled error rate
  • LASSO Cox regression alone was used for final gene selection from the 193 candidate genes
    Could also: Use elastic net Cox regression (combining L1 and L2 penalties) as an alternative variable selection approach — When candidate predictors are correlated (as co-expression module genes tend to be), elastic net can select groups of correlated genes more stably than LASSO, which tends to pick one representative arbitrarily from a correlated cluster
  • Predictive performance was summarized using time-dependent AUC from tROC analysis
    Could also: Also report Harrell's C-index (concordance statistic) with a 95% confidence interval — The C-index is the standard summary measure of discrimination for survival models and is more directly comparable across published breast cancer prognostic signatures; reporting both tROC-AUC and C-index would facilitate cross-study comparison
Software: R 4.0.3 · R/survival 3.2-7 · R/glmnet 4.0-2 · R/survminer 0.4.3 · R/WGCNA · R/clusterProfiler 3.18.0 · R/timeROC 0.4 · R/meta 4.15-1 · R/rms · R/ggplot2 3.3.2 · R/pRRophetic 4.15-1 · R/estimate · ImmuCellAI (web tool)

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
14
Impact: medium
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-33986754

Paper: Peng Y, et al. Construction and Validation of an Immune Infiltration-Related Gene Signature for the Prediction of Prognosis and Therapeutic Response in Breast Cancer. Front Immunol 2021. PMID 33986754 / PMCID PMC8110914 / DOI 10.3389/fimmu.2021.666137.

"Code" link (brief): https://github.com/paulgeeleher/pRRophetic — this is a third-party drug-response prediction tool, NOT the authors' analysis pipeline. The authors released no code for the signature construction / survival validation. Per brief rule P16, applying a public tool/known signature to the paper's own public data is an equally valid reproduction.

Data: GEO GSE20685 (Kao et al. 2011; Taiwan breast cancer cohort, n=327, Affymetrix HG-U133 Plus 2.0 / GPL570) — the paper's external validation set #3. Public, fully obtainable via GEOquery.

Signature (fully specified, machine-readable): 15 immune-related genes + LASSO coefficients are published in Supplementary Table S2 (DataSheet_2.csv). IRS = Σ(coefficient × normalized mRNA expression); cohort cutoff -0.03 (derived by surv_cutpoint on TCGA, applied to all cohorts).

Pipeline-derived results, in scope (low-hanging, clearly specified)

  • R1 — GSE20685 Kaplan–Meier external validation (Fig 5F). Reported log-rank P = 0.0091; high-IRS group reported to have worse OS. Recompute IRS on GSE20685 from the 15 published coefficients, split at the paper's cutoff, run log-rank + Cox. This is the single cleanest, fully-public, paper-specific pipeline output. ATTEMPTED.
  • R2 — Direction of the IRS effect (meta-analysis HR = 2.72, Fig 5G; "high-IRS worse OS"). Checked as part of R1 (sign of the Cox HR). ATTEMPTED.

Out of scope (not attempted — the optional hard ~20%, with reasons)

  • Signature construction (ImmuCellAI 24-cell scores → univariate Cox → WGCNA red module → LASSO Cox on TCGA): no authors' code; needs exact TCGA-BRCA preprocessing + ImmuCellAI server output; not reproducible 1:1, and re-deriving it would not yield the same 15 genes/coefficients. The published signature is the documented entry point — we validate it, we don't re-derive it.
  • TCGA / METABRIC cohorts (Figs 4, 5D, 6): METABRIC is access-gated (cBioPortal /Synapse), TCGA needs the same custom normalization; GSE20685 is the clean public surrogate for the validation claim.
  • pRRophetic drug-IC50 ranking (CTRP/PRISM, Fig 9) and immunotherapy response 57%/90% (Fig 8G): both keyed off TCGA IRS deciles + external databases (CTRP/PRISM/TCIA), not reproducible without the TCGA IRS pipeline.
  • IRS↔immune-cell correlations (DataSheet_5/6, TIMER/TCIA): require recomputing IRS on TCGA + downloading TIMER/TCIA scores — out of the 80/20 low-hanging set.

Reproduction stance

Known published signature (15 coefficients, Suppl. Table S2) applied to the paper's own public validation cohort (GSE20685) at the documented entry point, on «our HPC». Deterministic given normalization choice → graded by tolerance on the p-value and by sign on the effect direction.

Figures / tables: Figure 5FFig 5FFig 3F
R1
Reported
0.0091 (GSE20685 KM log-rank p, Fig 5F)
Reproduced
0.0177 (fixed cutoff -0.03) / 0.0133 (median) / 0.0098 (continuous Cox)
within tolerance
R2
Reported
high-IRS WORSE OS (meta HR=2.72)
Reproduced
high-IRS BETTER OS, HR=0.58 (group) / 0.956 (continuous) -> direction INVERTED
did not match
R3
Reported
327 (validation n, OS>30d)
Reproduced
327
exact
F1
Reported
192 protective(neg)+1 risk(pos) coefficients (text/Fig 3F)
Reproduced
all 15 published coefficients POSITIVE (0.835-1.455)
did not match

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 51/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🔴3. Location of the main deviation
🔴4. Cause of the deviation
🔴5. Derivability / plausibility
🔴6. Severity of the deviation
🔴7. Core claim
🔴8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Main result did not reproduce
Decisive
From: Q5 · Derivability / plausibility 🔴
Main result did not reproduce
Decisive
From: Q7 · Core claim 🔴
Main result did not reproduce
Decisive
From: Q8 · Severity of the miss (overall human judgment) 🔴

The external-validation step reproduces well on the public side — n=327 matches exactly (R3) and the 15-gene signature still significantly stratifies OS (p=0.0177 vs reported 0.0091, within-tol R1). But using the exact published all-positive coefficients (Suppl Table S2) the prognostic direction inverts: high-IRS = BETTER OS (HR 0.58), opposite the paper's high-IRS-worse / meta HR=2.72 claim. This sits on the authors' side: the deposited coefficient signs contradict the paper's own text ('192 protective + 1 risk') and Fig 3F, an internal inconsistency that mechanically drives the inversion and is biologically consistent with our result (MS4A1/CD20, CCR9, KLRC3 immune-activation genes). Severe and fabrication-suspect on the directional claim, hence red on q5/q7/q8.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at [email protected].

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

120.2 k
tokens (I/O) · 7.9 M incl. cache
16 min
runtime · 0.02 CPU-h
2.5 GB
peak RAM
2 (1 failed)
HPC jobs
hummel
machine