Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Comparative analysis of molecular signatures reveals a hybrid approach in breast cancer: Combining the Nottingham Prognostic Index with gene expressions into a

PLoS One · 2022
L1 48/100 3/4
⚑ Flagged for review — a reproduced result did not match the reported value

Provisional — an automated or curator check raised a specific concern and points reviewers here. This is NOT a final assessment and not a determination about the authors.

Why this verdict

The main result did not reproduce in this reproduction attempt. Where our recomputation produced values that differ from the published ones, those discrepancies are listed below. This is a single automated attempt — not peer review and not a finding of error or misconduct — and differences can also arise from data access, undocumented parameters or the computing environment. The verdict can be contested via “report an error”.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Input / endpoint not comparable 1:1
+1 pts
From: Q1 · Data identity 🔴
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +7
✓ What held up
  • Nothing in this column.
What did not (or only partly)
  • 🔴Could not use the authors’ exact input data
  • 🟡Reported values were only indirectly comparable
  • 🔴A deviation arose in the data or preprocessing
  • 🔴A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
48/100
Reproducibility score
1.5 SD below mean
vs. all fields · 1173 studies
🎯 Scores higher than 6% of all assessed papers rank 1092 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough only in PART. The repo (DiTscho/HybridSignature @42a1464, 100% R) ships the analysis scripts plus four small Rdata artifacts (selections, randomSignatures, patient_ids.training/test) but NOT the input matrix. The MAIN pipeline (hybrid_signature_analysis.R -> Tables 1-3, Test Set 1) requires Metabric.data, which is access-restricted (METABRIC/EGA DAC, EGAC00001000484), so the headline C-index/IAUC/SSS/HR numbers could NOT be regenerated. The paper's GSE96058 (SCAN-B, Test Set 2) external validation has NO shipped code at all, so although GSE96058 is public on GEO it was not attempted (would require re-implementing the whole external pipeline + a METABRIC-trained model + the exact n=440 event-ratio downsample -- the hard >20%). What WAS reproduced 1:1 on «our HPC» (SLURM 2177279, R 4.3.3) from the public shipped data: signature gene composition and the 100-random-signature metric distribution (suppl_100_random_signatures.R). Results: OncotypeDx=21 genes (exact match), EndoPredict=15 (exact), Hybrid=14 genes+NPI (count consistent); random-sig mean C-index 0.6240 (95% CI 0.6198-0.6282), Brier 0.1614, R2 0.0557, LogRank 39.6. TWO NOTABLE MISMATCHES flagged for human review: (a) shipped patient_ids give a 703/131 train/test split (834 total), not the paper's stated 883/379 (1262 total); (b) shipped selections$hybrid genes are DISJOINT from the paper's Table 1a Hybrid signature (only NPI shared). These could be a stale/placeholder commit rather than misreporting and are not resolvable without the restricted METABRIC data; they do not by themselves prove the paper's tables are wrong. NOT attempted: anything needing METABRIC (all Test Set 1 tables, Table 1 HRs) and the GSE96058 external validation.

💻 Code ↗ 🗄 Data: GSE96058

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 48
    assessed: 2026-06-15 ⛓ ca910f356fe3
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-15
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

Can a signature be constructed whose prognostic performance is substantially better than competing breast cancer signatures and random gene sets, and what information gain do molecular signatures provide—specifically by combining the Nottingham Prognostic Index with gene expressions into a hybrid signature?

Core claims
  • A hybrid signature combining the Nottingham Prognostic Index with SIS-selected gene expressions can be built in a data-driven fashion (NPI treated as a gene expression during feature selection). method
  • Random sets of genes are significantly associated with disease outcome, so most signatures are not substantially better than random gene sets. finding
  • None of the evaluated signatures (Hybrid, OncotypeDxGL, EndoPredictGL, random) can be considered the main candidate for providing prognostic information. finding
  • Both the Hybrid signature and the OncotypeDx gene expression list identify patients who may not require adjuvant chemotherapy. finding
  • Combining signatures substantially improves identification of patients who do not need adjuvant chemotherapy. finding
  • The Signature Skill Score is introduced as a simple measure to assess improvement over random signatures. method
  • Overlap among the six commercially available breast cancer signatures is almost non-existent; not a single gene is shared by all six. finding
  • Sure Independence Screening (SIS) was used to rank and select the 15 most important genes including the NPI. method
Experimental setups
Assay System Perturbation Readout Platform
Gene expression microarray (bulk transcriptomics) METABRIC cohort, ER+/Her2- breast cancer patients (human tumor tissue) none (observational; patients who did not receive chemotherapy) log2-normalized expression of ~24,000 genes; long-term survival Illumina HT-12 v3
RNA-seq gene expression (external validation) SCAN-B / GSE96058 cohort, ER+/Her2- breast cancer patients <80 years (human) none (observational; no chemotherapy) log2-normalized expression values; overall survival
Feature selection (Sure Independence Screening) METABRIC training set (n=883) none 15 most important genes (Hybrid signature including NPI) ranked by importance for survival R package SIS (cox response, lasso penalty, bic tuning)
Cox proportional hazards regression METABRIC training set (n=883) none hazard ratios, 95% CI, p-values per signature gene R
Decision Trees and survival analysis METABRIC test set 1 (n=379) and test set 2 (downsampled GSE96058) none patient outcome stratification / survival prediction R
Batch effect adjustment METABRIC and GSE96058 expression data none batch-corrected expression matrices R package sva (ComBat)
Key results
  • NPI is a significant predictor in the Hybrid signature Cox regression HR=1.54
  • PREX1 strongly associated with outcome in Hybrid signature HR=0.62
  • NPI alone is a significant predictor of survival HR=1.65
  • BIRC5 significantly associated with outcome in EndoPredictGL HR=1.61
  • PGR significantly associated with outcome in OncotypeDxGL HR=0.83
  • Random signature genes (e.g. PLAU, SPON1) significantly associated with outcome, supporting that random gene sets predict outcome PLAU HR=1.38; SPON1 HR=0.72
  • Low agreement between MammaPrint and OncotypeDx reported in OPTIMA study κ=0.40 (95% CI 0.30–0.49)
  • Any set of 100 or more randomly selected genes has high probability of significant association with patient outcome probability=0.9
Key statistics
  • count 1262 selected ER+/Her2- samples (883 training / 379 test set 1) (METABRIC samples selected from initial 2136)
  • count 1381 selected ER+/Her2- samples (GSE96058/SCAN-B samples from initial 3273)
  • correlation κ = 0.40 (95% CI: 0.30–0.49) (agreement between MammaPrint and OncotypeDx (OPTIMA study))
  • other probability of 0.9 (≥100 random genes significantly associated with outcome (Venet et al.))
  • fold_change HR=1.54, p<0.001 (NPI in Hybrid signature Cox regression)
  • fold_change HR=1.65, p<0.001 (NPI alone Cox regression)
  • pvalue p<0.001 (PREX1 in Hybrid signature (HR=0.62))
  • count 15 genes (Hybrid signature size selected by SIS (including NPI))

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The paper uses the METABRIC cohort (n=1262 ER+/HER2- patients not receiving chemotherapy) split 70/30 into training and internal test sets; an independent external cohort (SCAN-B/GSE96058, n=1381 selected patients) serves as test set 2 via 1000 random downsamplings matched on event-to-patients-at-risk ratio. Sure Independence Screening (SIS) with a Cox-LASSO penalty and BIC tuning selected 15 features (genes plus NPI) from the training set to form the 'Hybrid signature', which was then compared with OncotypeDxGL, EndoPredictGL, NPI alone, Random, and NPI+Random signatures via Cox proportional hazards regression. Results are reported as per-gene hazard ratios with 95% confidence intervals and p-values; patient stratification is additionally assessed using decision trees and survival analysis.

Replicationbiological Sample size1262 METABRIC patients (883 training / 379 test set 1, 70/30 random split); 1381 SCAN-B patients for test set 2, downsampled 1000 times to match training set event-to-patients-at-risk ratio GroupsSix signatures: Hybrid, EndoPredictGL, OncotypeDxGL, NPI alone, Random (15 genes), NPI+Random; random and NPI+Random serve as negative controls Pairingunpaired Randomization/blindingnot stated Dispersionnone Exact p-valuesyes Effect sizesyes Confidence intervalsyes Multiplicity correctionnone stated
Statistical tests used
Test Applied to n Assumptions
Cox proportional hazards regression (per-gene and signature-level) All six signatures (Hybrid, EndoPredictGL, OncotypeDxGL, NPI, Random, NPI+Random) fitted in METABRIC training set (Table 1) and applied to both test sets 883 (training); 379 (test set 1); downsampled subset of 1381 repeated 1000 times (test set 2) not stated
Sure Independence Screening (SIS) with Cox response, LASSO penalty, BIC tuning Feature selection of 15 most important genes/covariates from training set 883 not stated
Decision Trees (survival stratification) Patient risk-group stratification and survival outcome assessment across signatures na
Survival analysis (specific test not named in available text) Comparison of patient outcome groups stratified by signature risk categories not stated
Approaches that could also have been used
  • Overall signature performance was assessed via per-gene Cox HRs in separate models rather than a single summary discrimination metric per signature
    Could also: Harrell's concordance index (C-statistic) or time-dependent AUC with bootstrap confidence intervals could also be computed per signature — A discrimination metric directly quantifies each signature's overall ability to rank patients by risk, enabling straightforward head-to-head comparison between signatures — which is the paper's central research question
  • Individual gene-level p-values are reported across multiple Cox models (up to 21 genes per model, six models) without stated multiple testing correction
    Could also: Benjamini-Hochberg FDR correction or Bonferroni correction applied within each signature model (or across all tests) could also be used — Adjusting for the number of simultaneous comparisons is a standard practice when reporting many p-values from the same dataset, and would make the gene-level thresholds more conservative
  • Feature selection used SIS with a single 70/30 training/test split
    Could also: Repeated k-fold cross-validation or bootstrap-based stability selection could also be used to assess feature selection robustness — These approaches provide internal estimates of both model performance and the consistency with which each feature is selected, reducing sensitivity to the particular random split
  • External validation on SCAN-B used random downsampling to match event-to-patients-at-risk ratio, repeated 1000 times
    Could also: Validation on the full external cohort using the C-statistic or calibration plots (e.g., observed vs. predicted survival curves) with bootstrap CIs could also be used — Using the complete external cohort avoids information loss from downsampling and provides a direct assessment of discrimination and calibration in an independent population
  • The proportional hazards assumption of the Cox models is not stated as having been verified
    Could also: Schoenfeld residual tests or log-log survival plots could also be used to assess whether the hazard ratio remains constant over time for each covariate — With follow-up extending up to 30 years, time-varying effects are plausible for some genes; verifying this assumption is a standard diagnostic step and informs whether time-varying coefficient extensions are warranted
  • Standard decision trees were used for patient survival stratification
    Could also: Survival-specific recursive partitioning (e.g., rpart with a survival splitting rule) or conditional inference trees could also be used — Survival-tailored tree methods optimize splits directly on the survival outcome and handle censoring natively, which may yield more accurate and statistically grounded risk strata than general-purpose decision trees
Software: R/SIS · R/limma (avereps) · R/sva (ComBat batch correction) · R/illuminaHumanv3.db

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
11
Impact: medium
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

GSE96058 GEO in Results (http://purl.org/orb/Results)
also used by 2 papers:
EGAS00000000083 EGA in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
NCT02306096 NCT in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet

Downstream reach in the literature

199 downstream papers · 2 datasets

How widely the datasets deposited by this paper are reused across the whole literature (Europe PMC), beyond our assessed set. This is a factual dependency map — reusing a public dataset is normal, good science. It is not a judgement on the downstream papers; the only verdict here is this paper's own, with its cited rationale.

This paper is currently under reproducibility review (see the verdict above). The map below shows where the data in question has propagated — so reuse can be traced, not so the downstream work is presumed affected.

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Figures / tables: Table
otdx_ngenes
Reported
21 genes
Reproduced
21
exact
endo_ngenes
Reported
15 genes
Reproduced
15
exact
hybrid_nfeat
Reported
15 genes + NPI
Reproduced
14 genes + NPI (15 Cox terms)
partial
hybrid_genes
Reported
CDC20,CLSPN,SYTL4,NUSAP1,MELK,CEP55,CKAP2L,PTTG1,IRF3,PREX1,OR5M11,CCNB2,NLRP1,BUB1,NPI (Table 1a)
Reproduced
GAS7,CCT6B,LAMA2,DNAJB9,PDCD4,CLIC6,ARHGEF5,SORCS2,ADAMTSL2,ACTN1,COL16A1,KLRB1,POLR2D,ENC1,NPI
did not match
train_n
Reported
883 (70%)
Reproduced
703
did not match
test1_n
Reported
379 (30%)
Reproduced
131
did not match
rand_cindex_mean
Reported
supp figure (not text-extractable)
Reproduced
0.6240 (95% CI 0.6198-0.6282)
partial
rand_brier_mean
Reported
supp figure (not text-extractable)
Reproduced
0.1614 (95% CI 0.1600-0.1628)
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 48/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🔴1. Data identity
🟡2. Endpoint comparability
🔴3. Location of the main deviation
🔴4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Input / endpoint not comparable 1:1
+1 pts
From: Q1 · Data identity 🔴
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +7

Only auxiliary claims were checkable from public shipped data and these reproduced exactly (OncotypeDx=21, EndoPredict=15 genes; random-signature mean C-index 0.6240). The study's headline result is untestable here because the main pipeline needs access-restricted METABRIC and the GSE96058 validation ships no code — that is a data-availability limitation, not proof the paper is wrong. However, the deposited artifacts internally contradict the paper (hybrid genes disjoint from Table 1a; 703/131 vs stated 883/379 split), which is a genuine flag on the authors'/deposit side that a human must resolve; it is plausibly a stale/placeholder commit and cannot be settled without the restricted data, so this is a flagged partial, not confirmed fabrication.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

96.9 k
tokens (I/O) · 6.1 M incl. cache
12 min
runtime · 0.01 CPU-h
2.5 GB
peak RAM
1
HPC jobs
hummel
machine