Corpus 1,286 assessed · 1,187 scored · 648 reproduced ≥75 · 174 flagged ·∅ 73.9/100
← New search

Integrative analyses reveal signaling pathways underlying familial breast cancer susceptibility.

Mol Syst Biol · 2016
L1 67/100 3/4
Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +7
✓ What held up
  • Reported values were directly comparable
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
67/100
Reproducibility score
0.4 SD below mean
vs. all fields · 1187 studies
🎯 Scores higher than 30% of all assessed papers rank 801 of 1187 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

DESCRIBED WELL ENOUGH: partly. The authors' repo (srp33/BCRiskPathways @ fa9cb60) ships the 932 pathway gene sets, the GSE17072 class labels, the ML-Flex2 experiment config, and the EXACT permutation-p R code - enough to reproduce the structure of the normal-breast (Lim/Visvader = GSE17072) validation. It does NOT ship matrices/Visvader.txt nor the SVM wrapper (Internals/ gitignored), so the exact ML-Flex2 Java/Weka/R SVM and the Illumina probe->gene matrix are under-specified. PER BRIEF (don't chase the 20%; third-party/own implementation on the paper's data is equally valid), we rebuilt the gene matrix from the GEO series matrix (GPL6884->Entrez) and ran the DESCRIBED classifier (SVM-RFE + radial SVM via LIBSVM/scikit-learn, LOOCV) on all 931 evaluable pathways, then scored them with the authors' verbatim permutation-p logic. RESULT = 1:1-in-method, near-1:1-in-value for the headline pathway: REACTOME Integrin Cell Surface Interactions reproduced P=0.030 vs reported 0.038 (within-tol); KEGG Small Cell Lung Cancer P=0.040 vs 0.007 (both significant, same direction, magnitude off = partial); KEGG Focal Adhesion strongly significant (P=0.008), consistent with the paper highlighting it. The GSE17072 cohort itself (20 samples, 15 familial/5 control) reproduced exactly. NOT ATTEMPTED: the '9 of 45 pathways replicate' joint claim (needs GSE47862 PBMC + GSE19383 Bellacosa + cross-dataset rankPvalue/metap), the PBMC 45-pathway result (needs a non-shipped clinical-variables file), and the exome 94.2% BRCA1/2 figure (dbGaP-restricted data, data_restricted). HONESTY FLAG: 316/931 pathways reach P<0.05 in this single small cohort, so single-cohort significance is liberal - the paper's rigor comes from cross-cohort replication we did not reproduce. No fabrication signal: the two named reported p-values are derivable from the shipped data+method.

💻 Code ↗ 🗄 Data: GSE17072

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 67
    assessed: 2026-06-16 ⛓ b44639637d9f
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-16
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-16
no human curator yet
Last updated
2026-09-19

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The authors hypothesized that a pathway-based strategy examining multiple omic data types across multiple tissues (peripheral blood and breast) could reveal biological signaling pathways underlying familial breast cancer (FBC) risk, since known genetic risk variants explain at most 30% of FBC cases.

Core claims
  • Cell adhesion pathways are significantly and consistently dysregulated in women who develop familial breast cancer (FBC) finding
  • Integrated pathway-based analysis of gene expression and exome-sequencing data from peripheral blood mononuclear cells (PBMCs) can identify pathways associated with FBC development method
  • Cell adhesion pathway dysregulation identified in PBMCs was also detected in normal breast tissue from two independent cohorts of high-risk/BRCA1/2 women finding
  • Genomic pathway findings were validated using cell-based functional assays, drug-response assays, fluorescence microscopy, and Western blotting in normal primary mammary epithelial cells method
  • Normal mammary epithelial cells from FBC women show decreased cell adhesion ability, increased cell death upon treatment with adhesion-modulating drugs, and aberrant morphology compared to controls finding
  • Vitronectin (VTN) protein levels are increased and F-actin levels are decreased in breast epithelial tissue from FBC/high-risk women compared to controls finding
  • 45 canonical signaling pathways consistently and significantly discriminated FBC women from controls across Utah and Ontario cohorts and both gene expression and mutation data types finding
  • FAK/PTK2 and PTEN protein levels trended lower in FBC patient samples compared to controls finding
Experimental setups
Assay System Perturbation Readout Platform
gene expression microarray peripheral blood mononuclear cells (PBMCs), Utah and Ontario cohorts none (FBC vs familial non-cancer controls) pathway-level classification accuracy (SVM) discriminating FBC vs control
exome sequencing PBMCs, Utah cohort (n=35) none (FBC vs control) pathway-level germline variant burden (Barnard's exact test)
gene expression microarray normal breast tissue, Lim et al. and Bellacosa et al. cohorts none (family history/BRCA1/2 carriers vs controls) pathway-level discrimination score (genomic model score)
Western blotting primary mammary epithelial cells, prophylactic surgery vs breast reduction cohort (n=27) none (high-risk vs control) protein levels of FAK/PTK2, ITGA6, ICAM2, p53, PTEN, ITGA IV, ITGA V, VTN, F-actin, β-actin
fluorescence microscopy primary mammary epithelial cells, 24 patient cultures none (high-risk prophylactic mastectomy vs breast reduction control) cell size, F-actin staining, focal adhesion staining, cell morphology
cell-based functional and drug-response assays primary mammary epithelial cells (high-risk vs control) drugs that modulate cell adhesion capacity cell adhesion ability and drug-induced cell death
Key results
  • REACTOME Integrin Cell Surface Interactions pathway was top-performing pathway in PBMC gene expression/mutation analysis P = 0.004
  • KEGG Small Cell Lung Cancer pathway was second top-performing pathway in PBMC analysis P = 0.005
  • KEGG Focal Adhesion and REACTOME Cell Surface Interactions at the Vascular Wall pathways also performed well in PBMC analysis P = 0.012 and P = 0.007
  • Integrin Cell Surface Interactions and Small Cell Lung Cancer pathways significantly discriminated high-risk/BRCA1/2 vs control women in normal breast tissue data sets P = 0.038/0.030 (Integrin); P = 0.007/0.003 (SCLC)
  • Vitronectin (VTN) protein levels increased and F-actin decreased in FBC/high-risk women vs controls by Western blot P-value < 0.05
  • FAK/PTK2 and PTEN protein levels trended lower in FBC samples vs controls (not stated as statistically significant)
  • Significant differences in cell size, F-actin staining, and focal adhesion staining between FBC and control mammary epithelial cell cultures P < 0.05
  • 45 pathways consistently and significantly discriminated FBC women from controls across Utah/Ontario cohorts and gene expression/mutation data rank P-value < 0.05
Key statistics
  • pvalue P = 0.004 (REACTOME Integrin Cell Surface Interactions pathway, PBMC analysis)
  • pvalue P = 0.005 (KEGG Small Cell Lung Cancer pathway, PBMC analysis)
  • pvalue P = 0.012 (KEGG Focal Adhesion pathway, PBMC analysis)
  • pvalue P = 0.007 (REACTOME Cell Surface Interactions at the Vascular Wall pathway, PBMC analysis)
  • pvalue P = 0.038 (Lim et al) and P = 0.030 (Bellacosa et al) (Integrin Cell Surface Interactions pathway in normal breast tissue data sets)
  • pvalue P = 0.007 and 0.003 (Small Cell Lung Cancer pathway in normal breast tissue data sets (Lim, Bellacosa))
  • count n = 35 (Utah patients with exome sequencing data)
  • count n = 27 (breast epithelial tissue samples used for Western blotting)
  • count 24 patient cultures; 10 patients quantified with ~10 microscope fields each (fluorescence microscopy of primary mammary epithelial cells)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The study used a multi-cohort, multi-omic design integrating gene expression microarray data from PBMCs and normal breast tissue with exome-sequencing data to identify signaling pathways dysregulated in women who developed familial breast cancer (FBC) versus unaffected high-risk controls. Pathway-level classification was performed with Support Vector Machines ranked by cross-validated accuracy, and germline variant burden was compared with Barnard's exact test; P-values were combined across cohorts via a rank-based approach and corrected for multiple testing with Storey's q-value in sensitivity analyses. Genomic findings were validated in cell-based functional assays, Western blotting, and fluorescence microscopy with significance reported at P < 0.05.

Replicationbiological Sample sizeUtah and Ontario PBMC cohorts (exact n not stated in provided text); exome-sequencing subset n=35 (Utah); Western blot cohort n=27; fluorescence microscopy n=24 patient cultures (10 selected for quantitative comparison); two independent breast tissue expression cohorts (Lim et al, Bellacosa et al; n not stated) GroupsFBC-affected vs. unaffected women from high-risk families (main comparison); BRCA1/2 mutation carriers vs. BRCAX (subgroup); high-risk prophylactic mastectomy vs. breast-reduction controls (cell-based assays) Pairingunpaired Randomization/blindingstated DispersionIQR Effect sizesno Confidence intervalsno Multiplicity correctionStorey's q-value (FDR)
Statistical tests used
Test Applied to n Assumptions
Support Vector Machines (SVM) classifier, accuracy-ranked cross-validation Pathway-level gene expression discrimination between FBC-affected and unaffected high-risk women across Utah (training), Ontario (test), Lim et al, and Bellacosa et al cohorts not stated
Barnard's exact test Pathway-level germline DNA variant burden comparison between FBC and control groups in the exome-sequencing subset 35 Utah patients not stated
Rank-based P-value combination (plus three additional unstated combination methods in sensitivity analyses) Aggregating pathway significance across multiple cohorts and omic data types; rank P-value < 0.05 threshold used for primary filtering not stated
Storey's q-value (FDR correction) Multiple-testing correction across 932 pathways in robustness/sensitivity analyses using alternative P-value combination methods not stated
Unspecified test, P < 0.05 threshold Western blot protein level comparison (VTN, F-actin, FAK/PTK2, PTEN) between FBC and control primary mammary epithelial cells 27 not stated
Unspecified test, P < 0.05 threshold Fluorescence microscopy quantification of estimated cell size, F-actin staining, and focal adhesion staining across microscopy fields from selected patients ~10 fields per patient across 10 selected patients (from 24 total cultures) not stated
Approaches that could also have been used
  • Pathway classification relied on SVMs ranked by cross-validated accuracy; the SVM itself does not produce a classical null-hypothesis P-value for pathway enrichment
    Could also: Gene Set Enrichment Analysis (GSEA), CAMERA, or fgsea could also be applied to test pathway-level enrichment in differential expression and return permutation- or rotation-based P-values with established FDR procedures — These methods are widely adopted for pathway analysis in transcriptomic studies and produce enrichment scores and FDR-adjusted statistics that align with standard reporting conventions, which may facilitate cross-study comparisons
  • P-values were combined across cohorts using a rank-based aggregation approach as the primary strategy
    Could also: Fisher's method for combining P-values (Fisher's combined probability test) or the Stouffer Z-score method are standard meta-analytic alternatives for aggregating evidence across independent cohorts — Fisher's and Stouffer's methods have well-characterized null distributions, are widely implemented, and are explicitly framed as meta-analytic procedures, which may offer clearer statistical properties and reproducibility than a rank-based threshold
  • Formal multiple-testing correction (Storey's q-value) was applied only in sensitivity analyses; the primary filter used an unadjusted rank P < 0.05 across 932 pathways
    Could also: Applying Benjamini-Hochberg FDR or Bonferroni correction as the primary multiplicity control across all 932 tested pathways would also explicitly control the false-discovery or family-wise error rate from the outset — Designating a formal FDR procedure as the primary analysis step makes the multiplicity strategy more transparent and consistent with common practice in large-scale genomic pathway studies
  • Barnard's exact test was used to compare pathway-level variant burden (proportion of samples with at least one variant) between FBC and control groups
    Could also: Fisher's exact test is another standard choice for 2×2 contingency tables comparing variant presence/absence between two groups; burden tests or sequence kernel association tests (SKAT) could also aggregate rare variant effects at the pathway level — Fisher's exact test is more universally implemented and familiar; SKAT-type approaches additionally account for variant frequency and functional annotation, which can increase power when multiple rare variants contribute to pathway burden
  • The Western blot and fluorescence microscopy group comparisons were reported only as P < 0.05 with no test name or exact P-values stated
    Could also: Explicitly stating the test used (e.g., Mann-Whitney U, Student's t-test, permutation test) and reporting exact P-values alongside effect sizes (e.g., fold-change or Cohen's d with 95% CI) would also fully characterize the inferential procedure — Naming the test and reporting exact P-values and effect sizes allows assessment of distributional assumptions, practical significance, and reproducibility — this is particularly informative at small n (n = 27 Western blot; 10 patients for microscopy quantification)
  • Dispersion in box plots was conveyed as IQR with whiskers to extreme values; no confidence intervals or variance estimates were reported for the SVM model scores or cell-based measurements
    Could also: Bootstrapped 95% confidence intervals around the cross-validated accuracy or around the median genomic model score per group could also be reported to convey estimation uncertainty — Confidence intervals communicate precision of the group-level estimates independently of sample size and would complement IQR summaries, especially given the small cohort sizes where interval width is informative
Software: SVM (Support Vector Machines; specific package or implementation not stated) · Storey's q-value method (specific R package or implementation not stated)

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
15
Impact: medium
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

GSE47862 GEO in Methods (http://purl.org/orb/Methods)
also used by 1 paper:
phs001044 dbGaP in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
rs80358061 RefSNP in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-26969729

Paper: Piccolo SR et al. "Integrative analyses reveal signaling pathways underlying familial breast cancer susceptibility." Mol Syst Biol 2016. PMID 26969729 · PMCID PMC4812528 · DOI 10.15252/msb.20156506 Authors' code: https://github.com/srp33/BCRiskPathways @ fa9cb603373a196d1370bd8ffd4b5d7d883f41de Assigned data accession: GEO GSE17072 = the Lim et al. (Visvader) normal breast tissue validation cohort (GPL6884 Illumina, 20 QC-passed samples).

The paper's computational components

# Result Pipeline / tool In scope? Why
1 PBMC pathway dysregulation → 45 pathways (rank P<0.05), Dataset EV1 SCAN normalization + ComBat + ML-Flex2 (SVM-RFE + radial SVM) + WGCNA rankPvalue, on Utah+Ontario PBMC (GSE47862) partial needs GSE47862 + a non-shipped clinical/demographic variables file (secure/epidemiologic.txt, used to exclude 334 genes) → exact PBMC result not reproducible. Not the assigned accession.
2 Normal-breast validation: 9/45 pathways replicate; e.g. REACTOME Integrin Cell Surface Interactions P=0.038 (Lim/GSE17072), KEGG Small Cell Lung Cancer P=0.007 (Lim) — Fig 3, Datasets EV1/EV12 ML-Flex2 per-pathway LOOCV SVM-RFE → AUC → permutation P (code/CalculatePredictionMetrics.R), on GSE17072 with shipped genesets + class labels YES (primary) assigned accession; data public; class labels + 932 gene sets + the exact permutation-P R code are all shipped. Matrix must be rebuilt from GEO.
3 Exome variant counts (941,507 loci → 6,908; 94.2% BRCA1/2 detection) BWA+GATK+snpEff pipeline NO data are dbGaP phs001044 + "require security approvals"data_restricted.
4 Wet-lab assays (adhesion, FAK inhibitor EC50, microscopy, Western blot) bench / ImageJ / GraphPad NO non_pipeline (wet-lab).

What we attempt (80/20)

Component 2 on GSE17072 — the clearly-specified, low-hanging, fully-shipped validation pipeline keyed to the assigned accession. Two pinnable claims: the Lim permutation P-value for (a) REACTOME Integrin Cell Surface Interactions (paper 0.038) and (b) KEGG Small Cell Lung Cancer (paper 0.007), plus the structure "≈9 of 45 / how many of 932 pathways are significant."

Honest fidelity note (the hard ~20%, not chased)

The permutation P-value is a deterministic function of the per-pathway LOOCV AUC (CalculatePredictionMetrics.R: shuffle probabilities vs fixed labels, 1000 fixed seeds, P=(#≥AUC)/1000 + 1/1000). We reproduce that R step verbatim (authors' own logic). The only reimplemented piece is the LOOCV SVM probability vector: the authors' ML-Flex2 SVM wrapper (Internals/R/Predict.R) is .gitignored, and matrices/Visvader.txt is not shipped, so the exact SVM hyper-parameters and the Illumina probe→gene matrix construction are under-specified. We therefore apply the described method (SVM-RFE + radial SVM via scikit-learn's LIBSVM — the same LIBSVM the paper cites — feature scaling on, gamma=1/n_features, nested-CV C tuning) to the paper's own data (P16: a faithful method re-run on the paper's data is equally valid). Exact bit-for-bit P-values are NOT expected; we report ours beside the paper's and let the human auditor judge.

Figures / tables: Fig 3Fig 2B
C1
Reported
REACTOME Integrin Cell Surface Interactions P=0.038 (Lim/GSE17072)
Reproduced
P=0.030 (AUC 0.813)
within tolerance
C2
Reported
KEGG Small Cell Lung Cancer P=0.007 (Lim/GSE17072)
Reproduced
P=0.040 (AUC 0.787)
partial
C3
Reported
KEGG Focal Adhesion highlighted (no exact Lim P)
Reproduced
P=0.008 (AUC 0.867)
partial
C4
Reported
GSE17072 validation cohort = 20 samples
Reproduced
20 samples, 15 familial vs 5 control
exact
C5
Reported
9 of 45 PBMC pathways replicate in normal breast
Reproduced
not attempted (out of scope)
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 67/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟢2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +7

In-method, near-1:1-in-value reproduction of the GSE17072 (Lim/Visvader) normal-breast validation: the cohort (20 samples, 15/5) and the authors' exact permutation-p R code were reused, and the headline REACTOME Integrin pathway reproduced P=0.030 vs 0.038 (within-tol). Deviations are on our side — the unshipped SVM wrapper and expression matrix forced a self-chosen probe->gene collapse and approximate hyperparameters, producing one ~6x gap (KEGG Small Cell Lung Cancer 0.040 vs 0.007) while keeping direction and significance. No fabrication signal: both named p-values are derivable from the shipped data+method. Confirmation is limited, not full, because the paper's actual stringency (9/45 pathways replicating across 5 datasets) was out of scope and single-cohort significance is liberal (34% of pathways reach P<0.05).

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at [email protected].

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

250.3 k
tokens (I/O) · 25.6 M incl. cache
31 min
runtime · 0.86 CPU-h
0.4 GB
peak RAM
1 (1 failed)
HPC jobs
hummel
machine