Corpus 1,273 assessed · 1,174 scored · 643 reproduced ≥75 · 169 flagged ·∅ 74.1/100
← New search

Integrative analyses reveal signaling pathways underlying familial breast cancer susceptibility.

Mol Syst Biol · 2016
L1 67/100 3/4
Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +7
✓ What held up
  • Reported values were directly comparable
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
67/100
Reproducibility score
0.4 SD below mean
vs. all fields · 1174 studies
🎯 Scores higher than 29% of all assessed papers rank 795 of 1174 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

DESCRIBED WELL ENOUGH: partly. The authors' repo (srp33/BCRiskPathways @ fa9cb60) ships the 932 pathway gene sets, the GSE17072 class labels, the ML-Flex2 experiment config, and the EXACT permutation-p R code - enough to reproduce the structure of the normal-breast (Lim/Visvader = GSE17072) validation. It does NOT ship matrices/Visvader.txt nor the SVM wrapper (Internals/ gitignored), so the exact ML-Flex2 Java/Weka/R SVM and the Illumina probe->gene matrix are under-specified. PER BRIEF (don't chase the 20%; third-party/own implementation on the paper's data is equally valid), we rebuilt the gene matrix from the GEO series matrix (GPL6884->Entrez) and ran the DESCRIBED classifier (SVM-RFE + radial SVM via LIBSVM/scikit-learn, LOOCV) on all 931 evaluable pathways, then scored them with the authors' verbatim permutation-p logic. RESULT = 1:1-in-method, near-1:1-in-value for the headline pathway: REACTOME Integrin Cell Surface Interactions reproduced P=0.030 vs reported 0.038 (within-tol); KEGG Small Cell Lung Cancer P=0.040 vs 0.007 (both significant, same direction, magnitude off = partial); KEGG Focal Adhesion strongly significant (P=0.008), consistent with the paper highlighting it. The GSE17072 cohort itself (20 samples, 15 familial/5 control) reproduced exactly. NOT ATTEMPTED: the '9 of 45 pathways replicate' joint claim (needs GSE47862 PBMC + GSE19383 Bellacosa + cross-dataset rankPvalue/metap), the PBMC 45-pathway result (needs a non-shipped clinical-variables file), and the exome 94.2% BRCA1/2 figure (dbGaP-restricted data, data_restricted). HONESTY FLAG: 316/931 pathways reach P<0.05 in this single small cohort, so single-cohort significance is liberal - the paper's rigor comes from cross-cohort replication we did not reproduce. No fabrication signal: the two named reported p-values are derivable from the shipped data+method.

💻 Code ↗ 🗄 Data: GSE17072

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 67
    assessed: 2026-06-16 ⛓ b44639637d9f
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-16
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-16
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

The authors hypothesized that comparing molecular profiles of women from high-risk families who developed familial breast cancer (FBC) against high-risk women who did not, using a pathway-based multi-omic strategy across multiple tissues, could identify biological signaling pathways underlying FBC susceptibility.

Core claims
  • Cell adhesion (cell-cell and cell-ECM) pathways are significantly and consistently dysregulated in women who develop familial breast cancer across multiple omic data types and tissues. finding
  • A pathway-based integrative strategy combining gene expression and exome-sequencing data can elucidate the biological basis of FBC where single risk variants cannot. method
  • REACTOME Integrin Cell Surface Interactions and KEGG Small Cell Lung Cancer pathways are the top-ranked pathways discriminating FBC women from controls in both PBMC and normal breast tissue. finding
  • Normal mammary epithelial cells from FBC women show decreased cell adhesion, increased cell death with adhesion-modulating drugs, and aberrant morphology (larger spread cells, fewer cell-cell contacts). finding
  • Disrupted cell adhesion processes in non-malignant cells may play a role in FBC development and serve as an indicator of breast cancer risk. mechanism
  • Gene and protein expression for a subset of adhesion-related genes (VTN up, F-actin down) are concordantly dysregulated in FBC women. finding
  • Pathway-level germline DNA variation differs between FBC women and controls consistent with expression-based dysregulation. finding
Experimental setups
Assay System Perturbation Readout Platform
Gene expression microarray (transcriptomic profiling) Peripheral blood mononuclear cells (PBMCs) from Utah and Ontario cohorts of women with breast cancer family history none (FBC-affected vs unaffected high-risk women) Pathway-level multigene classifier accuracy discriminating FBC vs controls
Exome sequencing PBMCs from 35 Utah patients none (FBC vs control) Per-pathway germline DNA variant counts (likely pathogenic variants)
Gene expression profiling (pathway analysis) Normal breast tissue from two independent cohorts (Lim et al, Bellacosa et al) none (family history/BRCA1/2 carriers vs controls) Genomic model score discriminating high-risk vs control
Western blotting Snap-frozen breast epithelial tissue / primary mammary epithelial cell lysates from prophylactic surgery (high-risk) and breast reduction (control) women (n=27) none (FBC vs control) Protein levels of VTN, F-actin, FAK/PTK2, PTEN, ITGA6, ICAM2, p53, integrins; β-actin loading control
Fluorescence microscopy / immunostaining Normal primary mammary epithelial cells from 24 patient cultures (prophylactic mastectomy FBC vs breast reduction controls) none Cell size, F-actin staining, focal adhesion staining, cell morphology
Cell-based functional adhesion assays and drug-response assays Normal primary mammary epithelial cells from high-risk (prophylactic mastectomy) and control (breast reduction) women drugs modulating cell adhesion capacity Cell adhesion ability and cell death
Key results
  • 45 pathways consistently and significantly discriminated FBC women from controls across Utah and Ontario cohorts and both omic data types (rank P < 0.05). 45 pathways
  • REACTOME Integrin Cell Surface Interactions was a top discriminating pathway in PBMCs. P = 0.004
  • KEGG Small Cell Lung Cancer pathway discriminated FBC women from controls. P = 0.005
  • Of the 45 pathways, 9 showed significant differences in two normal breast tissue data sets and when combining all 5 data sets (P < 0.05). 9 of 45 pathways
  • Integrin Cell Surface Interactions and Small Cell Lung Cancer pathways remained top-ranked in normal breast tissue cohorts. Integrin P = 0.038 (Lim), 0.030 (Bellacosa); SCLC P = 0.007 (Lim), 0.003 (Bellacosa)
  • Vitronectin (VTN) protein increased and F-actin protein decreased in FBC women vs controls by Western blotting.
  • FBC normal mammary epithelial cells showed significant differences in cell size, F-actin staining, and focal adhesion staining vs controls.
  • FAK/PTK2 and PTEN protein levels trended lower in FBC patient samples compared to controls.
Key statistics
  • pvalue P = 0.004 (REACTOME Integrin Cell Surface Interactions pathway, PBMC analysis)
  • pvalue P = 0.005 (KEGG Small Cell Lung Cancer pathway, PBMC analysis)
  • pvalue P = 0.012 (KEGG Focal Adhesion pathway)
  • pvalue P = 0.007 (REACTOME Cell Surface Interactions at the Vascular Wall pathway)
  • pvalue P = 0.030–0.038 (Integrin Cell Surface Interactions in Lim (0.038) and Bellacosa (0.030) normal breast cohorts)
  • count 45 pathways (Pathways consistently significant across Utah/Ontario cohorts and both omic types)
  • count n = 27 (Breast epithelial tissue cohort for Western blotting)
  • count 35 patients (Utah patients whose exomes were sequenced)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The study used a multi-cohort, multi-omic design integrating gene expression microarray data from PBMCs and normal breast tissue with exome-sequencing data to identify signaling pathways dysregulated in women who developed familial breast cancer (FBC) versus unaffected high-risk controls. Pathway-level classification was performed with Support Vector Machines ranked by cross-validated accuracy, and germline variant burden was compared with Barnard's exact test; P-values were combined across cohorts via a rank-based approach and corrected for multiple testing with Storey's q-value in sensitivity analyses. Genomic findings were validated in cell-based functional assays, Western blotting, and fluorescence microscopy with significance reported at P < 0.05.

Replicationbiological Sample sizeUtah and Ontario PBMC cohorts (exact n not stated in provided text); exome-sequencing subset n=35 (Utah); Western blot cohort n=27; fluorescence microscopy n=24 patient cultures (10 selected for quantitative comparison); two independent breast tissue expression cohorts (Lim et al, Bellacosa et al; n not stated) GroupsFBC-affected vs. unaffected women from high-risk families (main comparison); BRCA1/2 mutation carriers vs. BRCAX (subgroup); high-risk prophylactic mastectomy vs. breast-reduction controls (cell-based assays) Pairingunpaired Randomization/blindingstated DispersionIQR Effect sizesno Confidence intervalsno Multiplicity correctionStorey's q-value (FDR)
Statistical tests used
Test Applied to n Assumptions
Support Vector Machines (SVM) classifier, accuracy-ranked cross-validation Pathway-level gene expression discrimination between FBC-affected and unaffected high-risk women across Utah (training), Ontario (test), Lim et al, and Bellacosa et al cohorts not stated
Barnard's exact test Pathway-level germline DNA variant burden comparison between FBC and control groups in the exome-sequencing subset 35 Utah patients not stated
Rank-based P-value combination (plus three additional unstated combination methods in sensitivity analyses) Aggregating pathway significance across multiple cohorts and omic data types; rank P-value < 0.05 threshold used for primary filtering not stated
Storey's q-value (FDR correction) Multiple-testing correction across 932 pathways in robustness/sensitivity analyses using alternative P-value combination methods not stated
Unspecified test, P < 0.05 threshold Western blot protein level comparison (VTN, F-actin, FAK/PTK2, PTEN) between FBC and control primary mammary epithelial cells 27 not stated
Unspecified test, P < 0.05 threshold Fluorescence microscopy quantification of estimated cell size, F-actin staining, and focal adhesion staining across microscopy fields from selected patients ~10 fields per patient across 10 selected patients (from 24 total cultures) not stated
Approaches that could also have been used
  • Pathway classification relied on SVMs ranked by cross-validated accuracy; the SVM itself does not produce a classical null-hypothesis P-value for pathway enrichment
    Could also: Gene Set Enrichment Analysis (GSEA), CAMERA, or fgsea could also be applied to test pathway-level enrichment in differential expression and return permutation- or rotation-based P-values with established FDR procedures — These methods are widely adopted for pathway analysis in transcriptomic studies and produce enrichment scores and FDR-adjusted statistics that align with standard reporting conventions, which may facilitate cross-study comparisons
  • P-values were combined across cohorts using a rank-based aggregation approach as the primary strategy
    Could also: Fisher's method for combining P-values (Fisher's combined probability test) or the Stouffer Z-score method are standard meta-analytic alternatives for aggregating evidence across independent cohorts — Fisher's and Stouffer's methods have well-characterized null distributions, are widely implemented, and are explicitly framed as meta-analytic procedures, which may offer clearer statistical properties and reproducibility than a rank-based threshold
  • Formal multiple-testing correction (Storey's q-value) was applied only in sensitivity analyses; the primary filter used an unadjusted rank P < 0.05 across 932 pathways
    Could also: Applying Benjamini-Hochberg FDR or Bonferroni correction as the primary multiplicity control across all 932 tested pathways would also explicitly control the false-discovery or family-wise error rate from the outset — Designating a formal FDR procedure as the primary analysis step makes the multiplicity strategy more transparent and consistent with common practice in large-scale genomic pathway studies
  • Barnard's exact test was used to compare pathway-level variant burden (proportion of samples with at least one variant) between FBC and control groups
    Could also: Fisher's exact test is another standard choice for 2×2 contingency tables comparing variant presence/absence between two groups; burden tests or sequence kernel association tests (SKAT) could also aggregate rare variant effects at the pathway level — Fisher's exact test is more universally implemented and familiar; SKAT-type approaches additionally account for variant frequency and functional annotation, which can increase power when multiple rare variants contribute to pathway burden
  • The Western blot and fluorescence microscopy group comparisons were reported only as P < 0.05 with no test name or exact P-values stated
    Could also: Explicitly stating the test used (e.g., Mann-Whitney U, Student's t-test, permutation test) and reporting exact P-values alongside effect sizes (e.g., fold-change or Cohen's d with 95% CI) would also fully characterize the inferential procedure — Naming the test and reporting exact P-values and effect sizes allows assessment of distributional assumptions, practical significance, and reproducibility — this is particularly informative at small n (n = 27 Western blot; 10 patients for microscopy quantification)
  • Dispersion in box plots was conveyed as IQR with whiskers to extreme values; no confidence intervals or variance estimates were reported for the SVM model scores or cell-based measurements
    Could also: Bootstrapped 95% confidence intervals around the cross-validated accuracy or around the median genomic model score per group could also be reported to convey estimation uncertainty — Confidence intervals communicate precision of the group-level estimates independently of sample size and would complement IQR summaries, especially given the small cohort sizes where interval width is informative
Software: SVM (Support Vector Machines; specific package or implementation not stated) · Storey's q-value method (specific R package or implementation not stated)

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
15
Impact: medium
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

GSE47862 GEO in Methods (http://purl.org/orb/Methods)
also used by 1 paper:
phs001044 dbGaP in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
rs80358061 RefSNP in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-26969729

Paper: Piccolo SR et al. "Integrative analyses reveal signaling pathways underlying familial breast cancer susceptibility." Mol Syst Biol 2016. PMID 26969729 · PMCID PMC4812528 · DOI 10.15252/msb.20156506 Authors' code: https://github.com/srp33/BCRiskPathways @ fa9cb603373a196d1370bd8ffd4b5d7d883f41de Assigned data accession: GEO GSE17072 = the Lim et al. (Visvader) normal breast tissue validation cohort (GPL6884 Illumina, 20 QC-passed samples).

The paper's computational components

# Result Pipeline / tool In scope? Why
1 PBMC pathway dysregulation → 45 pathways (rank P<0.05), Dataset EV1 SCAN normalization + ComBat + ML-Flex2 (SVM-RFE + radial SVM) + WGCNA rankPvalue, on Utah+Ontario PBMC (GSE47862) partial needs GSE47862 + a non-shipped clinical/demographic variables file (secure/epidemiologic.txt, used to exclude 334 genes) → exact PBMC result not reproducible. Not the assigned accession.
2 Normal-breast validation: 9/45 pathways replicate; e.g. REACTOME Integrin Cell Surface Interactions P=0.038 (Lim/GSE17072), KEGG Small Cell Lung Cancer P=0.007 (Lim) — Fig 3, Datasets EV1/EV12 ML-Flex2 per-pathway LOOCV SVM-RFE → AUC → permutation P (code/CalculatePredictionMetrics.R), on GSE17072 with shipped genesets + class labels YES (primary) assigned accession; data public; class labels + 932 gene sets + the exact permutation-P R code are all shipped. Matrix must be rebuilt from GEO.
3 Exome variant counts (941,507 loci → 6,908; 94.2% BRCA1/2 detection) BWA+GATK+snpEff pipeline NO data are dbGaP phs001044 + "require security approvals"data_restricted.
4 Wet-lab assays (adhesion, FAK inhibitor EC50, microscopy, Western blot) bench / ImageJ / GraphPad NO non_pipeline (wet-lab).

What we attempt (80/20)

Component 2 on GSE17072 — the clearly-specified, low-hanging, fully-shipped validation pipeline keyed to the assigned accession. Two pinnable claims: the Lim permutation P-value for (a) REACTOME Integrin Cell Surface Interactions (paper 0.038) and (b) KEGG Small Cell Lung Cancer (paper 0.007), plus the structure "≈9 of 45 / how many of 932 pathways are significant."

Honest fidelity note (the hard ~20%, not chased)

The permutation P-value is a deterministic function of the per-pathway LOOCV AUC (CalculatePredictionMetrics.R: shuffle probabilities vs fixed labels, 1000 fixed seeds, P=(#≥AUC)/1000 + 1/1000). We reproduce that R step verbatim (authors' own logic). The only reimplemented piece is the LOOCV SVM probability vector: the authors' ML-Flex2 SVM wrapper (Internals/R/Predict.R) is .gitignored, and matrices/Visvader.txt is not shipped, so the exact SVM hyper-parameters and the Illumina probe→gene matrix construction are under-specified. We therefore apply the described method (SVM-RFE + radial SVM via scikit-learn's LIBSVM — the same LIBSVM the paper cites — feature scaling on, gamma=1/n_features, nested-CV C tuning) to the paper's own data (P16: a faithful method re-run on the paper's data is equally valid). Exact bit-for-bit P-values are NOT expected; we report ours beside the paper's and let the human auditor judge.

Figures / tables: Fig 3Fig 2B
C1
Reported
REACTOME Integrin Cell Surface Interactions P=0.038 (Lim/GSE17072)
Reproduced
P=0.030 (AUC 0.813)
within tolerance
C2
Reported
KEGG Small Cell Lung Cancer P=0.007 (Lim/GSE17072)
Reproduced
P=0.040 (AUC 0.787)
partial
C3
Reported
KEGG Focal Adhesion highlighted (no exact Lim P)
Reproduced
P=0.008 (AUC 0.867)
partial
C4
Reported
GSE17072 validation cohort = 20 samples
Reproduced
20 samples, 15 familial vs 5 control
exact
C5
Reported
9 of 45 PBMC pathways replicate in normal breast
Reproduced
not attempted (out of scope)
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 67/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟢2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +7

In-method, near-1:1-in-value reproduction of the GSE17072 (Lim/Visvader) normal-breast validation: the cohort (20 samples, 15/5) and the authors' exact permutation-p R code were reused, and the headline REACTOME Integrin pathway reproduced P=0.030 vs 0.038 (within-tol). Deviations are on our side — the unshipped SVM wrapper and expression matrix forced a self-chosen probe->gene collapse and approximate hyperparameters, producing one ~6x gap (KEGG Small Cell Lung Cancer 0.040 vs 0.007) while keeping direction and significance. No fabrication signal: both named p-values are derivable from the shipped data+method. Confirmation is limited, not full, because the paper's actual stringency (9/45 pathways replicating across 5 datasets) was out of scope and single-cohort significance is liberal (34% of pathways reach P<0.05).

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

250.3 k
tokens (I/O) · 25.6 M incl. cache
31 min
runtime · 0.86 CPU-h
0.4 GB
peak RAM
1 (1 failed)
HPC jobs
hummel
machine