Increased prevalence of hybrid epithelial/mesenchymal state and enhanced phenotypic heterogeneity in basal breast cancer.
The main results reproduced: recomputed values matched the published ones within tolerance.
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
REPRODUCED 1:1 (WT landscape AND over-expression). RACIPE-1.0 (commit b4483533e69b1f12e11c3e0d0738d2d094e4ab6f) built on «our HPC»; topology re-derived from Supplementary Table S4 (mmc5.xlsx): WT 11-gene circuit = 38 edges (EXACT match to vetted reference), signals circuit = 46 edges over EXACTLY 13 nodes (11 genes + Tgfb + Il1b). KEY CORRECTION vs the two prior attempts: S4 contains NO Notch node, so the Fig 5F over-expression network IS fully specified and the OE experiment IS reproducible (prior rooms wrongly dropped it as out-of-scope). OE done with RACIPE's native -OEID 13(Il1b)/12(Tgfb) -OEFD 100 (paper Methods: 'overexpression (OE; 100x)'); gene IDs confirmed identical to the authors' core_OE_13/core_OE_12 files. Production «job» COMPLETED (1h21m): full triplicate, 12 runs (WT + signals-base + IL1B-OE + TGFB1-OE, each 10,000 parameter sets x 100 ICs, Euler, up to 6 states), every replicate stored exactly 10,000 models (integrity-checked). Classification follows the authors' Untitled.ipynb (11-node landscape) and signals.ipynb (OE: z-score OE states against the base mean/std, composite LBT/EMT scores on the 11 phenotype genes, authors' hardcoded per-replicate signals thresholds + KDE-minima re-derivation as robustness). RESULTS: C1 WT 5-category fractions reproduce to within 0.004 (signals-base, which is what the reported numbers actually come from); C2 luminal:basal = 0.381:0.619 (reported 0.380:0.620, exact); C3 luminal almost exclusively epithelial (hybrid 15.4% vs 14.6%) and basal spans the full E/M axis (non-epithelial 73.4% vs 73.7%); C4 multistability up-to-6 and 97% multistable reproduced; OE1 IL1B-OE (+Bas_epi, +Bas_H, -Lum_H) and TGFB1-OE (+Lum_H +91%->reproduced +77%, +Bas_mes, -Bas_epi, -Bas_H) reproduce every direction with absolute OE fractions within 0.005 of reported. The 11-node core landscape (Fig 5B) classified with the core hardcoded thresholds shows a modest ~0.03-0.05 mass shift Lum->Bas_epi/Bas_H (qualitative landscape figure; Bas_mes still near-exact) -- reported as an honest nuance. All grades PROVISIONAL for human review. OUT OF SCOPE (not attempted): ssGSEA 80-dataset meta-analysis, scRNA/spatial, survival (Figs 1-4), GSE42944 methylation (Fig 2A iii) -- transcriptomic/ancillary, not the mechanistic RACIPE pipeline.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 64assessed: 2026-06-19 ⛓ 1ed81710646f
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-22
- Rubric version
- not recorded
- Assessed by
- —
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetThe paper tests whether luminal-basal molecular subtype and epithelial-mesenchymal transition (EMT) status are equivalent axes of breast cancer heterogeneity, hypothesizing instead that basal breast cancer is specifically associated with a hybrid/partial-EMT (pEMT) state and greater phenotypic heterogeneity rather than a fully mesenchymal state.
- ★ Luminal breast cancer gene expression signature is closely/positively associated with an epithelial signature finding
- ★ Basal breast cancer signature correlates more strongly with a hybrid epithelial/mesenchymal (pEMT) signature than with a fully mesenchymal signature finding
- ★ Basal breast cancer exhibits higher phenotypic heterogeneity along the EMT spectrum than luminal breast cancer finding
- ★ Epigenetic patterns (DNA methylation, H3K27ac/H3K27me3) at epithelial and mesenchymal gene promoters underlie the luminal-epithelial and basal-pEMT associations mechanism
- ★ Mathematical modeling of a gene regulatory network coupling EMT and luminal-basal differentiation axes recapitulates the observed transcriptomic associations and heterogeneity mechanism
- ★ Spatial transcriptomics reveals intra-patient variability in EMT phenotypes specifically within basal breast cancer subtypes finding
- Epithelial-luminal joint classification stratifies TCGA breast cancer patient survival better than epithelial-mesenchymal classification alone finding
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| single-sample gene set enrichment analysis (ssGSEA) of bulk transcriptomic data | 80 breast-cancer-specific datasets (cell lines and primary tumors, meta-analysis) | none | correlation between luminal, basal, epithelial, mesenchymal, and pEMT gene signature scores | — |
| survival/prognosis analysis | TCGA breast cancer patient cohort | none | hazard ratio for survival stratified by epithelial/mesenchymal and epithelial/luminal group combinations | — |
| genome-wide DNA methylation profiling | CCLE breast cancer cell lines (GSE42944), Luminal/Basal A/Basal B subtypes | none | CpG island methylation at promoters of epithelial and mesenchymal genes, correlated with gene expression | — |
| bulk RNA sequencing | four breast cancer cell lines (MCF7, ZR-75 luminal; HCC38, HMLER basal) with CD44-high/low subpopulations (GSE184647) | none | ssGSEA scores for basal, mesenchymal, and pEMT gene expression programs | — |
| MINT-ChIP (histone modification profiling) | same luminal and basal breast cancer cell lines (e.g., HMLER CD44-high/low) | none | percentage of H3K27ac activation and H3K27me3 suppressive marks at epithelial/mesenchymal gene promoters | — |
| spatial transcriptomics | breast cancer tissue sections (patient tumor samples) | none | intra-patient spatial heterogeneity of EMT phenotype | — |
| mathematical/computational modeling of gene regulatory network | in silico model of EMT-luminal/basal crosstalk network | simulated network parameters | emergent phenotypic states and heterogeneity compared to transcriptomic data | — |
- ▲ Luminal-epithelial signatures showed significant positive correlation in 36 of 80 datasets vs. only 4 significantly negative r>0.3, p<0.05
- – Basal-mesenchymal correlation was only weakly skewed positive (11 positive vs. 7 negative datasets) r>0.3/r<-0.3, p<0.05
- ▲ Basal-pEMT signature showed significant positive correlation in 21 of 80 datasets vs. only 3 negative r>0.3, p<0.05
- ▲ Paired comparison showed basal-pEMT correlation significantly greater than basal-mesenchymal correlation T=2.9, p<0.01; T=4.0, p<0.001 (EMT_partial vs EMT_up)
- ▼ Patients with high mesenchymal (EPI-MES+) tumors had worse prognosis HR=1.5, p<0.05
- ▼ Basal score negatively correlated with luminal score in CCLE cell lines r=-0.34, p<0.05
- – Basal A cell lines showed low methylation of both epithelial and mesenchymal genes, consistent with a pEMT-like co-active state, while Basal B lines resembled a fully mesenchymal methylation pattern
- – CD44-high basal subpopulations clustered as more mesenchymal, while CD44-low basal subpopulations occupied an intermediate (pEMT) region of the EMT plane
- correlation r>0.3, p<0.05 in 36/80 datasets (luminal-epithelial) (meta-analysis of breast cancer datasets)
- correlation r>0.3, p<0.05 in 21/80 datasets vs 3 negative (basal-pEMT) (meta-analysis of breast cancer datasets)
- pvalue T=2.9, p<0.01 (basal-pEMT vs basal-mesenchymal, paired t test) (paired comparison of correlation coefficients)
- pvalue T=4.0, p<0.001 (basal-EMT_partial vs basal-EMT_up, paired t test) (paired comparison of correlation coefficients)
- other hazard ratio HR=1.5, p<0.05 (EPI-MES+ TCGA breast cancer patients survival vs reference group)
- correlation r=-0.34, p<0.05 (basal vs luminal ssGSEA score in CCLE breast cancer cell lines)
- count 40 vs 6 datasets (EMT_down-Luminal positive vs negative correlation) (breast-cancer-specific EMT gene set comparison)
- count 10 vs 5 datasets (Basal-EMT_up positive vs negative correlation) (breast-cancer-specific EMT gene set comparison)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
The paper integrates multi-modal transcriptomic data (bulk RNA-seq meta-analysis of 80 datasets, CCLE, TCGA-BRCA, single-cell RNA-seq, spatial transcriptomics, methylation) to characterize associations between luminal-basal and epithelial-mesenchymal (EMT) gene programs in breast cancer. The primary analytic strategy is single-sample gene set enrichment analysis (ssGSEA) to score five gene programs per sample, followed by Pearson correlation within each dataset and paired t-tests to compare correlation magnitudes. Survival differences were assessed with hazard ratios in the TCGA cohort. Results are reported as correlation coefficients, T-statistics, p-values, and dataset counts meeting a directional threshold (r > 0.3 or r < −0.3, p < 0.05).
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Pearson's correlation coefficient (with associated p-value) | Pairwise correlation of ssGSEA scores for each of five gene programs (luminal, basal, epithelial, mesenchymal, pEMT) within each of 80 breast cancer datasets; also in CCLE and TCGA-BRCA individually | 80 datasets (sample sizes per dataset not stated in excerpt); CCLE and TCGA-BRCA sample counts not stated in excerpt | not stated |
| Paired t-test | Comparison of per-dataset Pearson r values for basal-pEMT/basal-EMT_partial pairs vs. basal-mesenchymal/basal-EMT_up pairs across the 80 datasets (Figure 1E) | 80 paired observations (one r per dataset) | not stated |
| Survival analysis (hazard ratio reported, method not named in excerpt) | Prognostic capacity of epithelial-mesenchymal vs. epithelial-luminal classification in TCGA breast cancer cohort (Figure S1A) | null | not stated |
| ssGSEA (single-sample gene set enrichment analysis) | Quantification of gene program activity per sample across all datasets, CCLE, TCGA-BRCA, and spatial transcriptomics | null | na |
-
Pearson's correlation was used to quantify pairwise associations between ssGSEA scores across datasets↳ Could also: Spearman's rank correlation could also be used for the same comparisons — Spearman's correlation makes no assumption about the distributional shape of ssGSEA scores and is robust to outliers and non-linearity; for enrichment scores that may not be normally distributed, it is often reported alongside or instead of Pearson's r
-
A fixed threshold (r > 0.3, p < 0.05) was applied within each dataset independently to count 'significantly positive' or 'significantly negative' datasets across the 80-dataset meta-analysis↳ Could also: A random-effects meta-analytic model (e.g., DerSimonian-Laird) pooling the per-dataset r values into a single weighted summary estimate with a 95% CI and heterogeneity statistic (I²) could also be used — Formal meta-analysis provides a single pooled effect size with uncertainty quantification, explicitly accounts for between-study heterogeneity, and avoids the multiple-testing accumulation that arises from counting independent per-dataset p-values
-
A paired t-test was used to compare the distributions of per-dataset Pearson r values between two program pairs (e.g., basal-pEMT vs. basal-mesenchymal) across 80 datasets↳ Could also: A Wilcoxon signed-rank test could also be used for the same paired comparison — The Wilcoxon signed-rank test is the non-parametric counterpart that does not assume the within-pair differences of r values are normally distributed, which may be preferable given the bounded [-1, 1] range of correlation coefficients
-
No multiplicity correction was described despite testing multiple gene-program pairs across 80 datasets↳ Could also: A Benjamini-Hochberg false discovery rate (FDR) correction applied across all pairwise correlation tests within or across datasets could also be used — When many correlation tests are performed simultaneously, FDR control is a standard approach to limit the expected proportion of false discoveries; reporting q-values alongside p-values would let readers assess the robustness of the pattern counts reported
-
ssGSEA scores were used to reduce each sample's transcriptome to five scalar program scores before correlation analysis↳ Could also: GSVA (gene set variation analysis) or AUCell could also quantify per-sample gene-set activity, particularly for single-cell data where ssGSEA assumptions may be less appropriate — GSVA uses a different statistical model (kernel-based) and can behave differently at small n per cell in single-cell contexts; AUCell is specifically designed for sparse single-cell count matrices; comparing results across methods is a common robustness check
-
Survival group differences in TCGA were reported as hazard ratios with a p-value threshold, with the method of survival analysis not named in the excerpt↳ Could also: A multivariable Cox proportional-hazards model adjusting for known confounders (e.g., age, stage, PAM50 subtype) could also be used alongside or instead of unadjusted group comparisons — Unadjusted survival comparisons in heterogeneous clinical cohorts can conflate the program association with subtype or stage differences; multivariable Cox regression isolates the incremental prognostic contribution of the epithelial-luminal classification
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-38974967
Paper: Sahoo et al., "Increased prevalence of hybrid epithelial/mesenchymal state and enhanced phenotypic heterogeneity in basal breast cancer." iScience 2024. PMID 38974967 / PMC11225361 / DOI 10.1016/j.isci.2024.110116.
This is a systems-biology / mechanistic-modeling + transcriptomic meta-analysis paper. It combines (a) bulk/single-cell/spatial transcriptomic analyses of breast cancer and (b) a RACIPE gene-regulatory-network (GRN) simulation that is the paper's central mechanistic result (Figure 5).
Pipelines and in-scope vs out-of-scope
| Result | Pipeline | Scope | Reason |
|---|---|---|---|
| Fig 5B: RACIPE steady-state landscape of the luminal-basal × EMT GRN | RACIPE-1.0 (C) on the 11-node network (Table S4) | IN SCOPE — primary | Fully specified: topology in Table S4 (mmc5.xlsx), parameters in Methods, classifier in authors' notebook. Third-party tool (RACIPE) on the paper's own network = a valid reproduction (brief P16). |
| Fig 5B (i–iii): luminal states ~exclusively epithelial; basal states span E/hybrid/M | RACIPE + z-score scoring | IN SCOPE | Same simulation; the paper's headline mechanistic claim. |
| Fig 5F: IL-1β / TGF-β1 over-expression shifts (WT vs OE fractions) | RACIPE on the extended "signals" network | PARTIAL / 20% — not the primary target | The OE network folder is circuit_4_signals_notch; Table S4 lists Tgfb/Il1b edges but no Notch edges, and the shipped solution columns are 11 genes only. The exact OE base network is not fully specified by the shipped artifacts → treated as the hard 20%. The reported WT control fractions are still used as the comparison reference for the in-scope WT run. |
| Multistability (number of stable states per parameter set) | RACIPE | IN SCOPE (bonus) | Pure threshold-free RACIPE output. |
| Figs 1–4, S1–S6: ssGSEA scoring, correlations, meta-analysis of 80 GEO datasets, scRNA/spatial, survival | gseapy / pandas / Seurat-style | OUT OF SCOPE | Many version-dependent steps; data are large/various; not the mechanistic claim. Not attempted. |
| GSE42944 methylation (Fig 2A,iii) | beta-value download + plotting | OUT OF SCOPE | Ancillary; not the pipeline-defining result. |
| All wet-lab (qPCR, IHC, H3K27ac/me3) | experimental | OUT OF SCOPE | Non-computational. |
Primary reproduction target
Build RACIPE-1.0 and run it on the 11-node breast-cancer luminal-basal/EMT GRN
(38 edges, Table S4 "Base network" restricted to the 11 dynamical genes
[Slug, miR200, Zeb1, Cdh1, ERa66, ERa36, np63, Gata3, Foxa1, Pgr, Nrf2]) with the
paper's exact parameters (triplicate; 10,000 parameter sets; 100 initial conditions;
Euler; up to 6 stable states). Z-score the steady states and apply the authors'
composite scores (Epithelial, Mesenchymal, EMT, Luminal, Basal, LBT) with KDE-minima
thresholds (notebook Untitled.ipynb) to classify each state into
Lum / Lum_H / Bas_epi / Bas_H / Bas_mes, and compare the 5-category fractions and the
luminal:basal split to the reported WT values.
Code & data provenance (pointers)
- Tool:
github.com/simonhb1990/RACIPE-1.0@ commit b4483533e69b1f12e11c3e0d0738d2d094e4ab6f - Authors' analysis:
github.com/sarthak-sahoo-0710/luminal_basal_EMT_crosstalk(RACIPE_analysis_codes/) - Network: Supplementary Table S4 =
mmc5.xlsx(Elsevier CDN; NCBI mirror blocked) - Reported WT fractions: repo
RACIPE_analysis_codes/data_il1b.txt/data_tgfb1.txt(WT rows, triplicate) - «infra» work dir:
«path» - GEO accession in registry (GSE42944) is methylation data used only for Fig 2A,iii — not the pipeline reproduced here.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.