Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Increased prevalence of hybrid epithelial/mesenchymal state and enhanced phenotypic heterogeneity in basal breast cancer.

iScience · 2024
95/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
How its reproducibility compares
95/100
Reproducibility score
1.2 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 89% of all assessed papers rank 105 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

REPRODUCED 1:1 (WT landscape AND over-expression). RACIPE-1.0 (commit b4483533e69b1f12e11c3e0d0738d2d094e4ab6f) built on «our HPC»; topology re-derived from Supplementary Table S4 (mmc5.xlsx): WT 11-gene circuit = 38 edges (EXACT match to vetted reference), signals circuit = 46 edges over EXACTLY 13 nodes (11 genes + Tgfb + Il1b). KEY CORRECTION vs the two prior attempts: S4 contains NO Notch node, so the Fig 5F over-expression network IS fully specified and the OE experiment IS reproducible (prior rooms wrongly dropped it as out-of-scope). OE done with RACIPE's native -OEID 13(Il1b)/12(Tgfb) -OEFD 100 (paper Methods: 'overexpression (OE; 100x)'); gene IDs confirmed identical to the authors' core_OE_13/core_OE_12 files. Production «job» COMPLETED (1h21m): full triplicate, 12 runs (WT + signals-base + IL1B-OE + TGFB1-OE, each 10,000 parameter sets x 100 ICs, Euler, up to 6 states), every replicate stored exactly 10,000 models (integrity-checked). Classification follows the authors' Untitled.ipynb (11-node landscape) and signals.ipynb (OE: z-score OE states against the base mean/std, composite LBT/EMT scores on the 11 phenotype genes, authors' hardcoded per-replicate signals thresholds + KDE-minima re-derivation as robustness). RESULTS: C1 WT 5-category fractions reproduce to within 0.004 (signals-base, which is what the reported numbers actually come from); C2 luminal:basal = 0.381:0.619 (reported 0.380:0.620, exact); C3 luminal almost exclusively epithelial (hybrid 15.4% vs 14.6%) and basal spans the full E/M axis (non-epithelial 73.4% vs 73.7%); C4 multistability up-to-6 and 97% multistable reproduced; OE1 IL1B-OE (+Bas_epi, +Bas_H, -Lum_H) and TGFB1-OE (+Lum_H +91%->reproduced +77%, +Bas_mes, -Bas_epi, -Bas_H) reproduce every direction with absolute OE fractions within 0.005 of reported. The 11-node core landscape (Fig 5B) classified with the core hardcoded thresholds shows a modest ~0.03-0.05 mass shift Lum->Bas_epi/Bas_H (qualitative landscape figure; Bas_mes still near-exact) -- reported as an honest nuance. All grades PROVISIONAL for human review. OUT OF SCOPE (not attempted): ssGSEA 80-dataset meta-analysis, scRNA/spatial, survival (Figs 1-4), GSE42944 methylation (Fig 2A iii) -- transcriptomic/ancillary, not the mechanistic RACIPE pipeline.

💻 Code ↗ 🗄 Data: GSE42944

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 64
    assessed: 2026-06-19 ⛓ 1ed81710646f
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-22
Rubric version
not recorded
Assessed by
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The paper tests whether luminal-basal molecular subtype and epithelial-mesenchymal transition (EMT) status are equivalent axes of breast cancer heterogeneity, hypothesizing instead that basal breast cancer is specifically associated with a hybrid/partial-EMT (pEMT) state and greater phenotypic heterogeneity rather than a fully mesenchymal state.

Core claims
  • Luminal breast cancer gene expression signature is closely/positively associated with an epithelial signature finding
  • Basal breast cancer signature correlates more strongly with a hybrid epithelial/mesenchymal (pEMT) signature than with a fully mesenchymal signature finding
  • Basal breast cancer exhibits higher phenotypic heterogeneity along the EMT spectrum than luminal breast cancer finding
  • Epigenetic patterns (DNA methylation, H3K27ac/H3K27me3) at epithelial and mesenchymal gene promoters underlie the luminal-epithelial and basal-pEMT associations mechanism
  • Mathematical modeling of a gene regulatory network coupling EMT and luminal-basal differentiation axes recapitulates the observed transcriptomic associations and heterogeneity mechanism
  • Spatial transcriptomics reveals intra-patient variability in EMT phenotypes specifically within basal breast cancer subtypes finding
  • Epithelial-luminal joint classification stratifies TCGA breast cancer patient survival better than epithelial-mesenchymal classification alone finding
Experimental setups
Assay System Perturbation Readout Platform
single-sample gene set enrichment analysis (ssGSEA) of bulk transcriptomic data 80 breast-cancer-specific datasets (cell lines and primary tumors, meta-analysis) none correlation between luminal, basal, epithelial, mesenchymal, and pEMT gene signature scores
survival/prognosis analysis TCGA breast cancer patient cohort none hazard ratio for survival stratified by epithelial/mesenchymal and epithelial/luminal group combinations
genome-wide DNA methylation profiling CCLE breast cancer cell lines (GSE42944), Luminal/Basal A/Basal B subtypes none CpG island methylation at promoters of epithelial and mesenchymal genes, correlated with gene expression
bulk RNA sequencing four breast cancer cell lines (MCF7, ZR-75 luminal; HCC38, HMLER basal) with CD44-high/low subpopulations (GSE184647) none ssGSEA scores for basal, mesenchymal, and pEMT gene expression programs
MINT-ChIP (histone modification profiling) same luminal and basal breast cancer cell lines (e.g., HMLER CD44-high/low) none percentage of H3K27ac activation and H3K27me3 suppressive marks at epithelial/mesenchymal gene promoters
spatial transcriptomics breast cancer tissue sections (patient tumor samples) none intra-patient spatial heterogeneity of EMT phenotype
mathematical/computational modeling of gene regulatory network in silico model of EMT-luminal/basal crosstalk network simulated network parameters emergent phenotypic states and heterogeneity compared to transcriptomic data
Key results
  • Luminal-epithelial signatures showed significant positive correlation in 36 of 80 datasets vs. only 4 significantly negative r>0.3, p<0.05
  • Basal-mesenchymal correlation was only weakly skewed positive (11 positive vs. 7 negative datasets) r>0.3/r<-0.3, p<0.05
  • Basal-pEMT signature showed significant positive correlation in 21 of 80 datasets vs. only 3 negative r>0.3, p<0.05
  • Paired comparison showed basal-pEMT correlation significantly greater than basal-mesenchymal correlation T=2.9, p<0.01; T=4.0, p<0.001 (EMT_partial vs EMT_up)
  • Patients with high mesenchymal (EPI-MES+) tumors had worse prognosis HR=1.5, p<0.05
  • Basal score negatively correlated with luminal score in CCLE cell lines r=-0.34, p<0.05
  • Basal A cell lines showed low methylation of both epithelial and mesenchymal genes, consistent with a pEMT-like co-active state, while Basal B lines resembled a fully mesenchymal methylation pattern
  • CD44-high basal subpopulations clustered as more mesenchymal, while CD44-low basal subpopulations occupied an intermediate (pEMT) region of the EMT plane
Key statistics
  • correlation r>0.3, p<0.05 in 36/80 datasets (luminal-epithelial) (meta-analysis of breast cancer datasets)
  • correlation r>0.3, p<0.05 in 21/80 datasets vs 3 negative (basal-pEMT) (meta-analysis of breast cancer datasets)
  • pvalue T=2.9, p<0.01 (basal-pEMT vs basal-mesenchymal, paired t test) (paired comparison of correlation coefficients)
  • pvalue T=4.0, p<0.001 (basal-EMT_partial vs basal-EMT_up, paired t test) (paired comparison of correlation coefficients)
  • other hazard ratio HR=1.5, p<0.05 (EPI-MES+ TCGA breast cancer patients survival vs reference group)
  • correlation r=-0.34, p<0.05 (basal vs luminal ssGSEA score in CCLE breast cancer cell lines)
  • count 40 vs 6 datasets (EMT_down-Luminal positive vs negative correlation) (breast-cancer-specific EMT gene set comparison)
  • count 10 vs 5 datasets (Basal-EMT_up positive vs negative correlation) (breast-cancer-specific EMT gene set comparison)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The paper integrates multi-modal transcriptomic data (bulk RNA-seq meta-analysis of 80 datasets, CCLE, TCGA-BRCA, single-cell RNA-seq, spatial transcriptomics, methylation) to characterize associations between luminal-basal and epithelial-mesenchymal (EMT) gene programs in breast cancer. The primary analytic strategy is single-sample gene set enrichment analysis (ssGSEA) to score five gene programs per sample, followed by Pearson correlation within each dataset and paired t-tests to compare correlation magnitudes. Survival differences were assessed with hazard ratios in the TCGA cohort. Results are reported as correlation coefficients, T-statistics, p-values, and dataset counts meeting a directional threshold (r > 0.3 or r < −0.3, p < 0.05).

Replicationbiological Sample size80 publicly available breast cancer datasets enumerated in Table S3; four cell lines named (MCF7, ZR-75, HCC38, HMLER); n=6 patients for spatial transcriptomics; TCGA-BRCA and CCLE sample counts not stated in excerpt GroupsLuminal vs. Basal A vs. Basal B cell lines; CD44-low vs. CD44-high subpopulations within basal cell lines; ER+ vs. other subtypes in spatial transcriptomics; survival groups defined by combined epithelial/luminal or epithelial/mesenchymal ssGSEA status Pairingmixed Randomization/blindingnot stated Dispersionnone Exact p-valuesno Effect sizesyes Confidence intervalsno Multiplicity correctionnone stated
Statistical tests used
Test Applied to n Assumptions
Pearson's correlation coefficient (with associated p-value) Pairwise correlation of ssGSEA scores for each of five gene programs (luminal, basal, epithelial, mesenchymal, pEMT) within each of 80 breast cancer datasets; also in CCLE and TCGA-BRCA individually 80 datasets (sample sizes per dataset not stated in excerpt); CCLE and TCGA-BRCA sample counts not stated in excerpt not stated
Paired t-test Comparison of per-dataset Pearson r values for basal-pEMT/basal-EMT_partial pairs vs. basal-mesenchymal/basal-EMT_up pairs across the 80 datasets (Figure 1E) 80 paired observations (one r per dataset) not stated
Survival analysis (hazard ratio reported, method not named in excerpt) Prognostic capacity of epithelial-mesenchymal vs. epithelial-luminal classification in TCGA breast cancer cohort (Figure S1A) null not stated
ssGSEA (single-sample gene set enrichment analysis) Quantification of gene program activity per sample across all datasets, CCLE, TCGA-BRCA, and spatial transcriptomics null na
Approaches that could also have been used
  • Pearson's correlation was used to quantify pairwise associations between ssGSEA scores across datasets
    Could also: Spearman's rank correlation could also be used for the same comparisons — Spearman's correlation makes no assumption about the distributional shape of ssGSEA scores and is robust to outliers and non-linearity; for enrichment scores that may not be normally distributed, it is often reported alongside or instead of Pearson's r
  • A fixed threshold (r > 0.3, p < 0.05) was applied within each dataset independently to count 'significantly positive' or 'significantly negative' datasets across the 80-dataset meta-analysis
    Could also: A random-effects meta-analytic model (e.g., DerSimonian-Laird) pooling the per-dataset r values into a single weighted summary estimate with a 95% CI and heterogeneity statistic (I²) could also be used — Formal meta-analysis provides a single pooled effect size with uncertainty quantification, explicitly accounts for between-study heterogeneity, and avoids the multiple-testing accumulation that arises from counting independent per-dataset p-values
  • A paired t-test was used to compare the distributions of per-dataset Pearson r values between two program pairs (e.g., basal-pEMT vs. basal-mesenchymal) across 80 datasets
    Could also: A Wilcoxon signed-rank test could also be used for the same paired comparison — The Wilcoxon signed-rank test is the non-parametric counterpart that does not assume the within-pair differences of r values are normally distributed, which may be preferable given the bounded [-1, 1] range of correlation coefficients
  • No multiplicity correction was described despite testing multiple gene-program pairs across 80 datasets
    Could also: A Benjamini-Hochberg false discovery rate (FDR) correction applied across all pairwise correlation tests within or across datasets could also be used — When many correlation tests are performed simultaneously, FDR control is a standard approach to limit the expected proportion of false discoveries; reporting q-values alongside p-values would let readers assess the robustness of the pattern counts reported
  • ssGSEA scores were used to reduce each sample's transcriptome to five scalar program scores before correlation analysis
    Could also: GSVA (gene set variation analysis) or AUCell could also quantify per-sample gene-set activity, particularly for single-cell data where ssGSEA assumptions may be less appropriate — GSVA uses a different statistical model (kernel-based) and can behave differently at small n per cell in single-cell contexts; AUCell is specifically designed for sparse single-cell count matrices; comparing results across methods is a common robustness check
  • Survival group differences in TCGA were reported as hazard ratios with a p-value threshold, with the method of survival analysis not named in the excerpt
    Could also: A multivariable Cox proportional-hazards model adjusting for known confounders (e.g., age, stage, PAM50 subtype) could also be used alongside or instead of unadjusted group comparisons — Unadjusted survival comparisons in heterogeneous clinical cohorts can conflate the program association with subtype or stage differences; multivariable Cox regression isolates the incremental prognostic contribution of the epithelial-luminal classification
Software: ssGSEA implementation (package/tool not named in excerpt)

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-38974967

Paper: Sahoo et al., "Increased prevalence of hybrid epithelial/mesenchymal state and enhanced phenotypic heterogeneity in basal breast cancer." iScience 2024. PMID 38974967 / PMC11225361 / DOI 10.1016/j.isci.2024.110116.

This is a systems-biology / mechanistic-modeling + transcriptomic meta-analysis paper. It combines (a) bulk/single-cell/spatial transcriptomic analyses of breast cancer and (b) a RACIPE gene-regulatory-network (GRN) simulation that is the paper's central mechanistic result (Figure 5).

Pipelines and in-scope vs out-of-scope

Result Pipeline Scope Reason
Fig 5B: RACIPE steady-state landscape of the luminal-basal × EMT GRN RACIPE-1.0 (C) on the 11-node network (Table S4) IN SCOPE — primary Fully specified: topology in Table S4 (mmc5.xlsx), parameters in Methods, classifier in authors' notebook. Third-party tool (RACIPE) on the paper's own network = a valid reproduction (brief P16).
Fig 5B (i–iii): luminal states ~exclusively epithelial; basal states span E/hybrid/M RACIPE + z-score scoring IN SCOPE Same simulation; the paper's headline mechanistic claim.
Fig 5F: IL-1β / TGF-β1 over-expression shifts (WT vs OE fractions) RACIPE on the extended "signals" network PARTIAL / 20% — not the primary target The OE network folder is circuit_4_signals_notch; Table S4 lists Tgfb/Il1b edges but no Notch edges, and the shipped solution columns are 11 genes only. The exact OE base network is not fully specified by the shipped artifacts → treated as the hard 20%. The reported WT control fractions are still used as the comparison reference for the in-scope WT run.
Multistability (number of stable states per parameter set) RACIPE IN SCOPE (bonus) Pure threshold-free RACIPE output.
Figs 1–4, S1–S6: ssGSEA scoring, correlations, meta-analysis of 80 GEO datasets, scRNA/spatial, survival gseapy / pandas / Seurat-style OUT OF SCOPE Many version-dependent steps; data are large/various; not the mechanistic claim. Not attempted.
GSE42944 methylation (Fig 2A,iii) beta-value download + plotting OUT OF SCOPE Ancillary; not the pipeline-defining result.
All wet-lab (qPCR, IHC, H3K27ac/me3) experimental OUT OF SCOPE Non-computational.

Primary reproduction target

Build RACIPE-1.0 and run it on the 11-node breast-cancer luminal-basal/EMT GRN (38 edges, Table S4 "Base network" restricted to the 11 dynamical genes [Slug, miR200, Zeb1, Cdh1, ERa66, ERa36, np63, Gata3, Foxa1, Pgr, Nrf2]) with the paper's exact parameters (triplicate; 10,000 parameter sets; 100 initial conditions; Euler; up to 6 stable states). Z-score the steady states and apply the authors' composite scores (Epithelial, Mesenchymal, EMT, Luminal, Basal, LBT) with KDE-minima thresholds (notebook Untitled.ipynb) to classify each state into Lum / Lum_H / Bas_epi / Bas_H / Bas_mes, and compare the 5-category fractions and the luminal:basal split to the reported WT values.

Code & data provenance (pointers)

  • Tool: github.com/simonhb1990/RACIPE-1.0 @ commit b4483533e69b1f12e11c3e0d0738d2d094e4ab6f
  • Authors' analysis: github.com/sarthak-sahoo-0710/luminal_basal_EMT_crosstalk (RACIPE_analysis_codes/)
  • Network: Supplementary Table S4 = mmc5.xlsx (Elsevier CDN; NCBI mirror blocked)
  • Reported WT fractions: repo RACIPE_analysis_codes/data_il1b.txt/data_tgfb1.txt (WT rows, triplicate)
  • «infra» work dir: «path»
  • GEO accession in registry (GSE42944) is methylation data used only for Fig 2A,iii — not the pipeline reproduced here.
Figures / tables: Fig 5FFig 5BFigsFig 2A
C1
Reported
WT 5-phenotype fractions Lum 0.3247 / Lum_H 0.0553 / Bas_epi 0.1633 / Bas_H 0.1189 / Bas_mes 0.3378 (authors' data_il1b.txt WT triplicate = Fig 5F WT bars)
Reproduced
signals-base (13-node) triplicate mean Lum 0.3223 / Lum_H 0.0586 / Bas_epi 0.1645 / Bas_H 0.1209 / Bas_mes 0.3337 (every category within 0.004; per-rep SD<=0.0017)
exact
C2
Reported
luminal:basal landscape balance ~0.380:0.620 (Fig 5B i)
Reproduced
0.3809 luminal : 0.6191 basal
exact
C3
Reported
luminal states almost exclusively epithelial; basal states span epithelial/hybrid/mesenchymal (headline claim, Fig 5B i): luminal hybrid ~14.6%, basal non-epithelial ~73.7%
Reproduced
luminal hybrid 15.4% of luminal; basal hybrid+mes 73.4% of basal
exact
C4
Reported
enhanced heterogeneity = widespread multistability, up to 6 stable states per parameter set (Methods/Fig 5)
Reproduced
up to 6 confirmed; 97.18% multistable (>=2); modes 1:2.8% 2:14.3% 3:29.3% 4:29.0% 5:15.6% 6:9.0% (30000 param sets)
exact
OE1_IL1B
Reported
IL1B OE (100x) vs WT: Bas_epi +17.4% / Bas_H +11.5% / Bas_mes -5.1% / Lum -2.9% / Lum_H -28.3% (Fig 5F; data_percent.txt il1b)
Reproduced
Bas_epi +18.4% / Bas_H +9.9% / Bas_mes -5.4% / Lum -1.9% / Lum_H -30.9% (all 5 directions match; absolute OE fractions within 0.005); RACIPE -OEID 13 -OEFD 100
within tolerance
OE1_TGFB1
Reported
TGFB1 OE (100x) vs WT: Lum_H +82.5% / Bas_mes +24.5% / Bas_epi -48.4% / Bas_H -33.0% / Lum -3.1% (Fig 5F; data_percent.txt tgfb1)
Reproduced
Lum_H +77.5% / Bas_mes +26.0% / Bas_epi -49.3% / Bas_H -34.7% / Lum -2.9% (all 5 directions match, magnitudes within ~5 pp; absolute OE fractions within 0.004); RACIPE -OEID 12 -OEFD 100
within tolerance

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

No assessment has been recorded yet.
🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

1.4 M
tokens (I/O) · 96.7 M incl. cache
564 min
runtime · 15.52 CPU-h
0.3 GB
peak RAM
3
HPC jobs
hummel
machine