A mechanistic model captures the emergence and implications of non-genetic heterogeneity and reversible drug resistance in ER+ breast cancer cells.
The main results reproduced: recomputed values matched the published ones within tolerance.
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
REPRODUCED (described well enough; clean 1:1 on the headline RACIPE arm). Independent FULL-N re-run for a re-queued room (prior «infra» workdir reclaimed, so re-cloned+rebuilt from scratch). «our HPC» «job» COMPLETED exit 0 (6h49m): shipped core.cfg (100000 models) x3 replicates of RACIPE-1.0 (b4483533) on the TamRes (af08d7f) 5-node core network; rescore «job» after fixing a pandas tokenizer crash on ragged RACIPE rows. C1 Spearman EM-vs-resistance rho = 0.8073+/-0.0003 vs reported 0.806 (essentially exact, within RACIPE sampling noise; matches the earlier independent run's 0.8078). C2 multistability 66.0/22.2/7.8% bi/tri/mono vs reported ~65/20/8 (within-tol) -- taken from RACIPE's OWN in-log stability tally which sums to exactly 100000 models/seed (the per-model solution files have ~5% fewer physical rows because this RACIPE build occasionally merges two records onto one line; this affects no statistic). C3 ES & MR co-dominant (~42% each) and MS least (~0.75%) reproduced and stable across replicates (per-phenotype % provisional: Suppl Table S1 not pinned, repo fig_3_i broken). No fabrication indicators: every reported value is derivable from the shipped topology+tool. NOT attempted (separate pipelines, honestly deferred): StochasticSimulations (R/Euler-Maruyama, Fig2), PopulationDynamics (Julia, Figs5-6); GSE6532 patient correlation has no code in the repo (profiled-only: open, complete, grade A, N87 reported==observed).
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 85assessed: 2026-06-20 ⛓ 1c1179778143
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-23
- Rubric version
- not recorded
- Assessed by
- —
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetThe paper tests whether epithelial-mesenchymal transition (EMT) and tamoxifen resistance (TamR) in ER+ breast cancer cells are bidirectionally coupled, such that each process can drive the other to generate non-genetic phenotypic heterogeneity and reversible drug resistance.
- ★ EMT and tamoxifen-resistance (TamR) regulatory axes can drive one another, enabling non-genetic heterogeneity via six co-existing phenotypes (ES, ER, HS, HR, MS, MR) finding
- ★ A mechanism-based gene regulatory network model (ERα66, ERα36, SLUG, ZEB1, miR-200) reproduces multistable EMT/TamR phenotypes via RACIPE simulation method
- ★ Epithelial-sensitive (ES) and mesenchymal-resistant (MR) are the two dominant phenotypes among the six possible states finding
- ★ Phenotypic plasticity (drug-induced and intrinsic stochastic switching) and/or non-genetic heterogeneity promote population-level survival of a mixed sensitive/resistant cell population even without cell-cell cooperation finding
- ★ Combining mesenchymal-to-epithelial transition (MET) inducers with canonical anti-estrogen therapy is proposed as a strategy to limit emergence of reversible drug resistance mechanism
- ZEB1 expression increases monotonically across the six biologically defined phenotypes finding
- EM score distribution of RACIPE steady states is trimodal and resistance score distribution is bimodal finding
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| RACIPE (Random Circuit Perturbation) ODE-based network simulation | in silico gene regulatory network (ERα66, ERα36, SLUG, ZEB1, miR-200) | none (parameter ensemble sampling) | steady-state gene expression levels, EM score, resistance score | RACIPE |
| Stochastic simulation (Euler-Maruyama method) | same GRN, representative RACIPE parameter sets | noise term added to ODEs | probability density and pseudo-potential landscape of EM-resistance score pairs | — |
| Randomized network ensemble analysis | 100 randomized versions of the wild-type GRN | network topology randomization | distribution of Spearman correlation coefficients between EM and resistance scores | RACIPE |
| Microarray gene expression correlation analysis | publicly available GEO microarray datasets | none | Spearman correlation coefficients between genes | GEO |
| single-sample gene set enrichment analysis (ssGSEA) | bulk microarray datasets | none | hallmark EMT, early/late estrogen response, and TamRes signature scores | MSigDB gene sets |
| EMT scoring (76GS, KS, MLR methods) | bulk microarray datasets | none | EMT score | — |
| AUCell activity scoring | 10-cell and single-cell RNA-seq datasets | none | ESR1 regulon and EMT/MET signature activity | AUCell; BRCA ESR1 regulon from GRNDb |
| Agent-based population dynamics simulation | simulated population of sensitive/resistant cells | tamoxifen-like drug exposure, stochastic state switching, varying heterogeneity/plasticity parameters | population size, proportion of resistant vs sensitive cells over time | — |
- – RACIPE simulation of the EMT-ERα GRN yields six distinct phenotypes: epithelial-sensitive (ES), epithelial-resistant (ER), hybrid-sensitive (HS), hybrid-resistant (HR), mesenchymal-sensitive (MS), mesenchymal-resistant (MR)
- – ES and MR are the most prevalent/dominant phenotypes among the six
- – EM score histogram of RACIPE steady states is trimodal (epithelial, hybrid, mesenchymal peaks, hybrid smallest)
- – Resistance score histogram of RACIPE steady states is bimodal
- – UMAP visualization shows epithelial cluster more likely sensitive and mesenchymal cluster more likely resistant, while hybrid E/M cluster can be either sensitive or resistant
- ▲ Scatter plot of EM score vs resistance score shows positive association, quantified by Spearman correlation with reported P-value
- ▲ ZEB1 expression increases monotonically across the ordered set of six phenotypes
- – Correlation coefficient of the wild-type GRN differs from the distribution of correlation coefficients obtained from 100 randomized network topologies
- count 100,000 parameter sets (RACIPE simulation of the EMT-ERα GRN (Figure 1A/B))
- count 100 randomized network versions (ensemble used to compare wild-type vs randomized GRN correlation (Figure 1G))
- other proliferation rate = 0.91 (logistic growth model in population dynamics simulation)
- other carrying capacity = 10^5 cells (population dynamics simulation)
- other basal death probability = 0.1 (constant basal probability of cell death in population model)
- other c = 0.6 (constant in sigmoidal function mapping resistance score to drug-induced death probability)
- other resistance score range [-6, 6] (defined range representing a cell's fitness under anti-estrogen drug in population model)
- other sensitive/resistant score means at -2 and +2 (Gaussian distributions used to assign sensitive (S) and resistant (R) scores to cells at birth)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This is primarily a computational mechanistic modeling study using RACIPE (Random Circuit Perturbation) to simulate a five-node gene regulatory network (100,000 parameter sets) governing coupled EMT and tamoxifen resistance dynamics, yielding ensemble steady-state solutions analyzed by dimensionality reduction and clustering. Supplementary analysis of publicly available microarray and single-cell datasets was performed using ssGSEA and AUCell-based activity scoring. Statistical inference on correlations used Spearman's rho with corresponding P-values, and group comparisons used two-tailed Welch's t-tests; results were displayed as heatmaps, scatter plots, UMAP embeddings, and density histograms with SD error bars.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Spearman rank correlation | Quantifying association between EM Score and Resistance Score across all RACIPE steady-state solutions (Figure 1E); also applied to gene pairs in publicly available microarray datasets | 100,000 parameter sets (RACIPE solutions) for the GRN analysis; sample sizes for clinical microarray datasets not stated in excerpt | not stated |
| Two-tailed Student's t-test (Welch, unequal variances assumed) | Statistical comparisons between phenotypic groups (stated in Materials and Methods – Statistical testing section) | null | stated |
| Gaussian mixture model fitting (3-component and 2-component) | Density histogram of EM Score fitted to 3 Gaussians; Resistance Score fitted to 2 Gaussians (Figure 1C) | 100,000 RACIPE steady-state solutions | not stated |
| K-means clustering | Corroborating major expression patterns in RACIPE steady states (Supplementary Figure S1A) | null | not stated |
| Single-sample gene set enrichment analysis (ssGSEA) | Scoring hallmark EMT, early estrogen response, and late estrogen response signatures in bulk microarray datasets | null | na |
| AUCell activity scoring | Computing regulon and signature activity values in 10-cell and single-cell datasets | null | na |
-
Multiple Spearman correlations were computed across gene pairs and datasets without a stated multiple-testing correction procedure↳ Could also: Benjamini-Hochberg FDR correction (or Bonferroni for smaller families) applied across the full set of correlation tests — When many correlations are tested simultaneously, a FDR or FWER correction quantifies the expected proportion of false discoveries and is standard practice; reporting adjusted q-values alongside raw P-values would allow readers to assess which associations are robust after accounting for the number of tests
-
Two-tailed Welch's t-test was used for group comparisons across the six phenotypic groups↳ Could also: One-way ANOVA (or Kruskal-Wallis for non-parametric data) with a post-hoc correction such as Tukey HSD or Dunn's test — When more than two groups are compared, a single omnibus test followed by a corrected post-hoc procedure controls the family-wise error rate across all pairwise comparisons, which multiple independent t-tests do not; this is a widely recommended approach for multi-group designs
-
Gaussian mixture models with a fixed number of components (3 for EM Score, 2 for Resistance Score) were fit to the RACIPE output distributions↳ Could also: Model selection via BIC or AIC across a range of component numbers before fixing the mixture — Formally comparing models with different numbers of Gaussian components using an information criterion provides a data-driven justification for the chosen number of phenotypic states, complementing the visual inspection of histograms
-
K-means clustering was used to corroborate major expression patterns in the RACIPE steady-state solutions↳ Could also: Gaussian mixture model (GMM) clustering or hierarchical clustering with silhouette-score validation of the cluster count — GMM clustering provides probabilistic cluster assignments and naturally accommodates ellipsoidal cluster shapes; either approach paired with a cluster-validity index (silhouette, gap statistic) would give a more formal basis for the number of clusters chosen
-
Dispersion of gene expression levels across the six phenotypes was summarized with standard deviation (SD) error bars↳ Could also: 95% bootstrap confidence intervals or SEM alongside the mean, or box-and-whisker plots showing median and IQR — For distributions that may be non-Gaussian (as expected from ensemble ODE outputs spanning a wide parameter space), showing the median and IQR or a full distribution plot preserves more information about shape and skewness than SD alone
-
ssGSEA was used to score pathway activities (EMT, estrogen response) in bulk microarray data↳ Could also: GSVA (Gene Set Variation Analysis) or PROGENy for bulk transcriptomic pathway scoring — GSVA provides sample-level enrichment scores with somewhat different statistical properties than ssGSEA (it models the rank distribution within each sample), and PROGENy is specifically calibrated on perturbation data; using more than one scoring method and checking concordance would strengthen confidence in pathway-level conclusions
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-34316714
Paper: Sahoo et al. 2021, A mechanistic model captures the emergence and implications of non-genetic heterogeneity and reversible drug resistance in ER+ breast cancer cells. NAR Cancer 3(3):zcab027. PMCID PMC8271219. Code: https://github.com/csbBSSE/TamRes (HEAD af08d7f, 2021-03-15) — authors' own code (P16 N/A). Engine: RACIPE-1.0 (simonhb1990/RACIPE-1.0), referenced by the repo README.
The paper has three computational arms, each its own pipeline. RACIPE is the clearest, fully-pinned, deterministic one and is the focus of this reproduction.
IN SCOPE (pipeline-derived, attempted)
Pipeline RACIPE on the shipped 5-node core network
(RACIPE_analysis/topo_file/r1/core.{topo,cfg}; nodes ERa66, ERa36, ZEB1,
miR200, SLUG; 13 regulations). The cfg is fully pinned: 100 000 models, 100 RIVs,
Seed 1, default parameter ranges → deterministic. Steps: run RACIPE →
z-normalise steady states (repo Misc_code) → derive EM_score = ZEB1 − miR200,
resistance_score = ERa36 − ERa66 (repo fig_1_v).
| id | reported result | location | how reproduced |
|---|---|---|---|
| C1 | Spearman ρ(EM score, resistance score) = 0.806, P<0.01 | Results / Fig 1E | z-norm steady states → Spearman of the two scores |
| C2 | ~65% bistable, 20% tristable, 8% monostable parameter sets | Results text | count models per core_solution_N.dat / total |
| C3 | ES & MR phenotypes most dominant, MS least prevalent | Results / Fig 3, Suppl Table S1 | phenotype classification via score-distribution boundaries |
OUT OF SCOPE (not attempted, and why)
- Population-dynamics agent-based model (
PopulationDynamics/*.jl, Figs 5–6): separate stochastic Julia simulation arm; a different pipeline, not needed for the headline RACIPE claims. Deferred (80/20). - Stochastic ODE simulations (
StochasticSimulations/*.R, Euler–Maruyama, Fig 2): separate arm; deferred. - GSE6532 patient-data correlation (ESR1 vs ZEB1/SNAI2/EMT-hallmark ssGSEA, 87 ER+ patients): real microarray pipeline, in principle reproducible, but a distinct data+ssGSEA pipeline; deferred to keep to a few CLEAR RACIPE datapoints.
- All wet-lab measurements (qPCR, flow, viability) — non-computational.
Notes / gotchas
- Shipped
fig_3_ianalysis script has a bug:state_dataframe["Phenotype"] = phereferences an undefined namephe(the list built isphenotypes_array). The phenotype-count figure code therefore does not run as-is; C3 is reproduced by re-implementing the documented boundary logic. Flagged in AUDIT. - C3 phenotype boundaries depend on a GMM fit of score distributions
(
fig_1_iii); this is the harder ~20% and is reported at best partial.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.