Emergent dynamics of underlying regulatory network links EMT and androgen receptor-dependent resistance in prostate cancer.
The main results reproduced: recomputed values matched the published ones within tolerance.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Same input data as the authors
- ✓Reported values were directly comparable
- ✓No relevant deviation in data/preprocessing
- ✓No authors-side cause for any deviation
- ✓Reported values are derivable from the shared data
- ✓Any deviation was negligible
- ✓The central claim held under reproduction
- ✓Overall, the reproduction was clean
- Every checked point held up.
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
DESCRIBED WELL ENOUGH -> reproduced 1:1 (within stochastic tolerance). Per P16, ran the canonical RACIPE-1.0 tool (commit b4483533) on the paper's OWN published 9-node ESMR topology (ESMR.topo sha256 8f3c76d4...) with the Methods parameters (10000 parameter sets x 100 ICs, seed 1) on «our HPC» SLURM «job» (COMPLETED, 1.05h RACIPE runtime). The three quantitative paper claims all reproduce: C1 team strength 0.2297 -> 0.229716 (deterministic, topology-only, exact); C2 PCA of the RACIPE ensemble PC1 61.2%->63.38% / PC2 16.4%->15.30% on the all-stable-states ensemble (+2.2pp / -1.1pp; monostable-only ensemble does NOT match, confirming the paper pooled all states); C3 multistability 21.3/27.8/25.8/25.1% -> 21.06/27.88/25.37/25.69% (max deviation 0.59pp). C4 (bonus EMT-Resistance correlation) gives a strong, highly significant, structured correlation r+0.46 p~0, confirming the paper's qualitative EMT/AR-axis coupling claim; graded partial because no exact r is reported and the sign is definition-dependent. No fabrication concern: every reported number is directly derivable from the shipped topology + standard RACIPE. NOT ATTEMPTED (out of scope, see scope.md): GSE74685/clinical transcriptomic EMT-signature analyses (Fig 1,2D,6 - multi-signature ssGSEA across dozens of datasets, far past 80/20), the stochastic Euler-Maruyama landscape sims (qualitative, no single number), the extended PD-L1 10-node circuit (Fig 4-5), and all wet-lab validation.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 78assessed: 2026-06-15 ⛓ 7e3a689972bd
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-15
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15👤 1 human curator(s) · Level L2 2026-06-15
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetThe paper tests whether the underlying regulatory network connecting EMT (miR-200/ZEB1/SLUG/SNAIL) and androgen receptor/AR-V7-driven androgen independence generates emergent, coordinated phenotypic dynamics that link EMT to hormone therapy resistance in prostate cancer.
- ★ Simulations of the EMT-AR crosstalk network reveal four possible phenotypes: epithelial-sensitive (ES), epithelial-resistant (ER), mesenchymal-resistant (MR), and mesenchymal-sensitive (MS), with MS occurring rarely finding
- ★ Two antagonistic 'teams' of regulators (ZEB1, SLUG, LIN28, AR, AR-V7, hnRNPA1 vs. let-7, miR-200, SNAIL) act as a toggle switch and drive coordinated phenotype emergence mechanism
- ★ The 'team' structure is a property of the network topology itself, not of specific parameter choices, since randomized networks show near-zero team strength versus 0.2297 for the biological network finding
- ★ Model predictions of EMT-androgen independence coupling are supported by multiple bulk and single-cell transcriptomic datasets, including in vitro EMT induction models and clinical samples finding
- ★ Simulations reveal spontaneous stochastic switching between the ES and MR states finding
- ★ Addition of PD-L1 to the network captures interactions among AR, PD-L1, and SNAIL, confirmed via quantitative experiments finding
- RACIPE (Random Circuit Perturbation) was used to convert the network topology into an ensemble of coupled ODEs and simulate steady-state phenotypes across randomized kinetic parameters method
- Influence matrix and PCA analyses independently confirm the same two-team composition as the RACIPE co-expression/correlation analysis finding
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| RACIPE (mathematical/ODE-based network simulation) | in silico gene regulatory network model (EMT-AR crosstalk) | randomized kinetic parameter ensembles and initial conditions | steady-state expression levels of network nodes (phenotypes) | RACIPE (Random Circuit Perturbation) |
| Influence matrix / network topology analysis | in silico gene regulatory network (biological vs. 100 randomized networks) | edge-shuffling to generate random networks | team strength (0-1 scale) quantifying antagonism between node groups | — |
| Principal component analysis (PCA) | RACIPE steady-state solution ensemble | none | explained variance and clustering of nodes/samples along PC1/PC2 | — |
| ssGSEA-based transcriptomic correlation analysis | clinical CRPC patient cohort, GSE74685 (n=62) | none (observational) | correlation of Wang Androgen Independent ssGSEA score with ZEB1, SNAI2, KS score, 76GS score | — |
| ssGSEA-based transcriptomic correlation analysis | TCGA prostate cancer cohort | none (observational) | correlation of androgen independence ssGSEA score with EMT metrics (ZEB1, SNAI2, KS, 76GS) | — |
| Pairwise correlation of EMT and androgen-independence gene signatures | multiple transcriptomic datasets (GSE77959 n=30, GSE80042 n=44, GSE67681 n=6, GSE22010 n=12) | none (observational) | Spearman correlation among ssGSEA, ZEB1 expression, KS score, EMT Hallmark ssGSEA, 76GS score | — |
| Quantitative experiment (unspecified assay type) validating AR-PD-L1-SNAIL interaction | not specified in provided text | not specified | expression/interaction of AR, PD-L1, SNAIL | — |
- – Network simulation yields four phenotypes (ES, ER, MR, MS), with MS rare and ER/MR often clustering together by K-means (best K=2)
- ▲ Team strength of the experimentally-derived network was 0.2297, versus values centered around 0 for 100 randomized networks 0.2297
- ▲ Higher team strength in randomized networks correlated with stronger EMT-Resistance score correlation
- – PCA of RACIPE solutions: PC1 and PC2 together explain ~80% of variance PC1=61.2%, PC2=16.4%
- – ssGSEA androgen-independence score correlates positively with ZEB1 and SNAI2 expression and with KS score, and negatively with 76GS score, in both GSE74685 and TCGA
- – Bimodal distributions observed for ZEB1, SLUG, LIN28, let-7, and hnRNPA1 steady-state levels; less clear bimodality for miR-200, SNAIL, AR, AR-V7
- other team strength = 0.2297 (biological network) (compared to ~0 for 100 randomized networks)
- other PC1 = 61.2% variance, PC2 = 16.4% variance (PCA scree plot of RACIPE solutions)
- count n = 62 (GSE74685 castration-resistant prostate cancer patient cohort)
- count n = 149 (TCGA cohort used for ssGSEA correlation analysis (text); figure legend separately lists TCGA n=551)
- count n = 30, 44, 6, 551, 12 (sample sizes for GSE77959, GSE80042, GSE67681, TCGA, and GSE22010 respectively in pairwise correlation analysis (Fig. 2E))
- correlation Spearman's correlation (values not numerically stated in provided text) (correlation between EMT metrics (KS, 76GS, ZEB1, SNAI2) and androgen independence ssGSEA score across datasets)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
The paper is primarily a computational systems biology study that uses RACIPE (Random Circuit Perturbation) ODE-ensemble simulations to characterize the multi-stable dynamics of a coupled EMT–androgen receptor gene regulatory network in prostate cancer. Simulation outputs were characterized with K-means clustering, UMAP, PCA, and bimodality coefficients. Clinical validation was performed by computing Spearman correlations between multiple EMT scoring metrics and ssGSEA-derived androgen independence scores across five independent publicly available transcriptomic cohorts, with p >= 0.05 flagged as non-significant via figure markers. A subset of predictions was additionally supported by quantitative experiments (details not fully provided in the supplied text).
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Spearman's rank correlation | Correlation between EMT metrics (KS score, 76GS score, ZEB1 expression, SNAI2 expression) and ssGSEA androgen independence scores in TCGA and GSE74685 (Fig. 2D); pairwise correlations across five additional cohorts (Fig. 2E) | GSE77959 n=30, GSE80042 n=44, GSE67681 n=6, TCGA n=551, GSE22010 n=12, GSE74685 n=62 | not stated |
| Pairwise correlation (type unspecified) of RACIPE simulation steady-state values | Correlation matrix across all network node pairs from RACIPE ensemble steady-state solutions (Fig. 1B,ii); p >= 0.05 marked non-significant | null | not stated |
| K-means clustering with silhouette score for K selection | Partitioning of EM score vs. Resistance score scatter plot from RACIPE simulation outputs; K=2 chosen by highest silhouette score (Fig. 1C,iii; Fig. S1A) | null | na |
| Principal Component Analysis (PCA) | Dimensionality reduction of RACIPE simulation ensemble; PC1 61.2%, PC2 16.4% variance explained; correlation circle used to visualize team membership (Fig. 2C) | null | na |
| ssGSEA (single-sample Gene Set Enrichment Analysis) | Per-sample quantification of androgen independence (Wang Prostate Cancer Androgen Independent geneset, MSigDB) and Hallmark EMT pathway activity across all clinical cohorts (Fig. 2D, 2E) | Multiple cohorts; largest is TCGA n=551 | not stated |
| Bimodality Coefficient (BC) | Assessment of bimodality in RACIPE steady-state node expression distributions; KDE with Gaussian fits shown for each node (Fig. 1B,i; Tables S2, S3) | null | not stated |
-
Multiple Spearman correlations were computed across metric pairs and cohorts without a stated multiple-testing correction↳ Could also: A Benjamini-Hochberg FDR correction (or Bonferroni for a more conservative bound) could also be applied across the family of pairwise correlation tests within each dataset — With several metric pairs tested simultaneously per cohort, reporting FDR-adjusted q-values alongside raw p-values would provide a complementary perspective on which associations are most robust against inflation of the type I error rate
-
Bimodality in RACIPE output distributions was characterized with the Bimodality Coefficient (BC) and visual KDE-with-Gaussian-fit inspection↳ Could also: Hartigan's dip test for unimodality or a Gaussian mixture model (GMM) with BIC-based component selection could also be applied — These approaches provide either a formal hypothesis test for unimodality or a model-selection criterion for the number of modes, offering a more statistically explicit complement to the descriptive BC
-
K-means clustering with K selected by silhouette score was used to partition RACIPE phenotype space↳ Could also: Gaussian mixture models or hierarchical clustering with gap statistic or within-cluster sum-of-squares elbow assessment could also be used — Model-based clustering supplies probabilistic cluster memberships and uncertainty estimates, and multiple K-selection criteria (gap statistic, BIC) allow cross-validation of the chosen K=2
-
Network topology robustness was evaluated by comparing the observed team strength (0.2297) visually against a histogram of team strengths from 100 randomly shuffled networks↳ Could also: A formal permutation test yielding an explicit p-value (e.g., proportion of 100 or 1000 random networks with team strength >= 0.2297) could also be reported — An empirical permutation p-value would provide a quantitative statement of the rarity of the observed team strength under the null of random topology, complementing the visual histogram comparison
-
ssGSEA was used as the sole method for per-sample pathway scoring↳ Could also: GSVA or PAGE (parametric analysis of gene set enrichment) could also be applied as alternative single-sample enrichment methods — Different single-sample methods make different distributional assumptions; running one as a sensitivity analysis would help assess whether the observed EMT–AR correlations are robust to the choice of enrichment algorithm
-
Clinical associations between EMT and androgen independence were assessed with bivariate Spearman correlations, without adjustment for clinical covariates↳ Could also: Partial Spearman correlation or multivariate regression controlling for available clinical variables (e.g., Gleason score, disease stage, prior treatment) could also be used where covariate data are available — Covariate adjustment would help characterize whether the EMT–AR association is independent of other clinical variables that co-vary with both axes, offering additional interpretive context
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
Assessed papers, coloured by verdict. Click a node to open it.
- No assessed neighbours yet — the network grows as more papers are assessed.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Reproduction scope — pmid-36851919
Paper: Jindal et al. 2023, Emergent dynamics of underlying regulatory network
links EMT and androgen receptor-dependent resistance in prostate cancer.
Comput Struct Biotechnol J. PMID 36851919 · PMCID PMC9957767 · DOI
10.1016/j.csbj.2023.01.031
Code: https://github.com/csbBSSE/AR_Prostate_Cancer (csbBSSE / Jolly lab)
Network: 9-node "ESMR" gene regulatory network — nodes ZEB1, miR-200, SNAI1,
AR, SLUG, let-7, LIN28, hnRNPA1, AR-v7; 35 signed edges (ESMR.topo).
Core computational pipeline
The paper's mechanistic results come from RACIPE (Random Circuit Perturbation;
Huang et al. 2017; tool simonhb1990/RACIPE-1.0) applied to the hand-curated ESMR
topology: 10,000 random ODE parameter sets × 100 initial conditions → ensemble of
stable steady states (log2 expression). Downstream analysis (z-score, PCA,
correlations, multistability counting) is done by the repo's Python/R scripts.
Per BRIEF rule P16, running the canonical RACIPE tool on the paper's own topology
is a fully valid reproduction.
IN SCOPE (pipeline-derived, attempted)
| # | Result | Paper location | Pipeline | How reproduced |
|---|---|---|---|---|
| C1 | Team strength of biological network = 0.2297 | Results text / Fig 2B | topology-only influence matrix (sum of normalized signed path products to length 10) over teams T1={AR,AR-v7,LIN28,SLUG,hnRNPA1,ZEB1}, T2={SNAI1,let-7,miR-200} | deterministic numpy reimpl of the repo's Group Strength Distribution Histogram.py; no random component |
| C2 | PCA variance: PC1 = 61.2%, PC2 = 16.4% (~80% total) | Fig 2C text | RACIPE ensemble → z-score → PCA (sklearn) | run RACIPE, replicate PCA_ScreePlot.py |
| C3 | Multistability: 21.3% mono / 27.8% bi / 25.8% tri / 25.1% >3 states | Fig 3A text | RACIPE per-model #stable-states | count ESMR_solution_k.dat rows |
| C4 (bonus) | EMT–Resistance correlation (emt=ZEB1−miR200, res=AR+AR-v7, z-scored Pearson) | Fig 2B related script | RACIPE ensemble | replicate repo script |
OUT OF SCOPE (not attempted) — and why
- GSE74685 / clinical transcriptomic analyses (Fig 1, parts of Fig 2D, Fig 6 meta-analysis of 70 datasets): require multiple EMT signatures (76GS, KS, MLR), ssGSEA gene lists, and curation across dozens of datasets — large, multi-input, signature-dependent; far past the 80/20 line. The mechanistic RACIPE results (C1–C4) are the paper's central, self-contained, clearly-specified pipeline.
- Stochastic (Euler–Maruyama) landscape simulations (Stochastic Simulations/):
custom R (
simFuncs.R) producing per-parameter trajectory/landscape plots — the outputs are qualitative landscapes (no single reported number to grade against); the 60 MB*_simDat.csvare already shipped. Not a crisp quantitative claim. - PD-L1 extended 10-node network (Fig 4–5,
TS_CircuitB): a second/extended circuit; we reproduce the primary 9-node ESMR network only. - Wet-lab validation (enzalutamide-resistant cell lines, qPCR): experimental, inherently out of scope.
Compute
RACIPE build + run on «our HPC» (SLURM, partition std), data + repo cloned inside
the compute job («infra»). Team strength (C1) is topology-only and deterministic.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
A clean, faithful reproduction: the canonical RACIPE tool run on the authors' own published ESMR topology reproduces every quantitative claim — C1 team strength ~exactly (0.2297→0.229716, deterministic), C3 multistability within 0.59pp, and C2 PCA within ~2pp (PC1 61.2→63.4, PC2 16.4→15.3) — and C4 confirms the qualitative EMT/AR coupling (r≈+0.46, p≈0). Deviations are all technical/stochastic (RACIPE re-samples by design) and sit on neither the authors' nor our 'error' side; no fabrication signal, all values derivable from shared data. Graded overall yellow rather than green only because the match is not bit-exact, the PCA ensemble choice had to be inferred from the authors' repo code (paper underspecifies it), and C4 is qualitative with a definition-dependent sign.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.