Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

A systems-level analysis of the mutually antagonistic roles of RKIP and BACH1 in dynamics of cancer cell plasticity.

J R Soc Interface · 2023
95/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
How its reproducibility compares
95/100
Reproducibility score
1.2 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 89% of all assessed papers rank 105 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

REPRODUCED via a clean fresh «our HPC» run (SLURM «job» on node n093, COMPLETED; this room had been requeued only because its live reproduction/ folder was archived and the «infra» workdir was janitor-reclaimed -- no scientific defect). Scope = the paper's RACIPE systems-modeling core (Figs 3-4). Tier A = deterministic re-analysis of the repo's OWN shipped RACIPE solutions with the paper's own pipeline: monostable 4.58% (paper 4.58%), multistable 95.42% (95.42%), PCA PC1 48.77%/PC2 19.47% (paper 48.8%/19.5%), Fig 4e BACH1(+ve) enrichment 74.94/69.64/63.47% (paper 74.94/69.64/63.48) -- exact to 2 decimals, a strong anti-fabrication signal. Tier B = independent de-novo RACIPE-1.0 simulation (10000 models, third-party tool simonhb1990/RACIPE-1.0@b4483533, wall 4537s): monostable 4.53%, multistable 95.47%, RKIP-BACH1 Spearman rho=-0.474, PC1 48.82%/PC2 19.23% -- the multistability, the negative RKIP-BACH1 correlation (the paper's central thesis), and the PCA structure all reproduce from the topology alone. No evidence of fabrication. NOT attempted (out of 80/20 scope): Fig 1/2 TCGA correlations (TCGA matrices not in repo), Fig 2E multi-GEO ssGSEA, Fig 5 survival.

💻 Code ↗ 🗄 Data: GSE17705

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 95
    assessed: 2026-06-20 ⛓ 57a46670da8e
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-23
Rubric version
not recorded
Assessed by
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

Whether RKIP and BACH1, previously shown to be mutually antagonistic regulators of EMT in breast cancer, display consistent antagonistic trends across cancer types with respect to EMT, metabolic reprogramming and immune evasion, and whether a mechanism-based gene regulatory network can explain this antagonism and its link to cancer cell plasticity and patient survival.

Core claims
  • RKIP and BACH1 are negatively correlated with each other across most cancer types in TCGA finding
  • BACH1 associates with a mesenchymal/pro-EMT phenotype while RKIP associates with an epithelial phenotype pan-cancer finding
  • RKIP correlates positively and BACH1 negatively with oxidative phosphorylation and fatty acid oxidation signatures finding
  • BACH1 correlates positively and RKIP negatively with PD-L1 (immune evasion) and ferroptosis-sensitivity signatures finding
  • A mechanism-based GRN incorporating RKIP, BACH1 and EMT/stemness/tamoxifen-resistance regulators organizes into two mutually antagonistic 'teams' mechanism
  • The wild-type network's team strength is significantly higher than that of 1000 randomized networks, indicating a non-random topological signature finding
  • RACIPE simulations show the GRN is predominantly multi-stable (95.42% of solutions), supporting coexistence of multiple plastic cell states including stem-like states linked to BACH1 finding
  • Low RKIP and high BACH1 levels associate with worse clinical outcomes across many cancer types finding
Experimental setups
Assay System Perturbation Readout Platform
Spearman correlation / ssGSEA scoring of transcriptomic signatures TCGA pan-cancer transcriptomic data (35 cancer types) none correlation of RKIP/BACH1 expression with each other, EMT (KS-Epi/KS-Mes), OXPHOS, FAO, glycolysis, PD-L1 and ferroptosis signatures
Spearman correlation / ssGSEA scoring TCGA breast cancer samples none correlation of RKIP/BACH1 with ESR1, ZEB1, SNAI2, VIM, OVOL2, RKIP pathway metastasis signature (RPMS), BACH1 pathway metastasis signature (BPMS)
Spearman correlation / ssGSEA scoring of microarray data 6 ER+ breast cancer datasets (GSE6532, GSE9195, GSE17705, GSE24202, GSE43495, GSE67916) none correlation of RKIP/BACH1 with RPMS, BPMS, ferroptosis, PD-L1, OXPHOS, FAO, glycolysis signatures
Correlation analysis of transcriptomic data CCLE cancer cell line data none correlation trends analogous to TCGA/ER+ breast cancer datasets
Gene regulatory network construction, adjacency and influence matrix analysis with hierarchical clustering computational in silico model of breast cancer EMT/stemness/tamoxifen-resistance regulators including RKIP and BACH1 none (topology analysis) clustering into antagonistic 'teams', team strength score
Network randomization (edge shuffling, 1000 networks) computational GRN model randomized edge rewiring distribution of team strength compared to wild-type network
RACIPE (ODE-based mathematical modeling) with ensemble kinetic parameters/initial conditions, followed by PCA and clustering computational GRN model of RKIP/BACH1 and EMT/stemness/tamoxifen resistance modules parameter and initial condition ensemble sampling steady-state gene expression profiles, number of stable states per parameter set, PC1/PC2 variance and loadings, EM score, SN score
Key results
  • RKIP and BACH1 negatively correlated in 31 of 35 cancer types 88.57%
  • KS-Mes ssGSEA score negatively correlated with RKIP in 20/35 cancer types, positively with BACH1 in 31/35 57.14% / 88.57%
  • OXPHOS ssGSEA scores correlated positively with RKIP in 31/35 cancer types and negatively with BACH1 in 32/35 88.57% / 91.14%
  • PD-L1 ssGSEA score correlated negatively with RKIP in 24/35 cancer types and positively with BACH1 in 31/35 68.57% / 88.57%
  • In TCGA breast cancer, ESR1 correlated positively with RKIP and negatively with BACH1 rho=0.11 (RKIP), rho=-0.29 (BACH1)
  • ZEB1 correlated positively with BACH1 and negatively with RKIP in TCGA breast cancer rho=0.52 (BACH1), rho=-0.32 (RKIP)
  • Wild-type network team strength (0.37) exceeded that of all 1000 randomized networks 0.37 (scale 0-1)
  • 95.42% of RACIPE parameter sets produced multi-stable (≥2 states) solutions, versus 4.58% mono-stable 95.42%
Key statistics
  • correlation negative correlation in 31/35 cancer types (88.57%) (RKIP vs BACH1 expression, TCGA pan-cancer)
  • correlation positive correlation in 31/35 cancer types (88.57%) (RKIP vs OXPHOS ssGSEA score)
  • correlation negative correlation in 32/35 cancer types (91.14%) (BACH1 vs OXPHOS ssGSEA score)
  • correlation rho = 0.11 (RKIP vs ESR1, TCGA breast cancer)
  • correlation rho = -0.29 (BACH1 vs ESR1, TCGA breast cancer)
  • correlation rho = 0.52 (BACH1 vs ZEB1, TCGA breast cancer)
  • other team strength = 0.37 (scale 0-1), WT network significantly greater than 1000 randomized networks (influence matrix team strength analysis)
  • count 95.42% of RACIPE solutions multi-stable; 4.58% mono-stable (stable steady-state distribution from RACIPE simulation, 3 replicates)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This computational and bioinformatics study used Spearman's rank correlation to characterise pan-cancer associations of RKIP and BACH1 expression with EMT, metabolic, and immune-evasion gene-set scores (ssGSEA) across 35 TCGA cancer types and six external ER+ breast cancer GEO datasets. A mechanism-based gene regulatory network (GRN) was constructed and simulated using the RACIPE framework (ensemble ODE modelling across randomised kinetic parameters), with outputs analysed via PCA, hierarchical clustering of the influence matrix, and an edge-shuffling permutation test of 'team strength'. Results were reported as Spearman's ρ at nominal p-value thresholds, and RACIPE simulation outputs were summarised as mean ± SD across three independent replicates.

Replicationtechnical Sample sizeRACIPE simulations run in three independent replicates (results shown as mean ± SD); transcriptomic analyses span 35 TCGA cancer types and 6 ER+ GEO datasets; per-dataset patient sample sizes not reported in the provided text GroupsRKIP-associated vs BACH1-associated expression patterns; epithelial vs hybrid E/M vs mesenchymal states; stem-like vs non-stem-like; multiple TCGA cancer types compared descriptively Pairingunpaired Randomization/blindingnot stated DispersionSD Exact p-valuesno Effect sizesyes Confidence intervalsno Multiplicity correctionnone stated
Statistical tests used
Test Applied to n Assumptions
Spearman's rank correlation Correlation of RKIP and BACH1 with each other and with ssGSEA scores (KS-Epi, KS-Mes, EMT, OXPHOS, FAO, glycolysis, PD-L1, ferroptosis, RPMS, BPMS) across 35 TCGA cancer types and 6 ER+ GEO datasets; also individual gene correlations (ESR1, ZEB1, SNAI2, VIM, OVOL2) 35 TCGA cancer types; 6 named GEO datasets (GSE6532, GSE9195, GSE17705, GSE24202, GSE43495, GSE67916); per-dataset sample sizes not stated in provided text not stated
ssGSEA (single-sample Gene Set Enrichment Analysis) score computation Per-sample pathway quantification for KS-Epi, KS-Mes, OXPHOS, FAO, glycolysis, PD-L1, ferroptosis, RPMS, and BPMS gene signatures in all transcriptomic datasets null na
Network randomization permutation test (edge-shuffling) Testing whether the wild-type GRN team strength (0.37) exceeds that expected by chance by comparison to 1000 randomly rewired networks 1000 random networks not stated
Principal component analysis (PCA) Dimensionality reduction of RACIPE steady-state solutions; PC1 explained 48.8% of variance, PC2 explained 19.5% null not stated
Hierarchical clustering Clustering of the influence matrix (computed up to path length 8) to identify mutually activating/inhibiting 'teams' of network nodes null not stated
Kernel density estimation with Gaussian fitting Segregation of RACIPE solutions into epithelial/hybrid/mesenchymal states (EM score) and stem-like/non-stem-like states (SN score) using minima of fitted Gaussians null not stated
Approaches that could also have been used
  • Individual Spearman correlations were tested across 35 cancer types and multiple gene signatures simultaneously, with results evaluated at nominal p < 0.05 or p < 0.1 thresholds
    Could also: Apply a false-discovery rate correction (e.g., Benjamini-Hochberg) across the family of parallel tests — With many simultaneous tests (35 cancer types × multiple signatures), FDR control explicitly quantifies the expected proportion of false positives, giving readers a calibrated sense of confidence in individual associations
  • Spearman's rank correlation was used for all gene-expression association analyses without adjustment for covariates
    Could also: Partial Spearman or partial Pearson correlation controlling for tumour purity, sample batch, or cancer subtype composition — Bulk RNA-seq expression values are influenced by tumour cell content and stromal composition; partial correlation can help distinguish associations driven by cell-intrinsic expression from those driven by varying cellular admixture across samples
  • Network dynamics were simulated using RACIPE, an ensemble ODE-based approach requiring continuous kinetic parameters sampled over wide ranges
    Could also: Boolean or probabilistic Boolean network modelling — Boolean frameworks require fewer free parameters and yield attractors directly interpretable as discrete cell states, providing a complementary lower-assumption view of multistability that can corroborate or stress-test ODE-based findings
  • RACIPE steady-state solutions were visualised and clustered using PCA
    Could also: UMAP or t-SNE for nonlinear dimensionality reduction — Nonlinear embedding methods can better separate clusters that are not linearly separable in high-dimensional expression space, potentially revealing sub-populations or continua of states not apparent in PCA projections
  • Phenotypic state boundaries in RACIPE output were defined by applying thresholds at the minima of Gaussian-fitted kernel density curves of one-dimensional EM and SN composite scores
    Could also: Gaussian mixture modelling (GMM) or k-means clustering applied jointly in the multidimensional score space — Multivariate clustering simultaneously accounts for covariation among all scores and provides probabilistic cluster membership with uncertainty estimates, rather than applying sequential univariate thresholds that may not capture correlated structure
  • Dispersion for RACIPE replicate results was reported as mean ± SD across three replicates
    Could also: Report 95% confidence intervals for the mean, or show individual replicate values alongside the mean — With only three replicates, a 95% CI explicitly reflects the uncertainty in the mean estimate; SD alone does not convey the small-n context to readers, and individual data points aid transparency for such low-n technical replication
Software: RACIPE (Random Circuit Perturbation) · ssGSEA (implementation/package not specified)

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-37963558

Paper: Shyam, Ramu, Sehgal, Jolly (2023). A systems-level analysis of the mutually antagonistic roles of RKIP and BACH1 in dynamics of cancer cell plasticity. J R Soc Interface. DOI 10.1098/rsif.2023.0389 · PMCID PMC10645512.

Code repo (pinned): https://github.com/saishyam1/RKIP_BACH1_Data_Codes commit de576fc1d1961a7d16fd77a8e2d832df20d5c536 (default branch main, pushed 2023-10-10). No license file. ~34 MB; ships input topology + RACIPE output + several GEO expression matrices.

Third-party simulation tool (P16-valid): RACIPE-1.0 (https://github.com/simonhb1990/RACIPE-1.0) — the Random Circuit Perturbation ODE engine used by the paper (Methods §4.4). Applying this established tool to the paper's own network topology is an equally valid reproduction.


Result inventory & classification

# Reported result (paper) Pipeline In scope? Why
R1 Multistability distribution: 4.58% monostable / 95.42% ≥2 states (Fig 4a), from 10 000 RACIPE models on the 13-node GRN RACIPE-1.0 simulation → notebook Fig_3_4_paper.ipynb stable-state count YES Topology (Network5_trial*.topo) + RACIPE output (*_solution.dat) + analysis ship in repo; tool is public
R2 PCA of RACIPE steady states: PC1 = 48.8%, PC2 = 19.5% of variance (Fig 3e) z-score + sklearn PCA on RACIPE solutions YES Deterministic given shipped _solution.dat; sklearn PCA
R3 RKIP–BACH1 mutually antagonistic (negative) correlation in RACIPE steady states (Fig 3Bii Spearman heatmap) — the paper's core thesis Spearman ρ on pooled steady states YES Deterministic given shipped data; reproducible de novo from topology
R4 BACH1(+) enrichment in mesenchymal/stem states (74.94%, 69.64%, 63.48% hybrid) (Fig 4e) EM/SN scoring + threshold classification partial / lower priority Depends on hand-picked score thresholds (0.63/-0.25 etc., hard-coded in notebook); reproducible but threshold-sensitive
F1 Pan-cancer TCGA: RKIP–BACH1 negatively correlated in 31/35 (88.57%) cancer types; OXPHOS/FAO/PD-L1/KS patterns (Fig 1) ssGSEA + Spearman on TCGA OUT TCGA pan-cancer matrices NOT in repo; not the registered accession; download/processing = the hard last 20%
F2 Breast-cancer TCGA: ESR1 ρ(RKIP)=0.11, ρ(BACH1)=−0.29; ZEB1 ρ(BACH1)=0.52, ρ(RKIP)=−0.32 (Fig 2a,b) Spearman on TCGA-BRCA OUT TCGA-BRCA not shipped; out of 80/20 budget
F3 Fig 2E gene-signature scores across 6 GEO datasets (GSE24202/27473/43495/6532/67916/9195) ssGSEA with shipped .gmt signatures OUT (secondary) Data ships, but ssGSEA scoring + many datasets is broader than the few-clear-points target; RACIPE chosen as the cleaner core
F4 Survival analysis (Fig 5, rkip_bach1_survival.R) R survival / KM OUT Wet-data dependent (METABRIC/cohort), separate R stack; out of scope budget

Reproduction plan (in scope: R1, R2, R3; R4 noted)

Two complementary tiers, all compute on «our HPC», data on «infra»:

  • Tier A — re-analysis of shipped RACIPE output (deterministic, exact target). Run the paper's own analysis on the three shipped Network5_trial{1,2,3}_solution.dat files → recompute monostable %, full stable-state distribution (R1), PCA explained variance (R2), and RKIP–BACH1 Spearman ρ (R3). This tests whether the reported numbers are re-derivable from the shipped data (anti-fabrication check).

  • Tier B — independent RACIPE simulation (statistical target). Compile RACIPE-1.0, run 10 000 models on Network5_trial1.topo with default settings → recompute monostable % and RKIP–BACH1 Spearman ρ on a freshly simulated ensemble. RACIPE samples parameters randomly, so this is a statistical (not bit-exact) reproduction; agreement in sign + approximate magnitude confirms the result follows from the topology, not just from the shipped numbers.

Not attempted: F1, F2 (TCGA downloads), F3 (mult

Figures / tables: Fig 4aFig 3eFig 3Fig 4e
R1a
Reported
4.58% monostable models (Fig 4a)
Reproduced
4.58% (Tier A 3-trial mean 4.79/4.75/4.21)
exact
R1b
Reported
95.42% multistable (>=2 states) (Fig 4a)
Reproduced
95.42% (Tier A 3-trial mean 95.21/95.25/95.79)
exact
R2a
Reported
PCA PC1 = 48.8% variance (Fig 3e)
Reproduced
48.77% (Tier A trial3)
exact
R2b
Reported
PCA PC2 = 19.5% variance (Fig 3e)
Reproduced
19.47% (Tier A trial3)
within tolerance
R3
Reported
RKIP-BACH1 mutually antagonistic / negatively correlated (Fig 3Bii, core thesis)
Reproduced
Spearman rho=-0.48 (Tier A) / -0.474 (Tier B de-novo), p~0
within tolerance
R4a
Reported
BACH1(+ve) enrichment Mesenchymal 74.94% (Fig 4e)
Reproduced
74.94% (3-trial mean 74.55/75.41/74.85, sd 0.44)
exact
R4b
Reported
BACH1(+ve) enrichment Stem-like 69.64% (Fig 4e)
Reproduced
69.64% (3-trial mean 68.98/70.60/69.34, sd 0.85)
exact
R4c
Reported
BACH1(+ve) enrichment Hybrid 63.48% (Fig 4e)
Reproduced
63.47% (3-trial mean 63.57/64.13/62.71, sd 0.72)
exact
R5
Reported
Independent RACIPE-1.0 simulation consistent with R1/R3 (statistical)
Reproduced
monostable 4.53% / multistable 95.47% / rho=-0.474 / PC1 48.82% PC2 19.23% (10000 models, fresh «job»)
within tolerance

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

No assessment has been recorded yet.
🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

506.9 k
tokens (I/O) · 38.1 M incl. cache
245 min
runtime · 1.26 CPU-h
0.3 GB
peak RAM
1
HPC jobs
hummel
machine