Topological signatures in regulatory network enable phenotypic heterogeneity in small cell lung cancer.
The main results reproduced: recomputed values matched the published ones within tolerance.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Same input data as the authors
- ✓Reported values were directly comparable
- ✓No relevant deviation in data/preprocessing
- ✓No authors-side cause for any deviation
- ✓Reported values are derivable from the shared data
- ✓Any deviation was negligible
- ✓The central claim held under reproduction
- ✓Overall, the reproduction was clean
- Every checked point held up.
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Described well enough -> faithful 1:1. Reproduced the paper's headline Boolean result (Fig 1B) by running the authors' own Fast-Bool tool (github.com/uday2607/CSB-SCLC @ d1a885f) with their shipped bool.in (2^20 initial conditions, 5000 steps, async Ising) on the shipped 33-node WT-SCLC network, freshly on «our HPC» and independent of the committed output. Got exactly 10 stable steady states: 4 dominant (24.32-25.63% basins, frustration ~0.14) and 6 rare (<0.05%, frustration ~0.37-0.39), matching paper Fig 1B (10 states; 4 @ 24.3-25.5%; 6 @ <0.1%; frustration 0.14 vs 0.37) and the repo's committed Summary_Async.xlsx. Stochastic basin sizes agree to ~3 sig figs (re-seeded run); steady-state count and per-state frustration are exact. NOT attempted (optional hard 20%): RACIPE continuous ODE ensemble (external RACIPE-1.0 C tool + multi-GB Google-Drive data) and the CCLE/GSE73160 expression-clustering validation. No completeness claim; no fabrication signal.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 94assessed: 2026-06-15 ⛓ 79a0095a5e34
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-15
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusThe authors hypothesize that the underlying network topology of the small cell lung cancer (SCLC) regulatory network is fundamental to the emergence of multistability and the co-existence of distinct phenotypes, thereby explaining phenotypic (non-genetic) heterogeneity in SCLC.
- ★ Discrete (Boolean/Ising) and continuous (RACIPE) simulations of the SCLC regulatory network yield similar multistable phenotypic distributions, with four dominant steady states (X1-X4) that map onto experimentally observed SCLC molecular subtypes. finding
- ★ Multistability in the SCLC network emerges from two mutually inhibiting 'teams' of players whose members activate one another within a team, forming an effective toggle switch between the teams. mechanism
- ★ The two-team topological signature is specific to the wild-type SCLC network; randomizing or perturbing the topology disrupts the signature and the four phenotypes disappear. finding
- ★ The four predominant phenotypes map onto SCLC subtypes ASCL1high/NEUROD1low, ASCL1low/NEUROD1high, ASCL1high/NEUROD1high, and ASCL1low/NEUROD1low; the last can be sub-classified by YAP1/POU2F3 status. finding
- The SCLC network is resilient to single-edge perturbations, with NEUROD1 incoming edges identified as key to maintaining robustness of network dynamics. finding
- An influence matrix metric quantifies topological influence between nodes and recapitulates the two-team correlation structure, demonstrating that team members effectively activate one another while opposing teams inhibit each other. method
- The J metric quantifies the cumulative strength of within-team similarity and between-team competition, distinguishing wild-type from randomized networks. method
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Boolean modeling (Ising model formalism, asynchronous update) | SCLC master regulatory network (33 nodes, 357 edges) | none (WT); random initial conditions | ensemble of steady states and their frequencies; frustration | — |
| RACIPE (Random Circuit Perturbation, coupled ODE ensemble) | SCLC master regulatory network (33 nodes, 357 edges) | parameter sampling over biologically relevant range | steady-state distribution (discretized) and pairwise node correlations | — |
| Network randomization via edge swapping | 1000 randomized networks derived from SCLC WT network | random edge-pair swapping | number of steady states; J metric distribution | — |
| Single-edge perturbation (deletion or sign reversal) | 714 'mutant' SCLC networks | edge deletion or activation/inhibition sign reversal | Jensen-Shannon divergence vs WT steady-state distribution | — |
| Experimental gene expression correlation analysis | SCLC cell lines/tumors (CCLE n=52; GSE73160 n=63) | none | pairwise node correlation matrix; J metric | — |
| Influence matrix computation (path lengths up to lmax=10) | SCLC WT network | none | net topological influence (Inf_ij) between nodes; R1/R2 correlation with RACIPE correlations | — |
- – Boolean simulation yielded 10 unique steady states; four dominant states (X1-X4) each at ~24.3-25.5% frequency, remaining six states <0.1% 24.3-25.5% each (X1-X4)
- – High-frequency states had low frustration while low-frequency states had higher frustration 0.14 vs 0.37
- – Net frequency of X1-X4 states in RACIPE; top 20 RACIPE states (within 2 nodes of X1-X4) contributed substantial cumulative frequency 22% (X1-X4); top 20 states = 54%
- ▲ Random networks produced far more steady states than WT and had no overlap with WT distribution 10^4-10^6 states
- – Most single-edge mutations had negligible effect (JSD<0.01); of 36 high-JSD perturbations, 24 were incoming edges for NEUROD1 678 of 714 with JSD<0.01; 24 of 36
- – 32 of 33 nodes (all except NEUROD1) split into two groups positively correlated within and negatively across groups 32 of 33 nodes
- ▲ J metric much higher for WT than random networks, distinguishing WT topology Boolean J=496, RACIPE J=373.05 vs random mean 11 (Boolean) / 14.16 (continuous)
- ▲ Influence matrix (lmax=10) was highly similar to RACIPE correlation matrix; experimental datasets had J significantly larger than random networks p<0.01 (two-tailed z-test)
- correlation 0.851 ± 0.003 (Pearson correlation between JSD for edge deletion vs edge sign reversal, mean±SD over 3 replicates)
- count 33 nodes and 357 edges (size of SCLC master regulatory network)
- other J=496 (J metric for WT Boolean simulations)
- other J=373.05 (J metric for WT RACIPE simulations)
- mean 11 (average J across 1000 random networks (Boolean))
- mean 14.16 (mean J across 1000 random networks (continuous Pij sampling))
- pvalue p<0.01 (two-tailed z-test, WT/experimental J vs random networks)
- count n=52 (CCLE), n=63 (GSE73160) (experimental SCLC datasets used for correlation analysis)
Statistical methods review
Model: opusA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This is a computational/systems-biology study that characterizes a 33-node, 357-edge small cell lung cancer regulatory network using two simulation frameworks: a parameter-independent asynchronous Boolean (Ising) model and RACIPE, an ODE-based parameter-agnostic ensemble approach. Outputs are steady-state frequency distributions, pairwise Pearson node correlations, an information-theoretic Jensen-Shannon divergence between perturbed and wild-type distributions, and a custom 'J' team-strength metric; comparisons against 1000 edge-swapped random networks and two experimental expression datasets establish the topological signature. Quantitative significance of the J metric is reported with a two-tailed z-test, and simulation summaries are reported as mean ± standard deviation over three replicates.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| two-tailed z-test | comparison of J metric for WT Boolean and RACIPE simulations (and CCLE/GSE73160 datasets) versus the distribution of J across 1000 random networks (Figure 2B,ii) | 1000 random networks (reference distribution) | not stated |
| Pearson's correlation coefficient | pairwise node-node correlations across steady states (Figure 2A,B,C) and regression of correlation vs influence matrix (R1, R2) | — | not stated |
| Jensen-Shannon divergence (information-theoretic distance, not a hypothesis test) | similarity between WT and 714 single-edge perturbed network steady-state distributions, and 1000 random networks (Figure 1D,ii) | 2^20 initial conditions per network | na |
-
Significance of the J metric was assessed with a two-tailed z-test comparing the wild-type value against the distribution of J from 1000 random networks.↳ Could also: An empirical (permutation/bootstrap) p-value derived directly from the rank of the WT value within the 1000-network null distribution could also be reported. — An empirical null does not rely on an assumption of normality for the random-network J distribution and naturally accommodates skew or bounded ranges, which one might prefer when the null is generated by simulation.
-
p-values were reported as thresholds (p<0.01).↳ Could also: Exact p-values alongside the observed effect size (e.g., standardized distance of WT J from the null mean) could also be presented. — Exact values and effect magnitudes convey how far the observation lies from the null beyond a pass/fail cutoff, aiding cross-study comparison.
-
Variability across simulation runs was summarized as mean ± standard deviation over three replicates.↳ Could also: A 95% confidence interval or the full range/individual replicate values could also be shown. — With only three replicates, displaying individual points or a CI can communicate the precision of the estimate transparently and is often favored for very small n.
-
Distribution similarity was quantified with Jensen-Shannon divergence.↳ Could also: Complementary metrics such as Kullback-Leibler divergence, total variation distance, or earth-mover's distance could also be used. — Reporting an additional metric can show that conclusions about distributional overlap are robust to the choice of divergence measure.
-
Node relationships were assessed with Pearson's correlation coefficient.↳ Could also: Spearman's rank correlation could also be computed, particularly for the continuous (non-discretized) RACIPE outputs. — A rank-based correlation captures monotonic but nonlinear associations and is less sensitive to outliers, offering a useful cross-check on linear correlations.
-
Statistical comparisons treated the 1000 random networks as the reference family without a stated multiplicity adjustment across the several z-test comparisons.↳ Could also: A family-wise (e.g., Bonferroni) or FDR (Benjamini-Hochberg) adjustment across the related comparisons could also be applied. — When several significance statements are made against the same null, a correction documents control of the overall error rate across the family of tests.
Result convergence & founder nodes
Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.
-
32 of 33 SCLC regulatory network nodes partition into two internally positively correlated, mutually anti-correlated gene modules, reflecting an underlying bimodal phenotypic dichotomy.other sclc regulatory network 2021×1papers★ This paper is the founder (earliest)
-
The J metric (bimodality index) of the WT 33-node SCLC regulatory network far exceeds randomized edge-swap networks (Boolean J=496 vs mean 11; RACIPE J=373 vs mean 14), demonstrating that WT topology non-randomly encodes phenotypic heterogeneity.other sclc regulatory network up 2021×1papers★ This paper is the founder (earliest)
-
Dominant SCLC Boolean steady states have low Ising frustration (~0.14) while rare states have high frustration (~0.37), linking thermodynamic stability to phenotypic dominance.other sclc regulatory network mixed 2021×1papers★ This paper is the founder (earliest)
-
Single-edge perturbation of the SCLC network shows that 24 of 36 high-impact mutations (JSD>0.01) target incoming edges of NEUROD1, identifying it as the most topologically sensitive node governing network state distribution.other sclc regulatory network 2021×1papers★ This paper is the founder (earliest)
-
Boolean simulation of the 33-node SCLC master regulatory network yields 10 steady states; four dominant phenotypic states (X1–X4) each occur at ~24–25% frequency, encoding four near-equiprobable phenotypes.other sclc regulatory network 2021×1papers★ This paper is the founder (earliest)
-
Randomized SCLC networks (edge-swap) yield 10^4–10^6 steady states with no overlap with the WT distribution, demonstrating that WT topology uniquely constrains the network to a small, specific set of phenotypic attractors.other sclc regulatory network down 2021×1papers★ This paper is the founder (earliest)
-
J metric computed from experimental SCLC gene expression datasets (CCLE n=52; GSE73160 n=63) is significantly higher than randomized networks (p<0.01 by two-tailed z-test), validating model-predicted bimodal phenotypic heterogeneity in real tumors and cell lines.RNA-seq human sclc up 2021×1papers★ This paper is the founder (earliest)
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
Assessed papers, coloured by verdict. Click a node to open it.
- No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
- Generic injuries are sufficient to induce ecto... L1 100/100
- Competitive binding of STATs to receptor phosp... L1 89/100
- Co-regulation and function of <i>FOXM1</i>/<i>... L1 76/100
- Tbx5 drives <i>Aldh1a2</i> expression to regul... L1 84/100
- Disrupted PGR-B and ESR1 signaling underlies d... L1 78/100
- Firefly genomes illuminate parallel origins of... L1 93/100
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-33729159
Paper: Chauhan, Ram, Hari, Jolly. Topological signatures in regulatory network
enable phenotypic heterogeneity in small cell lung cancer. eLife 2021;10:e64522.
Code: https://github.com/uday2607/CSB-SCLC @ d1a885fc91fff11b9e8660c9767f69b55e8954de
Type: theory / computational-systems-biology paper (no wet lab to reproduce).
The repo wraps three computational pipelines:
| # | Pipeline | Tool | In scope? | Reason |
|---|---|---|---|---|
| 1 | Discrete (Boolean/Ising) modelling of the WT-SCLC network — asynchronous update → stable steady states, their frequency & frustration | Additional_Codes/Fast-Bool (pure Python + numba, self-contained, network .topo/.ids shipped in repo) |
YES — primary target | Fully specified: shipped bool.in (2^20 initial conditions, 5000 steps, Ising async). Produces the paper's headline Fig 1B numbers. Low-hanging, deterministic frustration, stochastic basin sampling. |
| 2 | Continuous (RACIPE) modelling — ensemble of ODE models of the WT-SCLC network | external RACIPE-1.0 C package (simonhb1990/RACIPE-1.0); simulation outputs hosted off-repo on Google Drive (authors say files are "quite huge") | NO — the optional hard 20% | Requires a separate external tool + multi-GB Drive downloads; stochastic ODE ensemble. Skipped per 80/20; would mainly re-confirm the same phenotype structure already covered by pipeline 1. |
| 3 | Experimental-data analysis — clustering/UMAP/correlation of CCLE & GSE73160 SCLC cell-line expression to match model phenotypes | Additional_Codes/BioData-Analysis (data mGSE73160.txt, CCLE.txt shipped) |
NO — secondary | Multi-step clustering/UMAP with several free parameters; validation layer rather than a single pinnable headline number. Out of 80/20 scope. |
What we reproduce (pipeline 1)
Run the authors' own Fast-Bool tool with their shipped bool.in on the shipped
33-node SCLC network, freshly on «our HPC» (independent of the committed output), and
compare the steady-state structure to paper Figure 1B (and the shipped
OUTPUT/sclcnetwork/Summary_Async.xlsx as a within-repo cross-check):
- number of unique stable steady states (paper: 10),
- count of dominant states at ~24.3–25.5% (paper: 4, "X1–X4"),
- count of rare states at <0.1% (paper: 6, "X5–X10"),
- frustration of dominant vs rare states (paper: 0.14 vs 0.37),
- network size 33 nodes / 357 edges (sanity check on the shipped
.topo).
Not attempted (honest)
RACIPE continuous ensemble (pipeline 2) and the CCLE/GSE73160 expression-clustering validation (pipeline 3). These are the optional hard tail; skipping them is recorded, not hidden. This is not a completeness claim.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
This is a clean 1:1 reproduction of the paper's headline Fig-1B Boolean result, run freshly on «our HPC» with the authors' own shipped code and 33-node network, independent of the committed output. All five claims hold: exact node/edge count (33/357), exact steady-state count (10), the 4-dominant/6-rare structure, and the deterministic frustration split (0.14 vs 0.37). The only deviations are stochastic basin-size differences at the 4th significant digit (e.g. 25.63% vs 25.5%), which are expected Monte-Carlo noise and on the technical/our-run side, not an authors' defect. The optional RACIPE and GSE73160 expression-clustering tails were out of scope but do not bear on the reproduced central claim.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.