Topological signatures in regulatory network enable phenotypic heterogeneity in small cell lung cancer.
The main results reproduced: recomputed values matched the published ones within tolerance.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Same input data as the authors
- ✓Reported values were directly comparable
- ✓No relevant deviation in data/preprocessing
- ✓No authors-side cause for any deviation
- ✓Reported values are derivable from the shared data
- ✓Any deviation was negligible
- ✓The central claim held under reproduction
- ✓Overall, the reproduction was clean
- Every checked point held up.
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Described well enough -> faithful 1:1. Reproduced the paper's headline Boolean result (Fig 1B) by running the authors' own Fast-Bool tool (github.com/uday2607/CSB-SCLC @ d1a885f) with their shipped bool.in (2^20 initial conditions, 5000 steps, async Ising) on the shipped 33-node WT-SCLC network, freshly on «our HPC» and independent of the committed output. Got exactly 10 stable steady states: 4 dominant (24.32-25.63% basins, frustration ~0.14) and 6 rare (<0.05%, frustration ~0.37-0.39), matching paper Fig 1B (10 states; 4 @ 24.3-25.5%; 6 @ <0.1%; frustration 0.14 vs 0.37) and the repo's committed Summary_Async.xlsx. Stochastic basin sizes agree to ~3 sig figs (re-seeded run); steady-state count and per-state frustration are exact. NOT attempted (optional hard 20%): RACIPE continuous ODE ensemble (external RACIPE-1.0 C tool + multi-GB Google-Drive data) and the CCLE/GSE73160 expression-clustering validation. No completeness claim; no fabrication signal.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 94assessed: 2026-06-15 ⛓ 79a0095a5e34
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-15
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15no human curator yet
- Last updated
- 2026-09-19
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetThe paper tests whether the topology of the SCLC regulatory network (33 nodes, 357 edges) is sufficient to generate multistability that explains the co-existence of four experimentally observed SCLC phenotypes, independent of genetic mutation.
- ★ The SCLC regulatory network is multistable and its steady states map onto four experimentally observed phenotypes (ASCL1high/NEUROD1low, ASCL1low/NEUROD1high, ASCL1high/NEUROD1high, ASCL1low/NEUROD1low) finding
- ★ Multistability arises from two 'teams' of nodes that mutually activate members within their own team but inhibit the opposing team, forming an effective toggle switch mechanism
- ★ Discrete Boolean (Ising) and continuous parameter-agnostic RACIPE simulations yield highly consistent steady-state distributions for the SCLC network finding
- ★ The specific steady-state distribution of the network is a property of its particular topology, since randomized (edge-swapped) networks produce vastly more states with no overlap to the wild-type distribution finding
- ★ The 'two-team' topological signature (quantified by metric J and the influence matrix) is specific to the wild-type SCLC network topology and disrupted by network perturbation mechanism
- NEUROD1 incoming edges are disproportionately important for maintaining robustness of the network's steady-state distribution finding
- The ASCL1low/NEUROD1low subtype can be further sub-classified into ASCL1low/NEUROD1low/YAP1low/POU2F3high and ASCL1low/NEUROD1low/YAP1high/POU2F3low classes finding
- A curated 33-node, 357-edge SCLC master regulatory network built from gene expression signatures resource
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Boolean/Ising model dynamical simulation | SCLC regulatory network (33 nodes, 357 edges) | none (2^20 random initial conditions) | steady-state frequency distribution, frustration | asynchronous Ising model formalism |
| RACIPE ODE ensemble simulation | SCLC regulatory network | none (parameter sampling, 10^6 models x 1000 initial conditions) | discretized steady-state distribution | RACIPE (Random Circuit Perturbation) |
| Network randomization via edge swapping | randomized SCLC network topologies (n=1000) | edge swap (topology randomization) | number of steady states vs wild-type | Ising model formalism |
| Single-edge mutant network simulation | SCLC network with one edge deleted or sign-reversed (n=714) | edge deletion or inhibitory/activating sign reversal | Jensen-Shannon divergence (JSD) vs wild-type steady-state distribution | Ising model formalism |
| Pairwise node correlation analysis of gene expression | CCLE cell lines (n=52) | none | Pearson correlation matrix, J metric | CCLE dataset |
| Pairwise node correlation analysis of gene expression | GSE73160 dataset (n=63) | none | Pearson correlation matrix, J metric | GEO dataset GSE73160 |
| Influence matrix (multi-path topological influence) computation | SCLC regulatory network topology | none | net activating/inhibiting influence between node pairs (Inf_ij) | computational path-length weighted-sum analysis (l_max=10) |
- – 10 unique steady states obtained; four dominant states (X1-X4) each with 24.3-25.5% frequency, remaining six states <0.1% each 24.3-25.5%
- – High-frequency states show low frustration while low-frequency states show high frustration 0.14 vs 0.37
- ▲ Randomized networks (n=1000) produced far more steady states than wild-type, with no overlap in steady-state distributions 10^4-10^6 vs 10 states
- – RACIPE top-20 states closely matched Boolean X1-X4 states (max 2/33 differing nodes); combined frequency of X1-X4-like states was 22% (top 4) and 54% (top 20) 22%/54%
- – Most single-edge perturbations had negligible effect on network dynamics; of 36 perturbations causing high JSD, 24 were incoming edges to NEUROD1 678/714 JSD<0.01; 24/36
- ▲ J metric for wild-type Boolean network (496) and RACIPE network (373.05) far exceeded mean J for random networks (11 and 14.16 respectively) 496 vs 11; 373.05 vs 14.16
- ▲ J metric for experimental datasets CCLE and GSE73160 was significantly larger than for random networks p<0.01
- – 32 of 33 network nodes form two anti-correlated teams (Group A: ASCL1, INSM1, FOXA1, FOXA2; Group B: REST, SMAD3, ZEB1)
- other J=496 (J metric for WT Boolean network correlation matrix)
- other J=373.05 (J metric for WT RACIPE correlation matrix)
- mean mean J=11 (average J metric across random (Boolean) networks)
- mean mean J=14.16 (average J metric across random networks (continuous correlation))
- pvalue p<0.01 (two-tailed z-test) (WT vs random network J metric significance)
- correlation r=0.851 ± 0.003 (Pearson) (JSD correlation between edge-deletion and sign-reversal perturbations)
- count n=52 (CCLE cell line samples used for correlation analysis)
- count n=63 (GSE73160 samples used for correlation analysis)
Statistical methods review
Model: opusA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This is a computational/systems-biology study that characterizes a 33-node, 357-edge small cell lung cancer regulatory network using two simulation frameworks: a parameter-independent asynchronous Boolean (Ising) model and RACIPE, an ODE-based parameter-agnostic ensemble approach. Outputs are steady-state frequency distributions, pairwise Pearson node correlations, an information-theoretic Jensen-Shannon divergence between perturbed and wild-type distributions, and a custom 'J' team-strength metric; comparisons against 1000 edge-swapped random networks and two experimental expression datasets establish the topological signature. Quantitative significance of the J metric is reported with a two-tailed z-test, and simulation summaries are reported as mean ± standard deviation over three replicates.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| two-tailed z-test | comparison of J metric for WT Boolean and RACIPE simulations (and CCLE/GSE73160 datasets) versus the distribution of J across 1000 random networks (Figure 2B,ii) | 1000 random networks (reference distribution) | not stated |
| Pearson's correlation coefficient | pairwise node-node correlations across steady states (Figure 2A,B,C) and regression of correlation vs influence matrix (R1, R2) | — | not stated |
| Jensen-Shannon divergence (information-theoretic distance, not a hypothesis test) | similarity between WT and 714 single-edge perturbed network steady-state distributions, and 1000 random networks (Figure 1D,ii) | 2^20 initial conditions per network | na |
-
Significance of the J metric was assessed with a two-tailed z-test comparing the wild-type value against the distribution of J from 1000 random networks.↳ Could also: An empirical (permutation/bootstrap) p-value derived directly from the rank of the WT value within the 1000-network null distribution could also be reported. — An empirical null does not rely on an assumption of normality for the random-network J distribution and naturally accommodates skew or bounded ranges, which one might prefer when the null is generated by simulation.
-
p-values were reported as thresholds (p<0.01).↳ Could also: Exact p-values alongside the observed effect size (e.g., standardized distance of WT J from the null mean) could also be presented. — Exact values and effect magnitudes convey how far the observation lies from the null beyond a pass/fail cutoff, aiding cross-study comparison.
-
Variability across simulation runs was summarized as mean ± standard deviation over three replicates.↳ Could also: A 95% confidence interval or the full range/individual replicate values could also be shown. — With only three replicates, displaying individual points or a CI can communicate the precision of the estimate transparently and is often favored for very small n.
-
Distribution similarity was quantified with Jensen-Shannon divergence.↳ Could also: Complementary metrics such as Kullback-Leibler divergence, total variation distance, or earth-mover's distance could also be used. — Reporting an additional metric can show that conclusions about distributional overlap are robust to the choice of divergence measure.
-
Node relationships were assessed with Pearson's correlation coefficient.↳ Could also: Spearman's rank correlation could also be computed, particularly for the continuous (non-discretized) RACIPE outputs. — A rank-based correlation captures monotonic but nonlinear associations and is less sensitive to outliers, offering a useful cross-check on linear correlations.
-
Statistical comparisons treated the 1000 random networks as the reference family without a stated multiplicity adjustment across the several z-test comparisons.↳ Could also: A family-wise (e.g., Bonferroni) or FDR (Benjamini-Hochberg) adjustment across the related comparisons could also be applied. — When several significance statements are made against the same null, a correction documents control of the overall error rate across the family of tests.
Result convergence & founder nodes
Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.
-
32 of 33 SCLC regulatory network nodes partition into two internally positively correlated, mutually anti-correlated gene modules, reflecting an underlying bimodal phenotypic dichotomy.other sclc regulatory network 2021×1papers★ This paper is the founder (earliest)
-
The J metric (bimodality index) of the WT 33-node SCLC regulatory network far exceeds randomized edge-swap networks (Boolean J=496 vs mean 11; RACIPE J=373 vs mean 14), demonstrating that WT topology non-randomly encodes phenotypic heterogeneity.other sclc regulatory network up 2021×1papers★ This paper is the founder (earliest)
-
Dominant SCLC Boolean steady states have low Ising frustration (~0.14) while rare states have high frustration (~0.37), linking thermodynamic stability to phenotypic dominance.other sclc regulatory network mixed 2021×1papers★ This paper is the founder (earliest)
-
Single-edge perturbation of the SCLC network shows that 24 of 36 high-impact mutations (JSD>0.01) target incoming edges of NEUROD1, identifying it as the most topologically sensitive node governing network state distribution.other sclc regulatory network 2021×1papers★ This paper is the founder (earliest)
-
Boolean simulation of the 33-node SCLC master regulatory network yields 10 steady states; four dominant phenotypic states (X1–X4) each occur at ~24–25% frequency, encoding four near-equiprobable phenotypes.other sclc regulatory network 2021×1papers★ This paper is the founder (earliest)
-
Randomized SCLC networks (edge-swap) yield 10^4–10^6 steady states with no overlap with the WT distribution, demonstrating that WT topology uniquely constrains the network to a small, specific set of phenotypic attractors.other sclc regulatory network down 2021×1papers★ This paper is the founder (earliest)
-
J metric computed from experimental SCLC gene expression datasets (CCLE n=52; GSE73160 n=63) is significantly higher than randomized networks (p<0.01 by two-tailed z-test), validating model-predicted bimodal phenotypic heterogeneity in real tumors and cell lines.RNA-seq human sclc up 2021×1papers★ This paper is the founder (earliest)
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
Assessed papers, coloured by verdict. Click a node to open it.
- No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
- Generic injuries are sufficient to induce ecto... L1 100/100
- Competitive binding of STATs to receptor phosp... L1 89/100
- Co-regulation and function of <i>FOXM1</i>/<i>... L1 76/100
- Tbx5 drives <i>Aldh1a2</i> expression to regul... L1 84/100
- Disrupted PGR-B and ESR1 signaling underlies d... L1 78/100
- Firefly genomes illuminate parallel origins of... L1 93/100
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-33729159
Paper: Chauhan, Ram, Hari, Jolly. Topological signatures in regulatory network
enable phenotypic heterogeneity in small cell lung cancer. eLife 2021;10:e64522.
Code: https://github.com/uday2607/CSB-SCLC @ d1a885fc91fff11b9e8660c9767f69b55e8954de
Type: theory / computational-systems-biology paper (no wet lab to reproduce).
The repo wraps three computational pipelines:
| # | Pipeline | Tool | In scope? | Reason |
|---|---|---|---|---|
| 1 | Discrete (Boolean/Ising) modelling of the WT-SCLC network — asynchronous update → stable steady states, their frequency & frustration | Additional_Codes/Fast-Bool (pure Python + numba, self-contained, network .topo/.ids shipped in repo) |
YES — primary target | Fully specified: shipped bool.in (2^20 initial conditions, 5000 steps, Ising async). Produces the paper's headline Fig 1B numbers. Low-hanging, deterministic frustration, stochastic basin sampling. |
| 2 | Continuous (RACIPE) modelling — ensemble of ODE models of the WT-SCLC network | external RACIPE-1.0 C package (simonhb1990/RACIPE-1.0); simulation outputs hosted off-repo on Google Drive (authors say files are "quite huge") | NO — the optional hard 20% | Requires a separate external tool + multi-GB Drive downloads; stochastic ODE ensemble. Skipped per 80/20; would mainly re-confirm the same phenotype structure already covered by pipeline 1. |
| 3 | Experimental-data analysis — clustering/UMAP/correlation of CCLE & GSE73160 SCLC cell-line expression to match model phenotypes | Additional_Codes/BioData-Analysis (data mGSE73160.txt, CCLE.txt shipped) |
NO — secondary | Multi-step clustering/UMAP with several free parameters; validation layer rather than a single pinnable headline number. Out of 80/20 scope. |
What we reproduce (pipeline 1)
Run the authors' own Fast-Bool tool with their shipped bool.in on the shipped
33-node SCLC network, freshly on «our HPC» (independent of the committed output), and
compare the steady-state structure to paper Figure 1B (and the shipped
OUTPUT/sclcnetwork/Summary_Async.xlsx as a within-repo cross-check):
- number of unique stable steady states (paper: 10),
- count of dominant states at ~24.3–25.5% (paper: 4, "X1–X4"),
- count of rare states at <0.1% (paper: 6, "X5–X10"),
- frustration of dominant vs rare states (paper: 0.14 vs 0.37),
- network size 33 nodes / 357 edges (sanity check on the shipped
.topo).
Not attempted (honest)
RACIPE continuous ensemble (pipeline 2) and the CCLE/GSE73160 expression-clustering validation (pipeline 3). These are the optional hard tail; skipping them is recorded, not hidden. This is not a completeness claim.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
This is a clean 1:1 reproduction of the paper's headline Fig-1B Boolean result, run freshly on «our HPC» with the authors' own shipped code and 33-node network, independent of the committed output. All five claims hold: exact node/edge count (33/357), exact steady-state count (10), the 4-dominant/6-rare structure, and the deterministic frustration split (0.14 vs 0.37). The only deviations are stochastic basin-size differences at the 4th significant digit (e.g. 25.63% vs 25.5%), which are expected Monte-Carlo noise and on the technical/our-run side, not an authors' defect. The optional RACIPE and GSE73160 expression-clustering tails were out of scope but do not bear on the reproduced central claim.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at [email protected].
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.