Corpus 1,286 assessed · 1,187 scored · 648 reproduced ≥75 · 174 flagged ·∅ 73.9/100
← New search

Topological signatures in regulatory network enable phenotypic heterogeneity in small cell lung cancer.

Elife · 2021
L1 94/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Every question reproduced
-1 pts
From: “every question reproduced”
Total score -7
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
  • No relevant deviation in data/preprocessing
  • No authors-side cause for any deviation
  • Reported values are derivable from the shared data
  • Any deviation was negligible
  • The central claim held under reproduction
  • Overall, the reproduction was clean
What did not (or only partly)
  • Every checked point held up.
How its reproducibility compares
94/100
Reproducibility score
1.1 SD above mean
vs. all fields · 1187 studies
🎯 Scores higher than 87% of all assessed papers rank 133 of 1187 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough -> faithful 1:1. Reproduced the paper's headline Boolean result (Fig 1B) by running the authors' own Fast-Bool tool (github.com/uday2607/CSB-SCLC @ d1a885f) with their shipped bool.in (2^20 initial conditions, 5000 steps, async Ising) on the shipped 33-node WT-SCLC network, freshly on «our HPC» and independent of the committed output. Got exactly 10 stable steady states: 4 dominant (24.32-25.63% basins, frustration ~0.14) and 6 rare (<0.05%, frustration ~0.37-0.39), matching paper Fig 1B (10 states; 4 @ 24.3-25.5%; 6 @ <0.1%; frustration 0.14 vs 0.37) and the repo's committed Summary_Async.xlsx. Stochastic basin sizes agree to ~3 sig figs (re-seeded run); steady-state count and per-state frustration are exact. NOT attempted (optional hard 20%): RACIPE continuous ODE ensemble (external RACIPE-1.0 C tool + multi-GB Google-Drive data) and the CCLE/GSE73160 expression-clustering validation. No completeness claim; no fabrication signal.

💻 Code ↗ 🗄 Data: GSE73160

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 94
    assessed: 2026-06-15 ⛓ 79a0095a5e34
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-15
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-09-19

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The paper tests whether the topology of the SCLC regulatory network (33 nodes, 357 edges) is sufficient to generate multistability that explains the co-existence of four experimentally observed SCLC phenotypes, independent of genetic mutation.

Core claims
  • The SCLC regulatory network is multistable and its steady states map onto four experimentally observed phenotypes (ASCL1high/NEUROD1low, ASCL1low/NEUROD1high, ASCL1high/NEUROD1high, ASCL1low/NEUROD1low) finding
  • Multistability arises from two 'teams' of nodes that mutually activate members within their own team but inhibit the opposing team, forming an effective toggle switch mechanism
  • Discrete Boolean (Ising) and continuous parameter-agnostic RACIPE simulations yield highly consistent steady-state distributions for the SCLC network finding
  • The specific steady-state distribution of the network is a property of its particular topology, since randomized (edge-swapped) networks produce vastly more states with no overlap to the wild-type distribution finding
  • The 'two-team' topological signature (quantified by metric J and the influence matrix) is specific to the wild-type SCLC network topology and disrupted by network perturbation mechanism
  • NEUROD1 incoming edges are disproportionately important for maintaining robustness of the network's steady-state distribution finding
  • The ASCL1low/NEUROD1low subtype can be further sub-classified into ASCL1low/NEUROD1low/YAP1low/POU2F3high and ASCL1low/NEUROD1low/YAP1high/POU2F3low classes finding
  • A curated 33-node, 357-edge SCLC master regulatory network built from gene expression signatures resource
Experimental setups
Assay System Perturbation Readout Platform
Boolean/Ising model dynamical simulation SCLC regulatory network (33 nodes, 357 edges) none (2^20 random initial conditions) steady-state frequency distribution, frustration asynchronous Ising model formalism
RACIPE ODE ensemble simulation SCLC regulatory network none (parameter sampling, 10^6 models x 1000 initial conditions) discretized steady-state distribution RACIPE (Random Circuit Perturbation)
Network randomization via edge swapping randomized SCLC network topologies (n=1000) edge swap (topology randomization) number of steady states vs wild-type Ising model formalism
Single-edge mutant network simulation SCLC network with one edge deleted or sign-reversed (n=714) edge deletion or inhibitory/activating sign reversal Jensen-Shannon divergence (JSD) vs wild-type steady-state distribution Ising model formalism
Pairwise node correlation analysis of gene expression CCLE cell lines (n=52) none Pearson correlation matrix, J metric CCLE dataset
Pairwise node correlation analysis of gene expression GSE73160 dataset (n=63) none Pearson correlation matrix, J metric GEO dataset GSE73160
Influence matrix (multi-path topological influence) computation SCLC regulatory network topology none net activating/inhibiting influence between node pairs (Inf_ij) computational path-length weighted-sum analysis (l_max=10)
Key results
  • 10 unique steady states obtained; four dominant states (X1-X4) each with 24.3-25.5% frequency, remaining six states <0.1% each 24.3-25.5%
  • High-frequency states show low frustration while low-frequency states show high frustration 0.14 vs 0.37
  • Randomized networks (n=1000) produced far more steady states than wild-type, with no overlap in steady-state distributions 10^4-10^6 vs 10 states
  • RACIPE top-20 states closely matched Boolean X1-X4 states (max 2/33 differing nodes); combined frequency of X1-X4-like states was 22% (top 4) and 54% (top 20) 22%/54%
  • Most single-edge perturbations had negligible effect on network dynamics; of 36 perturbations causing high JSD, 24 were incoming edges to NEUROD1 678/714 JSD<0.01; 24/36
  • J metric for wild-type Boolean network (496) and RACIPE network (373.05) far exceeded mean J for random networks (11 and 14.16 respectively) 496 vs 11; 373.05 vs 14.16
  • J metric for experimental datasets CCLE and GSE73160 was significantly larger than for random networks p<0.01
  • 32 of 33 network nodes form two anti-correlated teams (Group A: ASCL1, INSM1, FOXA1, FOXA2; Group B: REST, SMAD3, ZEB1)
Key statistics
  • other J=496 (J metric for WT Boolean network correlation matrix)
  • other J=373.05 (J metric for WT RACIPE correlation matrix)
  • mean mean J=11 (average J metric across random (Boolean) networks)
  • mean mean J=14.16 (average J metric across random networks (continuous correlation))
  • pvalue p<0.01 (two-tailed z-test) (WT vs random network J metric significance)
  • correlation r=0.851 ± 0.003 (Pearson) (JSD correlation between edge-deletion and sign-reversal perturbations)
  • count n=52 (CCLE cell line samples used for correlation analysis)
  • count n=63 (GSE73160 samples used for correlation analysis)

Statistical methods review

Model: opus

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This is a computational/systems-biology study that characterizes a 33-node, 357-edge small cell lung cancer regulatory network using two simulation frameworks: a parameter-independent asynchronous Boolean (Ising) model and RACIPE, an ODE-based parameter-agnostic ensemble approach. Outputs are steady-state frequency distributions, pairwise Pearson node correlations, an information-theoretic Jensen-Shannon divergence between perturbed and wild-type distributions, and a custom 'J' team-strength metric; comparisons against 1000 edge-swapped random networks and two experimental expression datasets establish the topological signature. Quantitative significance of the J metric is reported with a two-tailed z-test, and simulation summaries are reported as mean ± standard deviation over three replicates.

Replicationtechnical Sample sizeBoolean: 2^20 (and 2^25) initial conditions; RACIPE: 10^6 models each with 1000 initial conditions; 1000 randomized networks; 714 single-edge mutant networks; experimental datasets CCLE n=52 and GSE73160 n=63; simulation summaries over three replicates GroupsWT network vs random/perturbed networks; Boolean vs RACIPE vs experimental data Pairingna Randomization/blindingna DispersionSD Exact p-valuesno Effect sizesno Confidence intervalsno Multiplicity correctionnone stated
Statistical tests used
Test Applied to n Assumptions
two-tailed z-test comparison of J metric for WT Boolean and RACIPE simulations (and CCLE/GSE73160 datasets) versus the distribution of J across 1000 random networks (Figure 2B,ii) 1000 random networks (reference distribution) not stated
Pearson's correlation coefficient pairwise node-node correlations across steady states (Figure 2A,B,C) and regression of correlation vs influence matrix (R1, R2) not stated
Jensen-Shannon divergence (information-theoretic distance, not a hypothesis test) similarity between WT and 714 single-edge perturbed network steady-state distributions, and 1000 random networks (Figure 1D,ii) 2^20 initial conditions per network na
Approaches that could also have been used
  • Significance of the J metric was assessed with a two-tailed z-test comparing the wild-type value against the distribution of J from 1000 random networks.
    Could also: An empirical (permutation/bootstrap) p-value derived directly from the rank of the WT value within the 1000-network null distribution could also be reported. — An empirical null does not rely on an assumption of normality for the random-network J distribution and naturally accommodates skew or bounded ranges, which one might prefer when the null is generated by simulation.
  • p-values were reported as thresholds (p<0.01).
    Could also: Exact p-values alongside the observed effect size (e.g., standardized distance of WT J from the null mean) could also be presented. — Exact values and effect magnitudes convey how far the observation lies from the null beyond a pass/fail cutoff, aiding cross-study comparison.
  • Variability across simulation runs was summarized as mean ± standard deviation over three replicates.
    Could also: A 95% confidence interval or the full range/individual replicate values could also be shown. — With only three replicates, displaying individual points or a CI can communicate the precision of the estimate transparently and is often favored for very small n.
  • Distribution similarity was quantified with Jensen-Shannon divergence.
    Could also: Complementary metrics such as Kullback-Leibler divergence, total variation distance, or earth-mover's distance could also be used. — Reporting an additional metric can show that conclusions about distributional overlap are robust to the choice of divergence measure.
  • Node relationships were assessed with Pearson's correlation coefficient.
    Could also: Spearman's rank correlation could also be computed, particularly for the continuous (non-discretized) RACIPE outputs. — A rank-based correlation captures monotonic but nonlinear associations and is less sensitive to outliers, offering a useful cross-check on linear correlations.
  • Statistical comparisons treated the 1000 random networks as the reference family without a stated multiplicity adjustment across the several z-test comparisons.
    Could also: A family-wise (e.g., Bonferroni) or FDR (Benjamini-Hochberg) adjustment across the related comparisons could also be applied. — When several significance statements are made against the same null, a correction documents control of the overall error rate across the family of tests.
Software: RACIPE (Random Circuit Perturbation) · Asynchronous Boolean / Ising model formalism

Result convergence & founder nodes

Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
64
Impact: high
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

Assessed papers, coloured by verdict. Click a node to open it.

Built on (assessed references) (0)
  • No assessed neighbours yet — the network grows as more papers are assessed.
Cited by (assessed papers) (3)

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

GSE73160 GEO in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-33729159

Paper: Chauhan, Ram, Hari, Jolly. Topological signatures in regulatory network enable phenotypic heterogeneity in small cell lung cancer. eLife 2021;10:e64522. Code: https://github.com/uday2607/CSB-SCLC @ d1a885fc91fff11b9e8660c9767f69b55e8954de Type: theory / computational-systems-biology paper (no wet lab to reproduce).

The repo wraps three computational pipelines:

# Pipeline Tool In scope? Reason
1 Discrete (Boolean/Ising) modelling of the WT-SCLC network — asynchronous update → stable steady states, their frequency & frustration Additional_Codes/Fast-Bool (pure Python + numba, self-contained, network .topo/.ids shipped in repo) YES — primary target Fully specified: shipped bool.in (2^20 initial conditions, 5000 steps, Ising async). Produces the paper's headline Fig 1B numbers. Low-hanging, deterministic frustration, stochastic basin sampling.
2 Continuous (RACIPE) modelling — ensemble of ODE models of the WT-SCLC network external RACIPE-1.0 C package (simonhb1990/RACIPE-1.0); simulation outputs hosted off-repo on Google Drive (authors say files are "quite huge") NO — the optional hard 20% Requires a separate external tool + multi-GB Drive downloads; stochastic ODE ensemble. Skipped per 80/20; would mainly re-confirm the same phenotype structure already covered by pipeline 1.
3 Experimental-data analysis — clustering/UMAP/correlation of CCLE & GSE73160 SCLC cell-line expression to match model phenotypes Additional_Codes/BioData-Analysis (data mGSE73160.txt, CCLE.txt shipped) NO — secondary Multi-step clustering/UMAP with several free parameters; validation layer rather than a single pinnable headline number. Out of 80/20 scope.

What we reproduce (pipeline 1)

Run the authors' own Fast-Bool tool with their shipped bool.in on the shipped 33-node SCLC network, freshly on «our HPC» (independent of the committed output), and compare the steady-state structure to paper Figure 1B (and the shipped OUTPUT/sclcnetwork/Summary_Async.xlsx as a within-repo cross-check):

  • number of unique stable steady states (paper: 10),
  • count of dominant states at ~24.3–25.5% (paper: 4, "X1–X4"),
  • count of rare states at <0.1% (paper: 6, "X5–X10"),
  • frustration of dominant vs rare states (paper: 0.14 vs 0.37),
  • network size 33 nodes / 357 edges (sanity check on the shipped .topo).

Not attempted (honest)

RACIPE continuous ensemble (pipeline 2) and the CCLE/GSE73160 expression-clustering validation (pipeline 3). These are the optional hard tail; skipping them is recorded, not hidden. This is not a completeness claim.

Figures / tables: Fig 1B
C1
Reported
33 nodes, 357 edges
Reproduced
33 nodes, 357 edges
exact
C2
Reported
10 unique steady states (Fig 1B)
Reproduced
10
exact
C3
Reported
4 dominant states @ 24.3-25.5% (Fig 1B)
Reproduced
4 @ 24.32-25.63%
within tolerance
C4
Reported
6 rare states @ <0.1% (Fig 1B)
Reproduced
6 @ <0.05%
exact
C5
Reported
frustration 0.14 (dominant) vs 0.37 (rare), Suppl file 1a
Reproduced
0.143-0.146 vs 0.375-0.387
within tolerance

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 94/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟢3. Location of the main deviation
🟢4. Cause of the deviation
🟢5. Derivability / plausibility
🟢6. Severity of the deviation
🟢7. Core claim
🟢8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Every question reproduced
-1 pts
From: “every question reproduced”
Total score -7

This is a clean 1:1 reproduction of the paper's headline Fig-1B Boolean result, run freshly on «our HPC» with the authors' own shipped code and 33-node network, independent of the committed output. All five claims hold: exact node/edge count (33/357), exact steady-state count (10), the 4-dominant/6-rare structure, and the deterministic frustration split (0.14 vs 0.37). The only deviations are stochastic basin-size differences at the 4th significant digit (e.g. 25.63% vs 25.5%), which are expected Monte-Carlo noise and on the technical/our-run side, not an authors' defect. The optional RACIPE and GSE73160 expression-clustering tails were out of scope but do not bear on the reproduced central claim.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at [email protected].

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

95.4 k
tokens (I/O) · 8.5 M incl. cache
16 min
runtime · 1.08 CPU-h
2.1 GB
peak RAM
2 (1 failed)
HPC jobs
hummel
machine