Numb prevents a complete epithelial-mesenchymal transition by modulating Notch signalling.
The main result did not reproduce in this reproduction attempt. Where our recomputation produced values that differ from the published ones, those discrepancies are listed below. This is a single automated attempt — not peer review and not a finding of error or misconduct — and differences can also arise from data access, undocumented parameters or the computing environment. The verdict can be contested via “report an error”.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Same input data as the authors
- 🟡Reported values were only indirectly comparable
- 🔴A deviation arose in the data or preprocessing
- 🔴A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🔴The deviation was non-trivial in magnitude
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Described well enough to reproduce. Theory/modeling paper; key results are figure curves with no in-text scalars, so graded on direction/structure. (1) Model core (Fig 1 bifurcation) reproduced 1:1 by re-running the authors' OWN PyDSTool code on «our HPC»: Numb shifts the EMT bifurcation to higher external Jagged (6420->9982) and Delta (767->1192), i.e. a stronger stimulus is needed to complete EMT -- exactly the paper's central claim. (2) Survival validation (Fig 6a, GSE30219) reproduced for NUMBL (HR 1.39, p=0.020, predicted direction) but NOT for NUMB -- for the title gene the OS association is null-to-OPPOSITE (HR 0.89, n.s.; two probes significantly protective). Flagged as possible overstatement: the poor-survival signal in GSE30219 is carried by NUMBL, not NUMB; a human should confirm which gene Fig 6a actually plots (HR is inside the figure image). NOT attempted: wet-lab Fig 2; Fig 3/4 tissue percentages (no in-text scalar to compare); Fig 6 cohorts other than GSE30219.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 49assessed: 2026-06-14 ⛓ 1af9456d91ee
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-14
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15no human curator yet
- Last updated
- 2026-09-19
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetSince Notch–Jagged signalling stabilizes a hybrid epithelial/mesenchymal (E/M) phenotype, proteins that modulate Notch signalling—specifically Numb and Numb-like (Numbl), which inhibit Notch—may also modulate the stability of the hybrid E/M phenotype and thereby influence epithelial-mesenchymal transition (EMT) progression and CTC cluster formation.
- ★ Numb/Numbl acts as a 'phenotypic stability factor' (PSF) that inhibits a complete EMT by stabilizing the hybrid E/M phenotype mechanism
- ★ Mathematical modelling of the Notch-EMT-Numb axis predicts Numb enlarges the parameter range over which epithelial and hybrid E/M states are stable, delaying transition to a fully mesenchymal state finding
- ★ Knockdown of Numb or Numbl in stable hybrid E/M H1975 lung cancer cells drives them to a complete, fully mesenchymal EMT phenotype finding
- ★ At the tissue level, Numb alters the balance of hybrid E/M versus mesenchymal cells within simulated clusters of cells communicating via Notch, potentially affecting tumour-initiation ability finding
- ★ Higher Numb expression correlates with worse overall/relapse-free survival in multiple independent lung and ovarian cancer datasets finding
- Numb and Numbl form a mutually inhibitory feedback loop with Notch signalling (Numb inhibits Notch; NICD inhibits Numb) mechanism
- An ODE-based mathematical model coupling the EMT circuit (miR-34, miR-200, Snail, Zeb), Notch pathway components (Notch, Delta, Jagged, NICD) and Numb, implemented in PyDsTool method
- Source code for the model is freely available on GitHub resource
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| ODE-based mathematical modelling / bifurcation analysis | single-cell Notch-EMT-Numb circuit (computational) | varying external Jagged/Delta concentration, with/without Numb interactions | steady-state levels of miR-200 and other circuit species; phenotype (E, E/M, M) stability ranges | PyDsTool |
| Multi-cell lattice simulation | 2D lattice of 50x50 simulated cancer cells communicating via Notch | varying Jagged/Delta production rates, with/without Numb | fraction and spatial pattern of epithelial, hybrid E/M and mesenchymal cells | PyDsTool / Matplotlib |
| siRNA knockdown + bright-field microscopy | H1975 NSCLC cells (hybrid E/M) | siRNA against Numb or Numbl | cell morphology (spindle-shaped vs compact) | — |
| Immunofluorescence | H1975 NSCLC cells | siRNA against Numb or Numbl | CDH1 (E-cadherin) and VIM (vimentin) protein expression/localization | — |
| Transwell migration assay | H1975 NSCLC cells | siRNA against Numb or Numbl | collective vs individual cell migration | — |
| Proliferation assay | H1975 NSCLC cells | siRNA against Numb or Numbl | cell proliferation rate | — |
| RT-PCR | H1975 NSCLC cells | siRNA against Numb or Numbl | mRNA levels of CDH1, VIM, ZEB1, JAG1 | — |
| Western blot | H1975 NSCLC cells | siRNA against Numb or Numbl | protein levels of CDH1, VIM, ZEB1, JAG1 (quantified via ImageJ) | ImageJ |
- ▲ Numb widens the range of external Jagged/Delta concentrations over which epithelial and hybrid E/M states remain stable, requiring stronger ligand stimulus to reach a fully mesenchymal state
- – Numb/Numbl knockdown H1975 cells become spindle-shaped, lose E-cadherin staining, and stain only for vimentin, unlike control cells co-expressing both markers
- – Control H1975 cells show collective cell migration; Numb/Numbl-knockdown cells show individual cell migration
- ▼ Numb/Numbl knockdown inhibits H1975 cell proliferation
- – Numb/Numbl knockdown increases VIM, ZEB1 and JAG1 mRNA and protein levels while decreasing CDH1 mRNA and protein levels
- ▼ At low Jagged production (gJ=45 molecules h-1), Numb decreases the fraction of hybrid E/M and mesenchymal cells and increases the epithelial fraction in tissue simulations
- – At high Jagged production (gJ=80 molecules h-1), Numb reduces the fraction of fully mesenchymal cells while increasing the hybrid E/M fraction
- ▼ Higher Numb expression correlates with worse survival in multiple independent lung and ovarian cancer patient datasets
- pvalue p<0.05, p<0.005, p<0.001 (significance thresholds for RT-PCR/Western blot marker changes upon Numb/Numbl knockdown, two-tailed paired t-test)
- count N=5 (technical replicates per condition in proliferation assay)
- other gJ = 45 molecules h-1 (low Jagged production rate regime in tissue-level lattice simulation)
- other gJ = 80 molecules h-1 (high Jagged production rate regime in tissue-level lattice simulation)
- other gD = 20 molecules h-1 (fixed Delta production rate across tissue-level simulations)
- count 10 simulations (averaged replicate simulations for computing E/E-M/M cell fractions, each with different random initial conditions)
- other Next = 10000 molecules (fixed external Notch concentration used in all single-cell bifurcation simulations)
- count 50x50 cells (size of the two-dimensional cell lattice used for tissue-level Notch signalling simulations)
Statistical methods review
Model: opusA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This is an integrated computational–experimental study. A deterministic ODE model of the Notch–EMT–Numb circuit was analysed at single-cell (bifurcation/sensitivity analysis) and multi-cell (50×50 lattice) levels, solved numerically in Python. Experimental validation used siRNA knockdown of Numb/Numbl in H1975 cells assessed by imaging, RT-PCR and western blot, with group comparisons by two-tailed paired t-test and results shown as bar charts with s.e.m. Clinical association was assessed by median-split survival analysis across several lung and ovarian cancer datasets.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| two-tailed paired Student's t-test | Figure 2 RT-PCR and western blot quantifications (CDH1, VIM, ZEB1, JAG1) and proliferation comparisons after Numb/Numbl knockdown | N = 5 for each technical replicate (proliferation, fig. 2d) | not stated |
| survival comparison of below-median vs above-median Numb groups (test not explicitly named; performed in ProgGeneV2) | overall survival and relapse-free survival in multiple lung and ovarian cancer datasets | — | not stated |
-
Dispersion was summarized with the standard error of the mean (s.e.m.).↳ Could also: Reporting the standard deviation or a 95% confidence interval alongside the mean, and overlaying individual data points. — SD or a CI conveys the spread of the underlying data (rather than the precision of the mean estimate) and is often preferred for small n, helping readers gauge variability directly.
-
Group comparisons for several markers were made with two-tailed paired t-tests.↳ Could also: A single ANOVA (or mixed-effects model) with a post-hoc multiple-comparison correction such as Tukey HSD, or a nonparametric Mann–Whitney/Wilcoxon test. — An omnibus model with post-hoc correction controls the family-wise error rate across many simultaneous marker comparisons, and a nonparametric option avoids reliance on normality assumptions when n is small.
-
Significance was reported as thresholded categories (*, **, ***).↳ Could also: Reporting exact p-values together with effect-size estimates (e.g. fold-change with CI). — Exact p-values and effect sizes give readers a continuous sense of evidence strength and magnitude rather than discretized cut-offs.
-
Patients were split into two groups at the median Numb expression and survival curves were compared.↳ Could also: Treating expression as a continuous covariate in a Cox proportional-hazards model, or reporting the log-rank statistic with hazard ratios and CIs. — A continuous model uses all the information in the expression values and avoids dependence on a single dichotomization threshold, while hazard ratios with CIs quantify the magnitude of the association.
-
Multiple markers and multiple independent datasets were analysed without a stated multiplicity correction.↳ Could also: Applying a correction such as Benjamini–Hochberg FDR or Bonferroni across the family of tests. — A correction controls the chance of false positives when many comparisons are performed across markers and cohorts.
-
Model robustness was probed via a local sensitivity analysis around chosen parameters.↳ Could also: A global sensitivity analysis (e.g. Latin-hypercube/Sobol sampling) or Bayesian parameter exploration. — Global methods sample the full parameter space simultaneously and can capture interactions and robustness beyond small variations around a single point.
Result convergence & founder nodes
Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.
-
Numb/Numbl knockdown drives hybrid E/M H1975 NSCLC cells to a fully mesenchymal morphology with loss of E-cadherin and gain of vimentin.imaging h1975 2017×1papers★ This paper is the founder (earliest)
-
Numb/Numbl knockdown switches H1975 cells from collective to individual cell migration in transwell assays.other h1975 2017×1papers★ This paper is the founder (earliest)
-
Numb/Numbl knockdown inhibits proliferation of H1975 NSCLC cells.other h1975 down 2017×1papers★ This paper is the founder (earliest)
-
Higher NUMB/NUMBL expression correlates with worse overall and relapse-free survival across multiple lung and ovarian cancer patient datasets.other human-cancer up 2017×1papers★ This paper is the founder (earliest)
-
At low Jagged production in a tissue-level lattice simulation, Numb decreases hybrid and mesenchymal cell fractions and increases the epithelial fraction.other in-silico down 2017×1papers★ This paper is the founder (earliest)
-
In a single-cell ODE bifurcation model, Numb enlarges the parameter range stabilizing epithelial and hybrid E/M states and inhibits transition to complete EMT.other in-silico 2017×1papers★ This paper is the founder (earliest)
-
At high Jagged production in a tissue-level lattice simulation, Numb reduces the fully mesenchymal cell fraction and increases the hybrid E/M fraction.other in-silico down 2017×1papers★ This paper is the founder (earliest)
-
Numb/Numbl knockdown increases VIM, ZEB1, and JAG1 and decreases CDH1 mRNA and protein levels in H1975 cells.western-blot h1975 mixed 2017×1papers★ This paper is the founder (earliest)
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-29187638
Numb prevents a complete EMT by modulating Notch signalling. Bocci et al., J R Soc Interface 2017. Code: https://github.com/federicobocci91/Numb_project (authors' own, Python2 + PyDSTool). Data: GEO GSE30219 (n=282 lung adenocarcinoma, Affymetrix GPL570).
Nature of the paper
Theory/modeling paper. Core results = an ODE model of the Numb–Notch–EMT circuit: single-cell bifurcation diagrams (Fig 1) + multicell tissue simulations (Fig 3/4), plus a clinical-validation arm: survival stratified by NUMB/NUMBL expression in 6 public cohorts (Fig 6). Most reported numbers live INSIDE figure panels, not as in-text values — so precise scalar claims are sparse.
In scope (pipeline-derived, attempted)
- A. GSE30219 survival (Fig 6a). Public expression+clinical data named in the RU. Split n=282 patients by median NUMB expression; high NUMB → shorter overall survival. Reproduce direction + significance (log-rank p, Cox HR). This is the one claim that binds the named DATA to an expected result → primary data point.
- B. single_cell.py bifurcation (Fig 1c–f). Authors' own ODE code. Reproduce the qualitative claim: with Numb the hybrid-E/M (sender/receiver) state is stable over a WIDER range of external Jagged/Delta than without Numb. Run the shipped script on «our HPC». Qualitative (curve-shape) match; the paper gives no scalar.
Out of scope (not attempted, why)
- Fig 2 knockdown / morphology / E-cadherin / vimentin / proliferation = wet-lab.
- Fig 3/4 tissue-level percentages = in repo (multi_cell.py) but the paper reports NO in-text numbers to compare against (values only in panels) → no pinnable expected result; secondary, attempt only if time, not chased (80/20).
- Fig 6 cohorts other than GSE30219 (TCGA-LUAD, GSE41271, GSE73614, GSE9891) = same analysis on other accessions; out of this RU's named dataset → not attempted.
Pipelines named
- A: GEOquery (download) + survival (KM/log-rank/Cox) in R on GSE30219.
- B: authors' single_cell.py (PyDSTool AUTO continuation), Python.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
Modeling core reproduces 1:1 by re-running the authors' own PyDSTool code: Numb shifts the EMT bifurcation to higher external Jagged (6420->9982) and Delta (767->1192), exactly the paper's central mechanistic claim. The clinical validation (Fig 6a, GSE30219) is a split decision: NUMBL reproduces the predicted poor-survival direction (HR 1.39, p=0.020), but NUMB — the title gene — is null-to-opposite (HR 0.89 n.s.; all 4 probes HR<=1.0, two significantly protective). This is a robust direction/significance flip flagged as possible overstatement, tempered by genuine ambiguity that Fig 6a's HR is inside the figure image and may actually plot NUMBL. No fabrication, but a clear human-audit target: confirm which gene Fig 6a plots — if NUMB, this escalates to a red critical discrepancy.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at [email protected].
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.