Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Numb prevents a complete epithelial-mesenchymal transition by modulating Notch signalling.

J R Soc Interface · 2017
L1 49/100 3/4
Why this verdict

The main result did not reproduce in this reproduction attempt. Where our recomputation produced values that differ from the published ones, those discrepancies are listed below. This is a single automated attempt — not peer review and not a finding of error or misconduct — and differences can also arise from data access, undocumented parameters or the computing environment. The verdict can be contested via “report an error”.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +5
✓ What held up
  • Same input data as the authors
What did not (or only partly)
  • 🟡Reported values were only indirectly comparable
  • 🔴A deviation arose in the data or preprocessing
  • 🔴A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🔴The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
49/100
Reproducibility score
1.4 SD below mean
vs. all fields · 1173 studies
🎯 Scores higher than 7% of all assessed papers rank 1081 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough to reproduce. Theory/modeling paper; key results are figure curves with no in-text scalars, so graded on direction/structure. (1) Model core (Fig 1 bifurcation) reproduced 1:1 by re-running the authors' OWN PyDSTool code on «our HPC»: Numb shifts the EMT bifurcation to higher external Jagged (6420->9982) and Delta (767->1192), i.e. a stronger stimulus is needed to complete EMT -- exactly the paper's central claim. (2) Survival validation (Fig 6a, GSE30219) reproduced for NUMBL (HR 1.39, p=0.020, predicted direction) but NOT for NUMB -- for the title gene the OS association is null-to-OPPOSITE (HR 0.89, n.s.; two probes significantly protective). Flagged as possible overstatement: the poor-survival signal in GSE30219 is carried by NUMBL, not NUMB; a human should confirm which gene Fig 6a actually plots (HR is inside the figure image). NOT attempted: wet-lab Fig 2; Fig 3/4 tissue percentages (no in-text scalar to compare); Fig 6 cohorts other than GSE30219.

💻 Code ↗ 🗄 Data: GSE30219

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 49
    assessed: 2026-06-14 ⛓ 1af9456d91ee
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-14
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

The authors hypothesized that proteins affecting Notch signalling—specifically Numb and its homologue Numb-like, which inhibit Notch—modulate the stability of a hybrid epithelial/mesenchymal (E/M) phenotype, testing whether Numb prevents cells from undergoing a complete EMT.

Core claims
  • Numb (and Numbl) inhibits a full EMT by stabilizing a hybrid E/M phenotype, acting as a 'phenotypic stability factor'. finding
  • A mathematical model of the Notch-EMT-Numb axis predicts that Numb restricts progression to a complete EMT and enlarges the stability range of epithelial and hybrid E/M states. method
  • Knockdown of Numb or Numbl in stable hybrid E/M H1975 lung cancer cells drives them toward a fully mesenchymal phenotype, validating the model prediction. finding
  • Numb and Notch form a mutually inhibitory feedback loop; Numb inhibits Notch while NICD inhibits Numb, mediating Notch-driven EMT. mechanism
  • At a tissue/multi-cell level, Numb alters the balance of hybrid E/M versus mesenchymal cells in clusters, potentially increasing tumour-initiation ability and CTC cluster formation. finding
  • Higher Numb levels correlate with worse survival in multiple independent lung and ovarian cancer datasets, indicating association with cancer aggressiveness. finding
  • An open-source computational model (ODE-based, PyDsTool) of the Notch-EMT-Numb signalling axis is provided as a resource. resource
Experimental setups
Assay System Perturbation Readout Platform
Mathematical/ODE model (single-cell bifurcation analysis) In silico Notch-EMT-Numb regulatory circuit presence/absence of Numb; varying external Jagged (Jext) and Delta (Dext) steady-state level of miR-200 and EMT phenotype (E, E/M, M) PyDsTool python library; Matplotlib
Multi-cell tissue-level simulation In silico 50×50 two-dimensional lattice of cancer cells communicating via Notch signalling presence/absence of Numb; varying Jagged (gJ) and Delta (gD) production rates fraction and spatial patterning of E, E/M and M cells PyDsTool
siRNA knockdown with bright-field/immunofluorescence microscopy H1975 NSCLC lung cancer cells (stable hybrid E/M) Numb or Numbl knockdown (siRNA) cell morphology; E-cadherin (CDH1), Vimentin (VIM), DAPI staining
Transwell migration assay H1975 cells Numb or Numbl knockdown (siRNA) collective vs individual cell migration
Cell proliferation assay H1975 cells Numb or Numbl knockdown (siRNA) proliferation rate
RT-PCR H1975 cells Numb or Numbl knockdown (siRNA) mRNA levels of CDH1, VIM, ZEB1, JAG1
Western blot H1975 cells Numb or Numbl knockdown (siRNA) protein levels of CDH1, VIM, ZEB1, JAG1 (quantified in ImageJ) ImageJ quantification
Survival/clinical data analysis Multiple independent lung and ovarian cancer patient datasets none (patients stratified by Numb expression below/above median) overall survival and relapse-free survival ProgGeneV2
Key results
  • Knockdown of Numb or Numbl drove hybrid E/M H1975 cells to a spindle-shaped, vimentin-positive/E-cadherin-negative mesenchymal phenotype (full EMT).
  • Numb/Numbl knockdown increased mRNA and protein levels of Vimentin, ZEB1 and JAG1, while decreasing E-cadherin levels.
  • Numb/Numbl knockdown switched cells from collective to individual cell migration in transwell assays.
  • Numb/Numbl knockdown inhibited H1975 cell proliferation.
  • In the single-cell model, Numb enlarged the range of external ligand concentrations stabilizing epithelial and hybrid E/M states, inhibiting transition to complete EMT.
  • At low Jagged production (gJ=45 molecules/h), Numb decreased the fraction of both hybrid and mesenchymal cells and increased epithelial cells.
  • At high Jagged production (gJ=80 molecules/h) all cells underwent partial/complete EMT, but Numb reduced the fraction of fully mesenchymal cells and increased hybrid E/M cells.
  • Higher Numb/Numbl expression correlated with worse survival across multiple lung and ovarian cancer datasets.
Key statistics
  • count 50 × 50 cells (two-dimensional lattice size for tissue-level simulation)
  • count gJ = 45 molecules h−1 (low Jagged production rate, E-E/M crossing regime)
  • count gJ = 80 molecules h−1 (high Jagged production rate, full EMT regime)
  • count gD = 20 molecules h−1 (fixed Delta production rate in tissue simulations)
  • count Next = 10 000 molecules (fixed external Notch concentration in bifurcation simulations)
  • count transient of 120 h (5 days) (simulation time window before disruption of patterning)
  • count N = 5 (technical replicates for proliferation assay)
  • count 10 simulations (averaging cell-fraction over randomized initial conditions)

Statistical methods review

Model: opus

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This is an integrated computational–experimental study. A deterministic ODE model of the Notch–EMT–Numb circuit was analysed at single-cell (bifurcation/sensitivity analysis) and multi-cell (50×50 lattice) levels, solved numerically in Python. Experimental validation used siRNA knockdown of Numb/Numbl in H1975 cells assessed by imaging, RT-PCR and western blot, with group comparisons by two-tailed paired t-test and results shown as bar charts with s.e.m. Clinical association was assessed by median-split survival analysis across several lung and ovarian cancer datasets.

Replicationtechnical Sample sizeN = 5 stated for each technical replicate (proliferation assay); sample size/power not otherwise described, and survival cohort sizes not stated in text GroupssiRNA knockdown (Numb or Numbl) vs control/mock H1975 cells; below- vs above-median Numb expression in patient datasets Pairingpaired Randomization/blindingnot stated DispersionSEM Exact p-valuesno Effect sizesno Confidence intervalsno Multiplicity correctionnone stated
Statistical tests used
Test Applied to n Assumptions
two-tailed paired Student's t-test Figure 2 RT-PCR and western blot quantifications (CDH1, VIM, ZEB1, JAG1) and proliferation comparisons after Numb/Numbl knockdown N = 5 for each technical replicate (proliferation, fig. 2d) not stated
survival comparison of below-median vs above-median Numb groups (test not explicitly named; performed in ProgGeneV2) overall survival and relapse-free survival in multiple lung and ovarian cancer datasets not stated
Approaches that could also have been used
  • Dispersion was summarized with the standard error of the mean (s.e.m.).
    Could also: Reporting the standard deviation or a 95% confidence interval alongside the mean, and overlaying individual data points. — SD or a CI conveys the spread of the underlying data (rather than the precision of the mean estimate) and is often preferred for small n, helping readers gauge variability directly.
  • Group comparisons for several markers were made with two-tailed paired t-tests.
    Could also: A single ANOVA (or mixed-effects model) with a post-hoc multiple-comparison correction such as Tukey HSD, or a nonparametric Mann–Whitney/Wilcoxon test. — An omnibus model with post-hoc correction controls the family-wise error rate across many simultaneous marker comparisons, and a nonparametric option avoids reliance on normality assumptions when n is small.
  • Significance was reported as thresholded categories (*, **, ***).
    Could also: Reporting exact p-values together with effect-size estimates (e.g. fold-change with CI). — Exact p-values and effect sizes give readers a continuous sense of evidence strength and magnitude rather than discretized cut-offs.
  • Patients were split into two groups at the median Numb expression and survival curves were compared.
    Could also: Treating expression as a continuous covariate in a Cox proportional-hazards model, or reporting the log-rank statistic with hazard ratios and CIs. — A continuous model uses all the information in the expression values and avoids dependence on a single dichotomization threshold, while hazard ratios with CIs quantify the magnitude of the association.
  • Multiple markers and multiple independent datasets were analysed without a stated multiplicity correction.
    Could also: Applying a correction such as Benjamini–Hochberg FDR or Bonferroni across the family of tests. — A correction controls the chance of false positives when many comparisons are performed across markers and cohorts.
  • Model robustness was probed via a local sensitivity analysis around chosen parameters.
    Could also: A global sensitivity analysis (e.g. Latin-hypercube/Sobol sampling) or Bayesian parameter exploration. — Global methods sample the full parameter space simultaneously and can capture interactions and robustness beyond small variations around a single point.
Software: PyDsTool (Python numerical library) · Matplotlib (Python plotting) · ImageJ (blot quantification) · ProgGeneV2 (survival analysis)

Result convergence & founder nodes

Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
114
Impact: high
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-29187638

Numb prevents a complete EMT by modulating Notch signalling. Bocci et al., J R Soc Interface 2017. Code: https://github.com/federicobocci91/Numb_project (authors' own, Python2 + PyDSTool). Data: GEO GSE30219 (n=282 lung adenocarcinoma, Affymetrix GPL570).

Nature of the paper

Theory/modeling paper. Core results = an ODE model of the Numb–Notch–EMT circuit: single-cell bifurcation diagrams (Fig 1) + multicell tissue simulations (Fig 3/4), plus a clinical-validation arm: survival stratified by NUMB/NUMBL expression in 6 public cohorts (Fig 6). Most reported numbers live INSIDE figure panels, not as in-text values — so precise scalar claims are sparse.

In scope (pipeline-derived, attempted)

  • A. GSE30219 survival (Fig 6a). Public expression+clinical data named in the RU. Split n=282 patients by median NUMB expression; high NUMB → shorter overall survival. Reproduce direction + significance (log-rank p, Cox HR). This is the one claim that binds the named DATA to an expected result → primary data point.
  • B. single_cell.py bifurcation (Fig 1c–f). Authors' own ODE code. Reproduce the qualitative claim: with Numb the hybrid-E/M (sender/receiver) state is stable over a WIDER range of external Jagged/Delta than without Numb. Run the shipped script on «our HPC». Qualitative (curve-shape) match; the paper gives no scalar.

Out of scope (not attempted, why)

  • Fig 2 knockdown / morphology / E-cadherin / vimentin / proliferation = wet-lab.
  • Fig 3/4 tissue-level percentages = in repo (multi_cell.py) but the paper reports NO in-text numbers to compare against (values only in panels) → no pinnable expected result; secondary, attempt only if time, not chased (80/20).
  • Fig 6 cohorts other than GSE30219 (TCGA-LUAD, GSE41271, GSE73614, GSE9891) = same analysis on other accessions; out of this RU's named dataset → not attempted.

Pipelines named

  • A: GEOquery (download) + survival (KM/log-rank/Cox) in R on GSE30219.
  • B: authors' single_cell.py (PyDSTool AUTO continuation), Python.
Figures / tables: Fig 1cFig 1dFig 1eFig 1fFig 6a
bif_Jt
Reported
Fig 1c->1d: Numb widens the stable hybrid-E/M window along external Jagged (Jt); stronger stimulus needed to complete EMT (qualitative, no scalar printed)
Reproduced
authors' single_cell.py (commit 9316b35) re-run on «our HPC»: interior high-Jt bifurcation point 6420 -> 9982 molecules/h (control -> Numb)
partial
bif_Dt
Reported
Fig 1e->1f: same effect along external Delta (Dt) (qualitative)
Reproduced
interior high-Dt bifurcation point 767 -> 1192 molecules/h (control -> Numb)
partial
surv_NUMBL_OS
Reported
Fig 6a (GSE30219, n=282): high Numb/Numbl predicts shorter overall survival (HR>1)
Reproduced
NUMBL high vs low median split: Cox HR=1.39, p=0.020 (probe 242195_x_at); both NUMBL probes HR~1.40
within tolerance
surv_NUMB_OS
Reported
Fig 6a / Results: high NUMB (the title gene) predicts shorter overall survival in GSE30219
Reproduced
NUMB high vs low: HR=0.89, p=0.42 (mean of 4 probes); all 4 probes HR<=1.0; 2 probes significantly PROTECTIVE; opposite to reported direction
did not match

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 49/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟡2. Endpoint comparability
🔴3. Location of the main deviation
🔴4. Cause of the deviation
🟡5. Derivability / plausibility
🔴6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +5

Modeling core reproduces 1:1 by re-running the authors' own PyDSTool code: Numb shifts the EMT bifurcation to higher external Jagged (6420->9982) and Delta (767->1192), exactly the paper's central mechanistic claim. The clinical validation (Fig 6a, GSE30219) is a split decision: NUMBL reproduces the predicted poor-survival direction (HR 1.39, p=0.020), but NUMB — the title gene — is null-to-opposite (HR 0.89 n.s.; all 4 probes HR<=1.0, two significantly protective). This is a robust direction/significance flip flagged as possible overstatement, tempered by genuine ambiguity that Fig 6a's HR is inside the figure image and may actually plot NUMBL. No fabrication, but a clear human-audit target: confirm which gene Fig 6a plots — if NUMB, this escalates to a red critical discrepancy.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

225 k
tokens (I/O) · 17.4 M incl. cache
26 min
runtime · 0.08 CPU-h
1.3 GB
peak RAM
2
HPC jobs
hummel
machine