Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Yap1 safeguards mouse embryonic stem cells from excessive apoptosis during differentiation.

Elife · 2018
L1 50/100 3/4
Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +6
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
What did not (or only partly)
  • 🔴A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
50/100
Reproducibility score
1.4 SD below mean
vs. all fields · 1173 studies
🎯 Scores higher than 8% of all assessed papers rank 1026 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough to run, but exact counts NOT 1:1. First, the room's mined data accession (GSE69669) was wrong: it is an RNA-seq dataset reused from an earlier paper (PMID 26917425) and has no peaks to call; the paper's OWN generated Yap1 ChIP-seq is GSE112606 / SRP136948. Reproduced against the correct dataset using the paper's stated pipeline (Bowtie2 -> mm9 -> MACS2 2.2.7.1 callpeak -p 1e-5 -g mm); MACS2 is the room's third-party code pointer (taoliu/MACS), valid per P16. Outcome PARTIAL: reproduced peak counts are the same order of magnitude and reproduce the paper's central directional result (far more Yap1 binding in differentiating than self-renewal: reported 8453 vs 699 ~12x; ours 5580 vs 289 ~19x), but do not match the exact reported numbers. -LIF: pooled = 1064 (~8x low) yet the better single replicate (rep2) = 5580, within ~1.5x of 8453; rep1 nearly failed (57 peaks). +LIF: 289 vs 699 (~0.41x). The gap is attributable to MACS2 settings the paper left unspecified (version, --keep-dup given redundant rates 0.28-0.81, p- vs q-value mode, replicate pooling) plus a low-quality -LIF input control (37.2% mm9 alignment, 0.81 redundancy). No fabrication concern: rep2 alone already yields 5580 of the reported 8453, so the magnitude is derivable from the shipped data. NOT attempted: HOMER motif analysis and any parameter sweep to force-match counts (excluded per 80/20 / do-not-chase-last-20% rule).

💻 Code ↗ 🗄 Data: GSE69669

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 50
    assessed: 2026-06-14 ⛓ 12d4d0995d58
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-14
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

Does the Hippo pathway effector Yap1 regulate cell survival during the exit of mouse embryonic stem cells from self-renewal, and if so, does it protect differentiating ESCs from excessive apoptosis?

Core claims
  • ESCs lacking Yap1 undergo massive caspase-dependent cell death upon exit from self-renewal but not during self-renewal finding
  • Yap1 contextually protects differentiating (but not self-renewing) ESCs from hyperactivation of the apoptotic cascade finding
  • Yap1 directly activates anti-apoptotic genes (Bcl-2, Bcl-xL/Bcl2l1, Mcl-1) and mildly suppresses pro-apoptotic genes to moderate mitochondrial priming during differentiation mechanism
  • Loss of Yap1 leads to Casp9 hyperactivation that sustains elevated apoptosis during differentiation mechanism
  • Modulating expression of single Yap1-targeted apoptosis genes is sufficient to augment or hinder survival during differentiation finding
  • Yap1's pro-survival role is independent of differentiation lineage and acts directly after exit from self-renewal finding
  • Yap1 overexpression reduces cell death during differentiation finding
Experimental setups
Assay System Perturbation Readout Platform
LDH cytotoxicity assay J1, CJ7, E14 mouse ESCs (WT, Yap1 KO, KD, OE) Yap1 KO/KD/OE; LIF withdrawal; zVAD/necrostatin-1/verteporfin/IDE1/Casp9 KD cell death (% relative to fully lysed control)
Immunoblot / Western blot WT and Yap1 KO mouse ESCs during differentiation Yap1 KO; LIF withdrawal; STS treatment protein levels of Yap1, Casp8/9/3, cleaved Casp3, cleaved Parp1, Bcl-2, Bcl-xL, Mcl-1
Flow cytometry WT and Yap1 KO differentiating mouse ESCs (60 hr -LIF) Yap1 KO; LIF withdrawal annexin-V (CF594) and active Casp3 (NucView 488) positivity/intensity
Live fluorescence microscopy WT and Yap1 KO mouse ESCs LIF withdrawal active Casp3 (NucView 488 substrate)
Luminogenic caspase activity assay WT and Yap1 KO ESCs in ±LIF Yap1 KO; LIF withdrawal caspase activity (luminescence)
RT-qPCR WT, Yap1 KO/KD/OE mouse ESCs in various differentiation conditions Yap1 KO/KD/OE; N2B27, IDE1, EpiLC, LIF withdrawal mRNA of caspases, anti-/pro-apoptotic genes, lineage markers
Immunocytochemistry / confocal microscopy WT and Yap1 KO ESCs in -LIF (72 hr) Yap1 KO; LIF withdrawal Bcl-2, Mcl-1 expression and mitochondrial colocalization (MitoTracker) ImageJ; 63X oil objective confocal
ChIP-seq mouse ESCs none (Yap1 binding) Yap1 genomic binding at apoptosis-related cis-regulatory elements
Key results
  • Cell death rises from ~30% in WT to >70% in Yap1 KO ESCs 72 hr after LIF withdrawal >70% vs ~30%
  • Yap1 overexpression reduces cell death during differentiation to ~10% ~10%
  • All caspases tested ~two-fold more active in Yap1 KO than WT by 60 hr after LIF removal ~2-fold
  • Casp9 knockdown reduces Yap1 KO cell death back to WT levels during differentiation
  • Bcl2 upregulated 80–100-fold by 96 hr in WT during differentiation, blunted in Yap1 KO 80-100-fold
  • Bcl-2, Bcl-xL, and Mcl-1 anti-apoptotic proteins deficient in Yap1 KO cells after 72 hr LIF withdrawal
  • Bcl-2 and Mcl-1 strongly colocalize with mitochondria weighted colocalization coefficient ~0.7–0.9
  • Verteporfin before exit from self-renewal phenocopies Yap1 KO, but late treatment has modest effect as low as 1 μM
Key statistics
  • fold_change >70% (KO) vs ~30% (WT) cell death (LDH assay 72 hr after LIF withdrawal)
  • fold_change ~10% cell death (Yap1 OE cell lines during differentiation)
  • fold_change ~2-fold higher caspase activity (Yap1 KO vs WT at 60 hr -LIF)
  • fold_change 80-100-fold Bcl2 upregulation (WT ESCs by 96 hr differentiation relative to +LIF)
  • correlation weighted colocalization coefficient ~0.7–0.9 (Bcl-2/Mcl-1 colocalization with mitochondria)
  • other ~30% annexin V positive (human ESCs exiting self-renewal (cited prior work))
  • count 1 μM verteporfin (dose phenocopying Yap1 KO before exit from self-renewal)

Statistical methods review

Model: opus

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The study uses a controlled experimental design comparing Yap1 knockout (and knockdown/overexpression) mouse ESCs to wild-type cells across self-renewal and several differentiation conditions, with cell-death, caspase-activity, flow-cytometry, immunoblot, and RT-qPCR readouts. Group comparisons were made primarily with two-sample two-tailed t-tests (with paired t-tests used for some boxplot panels), and results were reported as mean ± standard deviation with significance indicated by p-value threshold bins (asterisks). Sample sizes were stated as numbers of independent samples (n = 4 for LDH assays, n = 3 for most other assays).

Replicationbiological Sample sizeStated as number of independent samples (n = 4 for LDH assays, n = 3 for most other assays, n = 8 for one verteporfin control); no power/sample-size calculation described GroupsYap1 KO/KD/OE or inhibitor-treated vs WT/control ESCs across self-renewal and multiple differentiation conditions Pairingmixed Randomization/blindingnot stated DispersionSD Exact p-valuesno Effect sizesno Confidence intervalsno Multiplicity correctionnone stated
Statistical tests used
Test Applied to n Assumptions
two-sample two-tailed Student's t-test LDH cell-death assays, caspase activity, flow cytometry fold enrichment, and RT-qPCR comparisons vs. WT throughout Figures 1, 2, 3 and supplements n = 4 independent samples for LDH assays; n = 3 for other experiments (n = 8 for one verteporfin positive control) not stated
paired t-test boxplots of pro- vs anti-apoptotic gene expression in Figure 3—figure supplement 1D (Yap1 OE vs BirA) and 1E (Yap1 KD vs empty KD) n = 3 independent samples not stated
Approaches that could also have been used
  • Multiple group comparisons (e.g. several apoptosis-related genes and several timepoints) were each evaluated with separate two-sample t-tests.
    Could also: A one-way or two-way ANOVA followed by a post-hoc test (e.g. Tukey HSD or Dunnett's vs WT), or Benjamini-Hochberg FDR / Bonferroni adjustment across the family of comparisons. — A grouped model with post-hoc correction also controls the family-wise or false-discovery rate when many comparisons share an experiment, and can borrow variance information across groups.
  • The t-test was selected for comparisons, which assumes approximate normality.
    Could also: A non-parametric test such as Mann-Whitney U (unpaired) or Wilcoxon signed-rank (paired) could also be used. — Non-parametric tests make fewer distributional assumptions and are often chosen for small n where normality is hard to assess.
  • Dispersion was reported as mean ± standard deviation.
    Could also: A 95% confidence interval could also be reported alongside or instead of SD. — A confidence interval conveys the precision of the estimated effect directly and is often preferred for communicating uncertainty, especially with small n.
  • Significance was reported using binned p-value thresholds (asterisks).
    Could also: Exact p-values together with effect-size estimates (e.g. mean difference or fold-change with CI) could also be reported. — Exact p-values and effect sizes give readers more information about magnitude and let them apply their own significance thresholds.
  • Sample sizes were stated as numbers of independent samples without a power analysis.
    Could also: An a priori power/sample-size justification could also be provided. — A stated power analysis documents the basis for the chosen n and the sensitivity to detect a given effect size.

Result convergence & founder nodes

Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
52
Impact: high
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

CVCL_C316 Cellosaurus in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GO:0006915 Gene Ontology (GO) in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GO:2001243 Gene Ontology (GO) in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GO:2001244 Gene Ontology (GO) in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE112606 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE69669 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSM1706488 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSM1706489 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSM1706495 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSM1706496 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
HPA007415 HPA in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
PD184352 ENA in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
R37117 ENA in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-30561326

Paper: LeBlanc et al. (2018), Yap1 safeguards mouse embryonic stem cells from excessive apoptosis during differentiation. eLife 7:e40167. PMID 30561326 / PMC6307859.

Accession correction (important)

The room metadata listed GSE69669 as the dataset. That is wrong for this paper: GSE69669 ("Roles of Yap1 in mouse embryonic stem cells", PMID 26917425, Chung et al. 2016) is a previously-published RNA-seq dataset merely reused by this paper, and RNA-seq has no peaks to call with MACS. The paper's own generated dataset is GSE112606 (SRA study SRP136948) — Yap1 ChIP-seq in mouse ES cells. We reproduce against GSE112606, the correct artifact.

The room "code" pointer is github.com/taoliu/MACS (MACS/MACS2). Per brief rule P16, applying this third-party peak caller to the paper's own ChIP-seq data per the paper's parameters is a fully valid reproduction.

ChIP-seq samples (GSE112606 / SRP136948), single-end 75 bp

Run GSM role condition
SRR8260152 GSM3495122 Input (BirA) +LIF (self-renewal)
SRR8260153 GSM3495123 Input (BirA) −LIF (differentiating)
SRR8260154 GSM3495124 Yap1 ChIP +LIF (self-renewal)
SRR8260155 GSM3495125 Yap1 ChIP −LIF rep1 (differentiating)
SRR8260156 GSM3495126 Yap1 ChIP −LIF rep2 (differentiating)

In scope (pipeline-derived, reproducing)

  • Yap1 ChIP-seq peak counts via the paper's stated pipeline: Bowtie2 → mm9 → MACS2 callpeak −p 1e-5 (treatment = Yap1 ChIP, control = matched input).
    • −LIF / differentiating: reported 8453 peaks (pool of rep1+rep2 vs −LIF input). [primary claim]
    • +LIF / self-renewal: reported 699 peaks (vs +LIF input). [primary claim]

Optional 20% (motif enrichment — may attempt, lower priority)

  • HOMER de-novo/known motif enrichment under the −LIF Yap1 peaks: paper reports Tead, Zic3, and AP-1 (JunB / Fra1/Fosl1) motifs enriched. Qualitative (motif presence + significance), not a single clean number → secondary.

Out of scope (wet-lab / not pipeline)

  • Cell-death rates (Fig 1A), caspase activity (Fig 1H), mitochondrial membrane potential (Fig 5), Bcl2 80–100-fold qPCR/protein (Fig 3E) — flow-cytometry, assays, microscopy. Not computational-pipeline outputs; not attempted.
  • RNA-seq DE gene lists — the paper gives thresholds (|log2|≥0.5) but no single pinned count in main text; not the cleanest target → not primary.

Pipeline parameters (from Methods)

  • Aligner: Bowtie2, reads 75 bp, genome mm9.
  • Peak caller: MACS2, p-value threshold 1e-5.
  • Motif: HOMER. Genome size for MACS2: -g mm (1.87e9).
Figures / tables: Figure 4Fig 4
C1
Reported
8453 Yap1 ChIP-seq peaks in differentiating (-LIF) ESCs (MACS2 p<1e-5)
Reproduced
1064 (pooled rep1+rep2); 5580 (rep2 SRR8260156 alone); 57 (rep1 SRR8260155 alone)
partial
C2
Reported
699 Yap1 ChIP-seq peaks in self-renewal (+LIF) ESCs (MACS2 p<1e-5)
Reproduced
289
partial
C3
Reported
HOMER motif enrichment (Tead/Zic3/AP-1) under -LIF Yap1 peaks
Reproduced
not attempted (optional 20%)
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 50/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🔴3. Location of the main deviation
🟡4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +6

Yap1 ChIP-seq peak counts do not reproduce exactly (reported 8453/699 vs 1064 pooled / 5580 best-rep / 289), but the paper's central directional result reproduces — far more Yap1 binding in differentiating (-LIF) than self-renewal (+LIF) ESCs (reported ~12x; reproduced ~19x best-rep). The gap is methodological under-specification: MACS2 version, --keep-dup (redundancy 0.28-0.81), p-vs-q mode and replicate pooling are all unstated, and the -LIF input control aligned only 37% (0.81 redundancy). The magnitude is derivable from shared data (rep2 alone gives 5580 of 8453), so this is not fabrication. The agent also corrected a mis-mined registry accession (GSE69669 RNA-seq -> the real GSE112606 ChIP-seq). A solid directional reproduction with quantitatively under-specified peak calling.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

130 k
tokens (I/O) · 8.8 M incl. cache
48 min
runtime · 25.94 CPU-h
12.1 GB
peak RAM
1
HPC jobs
hummel
machine