Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

A single cell characterisation of human embryogenesis identifies pluripotency transitions and putative anterior hypoblast centre.

Nat Commun · 2021
L1 90/100 PQI 97
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +1
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
  • No authors-side cause for any deviation
  • Reported values are derivable from the shared data
  • Any deviation was negligible
  • The central claim held under reproduction
What did not (or only partly)
  • 🟡A deviation arose in the data or preprocessing
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
90/100
Reproducibility score
0.9 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 79% of all assessed papers rank 211 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

DESCRIBED WELL ENOUGH for the shipped-data path; reproduced 1:1 on a clear self-contained slice. The repo ships the Fig 1d-f epiblast-staging count matrix (extended_data_table_4.csv, 262 cells: 165 own-data epiblast + 97 Stirparo reference, with per-cell developmental-stage annotation). Running the authors' OWN normalization (Seurat::NormalizeData) on it reproduced: (a) epiblast cell counts within 1-2 cells of the paper (165 vs 166; 106 vs 108 @day9; 59 vs 58 @day11); (b) all three Fig 1d-f directional claims -- naive pluripotency markers down post-implantation (all 7 named genes), primed markers up (4/5; SALL2 undetected), core POU5F1/NANOG/SOX2 steady upregulation (all 3). DID NOT attempt the headline Fig 1a-c clustering (4820 cells / 4 lineages Epi166-Hypo136-CTB2182-STB2336): the authors' own raw data E-MTAB-8060 is FASTQ-only (ENA ERP129702) with no processed matrices and the souporcell MPlex*/clusters.tsv intermediates are not deposited, so exact reconstruction needs cellranger+souporcell+Seurat integration = beyond the 80/20 line. Also skipped Fig 1g/ED Fig 3,5 logistic regression (sources logisticRegression.R/similarity.R that are absent from the repo) and Fig 2-3 (need unshipped integrated .Rdata). No fabrication signal: every reproduced value derives from shipped data+code; unverifiable cluster sizes reflect a data-deposition gap, not fabrication. NOTE: brief listed GSE136447 as 'the data' but that is the Xiang COMPARISON set; the own data is E-MTAB-8060.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 90
    assessed: 2026-06-15 ⛓ fd6a57d9e361
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-15
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

How is epiblast patterning and the anterior-posterior axis initiated in the human embryo between implantation and gastrulation, and what signalling interactions between embryonic and extra-embryonic tissues drive these morphogenetic transformations?

Core claims
  • The embryonic epiblast progressively transitions from a naïve to a primed pluripotent state as it develops from pre- to post-implantation. finding
  • The post-implantation epiblast acts as a source of FGF (FGF2/FGF4) signals that are required for proliferation of embryonic and extra-embryonic lineages. mechanism
  • A subset of asymmetrically positioned hypoblast cells expressing CER1 and other NODAL/BMP/WNT inhibitors constitutes a putative anterior signalling centre analogous to the mouse AVE. finding
  • FGF signalling is necessary for proliferation of epiblast, hypoblast and trophoblast lineages after implantation. finding
  • Single-cell RNA sequencing of post-implantation human embryos identifies four major lineages (epiblast, hypoblast, cytotrophoblast, syncytiotrophoblast). resource
  • Conventional primed human ESCs (H9) resemble the 11 d.p.f. post-implantation epiblast, while naïve H9-Reset ESCs resemble the 6-7 d.p.f. pre-implantation epiblast. finding
  • An open-source web server (www.humanembryo.org) is provided for the dataset. resource
Experimental setups
Assay System Perturbation Readout Platform
single-cell RNA-seq (scRNA-seq) in vitro cultured human embryos at 9 and 11 d.p.f. (16 embryos, 4820 cells) none transcriptional profiles / lineage clustering and marker gene expression 10x Genomics Chromium
logistic regression cell-matching / dataset integration human embryo epiblast/ICM datasets and human ESCs (H9, H9-Reset) none cell matching scores between pluripotency states
immunofluorescence in vitro cultured human embryos at 8 d.p.f. MEK inhibitor PD0325901 (3/1 μM), pan-FGFR inhibitor LY2874455 (1 μM/500 nM), FGF2/FGF4+heparan sulphate cell number of OCT4+ epiblast, SOX17+ hypoblast, and trophoblast (OCT4/SOX17 negative)
immunofluorescence in vitro cultured human embryos at 9 d.p.f. (19-28 embryos) none CER1+ cells among GATA6+ hypoblast cells and their angular/spatial distribution
Key results
  • Naïve pluripotency markers (KLF4, KLF17, PRDM14, DNMT3L, SOX15, TFCP2L1, ZFP42) downregulated post-implantation at 9 and 11 d.p.f.
  • Primed pluripotency markers (FGF2, DNMT3B, SOX11, SFRP2, SALL2) increased in post-implantation epiblast; core markers POU5F1, NANOG, SOX2 steadily upregulated.
  • PD (3 μM) and LY (1 μM and 500 nM) significantly reduced OCT4+ epiblast cell number; PD 1 μM non-significant for epiblast.
  • PD (3 and 1 μM) and LY (1 μM and 500 nM) significantly reduced SOX17+ hypoblast cell number.
  • Trophoblast cell number not significantly affected by PD; LY reduced trophoblast only at 500 nM.
  • 17-43% (IQR) of GATA6+ hypoblast cells in contact with epiblast expressed CER1 at 9 d.p.f. 17-43% IQR
  • CER1+ cells showed a significant localisation bias towards one side of the hypoblast in 10/28 embryos. 10/28 embryos
  • CER1 significantly co-expressed with NODAL/BMP/WNT antagonists LEFTY1, LEFTY2, HHEX, NOG.
Key statistics
  • count 16 embryos (8 at 9 d.p.f., 8 at 11 d.p.f.); 4820 cells total (final scRNA-seq dataset after excluding 13 of 29 embryos lacking ICM derivatives)
  • pvalue p = 0.0026 (control vs PD 3 μM, OCT4+ epiblast cell reduction (Kruskal-Wallis, Dunn's))
  • pvalue p = 0.0006 (control vs LY 1 μM); p = 0.0109 (control vs LY 500 nM) (OCT4+ epiblast cell reduction)
  • pvalue p = 0.0011 (control vs LY 1 μM); p = 0.0002 (control vs LY 500 nM) (SOX17+ hypoblast cell reduction)
  • pvalue p = 4.21E-13 (LEFTY1); 2.75E-10 (LEFTY2); 1.88E-07 (HHEX) (CER1 co-expression correlation, Benjamini-Hochberg corrected)
  • pvalue H9 vs H9-Reset: 5 d.p.f. p=0.0015, 6-7 d.p.f. p<0.0001, 9 d.p.f. p=0.0101, 11 d.p.f. p<0.0001 (Sidak's multiple comparisons, 2-way ANOVA, logistic regression matching)
  • count epiblast 166, hypoblast 136, cytotrophoblast 2182, syncytiotrophoblast 2336 cells (cell counts per lineage cluster)
  • count n = 19 (control), 12 (PD 3 μM), 15 (PD 1 μM), 8 (LY 1 μM), 8 (LY 500 nM), 7 (FGF2/4) (number of embryos per FGF treatment condition)

Statistical methods review

Model: opus

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This study combines single-cell RNA sequencing (10x Genomics) of 16 post-implantation human embryos with functional immunofluorescence-based cell-counting experiments. Cell identities were defined by clustering and differential gene expression, datasets were integrated and compared to published references via a logistic-regression projection, and functional perturbations (FGF pathway inhibitors/activators) were assessed by quantifying lineage cell numbers. Group comparisons used non-parametric tests (Kruskal–Wallis with Dunn's) and two-way ANOVA with Sidak's post-hoc, while gene co-expression used correlation with Benjamini–Hochberg correction; results were shown mainly as box plots and bar graphs with medians/means.

Replicationbiological Sample sizeReported as numbers of embryos per condition and numbers of cells per cluster, plus numbers of experimental replicates (e.g. 9 or 3); no formal power/sample-size calculation described GroupsDevelopmental stages (9 vs 11 d.p.f.; pre- vs post-implantation), lineages (epiblast/hypoblast/trophoblast), and drug treatments vs control Pairingunpaired Randomization/blindingnot stated Dispersionmixed Exact p-valuesyes Effect sizesno Confidence intervalsno Multiplicity correctionBenjamini–Hochberg FDR (correlations), Sidak's correction (within two-way ANOVA), Dunn's correction (within Kruskal–Wallis)
Statistical tests used
Test Applied to n Assumptions
Two-tailed Kruskal–Wallis test with Dunn's multiple-comparison correction Fig. 2d–f: quantification of epiblast (OCT4+), hypoblast (SOX17+) and trophoblast cell numbers across control and inhibitor/activator treatments embryos per group: control n=19, PD 3 µM n=12, PD 1 µM n=15, LY 1 µM n=8, LY 500 nM n=8, FGF2/4 n=7 not stated
Two-way ANOVA with Sidak's multiple-comparisons post-hoc test Fig. 1g: logistic-regression matching scores, H9 vs H9-Reset ESCs across developmental stages (5, 6–7, 9, 11 d.p.f.) not stated
Correlation analysis with Benjamini–Hochberg correction Fig. 3i: co-expression of CER1 with WNT/BMP/NODAL antagonists (LEFTY1, LEFTY2, HHEX, NOG) not stated
Differential gene expression analysis Fig. 1b and Supplementary Data 2: marker genes defining the four lineages na
Logistic regression model (cell-type projection/classification) Fig. 1c,g and Supplementary Fig. 3/5: projecting published embryo and ESC datasets onto the authors' clusters na
Angular-distribution / localisation-bias analysis Fig. 3h: angular distribution of CER1+ cells across the hypoblast hemisphere (bias in 10/28 embryos) n=28 embryos not stated
Approaches that could also have been used
  • Functional cell-count comparisons used the Kruskal–Wallis test with Dunn's correction on per-embryo counts
    Could also: A generalized linear model for count data (e.g. Poisson or negative-binomial regression), or an ANOVA on transformed counts with a post-hoc correction — A count-based model can directly accommodate the integer nature of cell numbers and yield effect-size estimates with confidence intervals alongside p-values
  • Bar graphs were summarized with mean ± SEM (Fig. 3e,f)
    Could also: Showing SD or a 95% confidence interval in addition to, or instead of, SEM — SD conveys the spread of the underlying data and a 95% CI conveys precision of the mean; both are often preferred, particularly with modest sample sizes, to make variability explicit
  • Reported significance relies primarily on p-values from the group comparisons
    Could also: Reporting effect sizes (e.g. differences in medians/means with confidence intervals, or rank-based effect sizes such as Cliff's delta) — Effect-size reporting communicates the magnitude and practical relevance of differences independently of sample size
  • Stage/cell-line matching scores were analyzed with two-way ANOVA and Sidak's post-hoc test
    Could also: A non-parametric or mixed-effects model that accounts for the nested structure of cells within embryos/cell lines — A mixed-effects approach can model within-sample correlation and pseudoreplication explicitly, which is useful when many cells derive from few embryos
  • Co-expression relationships were assessed with correlation and Benjamini–Hochberg FDR control
    Could also: Reporting the correlation coefficients/effect magnitudes alongside the FDR-adjusted p-values, or a regression-based co-expression model — Showing the strength of each association (not only its significance) gives a fuller picture of co-expression structure
  • Differential expression and marker identification were described without naming the specific test/model
    Could also: Explicitly stating the DE method (e.g. Wilcoxon rank-sum in Seurat, MAST, or DESeq2) and its multiplicity handling — Naming the DE test and correction makes the marker-selection criteria fully transparent and reproducible
Software: 10x Genomics Chromium (scRNA-seq platform) · Statistical software for tests (e.g. for ANOVA/Sidak's and Kruskal–Wallis/Dunn's) — not explicitly named in the provided text

Result convergence & founder nodes

Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
133
Impact: high
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

7dpf in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-34140473

Paper: Molè, Coorens et al. 2021, Nat Commun 12:3679. "A single cell characterisation of human embryogenesis identifies pluripotency transitions and putative anterior hypoblast centre." DOI 10.1038/s41467-021-23758-w. Code: https://github.com/TimCoorens/EarlyEmbryo_scRNA @ a8ad4e5 (only commit). Own data: ArrayExpress E-MTAB-8060 (10X v2, in-vitro cultured human embryos, day 9 & 11). Note: the brief listed GSE136447, but that is the Xiang et al. comparison dataset, not the authors' own data — corrected here.

Repo contents

Three R scripts (Seurat-based) + one shipped data file:

  • scRNA_embryo_data_process.R — cellranger → batch-correct (Seurat anchors) → cluster (res 0.05) → 4 lineages; Fig 1a-f, ED Fig 1.
  • scRNA_embryo_expression_patterns.R — CER1 co-expression, hypo/epiblast subclustering; Fig 2-3, ED Fig 6.
  • scRNA_embryo_logistic_regression.R — logistic-regression cross-dataset comparisons; Fig 1g, ED Fig 3/5.
  • extended_data_table_4.csv.zip (111 MB unzipped) — the integrated epiblast staging matrix = 262 cells × 69,476 genes (raw counts), row "Annotation" = developmental stage. 165 cells are own-data epiblast (embryo*); 97 are Stirparo et al. reference cells. This is the source data for Fig 1d-f.

In scope (attempted) — self-contained, deterministic

Using the authors' own shipped table + the authors' own normalization (Seurat::NormalizeData, LogNormalize), no external data required:

  1. Epiblast cell counts per stage (Fig 1 / main text): own-data epiblast = 166 cells (108 @9 d.p.f., 58 @11 d.p.f.).
  2. Fig 1d-f — median expression of naive / primed / core pluripotency gene panels across the 4 ordered epiblast stages (ICM → Pre-Epi → Peri-Epi → Post-Epi); the directional claims (naive down, primed up, core up).

Out of scope (NOT attempted) — with reason

  • Fig 1a-c headline clustering (4820 cells, 16 embryos, 4 lineages with sizes Epi 166 / Hypo 136 / CTB 2182 / STB 2336). The own raw data E-MTAB-8060 ships only FASTQ (ENA ERP129702) + metadata (IDF/SDRF) — no processed filtered_feature_bc_matrix. Full reproduction would need cellranger + souporcell deconvolution (the MPlex*/clusters.tsv intermediates are not shipped) + Seurat integration of 6 batches: days of compute, partly stochastic, and not exactly specified → beyond the 80/20 line.
  • Fig 1g, ED Fig 3/5 logistic regression — scripts source() logisticRegression.R / similarity.R, which are not in the repo (external Sanger helpers), and pull several large public datasets. Not attempted.
  • Fig 2-3 expression/co-expression — require the full integrated object (embryo_integrated_allembryos_filtered.Rdata), not shipped. Not attempted.
  • All wet-lab / imaging results — out of scope by definition.

Pipeline named per attempted result

Seurat v4 (LogNormalize) on the authors' shipped count table → per-stage medians. Faithful 1:1 to the mat <- GetAssayData(..., slot="data") + median(mat[Gene, epi_state==s]) computation in scRNA_embryo_data_process.R (Fig 1d-f block).

Figures / tables: Fig 1Fig 1dFig 1eFig 1f
C1
Reported
166 epiblast cells
Reproduced
165
within tolerance
C2
Reported
108 epiblast @9 d.p.f.
Reproduced
106 (Peri-Epi)
within tolerance
C3
Reported
58 epiblast @11 d.p.f.
Reproduced
59 (Post-Epi)
within tolerance
C4
Reported
Fig1d: naive pluripotency markers (KLF4,KLF17,PRDM14,DNMT3L,SOX15,TFCP2L1,ZFP42) downregulated post-implantation
Reproduced
all 7 named genes down; panel median 0.818->0.508
exact
C5
Reported
Fig1e: primed markers (FGF2,DNMT3B,SOX11,SFRP2,SALL2) increased post-implantation
Reproduced
4/5 up (SALL2 undetected); panel median 0.257->0.472
within tolerance
C6
Reported
Fig1f: core markers (POU5F1,NANOG,SOX2) steady upregulation pre->post
Reproduced
all 3 monotonic up; panel median 0.732->1.596
exact

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 90/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟡3. Location of the main deviation
🟢4. Cause of the deviation
🟢5. Derivability / plausibility
🟢6. Severity of the deviation
🟢7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +1

On the self-contained Fig 1d-f slice, reproduction is essentially 1:1: using the authors' own shipped staging matrix and their own Seurat normalization, epiblast counts match within 1-2 cells (165 vs 166) and all three pluripotency-panel directional claims (naive down, primed up, core up) hold. The only deviations are negligible count deltas (staging re-annotation) and SALL2 undetected in the shipped matrix — our-side/technical, not authors' fault. The limitation is data availability: the headline Fig 1a-c clustering could not be attempted because the own raw data is FASTQ-only with no processed matrices/intermediates deposited, so overall this is a solid-but-partial reproduction, not a clean full 1:1.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

140.3 k
tokens (I/O) · 9.9 M incl. cache
16 min
runtime · 0 CPU-h
1.3 GB
peak RAM
1
HPC jobs
hummel
machine