A single cell characterisation of human embryogenesis identifies pluripotency transitions and putative anterior hypoblast centre.
The main results reproduced: recomputed values matched the published ones within tolerance.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Same input data as the authors
- ✓Reported values were directly comparable
- ✓No authors-side cause for any deviation
- ✓Reported values are derivable from the shared data
- ✓Any deviation was negligible
- ✓The central claim held under reproduction
- 🟡A deviation arose in the data or preprocessing
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
DESCRIBED WELL ENOUGH for the shipped-data path; reproduced 1:1 on a clear self-contained slice. The repo ships the Fig 1d-f epiblast-staging count matrix (extended_data_table_4.csv, 262 cells: 165 own-data epiblast + 97 Stirparo reference, with per-cell developmental-stage annotation). Running the authors' OWN normalization (Seurat::NormalizeData) on it reproduced: (a) epiblast cell counts within 1-2 cells of the paper (165 vs 166; 106 vs 108 @day9; 59 vs 58 @day11); (b) all three Fig 1d-f directional claims -- naive pluripotency markers down post-implantation (all 7 named genes), primed markers up (4/5; SALL2 undetected), core POU5F1/NANOG/SOX2 steady upregulation (all 3). DID NOT attempt the headline Fig 1a-c clustering (4820 cells / 4 lineages Epi166-Hypo136-CTB2182-STB2336): the authors' own raw data E-MTAB-8060 is FASTQ-only (ENA ERP129702) with no processed matrices and the souporcell MPlex*/clusters.tsv intermediates are not deposited, so exact reconstruction needs cellranger+souporcell+Seurat integration = beyond the 80/20 line. Also skipped Fig 1g/ED Fig 3,5 logistic regression (sources logisticRegression.R/similarity.R that are absent from the repo) and Fig 2-3 (need unshipped integrated .Rdata). No fabrication signal: every reproduced value derives from shipped data+code; unverifiable cluster sizes reflect a data-deposition gap, not fabrication. NOTE: brief listed GSE136447 as 'the data' but that is the Xiang COMPARISON set; the own data is E-MTAB-8060.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 90assessed: 2026-06-15 ⛓ fd6a57d9e361
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-15
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusHow is epiblast patterning and the anterior-posterior axis initiated in the human embryo between implantation and gastrulation, and what signalling interactions between embryonic and extra-embryonic tissues drive these morphogenetic transformations?
- ★ The embryonic epiblast progressively transitions from a naïve to a primed pluripotent state as it develops from pre- to post-implantation. finding
- ★ The post-implantation epiblast acts as a source of FGF (FGF2/FGF4) signals that are required for proliferation of embryonic and extra-embryonic lineages. mechanism
- ★ A subset of asymmetrically positioned hypoblast cells expressing CER1 and other NODAL/BMP/WNT inhibitors constitutes a putative anterior signalling centre analogous to the mouse AVE. finding
- ★ FGF signalling is necessary for proliferation of epiblast, hypoblast and trophoblast lineages after implantation. finding
- ★ Single-cell RNA sequencing of post-implantation human embryos identifies four major lineages (epiblast, hypoblast, cytotrophoblast, syncytiotrophoblast). resource
- Conventional primed human ESCs (H9) resemble the 11 d.p.f. post-implantation epiblast, while naïve H9-Reset ESCs resemble the 6-7 d.p.f. pre-implantation epiblast. finding
- An open-source web server (www.humanembryo.org) is provided for the dataset. resource
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| single-cell RNA-seq (scRNA-seq) | in vitro cultured human embryos at 9 and 11 d.p.f. (16 embryos, 4820 cells) | none | transcriptional profiles / lineage clustering and marker gene expression | 10x Genomics Chromium |
| logistic regression cell-matching / dataset integration | human embryo epiblast/ICM datasets and human ESCs (H9, H9-Reset) | none | cell matching scores between pluripotency states | — |
| immunofluorescence | in vitro cultured human embryos at 8 d.p.f. | MEK inhibitor PD0325901 (3/1 μM), pan-FGFR inhibitor LY2874455 (1 μM/500 nM), FGF2/FGF4+heparan sulphate | cell number of OCT4+ epiblast, SOX17+ hypoblast, and trophoblast (OCT4/SOX17 negative) | — |
| immunofluorescence | in vitro cultured human embryos at 9 d.p.f. (19-28 embryos) | none | CER1+ cells among GATA6+ hypoblast cells and their angular/spatial distribution | — |
- ▼ Naïve pluripotency markers (KLF4, KLF17, PRDM14, DNMT3L, SOX15, TFCP2L1, ZFP42) downregulated post-implantation at 9 and 11 d.p.f.
- ▲ Primed pluripotency markers (FGF2, DNMT3B, SOX11, SFRP2, SALL2) increased in post-implantation epiblast; core markers POU5F1, NANOG, SOX2 steadily upregulated.
- ▼ PD (3 μM) and LY (1 μM and 500 nM) significantly reduced OCT4+ epiblast cell number; PD 1 μM non-significant for epiblast.
- ▼ PD (3 and 1 μM) and LY (1 μM and 500 nM) significantly reduced SOX17+ hypoblast cell number.
- – Trophoblast cell number not significantly affected by PD; LY reduced trophoblast only at 500 nM.
- – 17-43% (IQR) of GATA6+ hypoblast cells in contact with epiblast expressed CER1 at 9 d.p.f. 17-43% IQR
- – CER1+ cells showed a significant localisation bias towards one side of the hypoblast in 10/28 embryos. 10/28 embryos
- ▲ CER1 significantly co-expressed with NODAL/BMP/WNT antagonists LEFTY1, LEFTY2, HHEX, NOG.
- count 16 embryos (8 at 9 d.p.f., 8 at 11 d.p.f.); 4820 cells total (final scRNA-seq dataset after excluding 13 of 29 embryos lacking ICM derivatives)
- pvalue p = 0.0026 (control vs PD 3 μM, OCT4+ epiblast cell reduction (Kruskal-Wallis, Dunn's))
- pvalue p = 0.0006 (control vs LY 1 μM); p = 0.0109 (control vs LY 500 nM) (OCT4+ epiblast cell reduction)
- pvalue p = 0.0011 (control vs LY 1 μM); p = 0.0002 (control vs LY 500 nM) (SOX17+ hypoblast cell reduction)
- pvalue p = 4.21E-13 (LEFTY1); 2.75E-10 (LEFTY2); 1.88E-07 (HHEX) (CER1 co-expression correlation, Benjamini-Hochberg corrected)
- pvalue H9 vs H9-Reset: 5 d.p.f. p=0.0015, 6-7 d.p.f. p<0.0001, 9 d.p.f. p=0.0101, 11 d.p.f. p<0.0001 (Sidak's multiple comparisons, 2-way ANOVA, logistic regression matching)
- count epiblast 166, hypoblast 136, cytotrophoblast 2182, syncytiotrophoblast 2336 cells (cell counts per lineage cluster)
- count n = 19 (control), 12 (PD 3 μM), 15 (PD 1 μM), 8 (LY 1 μM), 8 (LY 500 nM), 7 (FGF2/4) (number of embryos per FGF treatment condition)
Statistical methods review
Model: opusA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This study combines single-cell RNA sequencing (10x Genomics) of 16 post-implantation human embryos with functional immunofluorescence-based cell-counting experiments. Cell identities were defined by clustering and differential gene expression, datasets were integrated and compared to published references via a logistic-regression projection, and functional perturbations (FGF pathway inhibitors/activators) were assessed by quantifying lineage cell numbers. Group comparisons used non-parametric tests (Kruskal–Wallis with Dunn's) and two-way ANOVA with Sidak's post-hoc, while gene co-expression used correlation with Benjamini–Hochberg correction; results were shown mainly as box plots and bar graphs with medians/means.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Two-tailed Kruskal–Wallis test with Dunn's multiple-comparison correction | Fig. 2d–f: quantification of epiblast (OCT4+), hypoblast (SOX17+) and trophoblast cell numbers across control and inhibitor/activator treatments | embryos per group: control n=19, PD 3 µM n=12, PD 1 µM n=15, LY 1 µM n=8, LY 500 nM n=8, FGF2/4 n=7 | not stated |
| Two-way ANOVA with Sidak's multiple-comparisons post-hoc test | Fig. 1g: logistic-regression matching scores, H9 vs H9-Reset ESCs across developmental stages (5, 6–7, 9, 11 d.p.f.) | — | not stated |
| Correlation analysis with Benjamini–Hochberg correction | Fig. 3i: co-expression of CER1 with WNT/BMP/NODAL antagonists (LEFTY1, LEFTY2, HHEX, NOG) | — | not stated |
| Differential gene expression analysis | Fig. 1b and Supplementary Data 2: marker genes defining the four lineages | — | na |
| Logistic regression model (cell-type projection/classification) | Fig. 1c,g and Supplementary Fig. 3/5: projecting published embryo and ESC datasets onto the authors' clusters | — | na |
| Angular-distribution / localisation-bias analysis | Fig. 3h: angular distribution of CER1+ cells across the hypoblast hemisphere (bias in 10/28 embryos) | n=28 embryos | not stated |
-
Functional cell-count comparisons used the Kruskal–Wallis test with Dunn's correction on per-embryo counts↳ Could also: A generalized linear model for count data (e.g. Poisson or negative-binomial regression), or an ANOVA on transformed counts with a post-hoc correction — A count-based model can directly accommodate the integer nature of cell numbers and yield effect-size estimates with confidence intervals alongside p-values
-
Bar graphs were summarized with mean ± SEM (Fig. 3e,f)↳ Could also: Showing SD or a 95% confidence interval in addition to, or instead of, SEM — SD conveys the spread of the underlying data and a 95% CI conveys precision of the mean; both are often preferred, particularly with modest sample sizes, to make variability explicit
-
Reported significance relies primarily on p-values from the group comparisons↳ Could also: Reporting effect sizes (e.g. differences in medians/means with confidence intervals, or rank-based effect sizes such as Cliff's delta) — Effect-size reporting communicates the magnitude and practical relevance of differences independently of sample size
-
Stage/cell-line matching scores were analyzed with two-way ANOVA and Sidak's post-hoc test↳ Could also: A non-parametric or mixed-effects model that accounts for the nested structure of cells within embryos/cell lines — A mixed-effects approach can model within-sample correlation and pseudoreplication explicitly, which is useful when many cells derive from few embryos
-
Co-expression relationships were assessed with correlation and Benjamini–Hochberg FDR control↳ Could also: Reporting the correlation coefficients/effect magnitudes alongside the FDR-adjusted p-values, or a regression-based co-expression model — Showing the strength of each association (not only its significance) gives a fuller picture of co-expression structure
-
Differential expression and marker identification were described without naming the specific test/model↳ Could also: Explicitly stating the DE method (e.g. Wilcoxon rank-sum in Seurat, MAST, or DESeq2) and its multiplicity handling — Naming the DE test and correction makes the marker-selection criteria fully transparent and reproducible
Result convergence & founder nodes
Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.
-
CER1+ hypoblast cells show significant spatial localisation bias toward one side of the hypoblast in a subset of human embryos at 9 dpf, consistent with a putative anterior hypoblast centreimaging human embryo 2021×1papers★ This paper is the founder (earliest)
-
MEK inhibitor PD0325901 and pan-FGFR inhibitor LY2874455 significantly reduce OCT4+ epiblast cell number in human embryos at 8 dpfimaging human embryo down 2021×1papers★ This paper is the founder (earliest)
-
MEK inhibitor PD0325901 and pan-FGFR inhibitor LY2874455 significantly reduce SOX17+ hypoblast cell number in human embryos at 8 dpfimaging human embryo down 2021×1papers★ This paper is the founder (earliest)
-
Trophoblast cell number is not significantly reduced by MEK inhibitor PD0325901 but is reduced by pan-FGFR inhibitor LY2874455 at 500 nM in human embryos at 8 dpfimaging human embryo mixed 2021×1papers★ This paper is the founder (earliest)
-
CER1+ hypoblast cells co-express NODAL/BMP/WNT antagonists LEFTY1, LEFTY2, HHEX, and NOG, identifying a putative anterior hypoblast centre at 9 dpfscRNA-seq human embryo up 2021×1papers★ This paper is the founder (earliest)
-
Naive pluripotency markers KLF4, KLF17, PRDM14, DNMT3L, SOX15, TFCP2L1, and ZFP42 are downregulated in post-implantation epiblast at 9-11 dpfscRNA-seq human embryo down 2021×1papers★ This paper is the founder (earliest)
-
Core pluripotency factors POU5F1, NANOG, and SOX2 are steadily upregulated in post-implantation epiblast alongside primed markers FGF2, DNMT3B, SOX11, SFRP2, and SALL2 at 9-11 dpfscRNA-seq human embryo up 2021×1papers★ This paper is the founder (earliest)
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-34140473
Paper: Molè, Coorens et al. 2021, Nat Commun 12:3679. "A single cell
characterisation of human embryogenesis identifies pluripotency transitions and
putative anterior hypoblast centre." DOI 10.1038/s41467-021-23758-w.
Code: https://github.com/TimCoorens/EarlyEmbryo_scRNA @ a8ad4e5 (only commit).
Own data: ArrayExpress E-MTAB-8060 (10X v2, in-vitro cultured human embryos,
day 9 & 11). Note: the brief listed GSE136447, but that is the Xiang et al.
comparison dataset, not the authors' own data — corrected here.
Repo contents
Three R scripts (Seurat-based) + one shipped data file:
scRNA_embryo_data_process.R— cellranger → batch-correct (Seurat anchors) → cluster (res 0.05) → 4 lineages; Fig 1a-f, ED Fig 1.scRNA_embryo_expression_patterns.R— CER1 co-expression, hypo/epiblast subclustering; Fig 2-3, ED Fig 6.scRNA_embryo_logistic_regression.R— logistic-regression cross-dataset comparisons; Fig 1g, ED Fig 3/5.extended_data_table_4.csv.zip(111 MB unzipped) — the integrated epiblast staging matrix = 262 cells × 69,476 genes (raw counts), row "Annotation" = developmental stage. 165 cells are own-data epiblast (embryo*); 97 are Stirparo et al. reference cells. This is the source data for Fig 1d-f.
In scope (attempted) — self-contained, deterministic
Using the authors' own shipped table + the authors' own normalization
(Seurat::NormalizeData, LogNormalize), no external data required:
- Epiblast cell counts per stage (Fig 1 / main text): own-data epiblast = 166 cells (108 @9 d.p.f., 58 @11 d.p.f.).
- Fig 1d-f — median expression of naive / primed / core pluripotency gene panels across the 4 ordered epiblast stages (ICM → Pre-Epi → Peri-Epi → Post-Epi); the directional claims (naive down, primed up, core up).
Out of scope (NOT attempted) — with reason
- Fig 1a-c headline clustering (4820 cells, 16 embryos, 4 lineages with sizes
Epi 166 / Hypo 136 / CTB 2182 / STB 2336). The own raw data E-MTAB-8060 ships
only FASTQ (ENA ERP129702) + metadata (IDF/SDRF) — no processed
filtered_feature_bc_matrix. Full reproduction would need cellranger + souporcell deconvolution (theMPlex*/clusters.tsvintermediates are not shipped) + Seurat integration of 6 batches: days of compute, partly stochastic, and not exactly specified → beyond the 80/20 line. - Fig 1g, ED Fig 3/5 logistic regression — scripts
source()logisticRegression.R/similarity.R, which are not in the repo (external Sanger helpers), and pull several large public datasets. Not attempted. - Fig 2-3 expression/co-expression — require the full integrated object
(
embryo_integrated_allembryos_filtered.Rdata), not shipped. Not attempted. - All wet-lab / imaging results — out of scope by definition.
Pipeline named per attempted result
Seurat v4 (LogNormalize) on the authors' shipped count table → per-stage medians.
Faithful 1:1 to the mat <- GetAssayData(..., slot="data") + median(mat[Gene, epi_state==s]) computation in scRNA_embryo_data_process.R (Fig 1d-f block).
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
On the self-contained Fig 1d-f slice, reproduction is essentially 1:1: using the authors' own shipped staging matrix and their own Seurat normalization, epiblast counts match within 1-2 cells (165 vs 166) and all three pluripotency-panel directional claims (naive down, primed up, core up) hold. The only deviations are negligible count deltas (staging re-annotation) and SALL2 undetected in the shipped matrix — our-side/technical, not authors' fault. The limitation is data availability: the headline Fig 1a-c clustering could not be attempted because the own raw data is FASTQ-only with no processed matrices/intermediates deposited, so overall this is a solid-but-partial reproduction, not a clean full 1:1.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.