A comparative analysis of blastoid models through single-cell transcriptomics.
The main results reproduced, with only marginal, non-material deviations.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Reported values are derivable from the shared data
- ✓Any deviation was negligible
- 🟡Could not use the authors’ exact input data
- 🟡Reported values were only indirectly comparable
- 🟡A deviation arose in the data or preprocessing
- 🟡A deviation was attributed to the published material
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
DESCRIBED WELL ENOUGH: yes for the QC step. The paper (Balubaid et al. 2024, iScience) + repo + Table S1 specify the QC filtering thresholds exactly, so the QC/cell-filtering pipeline output is reproducible without ambiguity. We reproduced post-QC cell counts faithfully (Seurat-equivalent metrics in Python) on a «our HPC» SLURM job. RESULT: 1:1-ish PARTIAL, not exact. (1) Petropoulos2016_ref (E-MTAB-3929, the one comparative row with a printed reference): reported 1081 cells, reproduced 1137 on the E5-E7 subset (+5.2%) using the PUBLISHED matrix. The residual gap is fully explained by documented method differences: the authors requantified with kallisto (so their NGenes=29613 vs the matrix's 26178), and the published matrix carries 0 mitochondrial genes so the documented mito<0.2 filter could not remove high-mito cells on our side -> we keep slightly more cells. This is consistent with a GENUINE, non-fabricated number, not a fabrication. (2) Founding cell lines GSE247758 (the dataset assigned to this RU = the authors' OWN data): reproduced post-QC counts 2628 EPSC / 1721 nPSC under the script's documented thresholds (mito<0.1, nCount>20000, nFeature>5000) -- but these are DERIVED values because the paper/Table S1 never print founding-line post-QC counts (a transparency gap, flagged). (3) The repo's shipped result CSVs match Table S1 EXACTLY, confirming the shipped tables are the reported values. NOT ATTEMPTED (the hard ~20%, per 80/20): full kallisto|bustools requantification from raw FASTQ for the 8 external datasets (this is what sets the exact NGenes), plus cell-type annotation, Harmony integration, lineage module scores, DEGs, and all figures. No fabrication signal found.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 60assessed: 2026-06-14 ⛓ a2c760fb5eb9
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-14
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetThe paper investigates how similar pluripotent-stem-cell-derived blastoids are to natural peri-implantation human blastocysts, and what drives transcriptomic and cell-type heterogeneity across different blastoid generation protocols/founding cell lines.
- ★ EPSC-derived blastoids are transcriptomically distinct from nPSC-derived blastoids, with nPSC-blastoids clustering closer to natural blastocysts. finding
- ★ EPSC-blastoids show a higher proportion of primitive endoderm (PE) cells and ambiguous/unknown cells with endoderm-like signatures. finding
- ★ Blastoid models can be categorized into three subtypes by cell-type composition: balanced, EPI_ICM-enriched, or PE-enriched. finding
- ★ PE cells split into two transcriptomically distinct subtypes (PE I and PE II); PE II is enriched in EPSC-based blastoids and bears amnion-like traits. finding
- ★ Cell-type composition is significantly associated with the source cell type (blastocyst, nPSC, EPSC, fibroblast). finding
- ★ Gene expression heterogeneity in founding pluripotent cell lines (higher heterogeneity in nPSCs, prevalent amnionic signatures in EPSCs) influences blastoid lineage differentiation outcomes. mechanism
- A consensus cell-type annotation approach combining three automatic annotation methods was used to label cells across integrated datasets. method
- A lineage module score framework with Kolmogorov-Smirnov distance was developed to quantify lineage disposition and score-distribution differences between datasets. method
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| scRNAseq (reanalysis/integration of published data) | human blastoid models (Yanagida et al., Yu et al. 5iLA, Yu et al. PXGL, Kagawa et al., Liu et al., Sozen et al., Fan et al.) | none (comparative computational reanalysis) | cell-type composition, cluster distribution | — |
| scRNAseq (reanalysis/integration of published data) | natural peri-implantation human blastocysts (Petropoulos et al., Xiang et al., and others) | none | cell-type composition, transcriptomic identity | — |
| scRNAseq data integration (UMAP/PCA clustering) | combined blastoid + blastocyst datasets, sequenced by 10x Genomics or SMART-Seq2 | none | shared transcriptomic landscape, cluster overlap | 10x Genomics; SMART-Seq2 |
| lineage module score analysis (marker gene expression scoring) | integrated blastoid/blastocyst scRNAseq datasets | none | EPI_ICM, PE, TE module score expression per cell | — |
| Kolmogorov-Smirnov distance (KSD) statistical comparison | lineage module score distributions across datasets | none | distributional similarity/difference of module scores between datasets | — |
| scRNAseq (newly generated) | founding pluripotent stem cell lines: EPSCs and nPSCs | none | transcriptomic heterogeneity, marker/signature expression (e.g., amnionic signatures) | — |
| chi-square association test | cell-type annotations vs. dataset source (blastocyst, nPSC, EPSC, fibroblast) | none | statistical association between primary cell-type composition and source cell type | — |
- – Significant association between primary cell-type composition (EPI_ICM, PE, TE) and source cell type when only annotated cells included. χ2=4494.5, df=6, p<2.2e-16
- – Significant association between cell-type composition and source cell type when unassigned cells included. χ2=12661, df=9, p<2.2e-16
- – Three cell-type composition clusters identified: balanced (blastocysts + Yanagida et al.), EPI_ICM-enriched (Yu et al. 5iLA/PXGL, Kagawa et al.), and PE-enriched (Liu et al., Fan et al., Sozen et al.).
- – PE cells separate into PE I and PE II subtypes in shared transcriptomic space; PE II particularly enriched in EPSC-based blastoids, with a small population also in Yu2021 blastoids.
- ▼ KLF17 (EPI_ICM marker) is not expressed in Sozen et al. and Fan et al. datasets.
- ▼ TACSTD2 (TE marker) is absent in multiple datasets.
- – Most unidentified/unknown cells represent intermediate states between EPI_ICM and TE (or EPI_ICM and PE), suggesting immature or plastic cell states.
- – Sozen et al. blastoid PE and unknown cells show minimal transcriptomic overlap with other blastocyst/blastoid cells, indicating a distinct transcriptomic profile.
- pvalue p<2.2e-16 (chi-square association between cell-type composition and source cell type (annotated cells only))
- other χ2 statistic = 4494.5, df = 6 (chi-square test, annotated cells, cell-type composition vs. source cell type)
- pvalue p<2.2e-16 (chi-square association between cell-type composition and source cell type (including unassigned cells))
- other χ2 statistic = 12661, df = 9 (chi-square test, including unassigned cells, cell-type composition vs. source cell type)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This computational study integrates scRNAseq data from seven published blastoid models and reference blastocyst datasets to compare cell-type composition, transcriptomic identity, and lineage module scores. Similarity between datasets is quantified using chi-squared association tests, Jensen-Shannon Distance (JSD), Pearson correlation distance (PCD), and Kolmogorov-Smirnov distance (KSD) with bootstrapping. Results are presented primarily through dimensionality reduction (UMAP, PCA), hierarchical clustering, and heatmaps of pairwise distances rather than inferential hypothesis tests with correction for multiple comparisons.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Chi-squared test of association | Association between primary cell-type composition (EPI_ICM, PE, TE) and source cell type (blastocyst, nPSC, EPSC, fibroblast) — run twice: once with annotated cells only, once including unassigned cells | χ2 = 4494.5, df = 6 (annotated only); χ2 = 12661, df = 9 (with unassigned) — cell-level counts, exact n not stated | not stated |
| Jensen-Shannon Distance (JSD) | Pairwise similarity of cluster distribution across all integrated datasets (Figure 1G) | null | na |
| Pearson correlation distance (PCD) | Pairwise similarity of cluster distribution across datasets restricted to shared clusters (Figure 1H) | null | na |
| Kolmogorov-Smirnov distance (KSD) with bootstrapping | Pairwise comparison of lineage module score distributions (EPI_ICM, PE, TE) between datasets (Figures 2C–2E) | null | not stated |
| Hierarchical clustering | Clustering datasets by cell-type proportions (Figure 1B) and by cluster distribution (Figure 1F) | 7 blastoid models plus reference blastocyst datasets | na |
| Consensus of three automatic cell-type annotation methods | Cell-type annotation of all integrated single cells (Figures 1A, S3A–S3D) | null | not stated |
-
Cell-type composition differences across source groups were assessed with a single chi-squared test on cell counts↳ Could also: A compositional analysis framework such as scCODA or a Dirichlet-multinomial regression could also be used, treating each sample/dataset as the unit of observation rather than each cell — Treating individual cells as independent observations inflates the effective sample size because cells from the same sample are correlated; sample-level compositional models account for this clustering and produce estimates with appropriate uncertainty
-
Pairwise dataset similarity was quantified with JSD and PCD on cluster distributions, displayed as heatmaps without formal inference↳ Could also: A permutation-based test or PERMANOVA on the cluster-proportion vectors could also be used to attach p-values to the observed distances — Formal inference would allow statements about whether observed inter-group distances exceed what would be expected by chance, complementing the descriptive heatmaps
-
Lineage module score distributions were compared between datasets using KSD with bootstrapping↳ Could also: A two-sample Kolmogorov-Smirnov test with Bonferroni or Benjamini-Hochberg correction across all pairwise comparisons could also be used — Reporting adjusted p-values alongside the KSD statistic would clarify which pairwise differences are statistically notable after accounting for the number of comparisons
-
Dispersion in boxplots is shown as the 10th–90th percentile range rather than SD, SEM, or a 95% CI↳ Could also: A 95% confidence interval on the median (or mean) could also be displayed, especially for between-dataset comparisons — CIs directly convey estimation uncertainty and facilitate visual inference about overlap between groups, which is useful when comparing a moderate number of datasets
-
Cell annotations were determined by consensus of three automatic methods without a stated held-out validation or inter-method agreement metric↳ Could also: Reporting pairwise agreement (e.g., Cohen's kappa or percent agreement) between the three annotation methods could also be included — A quantified agreement metric would allow readers to assess how often the three methods concurred and how much the consensus relies on a majority versus unanimous agreement
-
Batch effects across datasets from different labs and sequencing platforms were addressed through a preprocessing pipeline, but the integration method's performance was assessed qualitatively via UMAP↳ Could also: A quantitative batch-mixing metric such as kBET, LISI, or ASW (average silhouette width) could also be reported before and after integration — Numeric mixing metrics provide an objective, dataset-agnostic measure of how well technical variation was removed relative to biological signal preserved, supplementing visual UMAP inspection
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-39524369
Paper: Balubaid et al. 2024, iScience — "A comparative analysis of blastoid
models through single-cell transcriptomics." DOI 10.1016/j.isci.2024.111122.
Repo: github.com/balubao/Blastoid_scRNAseq_Comparison @ f8315528f9cd05c731aca4e46206eed01d5e37c4
Assigned data: GEO GSE247758 (the authors' OWN founding stem-cell lines).
What the paper/repo produce (pipeline-derived)
The study reprocesses 8 published scRNA-seq blastoid/embryo datasets plus the authors' own 2 founding cell lines, with one shared pipeline: kallisto|bustools requantification → Seurat QC → consensus cell-type annotation (SCINA / scType / SingleR / majority) → Harmony integration → cluster & lineage-module-score analysis.
The repo ships small result tables that equal the paper's Table S1:
02_preprocessing/results/ncells_ngenes_postqc_table.csv— post-QC NCells/NGenes02_preprocessing/results/filtering_thresholds_table.csv— per-dataset thresholds03_annotation/results/samples_by_celltypes_annot_*.csv— cell-type counts ×4 methods
In scope (attempted)
The QC / filtering step is the cleanest, fully-specified, deterministic pipeline output (thresholds are printed in Table S1 and in the QC scripts). Two targets:
- A — Founding cell lines (GSE247758, the assigned dataset). Authors' own
public 10x matrices (CellRanger 6.1.2). Thresholds from
07_src_analysis/source_line_analysis.Rmd:mitoRatio<0.1 & nCount_RNA>20000 & nFeature_RNA>5000. → reproduce post-QC cell counts per sample. - B — Petropoulos2016_ref comparative row (E-MTAB-3929). Published count
matrix + Table S1 thresholds
nFeature_RNA>7500 & mitoRatio<0.2. → reproduce post-QC NCells (reported 1081) and NGenes (reported 29613). This is a third-party-data variant: we apply the documented QC to the published matrix instead of re-running the authors' kallisto requant.
QC metrics re-implemented exactly per Seurat definitions (nCount=colSums, nFeature=#genes>0, mitoRatio=PercentageFeatureSet('^MT-')/100).
Out of scope / not attempted (the hard ~20%)
- Full kallisto|bustools requantification from raw FASTQ for the 8 external datasets. This is what sets the authors' exact NGenes (e.g. 29613 for Petropoulos vs 26178 in the published matrix). Requires tens–hundreds of GB of FASTQ + per-dataset kb-count runs — deliberately skipped per the 80/20 rule.
- Cell-type annotation (SCINA/scType/SingleR), Harmony integration, lineage module scores, DEG/GSEA, all figures — downstream of the requant + heavier and stochastic; not attempted.
Notes affecting comparison (recorded honestly)
- Founding-line post-QC counts are not printed anywhere in the paper or Table S1 → no reported reference for target A (transparency gap); we provide a derived value.
- The published E-MTAB-3929 matrix contains 0 mitochondrial genes, so the
mito<0.2filter is inapplicable for target B (the authors' kallisto reference retained MT genes). This is the main documented reason our NCells runs slightly high vs the reported 1081.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
This is a solid partial reproduction of the paper's QC/cell-filtering step with no fabrication signal: the one printed reference value (Petropoulos NCells=1081) is recovered to ~5% (1137) and the repo CSVs match Table S1 exactly (C5). The deviations sit on the input/preprocessing side and are on our methodology / data-version side — we used the public E-MTAB-3929 matrix (0 MT genes, 26178 genes) instead of the authors' kallisto requant (29613 genes) — not an authors' defect. Severity is low (cell counts same order of magnitude, fully explained), but confirmation is limited because only the deterministic QC step was reproduced and the downstream comparative-transcriptomics conclusions were out of scope; one minor authors'-side transparency gap is that founding-line post-QC counts (2628/1721) are unreported.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.