Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

A comparative analysis of blastoid models through single-cell transcriptomics.

iScience · 2024
L1 60/100 PQI 88
Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +5
✓ What held up
  • Reported values are derivable from the shared data
  • Any deviation was negligible
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🟡Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
60/100
Reproducibility score
0.8 SD below mean
vs. all fields · 1173 studies
🎯 Scores higher than 21% of all assessed papers rank 918 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

DESCRIBED WELL ENOUGH: yes for the QC step. The paper (Balubaid et al. 2024, iScience) + repo + Table S1 specify the QC filtering thresholds exactly, so the QC/cell-filtering pipeline output is reproducible without ambiguity. We reproduced post-QC cell counts faithfully (Seurat-equivalent metrics in Python) on a «our HPC» SLURM job. RESULT: 1:1-ish PARTIAL, not exact. (1) Petropoulos2016_ref (E-MTAB-3929, the one comparative row with a printed reference): reported 1081 cells, reproduced 1137 on the E5-E7 subset (+5.2%) using the PUBLISHED matrix. The residual gap is fully explained by documented method differences: the authors requantified with kallisto (so their NGenes=29613 vs the matrix's 26178), and the published matrix carries 0 mitochondrial genes so the documented mito<0.2 filter could not remove high-mito cells on our side -> we keep slightly more cells. This is consistent with a GENUINE, non-fabricated number, not a fabrication. (2) Founding cell lines GSE247758 (the dataset assigned to this RU = the authors' OWN data): reproduced post-QC counts 2628 EPSC / 1721 nPSC under the script's documented thresholds (mito<0.1, nCount>20000, nFeature>5000) -- but these are DERIVED values because the paper/Table S1 never print founding-line post-QC counts (a transparency gap, flagged). (3) The repo's shipped result CSVs match Table S1 EXACTLY, confirming the shipped tables are the reported values. NOT ATTEMPTED (the hard ~20%, per 80/20): full kallisto|bustools requantification from raw FASTQ for the 8 external datasets (this is what sets the exact NGenes), plus cell-type annotation, Harmony integration, lineage module scores, DEGs, and all figures. No fabrication signal found.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 60
    assessed: 2026-06-14 ⛓ a2c760fb5eb9
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-14
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The paper investigates how similar pluripotent-stem-cell-derived blastoids are to natural peri-implantation human blastocysts, and what drives transcriptomic and cell-type heterogeneity across different blastoid generation protocols/founding cell lines.

Core claims
  • EPSC-derived blastoids are transcriptomically distinct from nPSC-derived blastoids, with nPSC-blastoids clustering closer to natural blastocysts. finding
  • EPSC-blastoids show a higher proportion of primitive endoderm (PE) cells and ambiguous/unknown cells with endoderm-like signatures. finding
  • Blastoid models can be categorized into three subtypes by cell-type composition: balanced, EPI_ICM-enriched, or PE-enriched. finding
  • PE cells split into two transcriptomically distinct subtypes (PE I and PE II); PE II is enriched in EPSC-based blastoids and bears amnion-like traits. finding
  • Cell-type composition is significantly associated with the source cell type (blastocyst, nPSC, EPSC, fibroblast). finding
  • Gene expression heterogeneity in founding pluripotent cell lines (higher heterogeneity in nPSCs, prevalent amnionic signatures in EPSCs) influences blastoid lineage differentiation outcomes. mechanism
  • A consensus cell-type annotation approach combining three automatic annotation methods was used to label cells across integrated datasets. method
  • A lineage module score framework with Kolmogorov-Smirnov distance was developed to quantify lineage disposition and score-distribution differences between datasets. method
Experimental setups
Assay System Perturbation Readout Platform
scRNAseq (reanalysis/integration of published data) human blastoid models (Yanagida et al., Yu et al. 5iLA, Yu et al. PXGL, Kagawa et al., Liu et al., Sozen et al., Fan et al.) none (comparative computational reanalysis) cell-type composition, cluster distribution
scRNAseq (reanalysis/integration of published data) natural peri-implantation human blastocysts (Petropoulos et al., Xiang et al., and others) none cell-type composition, transcriptomic identity
scRNAseq data integration (UMAP/PCA clustering) combined blastoid + blastocyst datasets, sequenced by 10x Genomics or SMART-Seq2 none shared transcriptomic landscape, cluster overlap 10x Genomics; SMART-Seq2
lineage module score analysis (marker gene expression scoring) integrated blastoid/blastocyst scRNAseq datasets none EPI_ICM, PE, TE module score expression per cell
Kolmogorov-Smirnov distance (KSD) statistical comparison lineage module score distributions across datasets none distributional similarity/difference of module scores between datasets
scRNAseq (newly generated) founding pluripotent stem cell lines: EPSCs and nPSCs none transcriptomic heterogeneity, marker/signature expression (e.g., amnionic signatures)
chi-square association test cell-type annotations vs. dataset source (blastocyst, nPSC, EPSC, fibroblast) none statistical association between primary cell-type composition and source cell type
Key results
  • Significant association between primary cell-type composition (EPI_ICM, PE, TE) and source cell type when only annotated cells included. χ2=4494.5, df=6, p<2.2e-16
  • Significant association between cell-type composition and source cell type when unassigned cells included. χ2=12661, df=9, p<2.2e-16
  • Three cell-type composition clusters identified: balanced (blastocysts + Yanagida et al.), EPI_ICM-enriched (Yu et al. 5iLA/PXGL, Kagawa et al.), and PE-enriched (Liu et al., Fan et al., Sozen et al.).
  • PE cells separate into PE I and PE II subtypes in shared transcriptomic space; PE II particularly enriched in EPSC-based blastoids, with a small population also in Yu2021 blastoids.
  • KLF17 (EPI_ICM marker) is not expressed in Sozen et al. and Fan et al. datasets.
  • TACSTD2 (TE marker) is absent in multiple datasets.
  • Most unidentified/unknown cells represent intermediate states between EPI_ICM and TE (or EPI_ICM and PE), suggesting immature or plastic cell states.
  • Sozen et al. blastoid PE and unknown cells show minimal transcriptomic overlap with other blastocyst/blastoid cells, indicating a distinct transcriptomic profile.
Key statistics
  • pvalue p<2.2e-16 (chi-square association between cell-type composition and source cell type (annotated cells only))
  • other χ2 statistic = 4494.5, df = 6 (chi-square test, annotated cells, cell-type composition vs. source cell type)
  • pvalue p<2.2e-16 (chi-square association between cell-type composition and source cell type (including unassigned cells))
  • other χ2 statistic = 12661, df = 9 (chi-square test, including unassigned cells, cell-type composition vs. source cell type)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This computational study integrates scRNAseq data from seven published blastoid models and reference blastocyst datasets to compare cell-type composition, transcriptomic identity, and lineage module scores. Similarity between datasets is quantified using chi-squared association tests, Jensen-Shannon Distance (JSD), Pearson correlation distance (PCD), and Kolmogorov-Smirnov distance (KSD) with bootstrapping. Results are presented primarily through dimensionality reduction (UMAP, PCA), hierarchical clustering, and heatmaps of pairwise distances rather than inferential hypothesis tests with correction for multiple comparisons.

Replicationunclear Sample sizeSeven blastoid models from published datasets plus reference blastocyst datasets; the authors also generated their own scRNAseq of EPSC and nPSC founding cells; exact cell counts and dataset sizes not stated in the visible text GroupsSeven blastoid models (nPSC-derived, EPSC-derived, fibroblast-derived) vs. natural blastocysts; EPSC-blastoids vs. nPSC-blastoids Pairingna Randomization/blindingnot stated DispersionIQR Exact p-valuesno Effect sizesno Confidence intervalsno Multiplicity correctionnone stated
Statistical tests used
Test Applied to n Assumptions
Chi-squared test of association Association between primary cell-type composition (EPI_ICM, PE, TE) and source cell type (blastocyst, nPSC, EPSC, fibroblast) — run twice: once with annotated cells only, once including unassigned cells χ2 = 4494.5, df = 6 (annotated only); χ2 = 12661, df = 9 (with unassigned) — cell-level counts, exact n not stated not stated
Jensen-Shannon Distance (JSD) Pairwise similarity of cluster distribution across all integrated datasets (Figure 1G) null na
Pearson correlation distance (PCD) Pairwise similarity of cluster distribution across datasets restricted to shared clusters (Figure 1H) null na
Kolmogorov-Smirnov distance (KSD) with bootstrapping Pairwise comparison of lineage module score distributions (EPI_ICM, PE, TE) between datasets (Figures 2C–2E) null not stated
Hierarchical clustering Clustering datasets by cell-type proportions (Figure 1B) and by cluster distribution (Figure 1F) 7 blastoid models plus reference blastocyst datasets na
Consensus of three automatic cell-type annotation methods Cell-type annotation of all integrated single cells (Figures 1A, S3A–S3D) null not stated
Approaches that could also have been used
  • Cell-type composition differences across source groups were assessed with a single chi-squared test on cell counts
    Could also: A compositional analysis framework such as scCODA or a Dirichlet-multinomial regression could also be used, treating each sample/dataset as the unit of observation rather than each cell — Treating individual cells as independent observations inflates the effective sample size because cells from the same sample are correlated; sample-level compositional models account for this clustering and produce estimates with appropriate uncertainty
  • Pairwise dataset similarity was quantified with JSD and PCD on cluster distributions, displayed as heatmaps without formal inference
    Could also: A permutation-based test or PERMANOVA on the cluster-proportion vectors could also be used to attach p-values to the observed distances — Formal inference would allow statements about whether observed inter-group distances exceed what would be expected by chance, complementing the descriptive heatmaps
  • Lineage module score distributions were compared between datasets using KSD with bootstrapping
    Could also: A two-sample Kolmogorov-Smirnov test with Bonferroni or Benjamini-Hochberg correction across all pairwise comparisons could also be used — Reporting adjusted p-values alongside the KSD statistic would clarify which pairwise differences are statistically notable after accounting for the number of comparisons
  • Dispersion in boxplots is shown as the 10th–90th percentile range rather than SD, SEM, or a 95% CI
    Could also: A 95% confidence interval on the median (or mean) could also be displayed, especially for between-dataset comparisons — CIs directly convey estimation uncertainty and facilitate visual inference about overlap between groups, which is useful when comparing a moderate number of datasets
  • Cell annotations were determined by consensus of three automatic methods without a stated held-out validation or inter-method agreement metric
    Could also: Reporting pairwise agreement (e.g., Cohen's kappa or percent agreement) between the three annotation methods could also be included — A quantified agreement metric would allow readers to assess how often the three methods concurred and how much the consensus relies on a majority versus unanimous agreement
  • Batch effects across datasets from different labs and sequencing platforms were addressed through a preprocessing pipeline, but the integration method's performance was assessed qualitatively via UMAP
    Could also: A quantitative batch-mixing metric such as kBET, LISI, or ASW (average silhouette width) could also be reported before and after integration — Numeric mixing metrics provide an objective, dataset-agnostic measure of how well technical variation was removed relative to biological signal preserved, supplementing visual UMAP inspection
Software: not stated in visible text

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
10
Impact: medium
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

5iLA PDBe in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
E-MTAB-3929 ArrayExpress in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE136447 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE150578 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE156596 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE158971 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE17182 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE171820 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE177689 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE178326 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE247758 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-39524369

Paper: Balubaid et al. 2024, iScience — "A comparative analysis of blastoid models through single-cell transcriptomics." DOI 10.1016/j.isci.2024.111122. Repo: github.com/balubao/Blastoid_scRNAseq_Comparison @ f8315528f9cd05c731aca4e46206eed01d5e37c4 Assigned data: GEO GSE247758 (the authors' OWN founding stem-cell lines).

What the paper/repo produce (pipeline-derived)

The study reprocesses 8 published scRNA-seq blastoid/embryo datasets plus the authors' own 2 founding cell lines, with one shared pipeline: kallisto|bustools requantification → Seurat QC → consensus cell-type annotation (SCINA / scType / SingleR / majority) → Harmony integration → cluster & lineage-module-score analysis.

The repo ships small result tables that equal the paper's Table S1:

  • 02_preprocessing/results/ncells_ngenes_postqc_table.csv — post-QC NCells/NGenes
  • 02_preprocessing/results/filtering_thresholds_table.csv — per-dataset thresholds
  • 03_annotation/results/samples_by_celltypes_annot_*.csv — cell-type counts ×4 methods

In scope (attempted)

The QC / filtering step is the cleanest, fully-specified, deterministic pipeline output (thresholds are printed in Table S1 and in the QC scripts). Two targets:

  • A — Founding cell lines (GSE247758, the assigned dataset). Authors' own public 10x matrices (CellRanger 6.1.2). Thresholds from 07_src_analysis/source_line_analysis.Rmd: mitoRatio<0.1 & nCount_RNA>20000 & nFeature_RNA>5000. → reproduce post-QC cell counts per sample.
  • B — Petropoulos2016_ref comparative row (E-MTAB-3929). Published count matrix + Table S1 thresholds nFeature_RNA>7500 & mitoRatio<0.2. → reproduce post-QC NCells (reported 1081) and NGenes (reported 29613). This is a third-party-data variant: we apply the documented QC to the published matrix instead of re-running the authors' kallisto requant.

QC metrics re-implemented exactly per Seurat definitions (nCount=colSums, nFeature=#genes>0, mitoRatio=PercentageFeatureSet('^MT-')/100).

Out of scope / not attempted (the hard ~20%)

  • Full kallisto|bustools requantification from raw FASTQ for the 8 external datasets. This is what sets the authors' exact NGenes (e.g. 29613 for Petropoulos vs 26178 in the published matrix). Requires tens–hundreds of GB of FASTQ + per-dataset kb-count runs — deliberately skipped per the 80/20 rule.
  • Cell-type annotation (SCINA/scType/SingleR), Harmony integration, lineage module scores, DEG/GSEA, all figures — downstream of the requant + heavier and stochastic; not attempted.

Notes affecting comparison (recorded honestly)

  • Founding-line post-QC counts are not printed anywhere in the paper or Table S1 → no reported reference for target A (transparency gap); we provide a derived value.
  • The published E-MTAB-3929 matrix contains 0 mitochondrial genes, so the mito<0.2 filter is inapplicable for target B (the authors' kallisto reference retained MT genes). This is the main documented reason our NCells runs slightly high vs the reported 1081.
Figures / tables: Table
C1
Reported
Petropoulos2016_ref post-QC NCells = 1081
Reproduced
1137 (E5-E7 subset) / 1374 (all 1529 cells)
partial
C2
Reported
Petropoulos2016_ref post-QC NGenes = 29613
Reproduced
24288-24427 detected (26178 total genes in published E-MTAB-3929 matrix)
partial
C3
Reported
EPSCs (GSM7900397) post-QC NCells = NOT REPORTED in paper/Table S1
Reproduced
2628 (of 4105 CellRanger-called)
partial
C4
Reported
nPSCs (GSM7900398) post-QC NCells = NOT REPORTED in paper/Table S1
Reproduced
1721 (of 3280 CellRanger-called)
partial
C5
Reported
Repo shipped result CSVs vs paper Table S1 (mmc2.xlsx)
Reproduced
Identical for all 10 comparative datasets (NCells, NGenes, thresholds, annotation counts)
exact

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 60/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟡2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟢5. Derivability / plausibility
🟢6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +5

This is a solid partial reproduction of the paper's QC/cell-filtering step with no fabrication signal: the one printed reference value (Petropoulos NCells=1081) is recovered to ~5% (1137) and the repo CSVs match Table S1 exactly (C5). The deviations sit on the input/preprocessing side and are on our methodology / data-version side — we used the public E-MTAB-3929 matrix (0 MT genes, 26178 genes) instead of the authors' kallisto requant (29613 genes) — not an authors' defect. Severity is low (cell counts same order of magnitude, fully explained), but confirmation is limited because only the deterministic QC step was reproduced and the downstream comparative-transcriptomics conclusions were out of scope; one minor authors'-side transparency gap is that founding-line post-QC counts (2628/1721) are unreported.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

160.4 k
tokens (I/O) · 9.7 M incl. cache
16 min
runtime · 0 CPU-h
1.5 GB
peak RAM
1
HPC jobs
hummel
machine