Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Transcriptomics, regulatory syntax, and enhancer identification in mesoderm-induced ESCs at single-cell resolution.

Cell Rep · 2022
L1 63/100 3/4
⚑ Flagged for review — a reproduced result did not match the reported value

Provisional — an automated or curator check raised a specific concern and points reviewers here. This is NOT a final assessment and not a determination about the authors.

Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +5
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
  • The central claim held under reproduction
What did not (or only partly)
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
63/100
Reproducibility score
0.6 SD below mean
vs. all fields · 1173 studies
🎯 Scores higher than 24% of all assessed papers rank 875 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

PARTIAL — described well enough to reproduce the bulk pipeline; mix of strong reproductions and two honest mismatches. (1) Bulk RNA-seq RPKM reproduced from raw SRA with the NAMED tool TrimGalore -> STAR (mm10/GENCODE vM25) -> featureCounts (SLURM 2214242, 45 min, 70-84% unique): reproduced RPKM correlates with the deposited Partek RPKM at Pearson(log)0.94 / Spearman0.95 for ALL 5 deposited samples in their correct 1:1 biological mapping (Naive->s01-03, Instructed->s04-05) — strong cross-pipeline agreement. (2) DEG re-thresholding (FDR<0.05 & |FC|>=1.5) on the deposited Partek ANOVA reproduces DOWN genes to within one (1289 vs 1288) but gives more UP (2539 vs 2022) — same threshold, so down=exact while up does not match: FLAGGED possible discrepancy / unstated up-filter. (3) Multiome cell count 20000 deposited barcodes vs 19800 reported (~1%, post-QC) = consistent. (4) Deposited MACS peak counts = factual. (5) Enhancer 'instructed-specific' triple-overlap (H3K4me1+/H3K27ac+/ATAC+) gives 4836-10380 vs reported 1482 across defensible interpretations — NOT reproduced because the replicate-merge and specificity definition are under-documented. NOT attempted: myogenic cluster fraction (needs full Seurat WNN); wet-lab/qPCR/FACS claims (out of scope). The Partek DE/RPKM statistics themselves are closed-source; we reproduce the threshold call + an independent RPKM pipeline.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 63
    assessed: 2026-06-22 ⛓ d6d3233122c8
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-22
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-22
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The study tests whether stepwise exposure of mouse ESCs to defined factors (Bmp4, then RDL, then HIFLR media) recapitulates embryonic transitions from naive pluripotency through anterior presomitic mesoderm to myogenic and neurogenic lineages, and whether integrated bulk/single-cell transcriptomic, epigenomic, and chromatin accessibility/conformation data can identify the cis-regulatory elements (including those controlling Pax7) driving these transitions.

Core claims
  • Bmp4 treatment instructs ESCs to downregulate pluripotency genes and upregulate genes associated with formative pluripotency and fate specification finding
  • Chromatin accessibility changes are frequently but not always concordant with gene expression changes, and accessibility can precede transcription finding
  • Single-cell RNA-seq reveals heterogeneous, pseudotemporally distinct transcriptional clusters within naive and instructed ESC populations finding
  • Pax3-GFP+ aPSM cell populations are heterogeneous, containing endothelial- and neurogenic-like subpopulations in addition to bona fide aPSM cells finding
  • Sox2 regulatory element usage shifts from SRR enhancers in naive ESCs to N1/R1 enhancers in aPSM-derived neurogenic precursors, recapitulating embryonic neural plate Sox2 regulation finding
  • Multiome (snRNA-seq + scATAC-seq, WNN) profiling of HIFLR cells identifies a pseudotemporal trajectory rooted in aPSM-like cells that bifurcates into neurogenic and myogenic lineage clusters finding
  • Genomic regions regulating Pax7 expression progressively gain chromatin accessibility through ESC differentiation and are shared with bona fide muscle stem cells and embryonic somites finding
  • The study provides a multi-omic resource for ESC differentiation toward myogenic and neurogenic lineages, including identification of a previously unappreciated Pax7 enhancer resource
Experimental setups
Assay System Perturbation Readout Platform
bulk RNA-seq mouse ESCs (Pax3-GFP line) Bmp4 (10 ng/ml) + 1% KSR, 48h (instructed ESCs) vs naive differentially expressed genes
ATAC-seq naive and instructed mouse ESCs Bmp4 instruction chromatin accessibility at TSS/enhancers
ChIP-seq (H3K4me1, H3K27ac) naive and instructed mouse ESCs Bmp4 instruction active enhancer histone marks
scRNA-seq naive and instructed mouse ESCs Bmp4 instruction single-cell transcriptomic clusters, pseudotime
scRNA-seq + scATAC-seq (integrated) FACS-isolated Pax3-GFP+ aPSM cells (RDL medium, 96h) R-spondin3 + LDN193189 (RDL) cell clusters, chromatin accessibility at pPSM/aPSM/Sox2 loci
snRNA-seq + scATAC-seq multiome (WNN) FACS-isolated Pax3-GFP+ HIFLR cells HGF+IGF-1+FGF-2+LDN193189+Rspo-3 (HIFLR), 48h joint gene expression/chromatin accessibility, pseudotime trajectories, TF motif footprinting
ATAC-seq freshly FACS-isolated MuSCs and E12.5 Pax7-YFP+ somite cells (Pax7-Cre;Rosa26-YFP mice) none (in vivo isolated) chromatin accessibility at Pax7 locus
transcriptome correlation analysis HIFLR cells vs E12.5 Pax7-YFP+ somites none transcriptome correlation (R2)
Key results
  • RNA-seq identified 2,022 upregulated and 1,288 downregulated genes in instructed vs naive ESCs (1.5-fold change, adj. p<0.05) 1.5-fold, padj<0.05
  • 68% (1,385/2,022) of upregulated genes gained increased chromatin accessibility at their TSS 68% (1385/2022)
  • 37% (479/1,288) of downregulated genes displayed reduced chromatin accessibility 37% (479/1288)
  • 1,482 instructed ESC-specific enhancer regions (H3K4me1+/H3K27ac+/ATAC+) identified; 25% associated with increased transcription of nearest gene 25% of 1482 regions
  • Instructed ESC clusters shared ~28% of genes with E8.5/E9.5 neuromesodermal progenitor signatures ~28%
  • Sox2 SRR enhancers (~100 Kb from TSS) accessible in naive ESCs became inaccessible in aPSM cells, while N1 and R1 elements (~10 Kb from TSS) gained accessibility
  • Myogenesis cluster in HIFLR multiome trajectory comprised 1,127 of 19,800 total cells analyzed (5.6%) 5.6% (1127/19800)
  • Transcriptomes of HIFLR cells and E12.5 Pax7-YFP+ somites showed significant correlation R2 > 0.8
Key statistics
  • fold_change 1.5-fold change, adjusted p < 0.05 (RNA-seq DE threshold, instructed vs naive ESCs)
  • count 2,022 upregulated genes (instructed ESCs RNA-seq)
  • count 1,288 downregulated genes (instructed ESCs RNA-seq)
  • other 68% (1385/2022) (upregulated genes gaining TSS accessibility)
  • other 37% (479/1288) (downregulated genes with reduced accessibility)
  • count 1,482 enhancer regions (instructed ESC-specific enhancers (H3K4me1+/H3K27ac+/ATAC+))
  • correlation R2 > 0.8 (HIFLR cells vs E12.5 Pax7-YFP+ somites transcriptome correlation)
  • count 1,127/19,800 (5.6%) (myogenesis cluster size in multiome WNN pseudotime analysis)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The excerpt describes bulk RNA-seq, ATAC-seq, and ChIP-seq profiling of ESCs at successive differentiation stages (naive, instructed, aPSM, HIFLR) alongside single-cell RNA-seq, scATAC-seq, and multiome (snRNA-seq + scATAC-seq) analyses integrated via UMAP and weighted-nearest-neighbor (WNN) clustering and pseudotime ordering. Differential gene expression between conditions was reported using a fold-change and adjusted p-value threshold, and cell populations/lineages were characterized largely through unsupervised clustering, marker-gene expression, and transcription-factor motif enrichment/footprinting rather than through explicitly named formal hypothesis tests in this portion of the text.

Replicationunclear Sample sizeCell cluster sizes are described as counts/proportions of total cells analyzed (e.g., 1,127/19,800, 5.6% of cells in the myogenesis cluster); number of independent biological replicates per condition is not stated in this excerpt GroupsNaive vs. instructed ESCs; ESC-derived aPSM and HIFLR populations vs. FACS-isolated MuSCs and E12.5 Pax7-YFP+ somites; single-cell/nuclei clusters within each population Pairingunclear Randomization/blindingnot stated Dispersionnone Exact p-valuesno Effect sizesyes Confidence intervalsno Multiplicity correctionyes
Statistical tests used
Test Applied to n Assumptions
Differential expression analysis (specific statistical test/tool not named in this excerpt; reported as adjusted p-value with fold-change cutoff) RNA-seq comparison of instructed vs. naive ESCs (Figure 1B; Table S1) not stated
Correlation analysis (R^2) Comparison of HIFLR cell and E12.5 Pax7-YFP+ somite transcriptomes (Figure S4A) not stated
TF DNA-binding motif overrepresentation/footprinting analysis scATAC-seq-derived cell clusters (e.g., root, endothelial, neurogenic, myogenic clusters; Figures 3B–3E, S3B–S3E) not stated
Approaches that could also have been used
  • Differentially expressed genes were defined using a fold-change cutoff (1.5-fold) combined with an adjusted p-value threshold, without reporting exact adjusted p-values or effect-size confidence intervals.
    Could also: Reporting exact adjusted p-values alongside log2 fold-change estimates and their confidence intervals (e.g., via a volcano plot or supplementary table) — This would let readers gauge both the magnitude and the statistical precision of expression differences rather than relying solely on a binary threshold.
  • Similarity between HIFLR cell and somite transcriptomes was summarized with a single R^2 value.
    Could also: Reporting the underlying correlation coefficient (Pearson or Spearman r) together with a confidence interval or p-value — This would convey both the strength and the statistical uncertainty of the association, which R^2 alone does not fully capture.
  • Cell populations and lineages were defined primarily through unsupervised clustering (UMAP, WNN) and pseudotime ordering, described qualitatively via marker gene and motif enrichment.
    Could also: Applying quantitative cluster-stability or classification metrics (e.g., bootstrapped clustering, silhouette scores, or reference-based label transfer statistics) — This would provide a formal measure of confidence in cluster assignments and pseudotime trajectory placement, complementing the visual/qualitative characterization.
  • Comparisons across single-cell clusters (e.g., marker gene expression differences between clusters) appear to be presented descriptively (feature plots, heatmaps) rather than through a stated formal statistical test.
    Could also: Formal pseudobulk differential expression testing between clusters (e.g., aggregating counts per cluster/replicate and applying DESeq2 or edgeR) with multiple-testing correction — This would allow statistical quantification of which marker genes differ significantly between clusters, in addition to the visual/qualitative comparisons shown.
  • TF motif overrepresentation and footprinting in scATAC-seq clusters are described without a stated statistical test or multiple-testing correction across the many motifs/clusters examined.
    Could also: A motif enrichment test (e.g., hypergeometric or Fisher's exact test) with family-wise or FDR correction (e.g., Benjamini-Hochberg) across all tested motifs — Given many motifs and clusters are surveyed in parallel, this would help control for false positives arising from the large number of comparisons.
  • Biological replication (number of independent ESC differentiation experiments, animals, or embryos per condition) is not detailed in this excerpt.
    Could also: Explicitly stating the number of independent biological replicates and, where relevant, using mixed-effects or hierarchical models that account for replicate-to-replicate variation — This would clarify the basis for generalizing findings beyond the specific samples profiled and support assessment of reproducibility.

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — PMID 35977485 (Cell Reports 2022; GSE198730)

"Transcriptomics, regulatory syntax, and enhancer identification in mesoderm-induced ESCs at single-cell resolution." Multi-omic resource: bulk RNA-seq, ATAC-seq, H3K27ac/H3K4me1 ChIP-seq, Hi-C, scRNA-seq, scATAC-seq, 10x Multiome.

Named code (Key Resources / BRIEF)

  • TrimGalore 0.6.6 (github.com/FelixKrueger/TrimGalore) — read trimming. Third-party tool applied to the paper's own data; per project rules this is a fully valid reproduction.
  • Other tools named in Methods (not "the repo"): TopHat 2.1.1 (RNA align, mm10), Bowtie 1.1.1 (ChIP/ATAC), MACS 1.4.2 (peaks, p<1e-5), Partek Genomics Suite 7.18 (RPKM + ANOVA DE), Seurat 4.1.0, Signac 1.5.0, Monocle3, Juicer/HOMER/HiCExplorer (Hi-C).

IN SCOPE (pipeline-derived, reproducible from deposited data)

  1. DEG counts 2022 up / 1288 down (Instructed vs Naive, 1.5-FC, adj p<0.05) — re-apply thresholds to deposited Partek ANOVA table. The ANOVA stats are Partek's (closed); we reproduce the call/threshold logic on the deposited numbers.
  2. Bulk RNA-seq RPKM — reproduce gene-level RPKM from raw SRA with the NAMED tool (TrimGalore) -> STAR (mm10, modern stand-in for deprecated TopHat) -> featureCounts, and correlate vs deposited RPKM table. (Method card: STAR 91% clean, featureCounts 80%.)
  3. Enhancer count 1482 instructed-specific enhancers — bedtools intersect of the deposited H3K27ac/H3K4me1/ATAC MACS peak BEDs (specificity = minus Naive).
  4. Deposited peak counts per assay/condition — direct count (profiling + supports peak claims).
  5. Cell counts (multiome 19800; per-sample scRNA/scATAC) — count deposited barcodes.

PARTIAL / STRETCH

  • Myogenic cluster fraction (1127/19800 = 5.6%) — needs full Seurat WNN re-run of multiome.
  • 68% (1385/2022) genes gained / 37% (479/1288) lost accessibility — DEG-to-ATAC linkage.

OUT OF SCOPE (wet-lab / manual / not a pipeline output)

  • Pax3-GFP FACS %, En7 deletion qPCR (~70% Pax7 reduction), functional myogenic/neurogenic assays — wet-lab. Hi-C TAD/loop calls — large compute, low specificity to a single number.
  • Partek ANOVA statistics themselves are closed-source (we reproduce the threshold call, not the p-value computation).

Data

  • GEO GSE198730 (open). Bulk RNA SRA runs: SRR18364188/189/195/196/204/205/206 (single-end).
  • All compute on «our HPC»; data on «infra»: «path»
RPKM_bulk
Reported
deposited RPKM table (5 samples)
Reproduced
Pearson(log) 0.936-0.942 all 5 cols, correct 1:1 mapping
within tolerance
DEG_down
Reported
1288
Reproduced
1289
within tolerance
DEG_up
Reported
2022
Reproduced
2539
did not match
CELLS_multi
Reported
19800
Reproduced
20000
within tolerance
PEAKS_dep
Reported
deposited
Reproduced
counted per BED (see agreement.json)
exact
ENH_instr
Reported
1482
Reproduced
4836-10380 (interpretation-dependent)
did not match

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 63/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟢7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +5

This is a multi-omic resource paper whose central conclusion reproduces well: bulk RPKM correlates at Pearson(log)≈0.94 across all five deposited samples, DEG_down matches almost exactly (1289 vs 1288), cell counts are ~1% apart, and instructed-specific enhancers do exist. Two flagged mismatches remain: DEG_up is ~25% higher (2539 vs 2022) under the same stated threshold that reproduces down exactly — pointing to an unstated up-filter or an inconsistent reported figure on the authors' side — and the enhancer count (4836–10380 vs 1482) cannot be pinned because the replicate-merge/specificity rule is under-documented, a mix of paper gaps and our interpretation. No fabrication signal; deviations are quantitative and the direction/biology holds, so overall yellow (solid with explainable deviations).

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

373.4 k
tokens (I/O) · 33 M incl. cache
114 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.