Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Pre-implantation embryo metabolism identified by PEMA reveals endogenous lactate insufficiency contributes to pre-implantation development arrest.

Fundam Res · 2024
not yet assessed 3/4
Why this verdict

Part of the results reproduced; minor but material deviations remained.

Reproduced on the brainbox compute brainarbeit.com
✓ What held up
  • Same input data as the authors
What did not (or only partly)
  • 🔴Reported values were only indirectly comparable
  • 🔴A deviation arose in the data or preprocessing
  • 🔴A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
Reproduction agent’s raw note

DROP / env_unresolvable, reached honestly at the solver-license wall. PEMA is a thin wrapper around COMPASS (YosefLab/Compass) on the Recon2 genome-scale metabolic model plus a deterministic R post-processing with Ribo-seq reweighting. Both the code (PEMA + COMPASS, cloned & installed on «our HPC»/«infra») and the data (public processed RNA+Ribo FPKM matrices: GSE165782 mouse, GSE197265 human, downloaded) were obtained successfully — so this is NOT a code- or data-availability gap. The entire in-scope numeric claim set (Fig 1C/1D/1F/1G) routes through COMPASS, which hard-requires IBM CPLEX to solve the Recon2 LP (4849 metabolite constraints, ~7400 reactions). The only obtainable solver, pip CPLEX Community Edition (v22.2.0.0, 1000-constraint cap), fails at the first solve with 'CPLEX Error 1016: Community Edition. Problem size limits exceeded.' (evidence: outputs/COMPASS_cplex1016_error.txt, «job»). No licensed CPLEX/Gurobi module, system install, env var or license file exists on «our HPC» (verified), and a full IBM academic CPLEX license cannot be obtained autonomously inside a batch job. Fig 1B additionally needs benchmark metabolomics (ref [5]) that is not under a resolvable public accession. Work done & verified: cleared a COMPASS multiprocessing bug (SwigPyObject pickle) with --num-processes 1; COMPASS parses the FPKM input (mouse 20987x15), builds the Recon2 model, and reaches the solve stage before the license wall. NOT attempted: all wet-lab results (embryo culture, microinjection, IF/H3K18lac, Lac-CoA rescue, CUT&Tag) and HRA006017 controlled-access human PIDA data — out of scope. No value was fabricated. This is a license/packaging reproducibility barrier: a nominally 'available' tool whose advertised deps omit that its core solver is a paid commercial product, with no bundled/open solver, no example I/O, and a non-runnable correlation.R (undefined 'aaaaa').

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment
    assessed: 2026-06-14 ⛓ d218cdf63a0d
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-14
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

The paper asks how abnormal metabolism contributes to non-genetic pre-implantation development arrest (PIDA), hypothesizing that insufficiency of endogenous lactate during major zygotic genome activation (ZGA) drives human embryonic arrest, and develops a computational tool (PEMA) to characterize embryo metabolism.

Core claims
  • PEMA, a Ribo-seq-weighted computational framework, characterizes metabolic states of human and mouse pre-implantation embryos more accurately than Compass method
  • Pre-implantation embryos exhibit high lactate dehydrogenase (LDH) activity and lactate synthesis when major ZGA occurs in humans and mice finding
  • Human 8-cell PIDA and corrected tripronuclear (ch3PN) embryos display insufficiency of endogenous lactate, coordinated with failed H3K18lac and major ZGA finding
  • Human PIDA embryos transcriptionally resemble endogenous lactate-deprived mouse embryos finding
  • Lac-CoA (lactyl-CoA) supplementation promotes H3K18lac, major ZGA, and ch3PN pre-implantation development mechanism
  • Lactate-derived H3K18lac is enriched on promoter regions of major ZGA / 8-cell signature genes and correlates with their expression mechanism
  • PEMA is publicly available as a resource at https://github.com/summus-kong/PEMA resource
Experimental setups
Assay System Perturbation Readout Platform
Computational metabolic modeling (PEMA framework, FBA-based with Ribo-seq weighting) Human and mouse pre-implantation embryos (zygote, 2-cell, 4-cell, 8-cell, blastocyst) none Predicted metabolic reaction scores (e.g., glycolysis, LDH activity, lactate synthesis) Recon2 database / GSMM; RNA-seq + Ribo-seq inputs (GSE165782, GSE209648)
Single-embryo RNA-seq Human 8-cell PIDA embryos (n=18), ch3PN embryos (n=5); mouse embryos PIDA / ch3PN; lactate deprivation in mouse Transcriptome / differential gene expression, GSEA of major ZGA genes Illumina NovaSeq 6000; TruePrep DNA Library Prep Kit
CUT&Tag / single-cell CUT&Tag Human 8-cell & 16-cell PIDA (5 + 2) and ch3PN (7 + 6) embryos; mouse embryos PIDA / ch3PN H3K18lac genome-wide signal peaks Hyperactive Universal CUT&Tag Assay Kit (Vazyme TD903); Illumina NovaSeq 6000
Immunofluorescence staining Human (8-cell, 16-cell PIDA n=2 each; ch3PN n=2) and mouse embryos PIDA / ch3PN / inhibitor H3K18lac and LDH protein levels Leica TCS SP8 confocal; anti-H3K18lac PTM-1406RM, anti-LDH ab52488
Genetically-encoded fluorescent lactate biosensor (FiLa) imaging Human ch3PN embryos (n=2) and mouse zygotes none / lactate sensing Intracellular lactate levels pAAV-CMV-MCS-FiLa, pLVX-Nuc-FiLa mRNA microinjection
Inhibitor treatment / drug perturbation Mouse pre-implantation embryos in modified KSOM GNE-140 (5 µM), GSK2837808A (1 µM) LDH inhibitors; NMN (50 µM) Developmental arrest, ZGA, H3K18lac MedChemExpress reagents
Lac-CoA rescue microinjection Human ch3PN zygotes (53 used for development statistics) 1 mM Lacetyl-CoA microinjection vs water control H3K18lac, major ZGA, pre-implantation development rate Eppendorf FemtoJet 4i microinjector
Key results
  • PEMA predicted metabolic activity correlated more strongly with measured mouse metabolomic profiles than Compass
  • Embryos show high LDH activity and lactate synthesis at major ZGA in humans and mice
  • Human PIDA and ch3PN embryos show insufficiency of endogenous lactate
  • Lactate synthesis and LDH activity reaction scores correlate with major ZGA (mouse) / 8-cell signature (human) gene expression more than other coding genes
  • Lac-CoA addition promoted H3K18lac, major ZGA, and improved ch3PN pre-implantation development
  • LDH inhibition (GNE-140/GSK2837808A) causes mouse 2-cell developmental arrest, major ZGA failure, and loss of H3K18lac (prior/related finding)
Key statistics
  • count ~10% of human IVF/ICSI embryos arrested at cleavage stages (prevalence of pre-implantation arrest)
  • count 18 8-cell PIDA embryos for single-embryo RNA-seq (human PIDA RNA-seq sample)
  • count 53 ch3PN embryos for embryo development statistics (ch3PN development quantification)
  • pvalue P < 0.01; P < 0.001 (Student's t-test) (glycolysis reaction score dynamics and metabolic activity comparisons)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The paper introduces a custom computational framework (PEMA) that weights metabolic reaction scores from RNA-seq and Ribo-seq data, validated against a reference metabolome via Pearson correlation. Differential gene expression between embryo groups was assessed with DESeq2, pathway-level activity with GSEA, and pairwise comparisons of metabolic reaction scores across developmental stages with Student's t-test. Results are summarized with SEM error bars and significance-threshold annotations; the paper text provided is truncated before the clustering and later analytical sections.

Replicationbiological Sample sizePer-assay embryo counts stated explicitly (e.g., 18 8-cell PIDA embryos for RNA-seq, 53 ch3PN embryos for developmental outcome statistics, 2 embryos per condition for immunofluorescence); no formal power analysis is mentioned GroupsPIDA (8-cell/16-cell) vs. normally developing human embryos; ch3PN vs. normal human embryos; LDH/lactate-inhibitor-treated vs. control mouse embryos; Lac-CoA-injected vs. water-injected ch3PN embryos Pairingunpaired Randomization/blindingnot stated DispersionSEM Exact p-valuesno Effect sizesno Confidence intervalsno Multiplicity correctionDESeq2 applies Benjamini-Hochberg FDR adjustment by default; no correction strategy is explicitly described for the Student's t-tests or GSEA nominal p-values
Statistical tests used
Test Applied to n Assumptions
Student's t-test (tail direction not stated) Comparison of mean metabolic reaction scores (glycolysis, LDH activity, lactate synthesis) across pre-implantation developmental stages in mouse and human embryos (Figs. 1C–E) not stated
DESeq2 Wald test Identification of differentially expressed genes between groups (e.g., 8-cell PIDA vs. normal, ch3PN vs. normal, inhibitor-treated vs. control mouse embryos) 18 8-cell PIDA embryos and 5 8-cell ch3PN embryos used for single-embryo RNA-seq not stated
Gene Set Enrichment Analysis (GSEA) via R/clusterProfiler; reported as normalized enrichment scores and nominal p-values Enrichment of a published major ZGA gene set in transcriptomes of PIDA or lactate-deprived embryos not stated
Pearson correlation Validation of PEMA against reference metabolomics profiles (Fig. 1B); correlation of LDH activity and lactate synthesis reaction scores with individual gene expression in mouse 2-cell and human 8-cell embryos (Fig. 1G) not stated
Approaches that could also have been used
  • Pairwise comparisons of metabolic reaction scores across developmental stages were performed with Student's t-test
    Could also: A non-parametric Mann-Whitney U (Wilcoxon rank-sum) test could also be applied — With the small per-stage embryo counts typical in this domain, normality of the underlying distribution is difficult to assess; a rank-based test requires no distributional assumption and is commonly recommended when n per group is fewer than ~10–15
  • Multiple pairwise t-tests are applied across developmental stage comparisons without an explicitly stated correction for multiple comparisons
    Could also: A one-way ANOVA followed by a post-hoc procedure (e.g., Tukey HSD or Dunnett's test) or a Benjamini-Hochberg FDR correction applied across the family of t-tests could also be used — When many pairwise comparisons share a common data set, a family-wise error rate or FDR approach controls the expected rate of false positives across the full set of tests; reporting which correction (if any) was applied helps readers interpret the significance annotations
  • Dispersion around group means is summarized exclusively as SEM
    Could also: SD or 95% bootstrap confidence intervals could also be used to describe spread — SEM quantifies precision of the mean estimate and decreases as n grows, which can make biological variability appear smaller than it is; SD or a CI conveys the actual spread among individual embryos and is often preferred when sample sizes are small
  • P-values are reported as categorical threshold annotations only (**, ***) rather than exact numeric values
    Could also: Exact p-values (e.g., p = 0.004) could also be reported alongside or instead of asterisk notation — Exact p-values allow readers to apply their own decision thresholds and facilitate meta-analytic reuse; they carry more information than categorical bins while requiring no additional space
  • No standardized effect sizes accompany the significance tests
    Could also: Cohen's d (for t-tests) or log2 fold-change with 95% confidence intervals (for DESeq2 results) could also be reported — Effect sizes express the magnitude of differences independently of sample size, complement p-values by indicating practical relevance, and provide the input needed for future power calculations in this clinically important domain
  • Developmental outcome in the Lac-CoA rescue experiment is assessed across 53 ch3PN embryos (a proportion reaching each stage)
    Could also: A Fisher's exact test or chi-squared test of proportions with a reported odds ratio and 95% CI could also be applied — When the outcome variable is categorical (e.g., proportion reaching blastocyst vs. arrested), tests designed specifically for count/proportion data are the standard approach and directly yield an effect size estimate (odds ratio) with uncertainty bounds that are interpretable in an IVF clinical context
Software: DESeq2 (R package) · clusterProfiler (R package) · HISAT2 2.2.1 · featureCounts · fastp · MACS2 2.2.7.1 · Bowtie2 2.5.0 · deepTools / bamCoverage · ChIPseeker (R package) · scploid (R package) · KOBAS (online tool) · PEMA (custom tool, available on GitHub)

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
0
Impact: low
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

GSE101571 GEO in Data Availability (http://purl.obolibrary.org/obo/IAO_0000611)
no other assessed paper uses this yet
GSE165782 GEO in Data Availability (http://purl.obolibrary.org/obo/IAO_0000611)
no other assessed paper uses this yet
GSE197265 GEO in Data Availability (http://purl.obolibrary.org/obo/IAO_0000611)
no other assessed paper uses this yet
GSE209648 GEO in Data Availability (http://purl.obolibrary.org/obo/IAO_0000611)
no other assessed paper uses this yet
GSE234027 GEO in Data Availability (http://purl.obolibrary.org/obo/IAO_0000611)
no other assessed paper uses this yet
GSE36552 GEO in Data Availability (http://purl.obolibrary.org/obo/IAO_0000611)
no other assessed paper uses this yet

Downstream reach in the literature

116 downstream papers · 6 datasets

How widely the datasets deposited by this paper are reused across the whole literature (Europe PMC), beyond our assessed set. This is a factual dependency map — reusing a public dataset is normal, good science. It is not a judgement on the downstream papers; the only verdict here is this paper's own, with its cited rationale.

GSE209648 GEO reused by 1 papers in the literature
GSE234027 GEO reused by 1 papers in the literature

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-42272466 (PEMA)

Paper: Li et al., Pre-implantation embryo metabolism identified by PEMA reveals endogenous lactate insufficiency contributes to pre-implantation development arrest. Fundam Res 2024. DOI 10.1016/j.fmre.2024.10.005. Code: https://github.com/summus-kong/PEMA (commit 3cb96db, pushed 2024-08-18; no license).

What PEMA actually is (read from the repo)

PEMA is a thin wrapper around COMPASS (YosefLab/Compass, Wagner et al. Cell 2021) plus a deterministic R post-processing step:

  1. PEMA (python) calls: compass --data <rna_fpkm.tsv> --species <homo_sapiens|mus_musculus> --output-dir <out> --num-threads <n> → COMPASS produces per-reaction penalties (reactions.tsv) from the RNA-seq FPKM matrix using the Recon2 genome-scale metabolic model (GSMM).
  2. calculate.R / calculateal.R convert penalties → reaction-consistency scores: get_reaction_consistencies(): score = -log(penalty + 1), drop rows with range ≤ 1e-3, subtract global min. (This is verbatim COMPASS's standard post-processing.) Then weight each reaction by mean Ribo-seq translation level of its associated genes (from data/reaction_metadata.csv, the Recon2 reaction→gene map), per developmental stage → final reaction-score matrix R.
  3. correlation.R correlates LDH/lactate-synthesis reaction scores against gene expression (Fig 1G). NB: ships with an undefined input object aaaaa; not runnable as distributed — must reconstruct the input from steps 1–2.

The novelty over plain COMPASS is the Ribo-seq reweighting (step 2). The repo ships only reaction_metadata.csv (Recon2 map) — no example input, no expected output, no parameter file.

Data (publicly resolvable — processed matrices ship on GEO)

  • Mouse (framework dev): GSE165782 + GSE209648. GSE165782 ships processed FPKM incl. GSE165782_E5_Btg4_ribo_RNA_fpkm.txt.gz (RNA + Ribo FPKM). → PEMA input ready.
  • Human (Fig 1E/F): GSE197265 ships GSE197265_1C_4C_DMSO_CHX_merge_fpkm.txt.gz (RNA + Ribo FPKM). → PEMA input ready.
  • Other accessions (GSE101571, GSE36552, HRA003366, GSE234027) for downstream figs.
  • HRA006017 (GSA-Human): scRNA/scCUT&Tag of human PIDA embryos generated in this study — controlled-access GSA-Human, not attempted.

IN SCOPE (pipeline-derived, attempted)

id claim fig/loc pipeline
C1 PEMA reaction scores for glycolysis high in mouse 2-cell & blastocyst; LDH activity & lactate synthesis high in 2-cell Fig 1C/1D COMPASS+postproc on GSE165782 mouse FPKM
C2 LDH activity & lactate-synthesis scores high at human 8-cell Fig 1F COMPASS+postproc on GSE197265 human FPKM
C3 LDH/lactate-synthesis reaction scores more strongly correlated with major-ZGA genes (mouse) / 8-cell-signature genes (human) than with other coding genes (P<0.001) Fig 1G correlation.R on PEMA scores
C4 Across 29 metabolic pathways, PEMA correlates better than COMPASS with the metabolomic benchmark in 22/29 Fig 1B needs benchmark metabolomics (ref [5])

OUT OF SCOPE (not pipeline-reproducible / not attempted)

  • All wet-lab results: embryo culture, microinjection, IF/H3K18lac intensities, Lac-CoA rescue, CUT&Tag heatmaps (Figs 2–5 wet parts) — manual/experimental.
  • C4 benchmark (Fig 1B/S1B): the "previously published metabolomic profiles" (ref [5]) used as ground truth are not deposited under a resolvable accession → cannot compute the 22/29 PEMA-vs-COMPASS comparison. Recorded, not attempted.
  • HRA006017 human PIDA scRNA/scCUT&Tag (controlled-access, this-study data).
  • Dimensionality-reduction/clustering figures (qualitative, no pinnable number).

Feasibility gate

The entire in-scope set routes through COMPASS, which hard-requires IBM CPLEX (cplex>=12.7, installed separately; community edition caps at 1000 constraints, Recon2 has ~7400 reactions). Whether the in-scope claims can be reproduced at a

Figures / tables: Fig 1CFig 1FTableFig 1GFig 1BFig S1B
C1
Reported
glycolysis & LDH/lactate-synthesis reaction scores high in mouse 2-cell & blastocyst (Fig 1C/D)
Reproduced
NOT PRODUCED — COMPASS LP solve blocked by CPLEX Error 1016
partial
C2
Reported
LDH/lactate-synthesis reaction scores high at human 8-cell (Fig 1F)
Reproduced
NOT PRODUCED — same CPLEX blocker
partial
C3
Reported
LDH/lactate scores correlate more with major-ZGA (mouse)/8-cell-signature (human) genes than other coding genes, P<0.001 (Fig 1G)
Reproduced
NOT PRODUCED — needs PEMA scores; correlation.R also ships referencing undefined object 'aaaaa'
partial
C4
Reported
PEMA beats COMPASS in correlation with metabolomic benchmark in 22/29 pathways (Fig 1B)
Reproduced
NOT ATTEMPTED — benchmark metabolomics (ref [5]) not under a resolvable public accession
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 38/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🔴2. Endpoint comparability
🔴3. Location of the main deviation
🔴4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

157.5 k
tokens (I/O) · 16.1 M incl. cache
19 min
runtime · 0.02 CPU-h
2 GB
peak RAM
3
HPC jobs
hummel
machine