Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Developmental hematopoietic stem cell variation explains clonal hematopoiesis later in life.

Nat Commun · 2024
L1 90/100 PQI 93
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score 0
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
  • No relevant deviation in data/preprocessing
  • No authors-side cause for any deviation
  • Reported values are derivable from the shared data
  • Any deviation was negligible
  • The central claim held under reproduction
What did not (or only partly)
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
90/100
Reproducibility score
0.9 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 79% of all assessed papers rank 211 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough to reproduce 1:1 (authors' own Julia code @ b1dc0e7, self-contained, data bundled). All clearly-specified pipeline-derived numbers reproduced exact or within rounding: Fig 1 decline slopes (-0.0020/-0.0033/-0.0025 per year) and birth-correlation intercepts (0.87/0.72/0.61) via deterministic LsqFit on shipped data; Fig 4 weak<neutral<strong model-vs-data distance ordering (weak best fit) recomputed exactly from shipped simulation files; Fig 5 cached simulated MZ birth Pearson 0.877. An INDEPENDENT fresh stochastic re-run of the development model recovered the empirical group birth correlations (MZ 0.882 / DZ 0.729 / UR 0.569) from the fitted N_split values, confirming the central mechanism. NOT attempted (the hard ~20%): (1) the raw-GEO-beta-matrix -> Pearson preprocessing in the front half of data_processing.ipynb, because its bespoke input files (GSE36642_MZ.csv, names.csv, GSE154566_..._betas.csv) are NOT shipped in the repo; (2) regenerating the full 9x100 lifespan Markov simulations (heavy 36500-step x 5000-cell), for which the authors' cached outputs were used and validated. Status 'partial' reflects that reproduction is downstream of the shipped intermediate Pearson values rather than end-to-end from raw GEO; every value actually compared matched. No fabrication detected; shipped data is internally consistent and matches reported figures.

💻 Code ↗ 🗄 Data: GSE36642

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 90
    assessed: 2026-06-14 ⛓ 805575384d47
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-14
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

HSC variation inevitably arises during development (before birth), conferring subtle weak selective advantages that are protected from stochastic loss by developmental expansion and eventually lead to clonal hematopoiesis later in life; this is testable using monozygotic twins sharing prenatal circulation and fluctuating CpG methylation lineage markers.

Core claims
  • Weak selection conferred by stem cell variation created before birth can reliably yield clonal hematopoiesis later in life. finding
  • Blood fluctuating CpG methylation patterns show low correlation between unrelated individuals but are highly correlated between many elderly monozygotic twins, supporting dominance of clones created before birth. finding
  • Stochastic loss of weakly selected variants is naturally prevented by the expansion of stem cell lineages during development. mechanism
  • A single-cell-resolved mathematical model of HSC methylation dynamics (fluctuating methylation clock) tracks clonal evolution from embryonic development through a 100-year lifespan. method
  • Frequency-dependent (rather than uniform) clonal growth during development is required to reproduce the observed clonal variation between individuals at birth. finding
  • Fluctuating CpG methylation serves as a high-error-rate lineage marker that becomes polymorphic before birth, enabling lineage tracing across life. method
  • Clonal relatedness (Pearson correlation of methylation profiles) declines slowly with age and more sharply after 65 years, consistent with loss of HSC clonal diversity. finding
  • Variants with weak selective advantages become dominant given large population size, turnover, and enough time — conditions satisfied by human hematopoiesis. mechanism
Experimental setups
Assay System Perturbation Readout Platform
DNA methylation array (fluctuating CpG site profiling) Human HSCs / blood from monozygotic twins, dizygotic twins, and unrelated individuals (prenatal to 86 years) none Pearson correlation of average methylation profiles between individuals over time
Single-cell single-site stochastic Markov model simulation (fluctuating methylation clock) Simulated HSC population (N_cell=5×10^3 cells, N_sites=10^3 fCpG sites per cell) varied parameters: cell replacement rate α, (de)methylation rate γ, N_split, N_clones, fitness coefficients, growth mode (uniform vs frequency-dependent) Clonal dynamics, β/methylation distributions, Pearson correlation coefficients
Key results
  • Mean Pearson coefficient at birth highest for MZ twins, intermediate for DZ twins, lowest for unrelated individuals MZ ~0.87, DZ ~0.72, unrelated ~0.61
  • Clonal relatedness declines with age; slope of decline differs by group DZ slope −0.0033, unrelated −0.0025, MZ −0.0020
  • Elderly MZ twins show wide range of HSC clonal relatedness, with some pairs highly correlated in later life Pearson ~0.3 to >0.9
  • Uniform growth during development cannot reproduce observed birth variation (profiles too similar) Pearson ≥0.95 (uniform; e.g. 0.99)
  • Frequency-dependent growth produces clonal variation matching data at birth initial Pearson 0.87
  • Pearson coefficients decrease more sharply after 65 years of age
  • Approximately two-thirds of MZ twins share circulation during development (monochorionic placentas) 2/3
Key statistics
  • correlation ~0.87 (mean Pearson at birth, MZ twins) (MZ twin methylation profile similarity at birth)
  • correlation ~0.72 (mean Pearson at birth, DZ twins) (DZ twin methylation profile similarity at birth)
  • correlation ~0.61 (mean Pearson at birth, unrelated) (unrelated individuals methylation similarity at birth)
  • other slopes −0.0033 (DZ), −0.0025 (unrelated), −0.0020 (MZ) (slope of decline in Pearson correlation over age)
  • correlation 0.99 (uniform growth model) vs 0.87 (frequency-dependent model) (simulated initial Pearson at birth, Fig. 3)
  • count 10^3 simulations per parameter set (independent model simulations for parameter estimation)
  • other 2/3 (fraction of MZ twins sharing circulation during development)
  • count N_cells = 5×10^3; N_sites = 10^3 (model HSC population size and fCpG sites per cell)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The paper combines analysis of nine publicly available DNA methylation datasets (spanning prenatal to age 86, covering MZ twins, DZ twins, and unrelated individuals) with a stochastic single-cell mathematical model of HSC clonal dynamics. Pairwise similarity of individual methylation profiles was quantified by Pearson correlation coefficients; trends over the lifespan were summarized by lines of best fit with 90% confidence intervals through dataset-level means. Model parameters were estimated by matching simulation means (±1 SD over 10³ replicate runs) to observed Pearson coefficients at birth; no classical hypothesis tests with p-values are reported.

Replicationbiological Sample sizeSample sizes reported per dataset in Table 1 (17–426 MZ twin pairs, 8–306 DZ twin pairs); no formal a priori power analysis described GroupsMZ twins vs. DZ twins vs. unrelated individuals, pooled across nine age-stratified public methylation datasets Pairingpaired Randomization/blindingnot stated Dispersionmixed Exact p-valuesno Confidence intervalsyes
Statistical tests used
Test Applied to n Assumptions
Pearson correlation coefficient Pairwise comparison of average blood methylation profiles between MZ twins, DZ twins, and unrelated individuals across all datasets (Fig. 1C–F) Varies by dataset: 17–426 MZ twin pairs, 8–306 DZ twin pairs (Table 1) not stated
Linear regression (line of best fit) Age trend of mean Pearson coefficients for MZ, DZ, and unrelated groups; slopes reported numerically (−0.0020, −0.0033, −0.0025 respectively) (Fig. 1F) Dataset-level means used as observations (9 datasets, varying time points) not stated
Simulation-based parameter estimation (grid/visual matching) Fitting N_clones and N_split to observed initial Pearson coefficients at birth for all three groups (Fig. 2C) 10³ independent stochastic simulations per parameter combination not stated
Continuous-time Markov chain / stochastic process simulation Single-cell model of fCpG methylation and clonal dynamics across a simulated 100-year lifespan (Fig. 2A, Table 2) N_cells = 5×10³ simulated cells; N_sites = 10³ fCpG sites per cell per simulation run stated
Approaches that could also have been used
  • Pairwise methylation profile similarity was quantified with the Pearson correlation coefficient on bounded beta values
    Could also: Spearman rank correlation or Jensen-Shannon divergence could also quantify similarity between methylation beta-value distributions — Spearman is robust to non-normality and outliers common in bounded [0,1] data; Jensen-Shannon divergence directly compares full distributions rather than linear association, which may better capture the multimodal beta distributions described in the paper
  • The age trend of mean Pearson coefficients was estimated by a single line of best fit through dataset-level means, treating each dataset as one data point
    Could also: A linear mixed-effects model with individual pairs as observations and dataset as a random effect could also estimate the age slope — Mixed-effects models account for within-dataset clustering and differences in dataset size when pooling nine datasets, and would yield confidence intervals for the slope that reflect between-individual rather than between-dataset variability
  • Model parameters N_clones and N_split were estimated by visually matching simulation means to observed Pearson coefficients at birth
    Could also: Approximate Bayesian Computation (ABC) or a likelihood-based grid search with a formal goodness-of-fit criterion could also be used for parameter estimation — Formal inference methods would yield posterior distributions or confidence sets for the estimated parameters, making uncertainty in N_clones and N_split explicit and enabling sensitivity analysis
  • Variability in simulation outputs was reported as ±1 SD across 10³ replicate runs
    Could also: Percentile-based simulation intervals (e.g., 5th–95th or 2.5th–97.5th percentile) could also characterize simulation variability — Percentile intervals communicate the full shape of the simulation output distribution — informative when that distribution may be skewed — and would align with the 90% CIs used for the empirical lines of best fit elsewhere in the paper
  • Fitness coefficients for HSC clones were drawn from a Gamma distribution parameterized from the literature
    Could also: Exponential, log-normal, or truncated normal distributions could also parameterize fitness effect distributions, and sensitivity analyses across these choices could be reported — The shape of the fitness distribution governs the probability of generating clones with large selective advantages; demonstrating that key conclusions hold across plausible alternative distributions would strengthen the robustness of the modeling findings
  • Differences in Pearson coefficient level and slope between groups (MZ, DZ, unrelated) were described narratively without formal inference
    Could also: Permutation tests or bootstrap-based comparisons of group means or slopes at specified age windows could also be applied — Formal tests would distinguish age-related changes in inter-individual methylation correlation from sampling variability, providing statistical quantification of the observed between-group differences in trajectory

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
8
Impact: low
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

GSE61496 GEO in Results (http://purl.org/orb/Results)
also used by 1 paper:
GSE100227 GEO in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GSE105018 GEO in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GSE154566 GEO in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GSE36642 GEO in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GSE43975 GEO in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GSE73115 GEO in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GSE89093 GEO in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-39592593

Paper: Kreger J, Mooney JA, Shibata D, MacLean AL. Developmental hematopoietic stem cell variation explains clonal hematopoiesis later in life. Nat Commun 2024. DOI 10.1038/s41467-024-54711-2 · PMCID PMC11599844.

Repo: https://github.com/maclean-lab/scFMC-model @ commit b1dc0e7cb08724813fd0be0bd5a6a2f7f02636fe (authors' own code — Julia ≥1.8.1, two Jupyter notebooks).

Nature of the study

A stochastic computational-modeling paper. The "scFMC" model (single-cell resolved fluctuating-CpG-methylation) is a discrete Markovian simulation of HSC clonal dynamics from embryonic development through a 100-year lifespan. There are no wet-lab experiments; the only empirical inputs are public methylation (fCpG) beta-value matrices from 8 GEO series (Table 1) used to compute Pearson correlations between twins / unrelated individuals at several ages.

What the repo actually ships

  • data/*_data_all.csv, data/*_data_mean_all.csvalready-computed per-pair Pearson correlations (and per-age group means) for MZ twins, DZ twins, unrelated (ur). These are the outputs of data_processing.ipynb.
  • simulation files/pearson_cells_all_sims_{mz,dz,ur}_{neutral,weak,strong}.csvcached simulation trajectories (101 ages × 100 sims) for 3 selection regimes.
  • simulation files/{mz,dz,ur}_data_trajectory.csv — bootstrapped data trajectories (6 ages × 100) used as the comparison target.
  • data/TmobT.csv (fCpG site list), data/data.xlsx.

IN SCOPE (pipeline-derived, reproducible from shipped artifacts/code)

id result (figure) how reproduced confidence
C1 Empirical Pearson decline slope per group, Fig 1 re-run notebook cells 8–16: LsqFit.curve_fit(linear) on shipped *_data_mean_all.csv high (deterministic)
C2 Birth-Pearson regression intercept per group (the 0.855/0.754/0.597 constants hard-coded in cells 39–41), Fig 1/2 linear-fit intercept from same fit high (deterministic)
C3 Model-vs-data distance, neutral/weak/strong selection (MZ), Fig 4 / Table S2 re-run cells 72–75 on shipped simulation files/*.csv → weak should be smallest high (deterministic)
C4 Cached weak-selection simulated birth Pearson mean (MZ/DZ/UR), Fig 5 re-run cell 67 mean over 100 cached sims high (deterministic)
C5 Fresh stochastic re-run of the development model: birth Pearson vs N_split (=36/15/10 → MZ/DZ/UR), Fig 2/3 faithful Julia port of cells 22–47, seeded, 20 realizations/group medium (stochastic; tests mechanism + ordering)

OUT OF SCOPE (the hard ~20% — not attempted, with reason)

  • Raw-GEO → Pearson step (front half of data_processing.ipynb): the notebook reads bespoke per-series beta matrices (GSE36642_MZ.csv, names.csv, GSE154566_..._betas.csv, …) and sample-pairing files that are NOT shipped in the repo and would each need separate GEO download + bespoke preprocessing. We therefore take the shipped per-pair Pearson CSVs as the pipeline input and do not re-derive them from raw methylation. (Possible-fabrication check: shipped birth means are internally consistent with the hard-coded intercept constants — see C2.)
  • Full 9×100 lifespan simulations (≈900 runs of a 36,500-step × 5,000-cell × 1,000-site Markov sim) that produced the cached pearson_cells_all_sims_*.csv. We validate those cached outputs (C3, C4) and re-run only the cheap development-stage sim fresh (C5), rather than regenerate all trajectories.

All heavy/any compute runs on «our HPC» («infra») per project rule; «host» holds results only. Repo cloned + run on «infra».

Figures / tables: Fig 1CFig 1DFig 1EFig 4TableFig 5
C1_mz_slope
Reported
-0.0020 per year (Fig 1C)
Reproduced
-0.0020020
exact
C1_dz_slope
Reported
-0.0033 per year (Fig 1D)
Reproduced
-0.0032684
within tolerance
C1_ur_slope
Reported
-0.0025 per year (Fig 1E)
Reproduced
-0.0024989
within tolerance
C2_mz_birth
Reported
~0.87 (Fig 1C)
Reproduced
0.8662 (fit intercept; raw mean 0.8259)
within tolerance
C2_dz_birth
Reported
~0.72 (Fig 1D)
Reproduced
0.7182 (fit intercept; raw 0.7290)
within tolerance
C2_ur_birth
Reported
~0.61 (Fig 1E)
Reproduced
0.6076 (fit intercept; raw 0.6405)
within tolerance
C3_mz_distance
Reported
weak selection best fit, < neutral (Fig 4 / Table S2)
Reproduced
weak=0.13229 < neutral=0.13355 < strong=0.14264
exact
C4_mz_sim_birth
Reported
~0.87 simulated MZ birth (Fig 5)
Reproduced
0.8769 (cached weak sims)
within tolerance
C5_mz_birth_sim
Reported
~0.87 MZ birth via N_split=36 (Fig 2/3)
Reproduced
0.8821 +/- 0.0690 (fresh sim, 20 reps)
within tolerance
C5_nsplit_order
Reported
MZ>DZ>UR via N_split 36/15/10 (Fig 2)
Reproduced
0.882 > 0.729 > 0.569
exact

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 90/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟢3. Location of the main deviation
🟢4. Cause of the deviation
🟢5. Derivability / plausibility
🟢6. Severity of the deviation
🟢7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score 0

Strong computational reproduction. Using the authors' own Julia code @ b1dc0e7 on bundled data, every pinnable number reproduced exact or within rounding — Fig 1 decline slopes (-0.0020/-0.0033/-0.0025/yr) and birth intercepts (0.87/0.72/0.61), and the Fig 4 weak<neutral<strong distance ordering — and an independent fresh stochastic re-run recovered the empirical MZ>DZ>UR birth correlations from the fitted N_split, corroborating the central mechanism with no fabrication signal. The only limitation is on scope/data availability, not the authors' side: the raw-GEO->Pearson preprocessing step (inputs not shipped) and the heavy 9x100 lifespan simulations (cached outputs reused) were not reproduced end-to-end. Hence q8 yellow (solid-partial) while all substantive claims hold green.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

207.6 k
tokens (I/O) · 17.6 M incl. cache
26 min
runtime · 0.26 CPU-h
2.6 GB
peak RAM
2 (1 failed)
HPC jobs
hummel
machine