Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Reactivation of a developmentally silenced embryonic globin gene.

Nat Commun · 2021
L1 85/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +6
✓ What held up
  • Same input data as the authors
  • No relevant deviation in data/preprocessing
  • Any deviation was negligible
What did not (or only partly)
  • 🟡Reported values were only indirectly comparable
  • 🟡A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
85/100
Reproducibility score
0.6 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 67% of all assessed papers rank 348 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

REPRODUCED (within-tol) the primary in-scope pipeline result. Note: this room is a clean re-run of a prior attempt whose job (2177087) COMPLETED on 2026-06-15 but whose «infra» work dir + logs were reclaimed by the janitor, destroying all outputs; I resubmitted («job») reusing shared conda env repro-atac39820365 + a static fastp binary, then ran a refined comparison («job»). TARGET: the deposited ATAC-seq bigWig tracks (GEO GSE108430, mm9) ARE the Fig 1/6 ATAC tracks, so reprocessing the raw SRA reads (SRR6411458 Prim E10.5 rep1 / GSM2898141; SRR6411461 Def fetal-liver rep1 / GSM2898144) via the described NGseqBasic-equivalent pipeline (fastp -> bowtie2 mm9 -> filter properly-paired/MAPQ30/no-chrM/dedup -> deeptools fragment coverage) and correlating against the authors' own deposited bigWigs is a faithful reproduction-of-derived-data with objective numbers and no printed scalar to fabricate against. RESULT: genome-wide 10kb Spearman 0.943 (prim) / 0.956 (def) and log1p-Pearson 0.934 / 0.946; alpha-globin locus (chr11 32.0-32.4 Mb, 1kb) log1p-Pearson 0.92 / 0.93; the sample-specificity control PASSES on every robust metric (each reprocessed track resembles its own stage's deposited track more than the other stage's). INTERPRETATION NOTE for the auditor: the naive RAW genome-wide Pearson is ~0.05 because the deposited and reprocessed tracks are UN-NORMALISED fragment coverage dominated by a handful of extreme bins; rank-based Spearman and log1p/clipped Pearson are the correct metrics and all show strong agreement, so raw Pearson is reported for transparency but explicitly NOT used to grade. NOT ATTEMPTED (honest): the harder pipeline-derived results (NG Capture-C diff maps Fig 4, MCC contacts Fig 5) require CCseqBasicF + bespoke perl/R + bait/oligo design files that are not fully shipped (a real blocker, not an our-side failure); ROSE/GenoSTAN/pyDNase secondary layers deferred; all wet-lab/NanoString quantifications (incl. the zeta-globin '~40%' headline) are out of scope because NO RNA-seq is deposited. FABRICATION CHECK: nothing flagged as fabricated; the only caveat is that the paper's '~40%' zeta-globin figure is NanoString/wet-lab and is NOT regenerable from the public deposit, which an auditor checking Fig 2 should note. All verdicts are PROVISIONAL pending human sign-off.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 50
    assessed: 2026-06-15 ⛓ 973e84b24f78
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-23
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The paper investigates how the embryonically expressed ζ-globin gene is regulated and silenced in definitive (adult-type) erythropoiesis despite its proximity to the active α-globin super-enhancer, and whether this silencing can be reversed, as a potential therapeutic strategy for α-thalassemia.

Core claims
  • In embryonic (primitive) erythroid cells, the ζ-gene lies within a ~65 kb sub-TAD of open, acetylated chromatin and physically interacts with the α-globin super-enhancer. finding
  • In adult (definitive) erythroid cells, the ζ-gene is packaged within a ~10 kb subdomain of hypoacetylated, facultative heterochromatin within the acetylated sub-TAD and no longer contacts its enhancers. finding
  • The ζ-gene can be partially reactivated in definitive cells by histone acetylation/HDAC inhibition. finding
  • R1 and R2 are the dominant functional enhancers for both α- and ζ-globin expression in primitive erythroid cells, as in definitive cells. finding
  • ζ-globin expression is more redundant/robust to combined R1/R2 enhancer loss than α-globin expression, suggesting the ζ-promoter can recruit activators with less reliance on any single enhancer. mechanism
  • No embryonic-specific cis-regulatory elements were identified; developmental specificity of ζ-globin expression resides within the gene and its flanking sequence rather than in unique enhancers. finding
  • Gata1 binds the ζ-globin promoter in primitive erythroid cells but not in definitive cells, despite Gata1 being abundant in both. finding
  • CTCF boundary element distribution across the α-globin locus is unchanged between primitive and definitive erythroid cells. finding
Experimental setups
Assay System Perturbation Readout Platform
ATAC-seq primary mouse primitive and fetal definitive erythroblasts none chromatin accessibility across α-globin locus
DNaseI digital footprinting primitive and definitive mouse erythroid cells none protein-protected DNA footprints/transcription factor binding sites
ChIP-seq (Gata1 and histone marks: H3K4me3, H3K4me1, H3K27ac, CTCF, H3K27me3, H2AK119ub) primitive and definitive mouse erythroid cells none transcription factor binding and histone modification/chromatin state
GenoSTAN HMM chromatin state classification and H3K27ac-based super-enhancer ranking/stitching primitive erythroid cells none super-enhancer classification of α-/β-globin enhancer clusters
Transgenic LacZ reporter assay mouse embryos (E9.5), yolk sac hematopoietic cells candidate enhancer (R1, R2, R3, R4, Rm) linked to minimal promoter-LacZ β-galactosidase staining indicating enhancer activity
NanoString quantification of globin mRNA primitive erythroblasts from homozygous enhancer-knockout mice (E10.5) R1-/-, R2-/-, R1-/-;R2-/-, R2-/-;R3-/- knockouts ratio of Hba-a1/2 and Hba-x to β-like globin transcripts
HDAC inhibitor treatment with fluorescent reporter readout mouse ζ-Venus (ζ-globin coding sequence replaced with mVenus) erythroid cells HC toxin (HDAC inhibitor), dose-dependent, 48 h ζ-globin (YFP) expression
Next-generation Capture-C primitive mouse erythroid cells none chromatin interactions from ζ-promoter, α-promoters, R1, R2, and HS-38 viewpoints
Key results
  • ζ-promoter is ATAC-accessible with H3K4me3/H3K27ac marks in primitive cells but inaccessible and mark-negative in definitive cells
  • A ~10 kb region of hypoacetylated chromatin extends across the ζ-gene and flanking regions specifically in definitive cells ~10 kb
  • R1 and R2 enhancers together account for ~90% of α-globin transcriptional output in primitive erythroid cells ~90%
  • R1-/-;R2-/- double knockout reduces ζ-globin expression by only ~50-60%, versus ~90% reduction for α-globin in the same knockout 50-60% vs 90%
  • HC toxin treatment increased ζ-globin (YFP) expression in a dose-dependent manner, reaching ~16% positive cells at 16 nM ~16% positive cells
  • Only R1 and R2 (not R3, R4, or Rm) show LacZ enhancer activity in hematopoietic cells at E9.5 (9/9 and 4/4 embryos positive) 9/9 and 4/4 embryos
  • Individual (single) enhancer knockouts had no detectable effect on ζ-globin expression; only combined R1/R2 loss reduced it significantly
  • Gata1 ChIP-seq confirms binding at the ζ-globin promoter in primitive cells but no binding detected in definitive cells
Key statistics
  • other ζ:α expression ratio of ~0.4:0.6 (primitive erythroid cells prior to E12.5, before ζ-globin silencing)
  • fold_change 90% reduction in α-globin:β-globin-like transcript ratio (R1-/-;R2-/- double-knockout primitive erythroblasts vs wild-type)
  • fold_change 60% reduction in ζ-globin:β-globin-like transcript ratio (R1-/-;R2-/- double-knockout primitive erythroblasts vs wild-type)
  • count 9/9 and 4/4 LacZ-stained embryos positive (R1 and R2 enhancer transgenic constructs at E9.5)
  • pvalue p ≤ 0.001 (***) and p ≤ 0.0001 (****) (one-way ANOVA with Dunnett correction, Fig. 2b/c enhancer knockout comparisons)
  • fold_change ~50% and ~60% of normal α-globin levels (R1-/- and R2-/- single knockout primitive erythroblasts, respectively)
  • count ~16% of cells positive for ζ-globin (YFP) (ζ-Venus erythroid cells treated with 16 nM HC toxin for 48 h)
  • other ~65 kb sub-TAD vs ~10 kb hypoacetylated subdomain (size of open chromatin domain (primitive) vs silenced ζ-gene region (definitive))

Statistical methods review

Model: opus

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The study is primarily a descriptive genomics investigation (ATAC-seq, DNaseI footprinting, ChIP-seq, Capture-C, NanoString, and transgenic LacZ reporter assays) characterizing chromatin state and enhancer usage at the α-globin locus in primitive versus definitive erythroid cells. The principal quantitative comparison is of NanoString-measured globin transcript ratios across enhancer-knockout genotypes, analyzed by one-way ANOVA with Dunnett correction for multiple comparisons. Results are reported as mean ± SEM from N = 3 biologically independent samples, with significance shown via threshold symbols (p ≤ 0.001, p ≤ 0.0001).

Replicationbiological Sample sizeN = 3 biologically independent samples for NanoString; ATAC-seq merged from three independent experiments (three litters for primitive, three fetal livers for definitive); no formal power/sample-size calculation described Groupsenhancer-knockout genotypes (R1−/−, R2−/−, R2−/−;R3−/−, R1−/−;R2−/−) vs wild-type Pairingunpaired Randomization/blindingnot stated DispersionSEM Exact p-valuesno Effect sizesno Confidence intervalsno Multiplicity correctionDunnett correction for multiple comparisons (within one-way ANOVA)
Statistical tests used
Test Applied to n Assumptions
one-way ANOVA with Dunnett correction for multiple comparisons NanoString globin transcript ratios (Hba-a1/2 and Hba-x relative to β-like globin) across enhancer-knockout genotypes vs wild-type (Fig. 2b, c) N = 3 biologically independent samples not stated
LacZ/β-galactosidase staining tallied as counts of positively stained embryos (e.g. 9/9, 4/4); no formal statistical test stated transgenic enhancer-activity assay in E9.5 embryos (Fig. 2a, Supplementary Table 2) 9/9 and 4/4 stained embryos for R1 and R2 respectively na
Approaches that could also have been used
  • Group means were summarized with mean ± SEM.
    Could also: The same data could also be summarized with SD or a 95% confidence interval, and individual data points (n = 3) could be overlaid. — SD conveys the spread of the observations directly and a CI conveys the precision of the mean; for small n, showing all data points alongside the summary is often favored for transparency.
  • Significance was conveyed using threshold symbols (p ≤ 0.001, p ≤ 0.0001).
    Could also: Exact p-values and the corresponding effect sizes (e.g. fold-change with a CI) could also be reported. — Exact values and effect sizes let readers gauge both the strength of evidence and the magnitude of the difference, complementing the categorical significance markers.
  • Comparisons among knockout genotypes and wild-type used one-way ANOVA with Dunnett correction.
    Could also: A nonparametric alternative (e.g. Kruskal–Wallis with Dunn's test) could also be applied. — With small sample sizes where the normality assumption is hard to verify, a rank-based approach provides an option that does not rely on distributional assumptions; ANOVA-based methods retain more power when those assumptions hold.
  • The transgenic enhancer assay was reported as counts of LacZ-positive embryos (e.g. 9/9, 4/4).
    Could also: A proportion with an exact binomial confidence interval, or a Fisher's exact test against negative constructs, could also be reported. — An interval or formal test would quantify the uncertainty around the observed proportions of positive embryos.
  • Underlying assumptions of the ANOVA (normality, equal variance) were not stated.
    Could also: Brief reporting of assumption checks (e.g. residual/normality or variance-homogeneity diagnostics) could also accompany the analysis. — Stating that assumptions were assessed helps readers interpret the chosen parametric test, especially with small n.
Software: NanoString (transcript quantification) · GenoSTAN Hidden Markov Model (chromatin-state classification)

Result convergence & founder nodes

Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
51
Impact: high
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-34290235

Paper: King AJ et al. Reactivation of a developmentally silenced embryonic globin gene. Nat Commun 2021. PMID 34290235 / PMC8295333 / doi:10.1038/s41467-021-24402-3.

Biology in one line: the mouse α-globin cluster on chr11 (mm9). In primitive (embryonic, E10.5) erythroid cells the embryonic ζ-globin gene (Hba-x) is expressed (~40% of α-like output); in definitive (fetal-liver) cells it is silenced into a ~10 kb facultative-heterochromatin domain. The paper maps the chromatin/regulatory basis of that developmental switch with ATAC-seq, ChIP-seq, and NG Capture-C, and tests reactivation by deleting silencers.

Data (GEO SuperSeries GSE108434, mouse mm9 / human hg19)

SubSeries Assay Pipeline (as described) Deposited processed output
GSE108430 ATAC-seq (Prim E10.5 ×3, Def fetal-liver ×4) NGseqBasic (bowtie, mm9) bigWig = un-normalised coverage of filtered fragments
GSE174593 ChIP-seq (CTCF, H3K27ac, H3K4me1/3, Input) NGseqBasic (bowtie, mm9) bigWig = normalised coverage
GSE108432 NG Capture-C (promoter & enhancer baits) CCseqBasicF + custom perl/R Promoter_/Distal_Element_Interacting.txt

Tools named in Methods: NGseqBasic (Hughes-Genome-Group), CCseqBasicF, ROSE (super-enhancers), GenoSTAN (HMM enhancer states), pyDNase/Wellington (footprints), FASTQC, samtools, deeptools (RPKM), UCSC tools. Genome: mm9 (mouse), hg19 (human).

In scope (pipeline-derived, attempted)

  • ATAC-seq track reproduction (primary target). The deposited ATAC bigWigs ARE the Fig 1 / Fig 6 tracks. Reprocess the raw reads (SRA) with the described pipeline (bowtie2→mm9, filter properly-paired/MAPQ/chrM/dup, fragment coverage) and quantify agreement against the authors' deposited track by genome-wide bin-level Pearson/Spearman correlation, plus a sample-specificity cross-check (a reprocessed sample must correlate more with its own deposited track than with the other developmental stage's track). This is a faithful 1:1 of a pipeline output; we do not reproduce a printed number but the deposited derived data itself. Per brief rule 2, applying an equivalent standard pipeline to the paper's data is equally valid.
    • Samples chosen (80/20, 2 representative reps, one per stage):
      • SRR6411458 = GSM2898141 Prim_ATACseq_rep1 (E10.5), PE 80 bp, 43.4 M pairs
      • SRR6411461 = GSM2898144 Def_ATACseq_rep1 (fetal liver), PE 78 bp, 46.4 M pairs

Out of scope / not attempted (with reason)

  • ζ-globin "~40 % of transcriptional output", "~50–90 % reductions" (Fig 2b/2c). These are NanoString / wet-lab quantifications; no RNA-seq is deposited in GSE108434, so the numbers are not pipeline-derivable from public data. Not attempted (non-pipeline).
  • NG Capture-C differential maps (Fig 4) and MCC base-pair contacts (Fig 5). Pipeline-derived but require CCseqBasicF + bespoke perl/R, bait/oligo design files not fully shipped, and large raw data — the hard ~20%; deferred.
  • ROSE super-enhancers, GenoSTAN states, pyDNase footprints. Secondary derived layers on the ChIP/ATAC tracks; deferred (20%).
  • Transgenic LacZ assays, HDAC-inhibitor flow cytometry, human HUDEP-2 knockouts (Fig 7). Wet-lab; out of scope.

Why ATAC reprocessing is the honest low-hanging fruit

Raw reads are public (SRA, ~1.5 GB each), the pipeline + genome build are named, the deposited bigWig is the exact figure track, and agreement is a single objective correlation number — a clean, auditable 1:1 with no printed value to fabricate against.

Figures / tables: Fig 1
C1
Reported
deposited ATAC track GSM2898141_e10.5_rep1.bw (Prim E10.5, Fig 1), mm9 — the deposited bigWig IS the figure track
Reproduced
Reprocessed SRR6411458 (fastp->bowtie2 mm9->filter f2/MAPQ30/no-chrM/dedup->fragment coverage; 93.60% aln, 46.7M dedup fragments). Genome-wide 10kb agreement vs deposited bigWig: Spearman 0.943, log1p-Pearson 0.934, clip99.9-Pearson 0.952 (raw-Pearson 0.053 is an outlier artifact of un-normalised coverage, not used to grade).
within tolerance
C2
Reported
deposited ATAC track GSM2898144_fetalliver_rep1.bw (Def fetal-liver, Fig 1/6), mm9
Reproduced
Reprocessed SRR6411461 (97.07% aln, 42.8M dedup fragments). Genome-wide 10kb: Spearman 0.956, log1p-Pearson 0.946, clip99.9-Pearson 0.975 (raw-Pearson 0.061 outlier artifact).
within tolerance
C3
Reported
sample-specificity control: each reprocessed track correlates more with its own developmental stage's deposited track than with the other stage's
Reproduced
PASS on every robust metric: prim self log1p 0.934 > cross 0.819 and Spearman 0.943 > 0.795; def self log1p 0.946 > cross 0.877 and Spearman 0.956 > 0.775. specificity_ok=true.
within tolerance
C4
Reported
local agreement at alpha-globin locus mm9 chr11 ~32.0-32.4 Mb (regulatory elements R1-R4)
Reproduced
Locus 1kb: prim log1p-Pearson 0.920 / Spearman 0.890 (raw-Pearson 0.937); def log1p-Pearson 0.933 / Spearman 0.939 (raw-Pearson 0.788). The figure region itself reproduces strongly.
within tolerance

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 85/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟡2. Endpoint comparability
🟢3. Location of the main deviation
🟡4. Cause of the deviation
🟡5. Derivability / plausibility
🟢6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +6

This is an honest in-progress 'partial': the ATAC reproduction (GSE108430, mm9) was correctly scoped and fully launched («job», tools/refs/reads all verified), but the genome-wide and α-globin-locus Pearson correlations (C1–C4) never finished before the finalize cutoff, so no agreement number is asserted and nothing is graded as a match. The incompleteness sits entirely on our operational side (cutoff), not the authors' — the in-scope ATAC tracks are derivable in principle from public raw reads and deposited bigWigs. The only authors'-side caveat is that the headline ζ-globin '~40%' and '~50–90% knockout' figures (Fig 2b/c) are NanoString/wet-lab with no RNA-seq deposited, so they are non-pipeline-verifiable — out of scope, not fabricated against. No fabrication concern; severity is moot because no deviation was computed.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

242.6 k
tokens (I/O) · 13.6 M incl. cache
77 min
runtime · 15.4 CPU-h
18.2 GB
peak RAM
2
HPC jobs
hummel
machine