An Erg-driven transcriptional program controls B cell lymphopoiesis.
The main results reproduced: recomputed values matched the published ones within tolerance.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Same input data as the authors
- ✓No relevant deviation in data/preprocessing
- ✓No authors-side cause for any deviation
- ✓Reported values are derivable from the shared data
- ✓Any deviation was negligible
- ✓The central claim held under reproduction
- 🟡Reported values were only indirectly comparable
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Described well enough to reproduce the CENTRAL computational result 1:1, with honest caveats. Brief metadata had two harvesting errors I corrected: the code link (sjmgarnier/viridis) is the viridis colour package (false positive; paper ships no code), and the data accession GSE114793 is the scRNA-seq used only for Fig 4e, not the bulk DE data. The real bulk RNA-seq is GSE132854, which ships a gene-level count matrix. I ran the paper's exact described pipeline (edgeR filterByExpr -> TMM -> limma voom/lmFit/eBayes; R4.5.3/edgeR4.8.2/limma3.66.0) on the contrast Rag1Cre;ErgD/D pre-proB vs Ergfl/fl pre-proB (GSE132854). Result: every named B-lineage/program gene moves in the reported direction (DOWN) — 18/18 by sign, with Ebf1 (-6.1) and Pax5 (-6.1) the most extreme, exactly matching the paper's 'loss of Ebf1 and Pax5' headline; the explicit negative-control set Foxo1/Spi1/Ikzf1 reproduces 'maintained' exactly (all flat, ns); Erg itself is ~16x down (KO sanity). Graded PARTIAL rather than reproduced because (a) the paper states NO DE-gene count and NO FDR/logFC threshold, so no single headline NUMBER exists to match 1:1 (Fig 4b is a logFC-ordered plot, hence direction is the right metric — and direction is perfect), and (b) with only n=2 replicates per group, 5 huge-effect genes (Ebf1, Pax5, Igll1, Xrcc6, Lef1) fall just short of FDR<0.05 (0.07-0.16). No fabrication concern: all directional claims are derivable from the shipped data. NOT attempted (out of scope / 80-20): upstream Rsubread/mm10 alignment (started from deposited counts), and the ChIP-seq/ATAC-seq/Hi-C (Fig 5-6) and scRNA-seq (Fig 4e) pipelines.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 81assessed: 2026-06-15 ⛓ cfbb29344e39
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-15
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15no human curator yet
- Last updated
- 2026-09-19
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetThe paper tests whether the ETS-family transcription factor Erg is required for early B lymphoid differentiation and whether it functions as an upstream regulator that coordinates Ebf1 and Pax5 expression to control V(D)J recombination and pre-B cell receptor formation.
- ★ Erg is essential for early B lymphoid differentiation finding
- ★ Erg initiates a transcriptional network involving Ebf1 and Pax5 that directly promotes genes required for V(D)J recombination and B cell receptor formation mechanism
- ★ Erg deletion causes B lymphoid developmental arrest at the pre-proB stage and loss of VH-to-DJH immunoglobulin heavy chain recombination finding
- ★ Complementation with a productively rearranged IgH allele (IgH VH10tar) rescues B lymphopoiesis in Erg-deficient mice, indicating failure of pre-BCR formation underlies the developmental block finding
- ★ Erg binding to VH gene families and to the iEmu mu A enhancer element does not structurally account for the Igh recombination defect seen upon Erg loss mechanism
- Rag1Cre-mediated conditional deletion efficiently ablates Erg specifically in lymphoid progenitors while sparing other hematopoietic lineages method
- Erg expression is developmentally regulated, present in HSCs and early lymphoid/B progenitors and declining with B and T cell maturation finding
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| lacZ reporter flow cytometry | Erg KI mouse bone marrow and thymus | none (Erg-driven LacZ knock-in reporter) | Erg promoter transcriptional activity (MFI ratio) | — |
| RNA-seq | Erg fl/fl mouse pre-proB, proB, preB cells | none | Erg expression (FPKM) across B developmental stages | — |
| RNA-seq | Erg fl/fl vs Rag1Cre;Erg Δ/Δ mouse pre-proB cells | Erg conditional knockout | differential gene expression (Erg, Ebf1/Pax5 network targets) | edgeR analysis |
| flow cytometry / blood counts | Erg fl/fl vs Rag1Cre;Erg Δ/Δ mouse blood and bone marrow | Erg conditional knockout | B220+, CD3+, Gr1+Mac1+ cell counts and B lymphoid subset (Hardy fraction) proportions | — |
| genomic PCR (degenerate primers) | B220+ bone marrow cells, Erg fl/fl vs Rag1Cre;Erg Δ/Δ mice | Erg conditional knockout | VH-to-DJH and DH-to-JH Igh recombination | — |
| fluorescence in situ hybridisation (FISH) | proB (Erg fl/fl) and pre-proB (Rag1Cre;Erg Δ/Δ) cells | Erg conditional knockout | intra-chromosomal distance between VH J558 and VH7183 gene families (locus contraction) | — |
| in situ Hi-C | wild-type proB cells vs Rag1Cre;Erg Δ/Δ pre-proB cells | Erg conditional knockout | long-range chromatin interaction counts across Igh locus | — |
| ChIP-seq / ChIP-PCR | wild-type/proB cells | none | Erg binding sites, H3K4me3 and H3K27ac marks across Igh locus | — |
- ▼ B-lymphoid development markedly compromised in Rag1Cre;Erg Δ/Δ mice with block at pre-proB stage and near-absent proB, preB, immature and mature B cells P adj = 3.5e-6 to 4.0e-2 across B populations
- ▼ Erg RNA significantly reduced in Rag1Cre;Erg Δ/Δ pre-proB cells confirming knockout edgeR adjusted P = 1.41e-5
- – VH-to-DJH Igh recombination lost in Erg-deficient B220+ cells while DH-to-JH recombination relatively preserved
- ▲ Reduced Igh locus contraction (increased intra-chromosomal distance) in Erg-deficient pre-proB cells n=129 Igh alleles
- ▼ Reduced long-range chromatin interactions across Igh locus in Erg-deficient pre-proB cells by Hi-C
- – No Erg binding detected across VH gene families of Igh locus, arguing against a direct structural role for Erg
- – Deletion of mu A element of iEmu (mu A Δ/Δ) did not reduce circulating mature B cells or impair VH-to-DJH recombination, despite Erg binding there; deletion of core iEmu (cEmu Δ/Δ) markedly reduced B220+CD19+IgM+IgD+ B cells P adj = 2.6e-5 (B220+), 1.8e-5 (B220+CD19+), 9.0e-6 (IgM+IgD+)
- ▲ Complementation with rearranged IgH VH10tar allele rescued B220+IgM+ B cells and CD25+CD19+IgM- preB cells in Erg-deficient mice P adj = 3.1e-3 (B220+IgM+), 3.3e-2 (preB)
- pvalue P adj < 0.028 (Erg KI lacZ MFI vs C57BL/6 across BM/thymic populations)
- pvalue P = 6.6e-8 (B220+ blood cell counts, Erg fl/fl vs Rag1Cre;Erg Δ/Δ)
- pvalue P adj = 3.5e-6 to 4.0e-2 (multiple B populations) (BM B-lymphoid population differences, Erg fl/fl (n=9) vs Rag1Cre;Erg Δ/Δ (n=10))
- pvalue edgeR adjusted P = 1.41e-5 (Erg RNA-seq expression, Erg fl/fl vs Rag1Cre;Erg Δ/Δ pre-proB cells)
- pvalue P adj = 2.1e-9 (B220+IgM+), 3.5e-4 (preB) (Rag1Cre;Erg Δ/Δ vs Erg fl/fl BM B lymphoid populations)
- pvalue P adj = 3.1e-3 (B220+IgM+), 3.3e-2 (preB) (IgH VH10tar rescue vs Rag1Cre;Erg Δ/Δ BM B lymphoid populations)
- pvalue P adj = 2.1e-11 (B220+), 5.6e-11 (Fol), 3.3e-5 (MZ) (Splenic B lymphoid populations, Erg fl/fl (n=14) vs Rag1Cre;Erg Δ/Δ (n=10))
- pvalue P adj = 2.6e-5 (B220+), 1.8e-5 (B220+CD19+), 9.0e-6 (IgM+IgD+) (cEμ Δ/Δ (n=3) vs cEμ Δ/+ (n=8) peripheral blood B cell counts)
Statistical methods review
Model: opusA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This experimental mouse study of Erg in B lymphopoiesis combined targeted gene-deletion models with flow cytometry, genomic/imaging assays (RNA-seq, ChIP-seq, ATAC-seq, in situ Hi-C, FISH, genomic PCR) and complementation crosses. Group comparisons of cell counts/proportions and expression intensities were made with two-tailed unpaired Student's t-tests, with multiplicity handled by Holm's modification or Benjamini–Hochberg correction depending on the comparison; differential RNA-seq expression was assessed with edgeR using two-sided adjusted P values. Results were reported as mean ± SD with stated biological replicate numbers and selected exact adjusted P values.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Student's two-tailed unpaired t-test with Holm's modification for multiple testing | Erg KI vs C57BL/6 lacZ MFI ratios across BM/thymus populations (Fig. 1b) | n=4 Erg KI and n=4 C57BL/6 biologically independent samples | not stated |
| edgeR two-sided adjusted P value (differential expression) | Erg expression in Erg fl/fl vs Rag1Cre;Erg Δ/Δ pre–proB cells (Fig. 1e; Supplementary Data 1) | n=2 biologically independent samples | not stated |
| Student's two-tailed unpaired t-test | B220+ B-cell blood counts, Erg fl/fl (n=4) vs Rag1Cre;Erg Δ/Δ (n=7) (Fig. 1f top left) | n=4 vs n=7 biologically independent samples | not stated |
| Student's two-tailed unpaired t-test with Holm's modification for multiple testing | BM B-lymphoid population ratios, Erg fl/fl (n=9) vs Rag1Cre;Erg Δ/Δ (n=10) (Fig. 1f bottom left) | n=9 vs n=10 biologically independent samples | not stated |
| Student's two-tailed unpaired t-test | Igh intra-chromosomal distance by FISH between distal VHJ558 and proximal VH7183 (Fig. 2b) | n=129 Igh alleles | not stated |
| Student's two-tailed unpaired t-test with Benjamini–Hochberg correction | Peripheral blood counts comparing cEμΔ/Δ to cEμΔ/+ controls (Fig. 2d) | cEμΔ/+ n=8, cEμΔ/Δ n=3, μAΔ/Δ n=7 | not stated |
| Student's two-tailed unpaired t-test with Holm's modification for multiple testing | Splenic B-lymphoid population proportions across genotypes (Fig. 3b) | n=14 Erg fl/fl, n=10 Rag1Cre;Erg Δ/Δ, n=9 Rag1Cre;Erg Δ/Δ;IgH VH10tar/+ | not stated |
-
Cell counts and population proportions were compared with multiple two-tailed unpaired Student's t-tests across several populations within a panel, with Holm or Benjamini–Hochberg adjustment.↳ Could also: A single one-way or two-way ANOVA (or mixed model) with a post-hoc multiple-comparison procedure such as Tukey's HSD or Dunnett's test (for comparisons against a common control) could also be used. — An omnibus model would estimate a shared variance across groups and control the family-wise error rate within one analytical framework, which some workflows prefer when several genotypes are compared against shared controls.
-
Group differences were assessed with parametric Student's t-tests, with assumptions not explicitly stated.↳ Could also: A nonparametric Mann–Whitney U test, or a t-test with explicit reporting of normality/variance checks (e.g., Welch's t-test for unequal variances), could also be applied. — For the smaller groups (e.g., n=3) a rank-based test or Welch's correction relaxes the normal-distribution or equal-variance assumption, which can be informative when distributional assumptions are hard to verify at small n.
-
Spread was summarized as mean ± SD.↳ Could also: A 95% confidence interval for the mean (or showing individual data points alongside the mean) could also be reported. — A confidence interval directly conveys the precision of the estimated difference and complements the descriptive SD, which is often favored for small sample sizes.
-
Differences were communicated primarily through adjusted P values.↳ Could also: Effect-size estimates with intervals (e.g., fold-change differences in counts with 95% CIs, or standardized mean differences) could also be reported. — Effect sizes quantify the magnitude of biological differences independently of sample size, complementing significance testing.
-
Differential RNA-seq expression with n=2 biological replicates was analyzed using edgeR adjusted P values.↳ Could also: Alternative count-based frameworks such as DESeq2, or limma-voom with empirical Bayes moderation, could also be used. — These methods share information across genes to stabilize variance estimates at low replicate numbers and offer alternative normalization and shrinkage options, so reporting them is common for cross-validation of differential-expression calls.
Result convergence & founder nodes
Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.
-
Igh locus chromatin accessibility is unchanged in ERG-deficient B cell progenitors despite impaired locus contractionATAC-seq mouse b-cell-progenitor none 2020×1papers★ This paper is the founder (earliest)
-
Deletion of the μA enhancer does not impair circulating mature B cells or VH-to-DJH Igh recombinationflow-cytometry mouse blood none 2020×1papers★ This paper is the founder (earliest)
-
ERG deficiency causes a developmental block at the pre-proB stage with accumulation of pre-proB cells and near-absence of proB, preB, and mature B cells in bone marrowflow-cytometry mouse bone marrow mixed 2020×1papers★ This paper is the founder (earliest)
-
A pre-rearranged IgH VH10tar allele rescues preB and IgM+ B cell populations in ERG-deficient bone marrow, bypassing the VH-to-DJH recombination blockflow-cytometry mouse bone marrow up 2020×1papers★ This paper is the founder (earliest)
-
ERG deficiency reduces long-range chromatin interactions across the Igh locus in pre-proB cellsHi-C mouse pre-prob-cell down 2020×1papers★ This paper is the founder (earliest)
-
ERG deficiency reduces Igh locus contraction in pre-proB cells as shown by increased intra-chromosomal VHJ558-VH7183 distance by FISHimaging mouse pre-prob-cell down 2020×1papers★ This paper is the founder (earliest)
-
VH-to-DJH recombination at the Igh locus is abolished in ERG-deficient B220+ bone marrow cells while DH-to-JH recombination is preservedother mouse bone marrow down 2020×1papers★ This paper is the founder (earliest)
-
ERG mRNA is significantly reduced in Rag1Cre;ErgΔ/Δ pre-proB cells confirming conditional knockoutRNA-seq mouse pre-prob-cell down 2020×1papers★ This paper is the founder (earliest)
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-32541654
Paper: Ng AP et al. An Erg-driven transcriptional program controls B cell lymphopoiesis. Nat Commun 2020. PMID 32541654 / PMC7296042 / DOI 10.1038/s41467-020-16828-y. WEHI (Davis & Smyth Bioinformatics Divisions).
Accession / code-link corrections (harvesting errors in BRIEF)
- Code link
github.com/sjmgarnier/viridisis a FALSE POSITIVE — that is the viridis R colour-palette package, not this paper's analysis code. The paper ships no code-availability statement (P16 applies: we reproduce by running the described, standard third-party pipeline — limma/edgeR — on the paper's own deposited data; equally valid). - BRIEF's data accession
GSE114793is NOT the main dataset. GSE114793 ("Functional dissection of lymphoid progenitor compartments…") is the scRNA-seq dataset used only for Fig 4e (imputed single-cell Erg/Ebf1/Pax5, 3297 cells). The paper's actual data-availability statement lists:- GSE132854 — bulk RNA-seq ← our in-scope target (has a shipped count matrix)
- GSE132853 — ChIP-seq · GSE132852 — ATAC-seq · GSE133246 — Hi-C
In scope (pipeline-derived, reproduced here)
Bulk RNA-seq differential expression (Fig 4a/4b), GSE132854.
- Shipped processed data:
GSE132854_Ng_etal_2019_RNAseqCounts_allSamples.txt.gz= gene-level (Entrez) counts, 27,180 genes × 8 samples (GSM3895109–116): Ergfl/fl pre-proB ×2, proB ×2, preB ×2; Rag1Cre;ErgΔ/Δ pre-proB ×2. - Pipeline exactly as Methods describe: edgeR
filterByExpr→ TMM (calcNormFactors) →voom→lmFit/eBayes(limma) →topTable. - Contrast reproduced: Rag1Cre;ErgΔ/Δ pre-proB − Ergfl/fl pre-proB (= Fig 4b).
- Pinnable claims (directional, from Results + Fig 4 legend):
- Erg itself strongly DOWN in KO (knockout sanity / positive control).
- DOWN upon Erg loss: Ebf1, Pax5, Cd19, Cd22, Igll1, Vpreb1, Vpreb2, Cd79a, Cd79b, Rag1, Rag2, Xrcc6, Lig4, Tcf3, Bach2, Irf4, Myc, Pou2af1, Lef1, Myb.
- MAINTAINED (explicitly NOT significantly changed): Foxo1, Spi1, Ikzf1 → built-in negative controls.
Out of scope (not attempted, why)
- Raw alignment (Rsubread
alignto mm10): upstream of the shipped count matrix; reproducing it needs the FASTQs (SRP201633) and adds no value over starting from the deposited counts. 80/20. - ChIP-seq / ATAC-seq / Hi-C (Fig 5–6, Bowtie2/MACS2/HOMER/HiC pipelines): separate large pipelines; deferred (the hard last 20%).
- scRNA-seq imputation (Fig 4e, GSE114793): non-deterministic imputation; deferred.
- Wet-lab (flow cytometry, Western blot Fig 4d, mouse phenotyping): not computational.
Why this is a faithful, honest target
The paper under-specifies the DE call: no DE-gene count and no explicit FDR/logFC threshold are stated in text or Fig 4 legend. We therefore reproduce the directional, named-gene claims (the actual scientific content of Fig 4b) and report our derived logFC/FDR per gene transparently; the DE-gene total at a standard FDR<0.05 is reported as our derived value (flagged: not stated by the paper, so not a 1:1 number comparison).
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
The central computational claim of Fig 4a/4b reproduces 1:1 by direction from the authors' deposited GSE132854 count matrix using the paper's described edgeR/limma-voom pipeline: 18/18 named B-lineage genes are down, Ebf1/Pax5 are the two most extreme, and the Foxo1/Spi1/Ikzf1 negative control is flat. No deviation is on the authors' side and there is no fabrication concern — all values are derivable. It falls short of a clean green only because the paper specifies no DE count/threshold (no numeric anchor), the registry metadata (code_url and accession) were wrong and needed correction, and n=2 softens FDR for 5 large-effect genes. Net: a solid, honest reproduction with fully explainable caveats.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at [email protected].
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.