Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Temporal control of progenitor competence shapes maturation in GABAergic neuron development in mice.

Nat Neurosci · 2025
L1 82/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +6
✓ What held up
  • Any deviation was negligible
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🟡Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🔴A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
82/100
Reproducibility score
0.4 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 60% of all assessed papers rank 459 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

STRONG PARTIAL, described-well-enough for the DEPOSITED data. «our HPC» reachable (prior attempt failed only on VPN 2FA, now cleared); built Seurat 4.4.0/R 4.3.3 on «infra» and ran 2 SLURM jobs. The paper's code repo (mayer-lab/Bright-et-al-2025) ships ONLY downstream figure-plotting R scripts; every script load()s integrated/annotated objects from a Processed_Objects/ dir that is NOT deposited (not in repo, not in GEO). GEO GSE255455 deposits RAW.tar (18 processed per-sample Seurat objects) + ONE integrated object (Sample15_16_17_18, 3498 cells) that is loaded by no script and lacks the Fig1g columns. RESULT: every claim gradeable directly from the deposited data reproduced EXACTLY (4/4) -- Nfib/x KO total = 47,079, Nfib OE total = 30,019, KO sgNfib/sgNfix = 5,887, KO sgLacZ = 30,328 -- all matching the paper to the digit (strong integrity evidence, no fabrication signal on these). NOT reproducible 1:1: Fig 1g stage-correlation (its input integrated object is not deposited; code shipped, input not shipped) -- this is a data-availability gap, not a fabrication finding. NOT attempted (out of 80/20 scope): scATAC/ArchR/TOBIAS/SCENIC+/CUT&RUN on separate accessions, and UMAP/pseudotime visual panels. Honest verdict: the deposited scRNA-seq data delivers its headline KO/OE numbers exactly; the published figures are only partly reproducible from public data+code because the figure-input integrated objects are largely undeposited.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 82
    assessed: 2026-06-20 ⛓ 865d20d5d6df
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-20
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-20
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The paper tests how progenitor competence (maturation competence and differentiation competence) in ganglionic eminence progenitors shapes the maturation and diversification of GABAergic neurons, and how this compares to the temporally progressive competence changes seen in dorsal cortical progenitors.

Core claims
  • Ganglionic eminence (ventral) progenitors maintain stable differentiation competence throughout neurogenesis, generating a consistent set of postmitotic precursor states at all stages, unlike dorsal cortical progenitors whose differentiation competence changes gradually. finding
  • Maturation competence changes with developmental timing: late-born (e16.5) GE neurons mature faster than early-born (e12.5) neurons, reaching similar transcriptional states in less time. finding
  • Chromatin remodeling together with a regulatory module centered on the transcription factor NFIB and its target genes drives increased maturation competence in late-born progenitors. mechanism
  • Ventral GE progenitors show stable resting membrane potential across e13.5-e15.5, in contrast to dorsal cortical progenitors, which progressively hyperpolarize over the same period. finding
  • Clonal lineage tracing (TrackerSeq) shows similar proportions of dispersing (multi-fate) clones at e12.5 and e16.5, indicating differentiation competence is maintained at the clonal level across neurogenesis. finding
  • Heterochronic transplantation experiments show that maturation competence of progenitors is influenced by the extrinsic tissue environment. finding
  • An interactive web-based resource is provided for exploring scRNA-seq, scATAC-seq, CUT&RUN and eGRN datasets comparing GE- and cortex-derived neurogenesis. resource
  • A multi-modal approach combining scRNA-seq, TrackerSeq barcode lineage tracing, FlashTag birthdating, perturbation sequencing and CUT&RUN was used to dissect progenitor competence. method
Experimental setups
Assay System Perturbation Readout Platform
scRNA-seq Dlx5/6-Cre::tdTomato mouse embryos (cortex, striatum, GE), e12.5/e14.5/e16.5 none transcriptomic cell states along developmental pseudotime Seurat, Monocle3
TrackerSeq (DNA barcode lineage tracing + scRNA-seq) GE progenitors, mouse embryos, in utero electroporation at e12.5 or e16.5, collected 96h later none (lineage labeling) clonal relationships and dispersal across postmitotic branch tips
FlashTag (FT) birthdating + scRNA-seq GE progenitors, wild-type and Dlx5/6-Cre::tdTomato mouse embryos, e12.5 and e16.5, collected 6h or 96h later none (isochronic cohort labeling with CFSE) pseudotime/maturation state of isochronic cell cohorts
RNAscope in situ hybridization GE tissue sections, mouse embryos none expression of Ascl1 (intermediate progenitor) and Gad2 (postmitotic precursor) markers
Whole-cell patch-clamp electrophysiology Cortical and GE progenitors, mouse embryos, e13.5 and e15.5 none resting membrane potential
scATAC-seq / eGRN inference GE progenitors, mouse embryos none chromatin accessibility and enhancer-driven gene regulatory networks
CUT&RUN GE progenitors/neurons, mouse none NFIB genomic binding/target validation
Perturbation sequencing GE progenitors, mouse embryos NFIB perturbation transcriptomic effects on maturation competence
Key results
  • Ventral (GE) progenitors show higher Pearson correlation coefficients between successive neurogenesis stages than dorsal progenitors, indicating less transcriptomic change over time
  • Dorsal cortical progenitors progressively hyperpolarize between e12.5 and e15.5, while ventral GE progenitor membrane potential remains stable
  • Pseudotime score of FT e16.5+6h cohort is markedly higher than FT e12.5+6h despite both being collected 6h after labeling
  • Genes upregulated in FT e16.5+6h overlap substantially with genes upregulated in FT e12.5+96h, indicating late-born neurons reach a similar expression profile much faster
  • Similar proportion of dispersing (multi-branch-tip) clones observed in TrackerSeq e12.5+96h and TrackerSeq e16.5+96h
  • Pseudotime distributions of FT+ cells differ significantly across conditions (FT e12.5+6h, e16.5+6h, e12.5+96h) P = 3.46×10^-82, 2.87×10^-255, 1.06×10^-107
  • Mitotic progenitor cells from nondispersing clones show no stronger transcriptomic correlation with their postmitotic clonal progeny than randomly selected progenitors
Key statistics
  • pvalue P = 3.46 × 10^-82, 2.87 × 10^-255, 1.06 × 10^-107 (****P < 0.0001) (Wilcoxon rank-sum test comparing pseudotime distributions of FT+ cohorts)
  • pvalue *P < 0.05, **P < 0.01 (Two-sided t-test, Pearson correlation between dorsal and ventral progenitors across stages)
  • count n = 41,460 cells; n = 40 embryos (Combined scRNA-seq, TrackerSeq and FT dataset (UMAP))
  • count n = 25,297 cells; n = 20 embryos (scRNA-seq dataset across collection stages)
  • count n = 75,431 cells (Combined ventral (GABAergic) and dorsal (glutamatergic) telencephalon dataset)
  • count n = 9,938 cells (TrackerSeq barcoded cells across IUE stages)
  • count n = 6,225 cells; n = 19 embryos (FlashTag (FT) dataset)
  • count Cortex n = 14; GE n = 37 (Whole-cell patch-clamp membrane potential recordings at e13.5/e15.5)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The paper primarily reports single-cell transcriptomic and electrophysiological comparisons using nonparametric and parametric two-group tests (Wilcoxon rank-sum, two-sided t-tests) applied to Monocle3 pseudotime scores, Pearson correlation coefficients, and membrane potential recordings, with results shown as boxplots/violin plots annotated with significance asterisks or exact p-values. Sample sizes are reported at both the cell level (n cells) and embryo/animal level (n embryos or n recorded cells per group) depending on the analysis. Differential gene expression and clustering analyses were performed within the Seurat/Monocle3 single-cell analysis framework, and one comparison (membrane potential across cortex/GE and stages) explicitly applied a false discovery rate correction.

Replicationmixed Sample sizeSample sizes given as combinations of cell numbers and embryo counts per dataset (e.g., n = 41,460 cells, n = 40 embryos for the combined UMAP; n = 6,225 cells, n = 19 embryos for FT data), and as per-group counts for specific statistical tests (e.g., Cortex n=14, GE n=37; n=2,000 for pseudotime comparison) Groupscortical vs. ganglionic eminence (GE) progenitors; dorsal vs. ventral lineages; early- vs. late-born neuron cohorts (FT e12.5 vs. e16.5); dispersing vs. nondispersing clones Pairingunpaired Randomization/blindingnot stated DispersionIQR Exact p-valuesyes Multiplicity correctionfalse discovery rate correction (specific procedure, e.g., Benjamini-Hochberg, not specified in the excerpt)
Statistical tests used
Test Applied to n Assumptions
two-sided t-test Pearson's correlation comparison between dorsal and ventral progenitors across developmental stages (Fig. 1g) not stated
two-sided, unpaired t-test with false discovery rate correction membrane potential comparison between cortical and GE progenitors at e13.5/e15.5 (Fig. 1j) Cortex n=14; GE n=37 not stated
two-sided, unpaired Wilcoxon rank-sum test pseudotime score distributions across FT birthdating cohorts (Fig. 2c) n = 2,000 not stated
Approaches that could also have been used
  • Two-group comparisons (e.g., membrane potential in Fig. 1j, n=14 vs n=37) were tested with two-sided unpaired t-tests.
    Could also: A nonparametric test such as the Mann-Whitney U (Wilcoxon rank-sum) test could also be used — This avoids relying on the normality assumption, which can be harder to verify with smaller group sizes such as n=14.
  • A false discovery rate correction was applied to the membrane potential comparisons (Fig. 1j), while the Pearson correlation comparisons across stages (Fig. 1g) are annotated with significance thresholds without an explicitly stated correction in this excerpt.
    Could also: A multiple-testing correction such as Benjamini-Hochberg FDR or Bonferroni could also be applied uniformly across all multi-comparison panels — Extending a single correction scheme across all families of comparisons in a figure keeps the family-wise or false-discovery error rate consistent throughout the manuscript.
  • Dispersion in boxplots (Fig. 1j) is shown as median with 25th/75th percentiles and min/max whiskers.
    Could also: Reporting mean ± SD or a 95% confidence interval could also be shown alongside or instead — Mean-based summaries with CIs directly complement parametric t-test results and make the estimated effect size and its precision explicit.
  • Very large sample sizes were used for some comparisons (e.g., n=2,000 in the Wilcoxon rank-sum test of pseudotime scores, Fig. 2c), where p-values were extremely small (e.g., P = 3.46 × 10⁻⁸²).
    Could also: An effect size measure such as rank-biserial correlation or Cliff's delta could also be reported alongside the p-value — With very large n, p-values can become vanishingly small even for modest differences, so an effect size helps convey the practical magnitude of the difference.
  • Differences between dorsal and ventral progenitor correlation coefficients were assessed using a t-test (Fig. 1g).
    Could also: A Fisher r-to-z transformation test for comparing two correlation coefficients could also be used — This is a standard approach specifically designed for comparing Pearson correlation coefficients between two independent groups rather than comparing raw values.
  • Differential gene expression between FT cohorts (e.g., FT e12.5+6h vs FT e16.5+6h) was analyzed within the Seurat/Monocle3 workflow.
    Could also: A pseudobulk approach with DESeq2 or edgeR, or a mixed-model method such as MAST, could also be applied — These methods explicitly model cell-to-cell correlation within samples or account for the zero-inflated nature of single-cell count data, which can complement cluster-based differential expression testing.
Software: Seurat · Monocle3

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-40629142 (Bright et al. 2025, Nat Neurosci)

"Temporal control of progenitor competence shapes maturation in GABAergic neuron development in mice." DOI 10.1038/s41593-025-01999-y. Code: https://github.com/mayer-lab/Bright-et-al-2025 (pushed 2025-04-11, public, MIT-ish, not archived). Data: GEO GSE255455 (scRNA-seq) · GSE255104 (scATAC) · GSE255103 (TotalRNA + CUT&RUN).

What the repo actually contains

  • Scripts/*.Rfigure-plotting scripts only (one per figure/panel group). Each source("Scripts/lib.R") then loads a pre-computed object from a Processed_Objects/ directory and writes a per-panel CSV into Results/source_data/.
  • Results/source_data/*.xlsx — the bundled reported figure values (the numbers plotted in the paper). These are the ground-truth "reported" values.
  • Scripts/create_source_data_all.R — just bundles the per-panel CSVs into the xlsx (and hard-codes Fig 1j ephys values inline).
  • Processed_Objects/ is NOT in the repo and NOT a GEO supplementary. This directory holds every integrated/annotated Seurat object (Inhibitory_datasets.Rdata, EXCIT_INHIBIT_cleaned[_sub], CFSE_sub.rds, gNFI_*, nfib_oe_* …), the ArchR peak GRanges (*-reproduciblePeaks.gr.rds), the TOBIAS footprint/BINDetect outputs (bindetect_results.txt, NFIB_1_MA1643.1_all.bed), and the SCENIC+ eRegulon tables (eRegulon_metadata_filtered.tsv, *_eRegulon_AUC_*).
  • The preprocessing / integration / annotation / pseudotime pipeline that PRODUCES Processed_Objects/ is not shipped at all. Only the downstream plotting code is public.

What GEO deposits

  • GSE255455_RAW.tar (2.29 GB) → contains per-sample processed *.Rdata.gz (Sample1–10 = reference + FlashTag + TrackerSeq; Sample41–48 = Nfib/Nfix KO + Nfib OE). These are per-library Seurat objects, NOT the merged/annotated ones.
  • GSE255455_Sample15_16_17_18.Rdata.gz (114 MB) — one integrated object (internal sample numbering; not in the GSM list). This is the only deposited integrated object → the single realistic entry point for a 1:1 panel.

In scope (pipeline-derived, attemptable)

panel script computation determinism
Fig 1g / SF3e Fig1g_SF3e.R 6×6 stage×stage Pearson correlation of AverageExpression (scale.data) over apical progenitors, after NormalizeData→FindVariableFeatures→CellCycleScoring→ScaleData(regress nFeature/nCount/percent.mt/CC) deterministic given the object + Seurat version
Fig 2b Fig2ab.R cell counts/proportions per stage from Inhibitory_datasets deterministic
Fig 2c–e / EDF3 Fig2c_e_*.R CFSE marker DE tables, scaled-expression heatmap matrix deterministic given object
Fig 4h/i, EDF7 Fig4*.R cluster proportion tables from KO/OE Seurat objects deterministic given object

All of the above are gated on having the right Processed_Objects/ object. The deposited Sample15_16_17_18.Rdata is the only candidate; whether it carries the needed Annotated2/Stage_DV2 metadata decides which (if any) panel is reproducible 1:1. Priority target = Fig 1g (smallest, fully scripted, scalar output, clean comparison vs the shipped xlsx).

Out of scope (not attempted; reason)

  • Fig 1j membrane potentials — wet-lab patch-clamp ephys, hard-coded inline. Not computational.
  • UMAP / DimPlot panels (Fig 1c–f, etc.) — coordinates depend on the exact stored reductions in non-deposited objects; visual, no scalar to grade 1:1.
  • scATAC peak calling (ArchR), TOBIAS footprinting, SCENIC+ GRN, CUT&RUN — their outputs are loaded but not deposited; reproducing them means rerunning large upstream pipelines on a different accession (GSE255104/GSE255103) with unscripted parameters → the hard ≥20%, explicitly skipped per 80/20.
  • Rebuilding Processed_Objects/ from RAW.tar — the integration + annotation pipeline is not provided and is stochastic/under-specified → cannot yiel
Figures / tables: Fig 4dFig 4fFig 1g
ko_total_cells
Reported
47079
Reproduced
47079
exact
oe_total_cells
Reported
30019
Reproduced
30019
exact
ko_sgNfib_sgNfix_cells
Reported
5887
Reproduced
5887
exact
ko_sgLacZ_cells
Reported
30328
Reproduced
30328
exact
fig1g_full_matrix
Reported
6x6 apical-progenitor stage correlation (e.g. E12_I x E14_I = 0.704154)
Reproduced
NOT-REPRODUCIBLE (input object EXCIT_INHIBIT_cleaned.Rdata not deposited)
did not match

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 82/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟡2. Endpoint comparability
🟡3. Location of the main deviation
🔴4. Cause of the deviation
🟡5. Derivability / plausibility
🟢6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +6

All four claims gradeable from the deposited scRNA-seq data reproduced exactly to the digit (KO 47,079; OE 30,019; sgNfib/x 5,887; sgLacZ 30,328), which is strong integrity evidence with no fabrication signal. The single non-reproducible item, Fig 1g's 6×6 stage-correlation matrix, fails purely because its input integrated object (Processed_Objects/EXCIT_INHIBIT_cleaned.Rdata) was never deposited in the repo or GEO — a data-availability/incomplete-deposit gap on the authors' side, not a numerical discrepancy. Severity of any measured deviation is zero; the limitation is that the core figure-level analyses cannot be independently verified from public data. Overall a solid partial with an explainable cause.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

323 k
tokens (I/O) · 21.3 M incl. cache
66 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.