Temporal control of progenitor competence shapes maturation in GABAergic neuron development in mice.
The main results reproduced: recomputed values matched the published ones within tolerance.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Any deviation was negligible
- 🟡Could not use the authors’ exact input data
- 🟡Reported values were only indirectly comparable
- 🟡A deviation arose in the data or preprocessing
- 🔴A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
STRONG PARTIAL, described-well-enough for the DEPOSITED data. «our HPC» reachable (prior attempt failed only on VPN 2FA, now cleared); built Seurat 4.4.0/R 4.3.3 on «infra» and ran 2 SLURM jobs. The paper's code repo (mayer-lab/Bright-et-al-2025) ships ONLY downstream figure-plotting R scripts; every script load()s integrated/annotated objects from a Processed_Objects/ dir that is NOT deposited (not in repo, not in GEO). GEO GSE255455 deposits RAW.tar (18 processed per-sample Seurat objects) + ONE integrated object (Sample15_16_17_18, 3498 cells) that is loaded by no script and lacks the Fig1g columns. RESULT: every claim gradeable directly from the deposited data reproduced EXACTLY (4/4) -- Nfib/x KO total = 47,079, Nfib OE total = 30,019, KO sgNfib/sgNfix = 5,887, KO sgLacZ = 30,328 -- all matching the paper to the digit (strong integrity evidence, no fabrication signal on these). NOT reproducible 1:1: Fig 1g stage-correlation (its input integrated object is not deposited; code shipped, input not shipped) -- this is a data-availability gap, not a fabrication finding. NOT attempted (out of 80/20 scope): scATAC/ArchR/TOBIAS/SCENIC+/CUT&RUN on separate accessions, and UMAP/pseudotime visual panels. Honest verdict: the deposited scRNA-seq data delivers its headline KO/OE numbers exactly; the published figures are only partly reproducible from public data+code because the figure-input integrated objects are largely undeposited.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 82assessed: 2026-06-20 ⛓ 865d20d5d6df
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-20
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-20no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetThe paper tests how progenitor competence (maturation competence and differentiation competence) in ganglionic eminence progenitors shapes the maturation and diversification of GABAergic neurons, and how this compares to the temporally progressive competence changes seen in dorsal cortical progenitors.
- ★ Ganglionic eminence (ventral) progenitors maintain stable differentiation competence throughout neurogenesis, generating a consistent set of postmitotic precursor states at all stages, unlike dorsal cortical progenitors whose differentiation competence changes gradually. finding
- ★ Maturation competence changes with developmental timing: late-born (e16.5) GE neurons mature faster than early-born (e12.5) neurons, reaching similar transcriptional states in less time. finding
- ★ Chromatin remodeling together with a regulatory module centered on the transcription factor NFIB and its target genes drives increased maturation competence in late-born progenitors. mechanism
- ★ Ventral GE progenitors show stable resting membrane potential across e13.5-e15.5, in contrast to dorsal cortical progenitors, which progressively hyperpolarize over the same period. finding
- ★ Clonal lineage tracing (TrackerSeq) shows similar proportions of dispersing (multi-fate) clones at e12.5 and e16.5, indicating differentiation competence is maintained at the clonal level across neurogenesis. finding
- ★ Heterochronic transplantation experiments show that maturation competence of progenitors is influenced by the extrinsic tissue environment. finding
- An interactive web-based resource is provided for exploring scRNA-seq, scATAC-seq, CUT&RUN and eGRN datasets comparing GE- and cortex-derived neurogenesis. resource
- A multi-modal approach combining scRNA-seq, TrackerSeq barcode lineage tracing, FlashTag birthdating, perturbation sequencing and CUT&RUN was used to dissect progenitor competence. method
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| scRNA-seq | Dlx5/6-Cre::tdTomato mouse embryos (cortex, striatum, GE), e12.5/e14.5/e16.5 | none | transcriptomic cell states along developmental pseudotime | Seurat, Monocle3 |
| TrackerSeq (DNA barcode lineage tracing + scRNA-seq) | GE progenitors, mouse embryos, in utero electroporation at e12.5 or e16.5, collected 96h later | none (lineage labeling) | clonal relationships and dispersal across postmitotic branch tips | — |
| FlashTag (FT) birthdating + scRNA-seq | GE progenitors, wild-type and Dlx5/6-Cre::tdTomato mouse embryos, e12.5 and e16.5, collected 6h or 96h later | none (isochronic cohort labeling with CFSE) | pseudotime/maturation state of isochronic cell cohorts | — |
| RNAscope in situ hybridization | GE tissue sections, mouse embryos | none | expression of Ascl1 (intermediate progenitor) and Gad2 (postmitotic precursor) markers | — |
| Whole-cell patch-clamp electrophysiology | Cortical and GE progenitors, mouse embryos, e13.5 and e15.5 | none | resting membrane potential | — |
| scATAC-seq / eGRN inference | GE progenitors, mouse embryos | none | chromatin accessibility and enhancer-driven gene regulatory networks | — |
| CUT&RUN | GE progenitors/neurons, mouse | none | NFIB genomic binding/target validation | — |
| Perturbation sequencing | GE progenitors, mouse embryos | NFIB perturbation | transcriptomic effects on maturation competence | — |
- ▲ Ventral (GE) progenitors show higher Pearson correlation coefficients between successive neurogenesis stages than dorsal progenitors, indicating less transcriptomic change over time
- – Dorsal cortical progenitors progressively hyperpolarize between e12.5 and e15.5, while ventral GE progenitor membrane potential remains stable
- ▲ Pseudotime score of FT e16.5+6h cohort is markedly higher than FT e12.5+6h despite both being collected 6h after labeling
- ▲ Genes upregulated in FT e16.5+6h overlap substantially with genes upregulated in FT e12.5+96h, indicating late-born neurons reach a similar expression profile much faster
- – Similar proportion of dispersing (multi-branch-tip) clones observed in TrackerSeq e12.5+96h and TrackerSeq e16.5+96h
- – Pseudotime distributions of FT+ cells differ significantly across conditions (FT e12.5+6h, e16.5+6h, e12.5+96h) P = 3.46×10^-82, 2.87×10^-255, 1.06×10^-107
- – Mitotic progenitor cells from nondispersing clones show no stronger transcriptomic correlation with their postmitotic clonal progeny than randomly selected progenitors
- pvalue P = 3.46 × 10^-82, 2.87 × 10^-255, 1.06 × 10^-107 (****P < 0.0001) (Wilcoxon rank-sum test comparing pseudotime distributions of FT+ cohorts)
- pvalue *P < 0.05, **P < 0.01 (Two-sided t-test, Pearson correlation between dorsal and ventral progenitors across stages)
- count n = 41,460 cells; n = 40 embryos (Combined scRNA-seq, TrackerSeq and FT dataset (UMAP))
- count n = 25,297 cells; n = 20 embryos (scRNA-seq dataset across collection stages)
- count n = 75,431 cells (Combined ventral (GABAergic) and dorsal (glutamatergic) telencephalon dataset)
- count n = 9,938 cells (TrackerSeq barcoded cells across IUE stages)
- count n = 6,225 cells; n = 19 embryos (FlashTag (FT) dataset)
- count Cortex n = 14; GE n = 37 (Whole-cell patch-clamp membrane potential recordings at e13.5/e15.5)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
The paper primarily reports single-cell transcriptomic and electrophysiological comparisons using nonparametric and parametric two-group tests (Wilcoxon rank-sum, two-sided t-tests) applied to Monocle3 pseudotime scores, Pearson correlation coefficients, and membrane potential recordings, with results shown as boxplots/violin plots annotated with significance asterisks or exact p-values. Sample sizes are reported at both the cell level (n cells) and embryo/animal level (n embryos or n recorded cells per group) depending on the analysis. Differential gene expression and clustering analyses were performed within the Seurat/Monocle3 single-cell analysis framework, and one comparison (membrane potential across cortex/GE and stages) explicitly applied a false discovery rate correction.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| two-sided t-test | Pearson's correlation comparison between dorsal and ventral progenitors across developmental stages (Fig. 1g) | — | not stated |
| two-sided, unpaired t-test with false discovery rate correction | membrane potential comparison between cortical and GE progenitors at e13.5/e15.5 (Fig. 1j) | Cortex n=14; GE n=37 | not stated |
| two-sided, unpaired Wilcoxon rank-sum test | pseudotime score distributions across FT birthdating cohorts (Fig. 2c) | n = 2,000 | not stated |
-
Two-group comparisons (e.g., membrane potential in Fig. 1j, n=14 vs n=37) were tested with two-sided unpaired t-tests.↳ Could also: A nonparametric test such as the Mann-Whitney U (Wilcoxon rank-sum) test could also be used — This avoids relying on the normality assumption, which can be harder to verify with smaller group sizes such as n=14.
-
A false discovery rate correction was applied to the membrane potential comparisons (Fig. 1j), while the Pearson correlation comparisons across stages (Fig. 1g) are annotated with significance thresholds without an explicitly stated correction in this excerpt.↳ Could also: A multiple-testing correction such as Benjamini-Hochberg FDR or Bonferroni could also be applied uniformly across all multi-comparison panels — Extending a single correction scheme across all families of comparisons in a figure keeps the family-wise or false-discovery error rate consistent throughout the manuscript.
-
Dispersion in boxplots (Fig. 1j) is shown as median with 25th/75th percentiles and min/max whiskers.↳ Could also: Reporting mean ± SD or a 95% confidence interval could also be shown alongside or instead — Mean-based summaries with CIs directly complement parametric t-test results and make the estimated effect size and its precision explicit.
-
Very large sample sizes were used for some comparisons (e.g., n=2,000 in the Wilcoxon rank-sum test of pseudotime scores, Fig. 2c), where p-values were extremely small (e.g., P = 3.46 × 10⁻⁸²).↳ Could also: An effect size measure such as rank-biserial correlation or Cliff's delta could also be reported alongside the p-value — With very large n, p-values can become vanishingly small even for modest differences, so an effect size helps convey the practical magnitude of the difference.
-
Differences between dorsal and ventral progenitor correlation coefficients were assessed using a t-test (Fig. 1g).↳ Could also: A Fisher r-to-z transformation test for comparing two correlation coefficients could also be used — This is a standard approach specifically designed for comparing Pearson correlation coefficients between two independent groups rather than comparing raw values.
-
Differential gene expression between FT cohorts (e.g., FT e12.5+6h vs FT e16.5+6h) was analyzed within the Seurat/Monocle3 workflow.↳ Could also: A pseudobulk approach with DESeq2 or edgeR, or a mixed-model method such as MAST, could also be applied — These methods explicitly model cell-to-cell correlation within samples or account for the zero-inflated nature of single-cell count data, which can complement cluster-based differential expression testing.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-40629142 (Bright et al. 2025, Nat Neurosci)
"Temporal control of progenitor competence shapes maturation in GABAergic neuron development in mice." DOI 10.1038/s41593-025-01999-y. Code: https://github.com/mayer-lab/Bright-et-al-2025 (pushed 2025-04-11, public, MIT-ish, not archived). Data: GEO GSE255455 (scRNA-seq) · GSE255104 (scATAC) · GSE255103 (TotalRNA + CUT&RUN).
What the repo actually contains
Scripts/*.R— figure-plotting scripts only (one per figure/panel group). Eachsource("Scripts/lib.R")then loads a pre-computed object from aProcessed_Objects/directory and writes a per-panel CSV intoResults/source_data/.Results/source_data/*.xlsx— the bundled reported figure values (the numbers plotted in the paper). These are the ground-truth "reported" values.Scripts/create_source_data_all.R— just bundles the per-panel CSVs into the xlsx (and hard-codes Fig 1j ephys values inline).Processed_Objects/is NOT in the repo and NOT a GEO supplementary. This directory holds every integrated/annotated Seurat object (Inhibitory_datasets.Rdata,EXCIT_INHIBIT_cleaned[_sub],CFSE_sub.rds,gNFI_*,nfib_oe_*…), the ArchR peak GRanges (*-reproduciblePeaks.gr.rds), the TOBIAS footprint/BINDetect outputs (bindetect_results.txt,NFIB_1_MA1643.1_all.bed), and the SCENIC+ eRegulon tables (eRegulon_metadata_filtered.tsv,*_eRegulon_AUC_*).- The preprocessing / integration / annotation / pseudotime pipeline that
PRODUCES
Processed_Objects/is not shipped at all. Only the downstream plotting code is public.
What GEO deposits
GSE255455_RAW.tar(2.29 GB) → contains per-sample processed*.Rdata.gz(Sample1–10 = reference + FlashTag + TrackerSeq; Sample41–48 = Nfib/Nfix KO + Nfib OE). These are per-library Seurat objects, NOT the merged/annotated ones.GSE255455_Sample15_16_17_18.Rdata.gz(114 MB) — one integrated object (internal sample numbering; not in the GSM list). This is the only deposited integrated object → the single realistic entry point for a 1:1 panel.
In scope (pipeline-derived, attemptable)
| panel | script | computation | determinism |
|---|---|---|---|
| Fig 1g / SF3e | Fig1g_SF3e.R | 6×6 stage×stage Pearson correlation of AverageExpression (scale.data) over apical progenitors, after NormalizeData→FindVariableFeatures→CellCycleScoring→ScaleData(regress nFeature/nCount/percent.mt/CC) |
deterministic given the object + Seurat version |
| Fig 2b | Fig2ab.R | cell counts/proportions per stage from Inhibitory_datasets |
deterministic |
| Fig 2c–e / EDF3 | Fig2c_e_*.R | CFSE marker DE tables, scaled-expression heatmap matrix | deterministic given object |
| Fig 4h/i, EDF7 | Fig4*.R | cluster proportion tables from KO/OE Seurat objects | deterministic given object |
All of the above are gated on having the right Processed_Objects/ object.
The deposited Sample15_16_17_18.Rdata is the only candidate; whether it carries
the needed Annotated2/Stage_DV2 metadata decides which (if any) panel is
reproducible 1:1. Priority target = Fig 1g (smallest, fully scripted, scalar
output, clean comparison vs the shipped xlsx).
Out of scope (not attempted; reason)
- Fig 1j membrane potentials — wet-lab patch-clamp ephys, hard-coded inline. Not computational.
- UMAP / DimPlot panels (Fig 1c–f, etc.) — coordinates depend on the exact stored reductions in non-deposited objects; visual, no scalar to grade 1:1.
- scATAC peak calling (ArchR), TOBIAS footprinting, SCENIC+ GRN, CUT&RUN — their outputs are loaded but not deposited; reproducing them means rerunning large upstream pipelines on a different accession (GSE255104/GSE255103) with unscripted parameters → the hard ≥20%, explicitly skipped per 80/20.
- Rebuilding
Processed_Objects/from RAW.tar — the integration + annotation pipeline is not provided and is stochastic/under-specified → cannot yiel
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
All four claims gradeable from the deposited scRNA-seq data reproduced exactly to the digit (KO 47,079; OE 30,019; sgNfib/x 5,887; sgLacZ 30,328), which is strong integrity evidence with no fabrication signal. The single non-reproducible item, Fig 1g's 6×6 stage-correlation matrix, fails purely because its input integrated object (Processed_Objects/EXCIT_INHIBIT_cleaned.Rdata) was never deposited in the repo or GEO — a data-availability/incomplete-deposit gap on the authors' side, not a numerical discrepancy. Severity of any measured deviation is zero; the limitation is that the core figure-level analyses cannot be independently verified from public data. Overall a solid partial with an explainable cause.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.