Sox8 remodels the cranial ectoderm to generate the ear.
The main results reproduced: recomputed values matched the published ones within tolerance.
- ✓Same input data as the authors
- ✓Reported values were directly comparable
- ✓No relevant deviation in data/preprocessing
- ✓No authors-side cause for any deviation
- ✓Reported values are derivable from the shared data
- ✓Any deviation was negligible
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
REPRODUCED (1:1, exact). Repo alexthiery/otic-reprogramming pinned at commit b28122f ships the bulk RNA-seq featureCounts matrices, enabling self-contained reproduction of the two DESeq2 DEA claims without re-aligning. In the authors' pinned environment (R 3.6.3, DESeq2 1.26.0, apeglm 1.8.0, Bioconductor 3.10) on a «our HPC» compute node (SLURM «job»), the deterministic DE core of bin/{sox8_dea.R,lmx1a_dea.R} (drop annotation cols, group by colname, DESeqDataSetFromMatrix ~Group, relevel ref, filter rowSums>=10, DESeq median-of-ratios, lfcShrink apeglm, count padj<0.05 & |log2FC|>1.5) reproduced EVERY reported number to the exact integer: Sox8-OE vs Control = 399 up / 112 down / 511 total (Dataset S5; '399 transcripts up', '112 down'); Lmx1aE1 vs Sox3U3 = 103 up-otic / 319 up-epibranchial / 422 total (Dataset S1; '103 and 319 genes up-regulated'). No fabrication concern: all six numbers are exactly derivable from the shipped data+code. Dataset GSE168089 profiled: GEO reports 801 samples for the SuperSeries (dominated by SmartSeq2 single cells); the 14 bulk RNA-seq samples used (3+3, 4+4; 24356 gene rows each) match the described designs and the shipped matrices fully deliver the DEA results (quality A, provisional). NOT attempted (out of quick scope, honestly declared): enhancer/motif/GO, peak annotation, and SmartSeq2 scRNA analyses (require non-shipped intermediates / full FASTQ re-alignment), and all wet-lab results.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessmentassessed: 2026-06-19 ⛓ 28eaa8ae0f28
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-26
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetThe paper investigates the transcriptional and epigenetic mechanisms that drive segregation of otic-epibranchial progenitors into distinct otic and epibranchial fates, testing whether a small transcriptional circuit—centered on Sox8—can act as a master regulator sufficient to impart ear identity on non-ear cranial ectoderm.
- ★ Sox8 sits at the top of a core transcriptional circuit (Sox8, Pax2, Lmx1a, Zbtb16) that determines otic identity mechanism
- ★ Sox8 loss-of-function abolishes expression of all other assayed core ear TFs, while Pax2/Lmx1a act in a positive feedback loop and Zbtb16 is required for Foxg1 finding
- ★ Misexpression of Sox8 alone in non-ear cranial ectoderm is sufficient to activate the ear program, forming ectopic otic vesicles with associated neurofilament-positive neurons finding
- ★ Otic-epibranchial progenitors transiently coexpress otic (Sox8, Lmx1a) and epibranchial (Foxi3, Tfap2e) markers before resolving into distinct branches finding
- Genome-wide profiling (ATACseq, H3K27ac/H3K4me3/H3K27me3 ChIPseq) in OEPs identifies 8,316 putative enhancers resource
- Motif enrichment across putative ear enhancers reveals overrepresentation of Sox, TEAD, Six, and Tfap2a binding sites finding
- The novel Lmx1aE1 enhancer drives EGFP expression specifically in the otic placode ectoderm resource
- ★ Combined misexpression of Sox8/Pax2/Lmx1a(/Zbtb16) activates ear-reporter expression and forms otic vesicles, but combinations lacking Sox8 fail to do so finding
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| bulk RNAseq | chick embryo, FACS-sorted Lmx1aE1-EGFP+ (otic) and Sox3U3-EGFP+ (epibranchial) cells | none (reporter labeling only) | differentially expressed genes between otic and epibranchial cell populations | — |
| single-cell RNAseq | chick embryo, FACS-sorted Pax2E1-EGFP+ cells at OEP (ss8-9), early-placode (ss11-12), and late-placode (ss14-15) stages | none | cell clustering, pseudotime trajectories, RNA velocity of otic/epibranchial fate segregation | Monocle2 (pseudotime), RNA velocity |
| ChIPseq | chick ss8-9 otic-epibranchial progenitors (OEPs) | none | genome-wide H3K27ac, H3K4me3, and H3K27me3 histone mark distribution | — |
| ATACseq | chick ss8-9 OEPs | none | chromatin accessibility / open regulatory regions | — |
| in vivo enhancer reporter electroporation | chick head-fold-stage embryos | electroporation of candidate enhancer-EGFP constructs with constitutive mCherry | EGFP fluorescence indicating enhancer activity in ear progenitors/otic placode | — |
| antisense oligonucleotide (aON) knockdown | chick OEP (head-fold-stage embryos) | unilateral knockdown of Sox8, Pax2, Zbtb16, or Lmx1a (with mRNA rescue) | in situ hybridization for otic markers (Foxg1, Soho1, and other core TFs) | — |
| gain-of-function electroporation / lineage conversion | chick non-ear head-fold cranial ectoderm | overexpression of Sox8, Pax2, Lmx1a, Zbtb16 alone or in combination (mCherry-tagged) | Lmx1aE1-EGFP reporter activation, Soho1+ ectopic otic vesicle formation, neurofilament+ neuron induction | — |
| bulk RNAseq | chick cranial ectoderm, FACS-sorted Sox8-mCherry/Lmx1aE1-EGFP double-positive cells vs constitutive mCherry/EGFP controls | Sox8 overexpression | differential gene expression (otic gene enrichment) | — |
- – 103 genes up-regulated in otic (Lmx1aE1-EGFP+) cells and 319 genes up-regulated in epibranchial (Sox3U3-EGFP+) cells log2FC > 1.5, adjusted P < 0.05
- ▼ Significantly more cells coexpress otic and epibranchial markers before the branch point than after, in both epibranchial and otic lineages W = 214, P = 0.0013 (epibranchial); W = 235, P < 0.0001 (otic)
- – 10,969 H3K27ac+/ATACseq+/H3K27me3-depleted genomic regions identified in OEPs, of which 8,316 are putative intergenic/intronic enhancers 8,316 enhancers (~70% of 10,969)
- ▲ Motif enrichment analysis of putative ear enhancers shows overrepresentation of Sox, TEAD, Six, and Tfap2a binding motifs
- ▼ Sox8 knockdown abolishes expression of all other core ear TFs (Pax2, Zbtb16, Lmx1a) and downstream otic markers, rescued by full-length Sox8 co-electroporation
- ▲ Sox8/Pax2/Lmx1a misexpression in non-ear ectoderm induces Soho1+ ectopic otic vesicles with neurofilament+ neurons; combinations lacking Sox8 fail to activate the reporter
- ▲ Sox8 alone activates Lmx1aE1-EGFP reporter and Pax2, Lmx1a, Soho1 expression, forming ectopic vesicles with neuronal projections
- – RNAseq of Sox8-overexpressing cranial ectoderm shows enrichment of otic genes relative to control 399 genes up-regulated, 112 down-regulated (log2FC > 1.5, adjusted P < 0.05)
- pvalue P = 0.0013 (W = 214) (Wilcoxon rank sum test: coexpression of otic/epibranchial markers before vs. after branch point, epibranchial branch)
- pvalue P < 0.0001 (W = 235) (Wilcoxon rank sum test: coexpression of otic/epibranchial markers before vs. after branch point, otic branch)
- fold_change log2FC > 1.5, adjusted P < 0.05 (differential expression significance threshold applied across RNAseq comparisons)
- count 103 genes up in otic cells; 319 genes up in epibranchial cells (bulk RNAseq differential expression, otic vs. epibranchial FACS-sorted cells)
- count 10,969 genomic regions (overlapping H3K27ac and ATACseq peaks depleted of H3K27me3 in OEPs)
- count 8,316 putative enhancers (intergenic/intronic subset (>70%) of the 10,969 candidate regulatory regions)
- count 399 up-regulated, 112 down-regulated genes (RNAseq of Sox8-overexpressing vs. control cranial ectoderm)
- other Sox8 expression detectable from 3 somite stage (ss3) (in situ hybridization timing of onset of core otic TFs relative to Pax2, Zbtb16, Lmx1a, Foxg1, Soho1, Sox10)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
The study combines bulk and single-cell RNA sequencing, epigenomic profiling (ChIPseq, ATACseq), and in vivo loss-/gain-of-function experiments in chick embryos to characterize ear versus epibranchial fate specification. Differential expression was assessed using a log2 fold-change and adjusted-P-value threshold, pseudotime trajectories were modeled with Monocle2/BEAM and cross-validated with RNA velocity, and a proportion comparison between cell populations was tested with a two-tailed Wilcoxon rank-sum test. Results are reported primarily through fold-change/adjusted-P thresholds, exact statistics for the Wilcoxon comparisons, and qualitative/image-based readouts (in situ hybridization, reporter activation) for most functional perturbation experiments.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Differential expression analysis (log2 fold change and adjusted P value threshold) | Lmx1aE1-EGFP+ (otic) vs Sox3U3-EGFP+ (epibranchial) bulk RNAseq (Fig. 1D); Sox8-overexpression vs control ectoderm RNAseq (Fig. 4F); Sox8-mCherry/Lmx1aE1-EGFP+ vs control cells (transcriptome comparison) | — | not stated |
| Two-tailed Wilcoxon rank-sum test | Proportion of cells coexpressing otic/epibranchial markers before vs after the branching point (Fig. 2B, C) | cell counts underlying W statistics (W=214, W=235); exact cell numbers not stated | not stated |
| Branch expression analysis modeling (BEAM, via Monocle2) | Identification of genes differentially regulated along otic vs epibranchial pseudotime branches (Fig. 2A) | — | not stated |
| RNA velocity analysis | Validation of pseudotime trajectory directionality (Fig. 1P) | — | not stated |
| Motif enrichment analysis | Transcription factor binding site enrichment across 8,316 putative ear enhancers (Fig. 3C) | — | not stated |
-
Differential expression between cell populations was reported using a log2 fold-change and adjusted P-value cutoff, without naming the specific multiple-testing correction method or the software/package used to generate adjusted P values.↳ Could also: Explicitly stating the correction method (e.g., Benjamini-Hochberg FDR) and the software/version (e.g., DESeq2, edgeR, limma-voom) used for differential expression testing — This would let readers evaluate the stringency of the multiple-comparisons control and reproduce the analysis pipeline exactly.
-
A two-tailed Wilcoxon rank-sum test was used to compare the proportion of coexpressing cells before versus after the branching point.↳ Could also: A generalized linear model for proportion/count data (e.g., logistic or beta-binomial regression) or a permutation test on the proportion difference — These approaches can incorporate additional covariates (e.g., stage, embryo of origin) and provide effect-size estimates with confidence intervals alongside the significance test.
-
Variability in Fig. 2B is displayed as error bars of ±1 SD.↳ Could also: Reporting a 95% confidence interval for the proportion estimates — A CI directly communicates the precision of the estimated difference in coexpression proportions, which can be more directly interpretable than SD for proportion data.
-
Most functional loss- and gain-of-function experiments (antisense knockdown, rescue, overexpression) are reported qualitatively via in situ hybridization/imaging rather than with a quantitative statistical test.↳ Could also: Quantifying the proportion of embryos showing the phenotype (e.g., contingency table with Fisher's exact test or chi-square test) across a stated number of biological replicates — Quantitative summary statistics with a formal test would allow readers to gauge the consistency and strength of the phenotypic effect across replicate embryos.
-
Significance for the Wilcoxon comparisons is also denoted by asterisk thresholds (**P<0.01, ***P<0.001) in the figure legend in addition to the exact P values in text.↳ Could also: Reporting exact P values consistently in the figure legend alongside or instead of threshold-based asterisks — Exact values throughout (rather than only in the main text) support more precise downstream reanalysis or meta-analysis.
-
Multiple distinct RNAseq comparisons (otic vs epibranchial; Sox8-OE vs control; Sox8/Lmx1aE1 double-positive vs control) each apply the same fold-change/adjusted-P threshold independently.↳ Could also: A unified linear-model framework with defined contrasts (e.g., limma or DESeq2 with a single design matrix) applied across all comparisons, with multiplicity control considered across the full set of tests performed in the study — This can help ensure consistent control of the family-wise or false-discovery rate across the entire set of comparisons made in the paper, not just within each individual comparison.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-35867760
Paper: Buzzi AL, Chen J, Thiery A, Delile J, Streit A. "Sox8 remodels the cranial
ectoderm to generate the ear." PNAS 2022. DOI 10.1073/pnas.2118938119.
Repo: https://github.com/alexthiery/otic-reprogramming (MIT, branch master,
last push 2023-05-12). Data: GEO GSE168089 (SuperSeries).
The repo is a set of Nextflow (DSL2) pipelines + R analysis scripts. Crucially it
ships the alignment outputs needed by the downstream analysis under
alignment_output/ (bulk RNA-seq count matrices and ChIP/ATAC peak files), so a
subset of the downstream results is reproducible without re-running alignment.
Pipelines in the repo
| Pipeline | Type | Input | Output |
|---|---|---|---|
| NF-sox8_alignment | bulk RNA-seq aln+quant (nf-core/rnaseq-like, STAR+featureCounts) | FASTQ (GEO) | featurecounts.merged.counts.tsv (shipped) |
| NF-lmx1a_alignment | bulk RNA-seq aln+quant | FASTQ (GEO) | featurecounts.merged.counts.tsv (shipped) |
| NF-ChIP_alignment | ChIP-seq (BWA + MACS2) | FASTQ | broadPeak (3 marks, shipped) |
| NF-ATAC_alignment | ATAC-seq (nf-core/atacseq) | FASTQ | narrowPeak (shipped) |
| NF-smartseq2_alignment | SmartSeq2 scRNA-seq aln+quant+velocyto | FASTQ | merged counts + loom (not shipped) |
| NF-downstream_analysis | R analysis (DESeq2, Antler, enhancer/motif/GO) | the above | DEA tables, figures |
In scope (pipeline-derived, attempted)
- sox8_dea — DESeq2 differential expression, Sox8_OE vs Control (6 samples:
3 control + 3 sox8_oe). Reads shipped
alignment_output/NF-sox8_alignment/featurecounts.merged.counts.tsv. ScriptNF-downstream_analysis/bin/sox8_dea.R. Self-contained. Reported (repo comment + Supplementary Data 5): 511 DE genes (399 up, 112 down) at padj<0.05 & |log2FC|>1.5. - lmx1a_dea — DESeq2 differential expression, Lmx1a_E1 vs Sox3U3 (8 samples:
4+4). Reads shipped
alignment_output/NF-lmx1a_alignment/featurecounts.merged.counts.tsv. ScriptNF-downstream_analysis/bin/lmx1a_dea.R. Self-contained. Reported (repo comment + Supplementary Data 1): 422 DE genes (103 up, 319 down) at padj<0.05 & |log2FC|>1.5.
These two are the QUICK MINIMUM (80%) target: precise, pinnable integer claims with shipped inputs and a pinned env (authors' Docker image rocker/tidyverse:3.6.3 + Bioconductor 3.10 → DESeq2 1.26, apeglm).
Stretch (attempt if quick targets land)
- enhancer_analysis — putative enhancers = ATAC∩H3K27Ac − H3K27me3, HOMER motif enrichment, gProfiler GO. Needs bigWig + peak dirs; repo ships only summary peaks, not the full bwa/mergedLibrary tree the downstream main.nf globs → likely needs re-running ChIP/ATAC alignment first. Lower priority.
- peak_annotations_frequency — ChIPpeakAnno feature distribution of peaks.
- smartseq_analysis — Antler scRNA clustering + velocyto. Needs merged_counts + loom which are NOT shipped → requires full smartseq2 alignment. Out of quick scope.
Out of scope (not pipeline-derived)
- All wet-lab results: in-situ hybridisation, immunostaining, electroporation phenotypes, ear-vesicle/neuron morphology, qPCR validation (qPCR/qpcr.R is a small barplot, manual Ct input — borderline, not attempted).
- Raw alignment re-run (NF-*_alignment) is in principle reproducible from GEO FASTQ but heavy and not needed for the DEA claims (outputs shipped); attempted only if time permits for the enhancer/smartseq results.
Pipeline naming per result
- sox8_dea / lmx1a_dea → DESeq2 (median-of-ratios norm, apeglm LFC shrinkage), threshold padj<0.05 & |log2FC|>1.5.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.