Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Phase transition specified by a binary code patterns the vertebrate eye cup.

Sci Adv · 2021
L1 74/100 3/4
Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
✓ What held up
  • Reported values were directly comparable
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
74/100
Reproducibility score
at the mean
vs. all fields · 1173 studies
🎯 Scores higher than 43% of all assessed papers rank 644 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

REPRODUCED (partial, honest). Balasubramanian 2021 Sci Adv, mouse optic-cup/ciliary-margin scRNA-seq, GSE139904 (one pooled GSM4148508). Ran Seurat v4 on the deposited CellRanger filtered matrix on «our HPC» (SLURM 2212619) and cross-checked against the authors' deposited graphclust + diffexp. CORE RESULTS REPRODUCE 1:1: C1 total cells 11239 vs reported 11235 (Delta 4, 0.04%); C2 mean genes/cell 2826.9 vs 2811 (0.57%); C4 CM markers Mitf/Wls/Msx1/Wfdc1 enriched in distinct CM clusters, confirmed both in our Seurat run AND in the authors' own deposited diffexp (exact). C3 clustering: all ~13 cell types recovered by canonical markers, exact cluster count differs (pooled-vs-per-genotype, Seurat v4-vs-v3) -> partial. NOT reproducible from the deposit: (a) the 6628/4607 control/mutant split (no genotype/demux label deposited), (b) the velocyto RNA-velocity trajectory (no BAM/loom deposited) -- the paper's headline 'binary code' model -> documented blocker, extendable via SRA re-derivation, not forced. Out of scope: all wet-lab/imaging. Dataset profiled: GEO open, N matches reported total, grade B (lacks genotype label + BAM/loom).

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 74
    assessed: 2026-06-21 ⛓ d6f62e10f506
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-21
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The paper tests whether the vertebrate eye cup (neural retina, RPE, ciliary margin) is specified by a combinatorial/binary code of FGF and Wnt signaling acting as a phase-transition mechanism, rather than by mutual inhibition of the two pathways as previously proposed.

Core claims
  • FGF signaling is required for ciliary margin (CM) development; loss of FGFRs in peripheral retina abolishes CM markers and causes aniridia finding
  • FGF signaling controls self-renewal versus differentiation and survival of CM progenitor cells, identified via single-cell RNA velocity analysis finding
  • Graded/nested expression of Fgf3, Fgf9, and Fgf15 patterns subdivision of the CM into distal, medial, and proximal domains in a dose-dependent manner mechanism
  • FGF signaling is required to maintain Wnt pathway activity (Lef1, Axin2) in the peripheral retina, contrary to prior models of mutual FGF-Wnt inhibition mechanism
  • Titrating FGF signaling strength in Fgf8-overexpressing retinas redirects RPE fate toward CM rather than NR finding
  • Single-cell RNA sequencing combined with RNA velocity analysis of Cre/GFP-sorted peripheral retinal cells method
  • Lineage tracing using Ai9 tdTomato reporter with Pax6 alpha-Cre and tamoxifen-inducible Msx1-CreERT2 to pulse-label CM progenitors method
Experimental setups
Assay System Perturbation Readout Platform
Immunostaining/in situ marker analysis Mouse embryonic/postnatal eye cup (E13.5-P7, adult) Fgfr1/Fgfr2 conditional knockout (Pax6 alpha-Cre) pERK, Etv5, Spry2, Vsx2, Sox2, Pax6, NICD, Gli1, Sfrp2, Atoh7, Mitf, Pcad, Cx43, Wfdc1 expression
Single-cell RNA sequencing (scRNAseq) Mouse E13.5 eye cup, Cre/GFP-sorted peripheral retinal cells Fgfr1/Fgfr2 conditional knockout vs control Transcriptomic clusters (RPC, neurogenic, RGC/AC/HC/PRC, CM subtypes), gene expression
RNA velocity / diffusion trajectory analysis Mouse E13.5 eye cup scRNAseq data Fgfr1/Fgfr2 conditional knockout vs control Cell differentiation trajectory, transition probabilities, cell cycle state
Lineage tracing (Cre-lox reporter) Mouse retina, Ai9 tdTomato reporter with Pax6 alpha-Cre or tamoxifen-induced Msx1-CreERT2 Fgfr1/Fgfr2 conditional knockout with pulse-labeling Persistence of tdTomato+ labeled cells (survival)
Conditional gene knockout/expression analysis Mouse retina Fgf9 or Fgf3/Fgf9 conditional knockout (Pax6 alpha-Cre) Mitf, Vsx2, Wfdc1, Msx1 expression domains
Ectopic overexpression/rescue immunostaining Mouse retina/RPE Fgf8 overexpression (R26 LSL-Fgf8) with or without Fgfr1/Fgfr2 deletion pERK, Spry2, Atoh7, Otx1, Msx1, Cdo expression (RPE-to-NR/CM conversion)
Key results
  • Fgfr deletion abolished pERK, Etv5, and Spry2 expression in distal retina
  • Fgfr mutants showed ectopic Mitf, Pcad, Cx43 expression with reduced Wfdc1, indicating loss of CM domain and aniridia in adults
  • scRNAseq identified three CM clusters (distal Wls/Otx2, medial Msx1, proximal Sox2/Cdo) with RNA velocity showing bidirectional self-renewal/differentiation capacity of CM progenitors; Fgfr mutants biased toward differentiation over self-renewal
  • Sequential loss of Fgf9 then Fgf3/Fgf9 caused progressive expansion of Mitf and reduction of Vsx2, Wfdc1, and Msx1
  • Lef1 and Axin2 (Wnt response genes) were down-regulated in Fgfr ΔRet mutants
  • Reducing FGF signaling strength in Fgf8-overexpressing RPE prevented Atoh7 induction but induced CM markers Otx1 and Msx1 instead
  • Msx1-CreERT2 pulse-labeled CM cells largely disappeared in Fgfr ΔMsx1 mutants by E18.5, indicating a cell survival defect
Key statistics
  • count 6628 control and 4607 Fgfr ΔRet mutant cells sequenced (scRNAseq of E13.5 eye cups)
  • mean mean depth of 2811 genes per cell (scRNAseq sequencing depth)
  • pvalue *P < 0.01, **P < 0.001, ***P < 0.0001 (one-way ANOVA) (Marker gene area quantification in Fgf9ΔRet/Fgf3/9ΔRet and Fgf8OE/FgfrΔRet;Fgf8OE mutants)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The paper uses single-cell RNA sequencing (scRNAseq) with unsupervised clustering, UMAP dimensionality reduction, diffusion-based pseudotime analysis, and RNA velocity to characterize cell populations and differentiation trajectories in the developing mouse eye cup (E13.5). Quantitative comparisons of immunostaining marker-gene expression domains across genotypes were evaluated by one-way ANOVA, with significance coded by threshold symbols. Findings were supported by multiple conditional knockout and overexpression mouse lines, each assessed at n = 3 animals for quantified area measurements. The provided text is truncated before the full Methods section, so additional statistical procedures may exist but are not visible here.

Replicationbiological Sample sizen = 3 animals stated for quantified area measurements (Fig. 4D); 6628 control and 4607 mutant cells stated for scRNAseq; no formal power calculation described in provided text GroupsControl vs. FgfrΔRet, Fgf9ΔRet, Fgf3/9ΔRet, Fgf8OE, FgfrΔRet;Fgf8OE, FgfrΔMsx1 conditional mouse genotypes Pairingunpaired Randomization/blindingnot stated Dispersionnone Exact p-valuesno Effect sizesno Confidence intervalsno Multiplicity correctionnone stated
Statistical tests used
Test Applied to n Assumptions
One-way ANOVA Relative area of marker-gene expression (Mitf, Vsx2, Wfdc1, Msx1, Atoh7, Otx1, Cdo) normalized to eye cup size, compared across genotypes (control, Fgf9ΔRet, Fgf3/9ΔRet, FgfrΔRet;Fgf8OE) — Fig. 4D n = 3 for all markers not stated
RNA velocity (stochastic/dynamical model; specific implementation not named in provided text) Direction and rate of transcriptomic change per cell cluster; self-renewal vs. differentiation bias — Fig. 2D, 2G 6628 control and 4607 FgfrΔRet mutant cells na
Diffusion/Markov-process analysis (forward and reverse) for pseudotime root and terminal-state identification Root and end of cell differentiation trajectories — Fig. 2D 6628 control and 4607 FgfrΔRet mutant cells na
Single-step transition probability comparison (method not formally named) Self-renewal versus differentiation bias of CM progenitors in control vs. FgfrΔRet — Fig. 2F null not stated
Unsupervised clustering (algorithm not specified in provided text) Identification of RPC, neurogenic, CM, and neuron clusters from scRNAseq — Fig. 2B, 2C 6628 control and 4607 FgfrΔRet mutant cells na
Approaches that could also have been used
  • One-way ANOVA was applied across multiple genotype contrasts and multiple marker genes, but no post-hoc correction method is described
    Could also: One-way ANOVA followed by Tukey's HSD or Dunnett's post-hoc test — When a single ANOVA covers several pairwise or treatment-vs-control contrasts, a post-hoc procedure such as Tukey's HSD or Dunnett's test controls the family-wise error rate, allowing each comparison to be interpreted at a stated α level
  • One-way ANOVA was used with n = 3 biological replicates per group, which assumes normally distributed residuals
    Could also: Kruskal-Wallis test with Dunn's post-hoc correction — At n = 3 per group, the normality assumption of ANOVA cannot be empirically verified; a nonparametric Kruskal-Wallis test makes no distributional assumption and is a widely used alternative for small-sample multi-group designs
  • P values were reported as threshold symbols (*, **, ***) rather than exact numeric values
    Could also: Report exact p values (e.g., P = 0.0038) in addition to or instead of symbols — Exact p values convey the precise strength of evidence rather than only whether a threshold was crossed, and facilitate downstream meta-analysis or replication assessment
  • Quantitative area measurements (Fig. 4D) are presented without any measure of spread
    Could also: Report mean ± SD or mean with 95% CI for each genotype group — Dispersion metrics communicate biological variability across animals and help readers assess consistency of group differences, which is especially informative at n = 3
  • Shifts in cell-state proportions (e.g., increased CM percentage at expense of RPC in mutants, fig. S3B) were described with reference to RNA velocity plots
    Could also: Apply a formal differential abundance method such as Milo (neighborhood-graph-based) or edgeR on pseudobulk cell-type counts — Formal differential abundance tests provide statistical uncertainty estimates (e.g., FDR-corrected p values) for changes in cell-state proportions between conditions, complementing visual interpretation of UMAP and velocity plots
  • Effect sizes are not reported alongside ANOVA p values for the marker-area comparisons
    Could also: Report eta-squared (η²) or partial η² as a standardized effect size — Effect sizes quantify the magnitude of group differences independently of sample size, providing information that p values alone do not convey and supporting cross-study comparisons
Software: Not specified in provided text (full Methods section not visible)

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-34757798

Paper: Balasubramanian et al. 2021, Sci Adv 7(46):eabj9846 — "Phase transition specified by a binary code patterns the vertebrate eye cup." (Xin Zhang lab, Columbia.)

Data: GEO GSE139904 (also Single-Cell Portal SCP1618) — one 10x Chromium v3 scRNA-seq run (GSM4148508, "FGF control and mutant"), Mus musculus, E13.5 peripheral retina + RPE, FACS-enriched Cre/GFP+ cells, NovaSeq 6000.

Code: Paper Methods cite velocyto.py for RNA velocity; no authors' own GitHub repo is given in the Data/Code Availability statement. The registry's code_url (github.com/velocyto-team/velocyto-notebooks) is the third-party tool itself. Per BRIEF P16, applying that existing tool (velocyto / scVelo) + the standard Seurat/Scanpy scRNA-seq pipeline to the paper's deposited data is an equally valid reproduction.

In scope (pipeline-derived, attempted)

# Reported result Pipeline Paper loc
C1 6,628 control + 4,607 mutant cells (=11,235) passing QC Cell Ranger v2.1.1 (mm10) → Seurat v3 QC (≥200 genes/cell, genes in >3 cells, mito <20%) Results / Fig 2
C2 Mean depth 2811 genes per cell Cell Ranger / Seurat summary Results
C3 Unsupervised clusters at resolution 0.7; cell-type set {RPC-1..4, Ngn-1, Ngn-2, boundary, RGC, AC/HC, PRC, 3×CM} (~13) Seurat v3 FindClusters res=0.7 + marker annotation Fig 2
C4 CM-specific markers Mitf, Wls, Msx1, Wfdc1 enriched in CM clusters Seurat marker DE Fig 2/3
C5 RNA velocity field over the RPC→neurogenic→CM transition (binary-code / phase-transition trajectory) velocyto.py (kNN imputation, 120 neighbors) / scVelo on spliced+unspliced Fig 2/3

Out of scope (wet-lab / manual / external — NOT attempted)

  • Mouse genetics, FACS, IHC/ISH, RNAscope, human iPSC organoid culture & quantification.
  • Any imaging-based quantification (MSX1+/CDO+ organoid areas).
  • The biological "binary code / phase transition" model itself (interpretive, not a pipeline number).

Reproducibility surface / risks

  • velocyto loom (spliced/unspliced counts) requires the Cell Ranger BAM, which is NOT in the standard 10x deposit. Must check whether GSE139904_analysis.tar.gz ships a precomputed .loom/velocyto output; if absent, full RNA-velocity (C5) may be only partially reproducible (we can still rerun clustering/UMAP).
  • Control vs mutant split (C1) is computational on a single pooled run — must find how cells were assigned (genotype/reporter label, likely in analysis.tar.gz).
C1
Reported
6628 control + 4607 mutant = 11235 cells passing QC
Reproduced
11239 cells in deposited CellRanger filtered matrix (Delta 4 = 0.04%); literal Methods QC ->10185. Control/mutant split not reconstructable (no demux label).
within tolerance
C2
Reported
mean 2811 genes/cell
Reproduced
2826.9 (median 2927)
within tolerance
C3
Reported
~13 cell types at res 0.7 (RPC-1..4, Ngn-1/2, boundary, RGC, AC/HC, PRC, 3x CM)
Reproduced
pooled Seurat res=0.7 -> 18 clusters (graphclust 16); all cell types recovered by canonical markers
partial
C4
Reported
CM markers Mitf/Wls/Msx1/Wfdc1 (distal Wls/Otx2, medial Msx1)
Reproduced
Mitf/Wls/Otx2 peak cl14, Msx1/Wfdc1 peak cl9; deposited graphclust diffexp: Mitf+Wls cl14 (log2FC 4.6/4.0), Msx1+Wfdc1 cl4 (3.68/1.94)
exact
C5
Reported
RNA velocity trajectory (velocyto.py, kNN 120)
Reproduced
not attempted (blocked: no BAM/loom deposited)
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 74/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟢2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

241.4 k
tokens (I/O) · 13.2 M incl. cache
78 min
runtime · 0.08 CPU-h
6.4 GB
peak RAM
1
HPC jobs
hummel
machine