Genetically engineered stem cell-derived retinal grafts for improved retinal reconstruction after transplantation.
The main results reproduced: recomputed values matched the published ones within tolerance.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓No authors-side cause for any deviation
- ✓Reported values are derivable from the shared data
- ✓Any deviation was negligible
- ✓The central claim held under reproduction
- 🟡Could not use the authors’ exact input data
- 🟡Reported values were only indirectly comparable
- 🟡A deviation arose in the data or preprocessing
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Described well enough for a faithful 1:1 reproduction of the ONE module with public data. The paper is overwhelmingly wet-lab: 8 of 9 analysis modules are Bayesian (Stan) re-analyses of measurement tables (cell counts, ERG/MEA, fluorescence, behaviour, human IHC) that are NOT deposited in the repo or GEO and so cannot be reproduced. The single module with public data is the mouse retinal-organoid microarray (GEO GSE178653, 9 samples = 3 lines x 3 differentiation days). Re-running the repo's described pipeline (Agilent normalized signal -> drop controls -> log2(KO/wt) per day -> 4-fold cut) on the deposited data reproduces the paper's reported marker-gene differences: Bhlhb4-/- shows lower photoreceptor markers (Cnga1/2/3, Gucy2d, Rho) at all days -> 5/5 exact; Islet1-/- shows higher markers (Cabp4, Cnga1, Cngb1, Gnat2, Gngt1, Rho, Rcvrn, Sag) at the mature day DD23 -> 8/8 (5/8 if naively averaged over all days, because they are lower at the immature DD10 -> a developmental-timing artifact, not a contradiction). The fourfold threshold and the 'largely unaltered' overall pattern both reproduce. NOT attempted: the GO-enrichment bar charts of SFig2 (need un-shipped helper files go.obo/take.R/genes.xlsx/Agilent annotation csv) and all non-microarray modules (no deposited data). No fabrication concerns: every checked microarray claim is derivable from the deposited data and reproduces in the reported direction.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 84assessed: 2026-06-15 ⛓ 094b5cfea54e
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-15
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusCan genetically engineered ESC/iPSC-retinal sheets with reduced secondary retinal neurons (rod/cone bipolar cells via Bhlhb4 or Islet1 knockout) but intact photoreceptor layers improve graft-host neural integration and visual recovery after subretinal transplantation into end-stage degenerate (rd1) mouse retinas?
- ★ Bhlhb4−/− and Islet1−/− grafts have markedly reduced bipolar cell populations after transplantation while retaining intact photoreceptor cell layers finding
- ★ Photoreceptors in bipolar-cell-KO organoids/grafts can functionally mature in vivo finding
- ★ Bhlhb4−/− and Islet1−/− grafts form more de novo synapses per host bipolar cell, improving neural integration finding
- ★ KO grafts better suppress heightened spontaneous spiking in the degenerated host retina finding
- ★ Mice transplanted with KO grafts perform better in light-guided behavior tests finding
- ★ Using genetically engineered cell lines (Bhlhb4−/−, Islet1−/−) to deplete bipolar cells is a strategy to enhance retinal sheet transplantation outcomes method
- Bhlhb4−/− and Islet1−/− cell lines differentiate into retinal organoids with similar potency, timing, and gene expression as wild-type finding
- ★ Reduced intragraft bipolar cells leave photoreceptor output directly available to contact host bipolar cells, reducing spontaneous activity and improving visual recovery mechanism
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Retinal organoid differentiation with Nrl-GFP reporter live fluorescence imaging | wt, Bhlhb4−/−, Islet1−/− mouse ESC/iPSC lines (Tg Nrl-GFP) | Bhlhb4 KO / Islet1 KO / none (wt) | Nrl-GFP fluorescence intensity over DD19–DD33; organoid morphology | — |
| RT-PCR / real-time PCR of key retinal genes | wt, Bhlhb4−/−, Islet1−/− iPSC- and ESC-derived organoids (DD10, DD16, DD23) | Bhlhb4 KO / Islet1 KO / none | mRNA expression of Bhlhb4, Islet1 and other retinal markers | — |
| Microarray gene expression analysis | Nrl-GFP;Ribeye-reporter iPSC-derived organoids (DD10, DD16, DD23) | Bhlhb4 KO / Islet1 KO / none | genome-wide expression patterns; differentially expressed genes / GO enrichment | — |
| Immunohistochemistry of in vitro organoids | wt, Bhlhb4−/−, Islet1−/− organoids DD15 and DD29 | Bhlhb4 KO / Islet1 KO / none | PKCα, Chx10, Rx, GS, Nrl-GFP marker expression | — |
| Subretinal transplantation followed by immunohistochemistry of grafts | rd1 mouse host retina with wt/Bhlhb4−/−/Islet1−/− grafts (4–5 weeks post-transplant) | Bhlhb4 KO / Islet1 KO / none graft genotype | counts of PKCα+, SCGN+, Chx10+, Calretinin+, Calbindin+ inner cells; rhodopsin+ photoreceptor content | — |
| Flat-mount confocal imaging with orthogonal reconstruction | transplanted rd1 retinas (wt, Bhlhb4−/−, Islet1−/− grafts) | graft genotype | number of photoreceptor cells per graft | — |
| De novo synapse quantification (immunostaining of pre/postsynaptic markers) | rd1;L7-GFP host mice transplanted with Ribeye-reporter-derived wt/Bhlhb4−/−/Islet1−/− grafts | graft genotype; host sex | number of host-graft synapses (RIBEYE + Cacna1s pairs) per L7-GFP host rod bipolar cell | — |
- ▼ PKCα+ rod bipolar cells consistently and markedly reduced in KO grafts vs wt wt 0.14 (95% CI 0.04–0.38) vs Bhlhb4−/− 0.04 (0.01–0.16) and Islet1−/− 0.02 (0.00–0.09) of inner cells
- ▼ Cone bipolar (SCGN+) cells decreased in Islet1−/− but not clearly in Bhlhb4−/− wt 0.28 (0.16–0.42) vs Islet1−/− 0.19 (0.09–0.30), Bhlhb4−/− 0.25 (0.14–0.36)
- ▲ Average number of de novo synapses per host bipolar cell notably increased in KO lines confidence in difference 98% Bhlhb4−/−-wt and 98% Islet1−/−-wt
- – Photoreceptor content not significantly different across genotypes ~0.6 rod photoreceptor content for all lines
- – Nrl-GFP differentiation time courses and maximum intensities broadly similar across wt and KO lines
- – Probability of host bipolar cells forming any synapse largely overlapping across genotypes (only modest improvement in Bhlhb4−/−) predicted probabilities 0.08–0.23; confidence in difference 84% wt-Bhlhb4−/−
- ▲ Host sex affects synapse outcome with more synapses per bipolar in females confidence in difference 91% male-female
- ▼ Chx10+ cells reduced in Bhlhb4−/− and slightly less in Islet1−/−; Islet1−/− Chx10+ cells showed weaker signal confidence in difference 99% Bhlhb4−/−-wt, 84% Islet1−/−-wt
- mean 0.14 (95% CI 0.04–0.38) PKCα+ fraction of inner cells in wt grafts (rod bipolar fraction wt grafts (3 wt, 3 Bhlhb4−/−, 4 Islet1−/− mice))
- mean 0.04 (95% CI 0.01–0.16) Bhlhb4−/−; 0.02 (0.00–0.09) Islet1−/− (PKCα+ rod bipolar fraction in KO grafts)
- other confidence 97% Bhlhb4−/−-wt, 99% Islet1−/−-wt, 84% Bhlhb4−/−-Islet1−/− (confidence in graft genotype effect on PKCα+ cells)
- other confidence 98% Bhlhb4−/−-wt, 98% Islet1−/−-wt, 64% Bhlhb4−/−-Islet1−/− (genotype effect on average synapses per host bipolar (9 wt, 9 Bhlhb4−/−, 9 Islet1−/− mice))
- count 10,354 L7-GFP bipolar cell observations from 27 mice (synapse quantification dataset)
- count 170 organoids (54 wt, 61 Bhlhb4−/−, 55 Islet1−/−) (Nrl-GFP growth curve analysis)
- mean ~0.6 rod photoreceptor content (photoreceptor ratio all genotypes (2 wt, 3 Bhlhb4−/−, 3 Islet1−/− mice))
- other confidence 91% male-female (effect of host sex on synapses per bipolar cell)
Statistical methods review
Model: opusA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
The study used a Bayesian modeling framework to characterize stem-cell-derived retinal grafts. Continuous and count outcomes were fit with explicit generative models (e.g., exponential growth curves for Nrl-GFP intensity, binomial models for marker-positive cell fractions, and a zero-inflated Poisson model for synapses per host bipolar cell), and effects of graft genotype (and host sex) were reported as posterior distributions. Results were summarized with modes and 95% compatibility (credible) intervals, and group differences were expressed as 'confidence in the difference' (the posterior probability/fraction of the difference distribution above zero) rather than as null-hypothesis p-values.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Bayesian exponential growth-curve model (Nrl-GFP fluorescence over differentiation days) | Figure 1B, 1C, S1; organoid GFP signal across wt/Bhlhb4-/-/Islet1-/- | 170 organoids (54 wt, 61 Bhlhb4-/-, 55 Islet1-/-) | stated |
| Bayesian binomial model of marker-positive fraction vs total inner cells | Figure 2B; PKCα+, SCGN+, Chx10+ (and Calbindin+/Calretinin+ in S2) cell counts by genotype | 11 mice PKCα (3 wt,3 Bhlhb4-/-,4 Islet1-/-); 11 SCGN (3,4,4); 9 Chx10 (2,3,4) | stated |
| Bayesian model of photoreceptor cell number as a function of graft cell number | Figure 2D; rod photoreceptor content by genotype | 8 mice (2 wt, 3 Bhlhb4-/-, 3 Islet1-/-) | stated |
| Zero-inflated Poisson model (probability of forming a synapse θ and mean synapses λ) | Figure 3J–L, S4; synapses per L7-GFP host bipolar cell by graft genotype and host sex | 10,354 L7-GFP bipolar cells from 27 mice (9 wt, 9 Bhlhb4-/-, 9 Islet1-/-) | stated |
| Microarray differential expression comparison with GO enrichment analysis | Figures 1E, 1F, S2; DD10/DD16/DD23 organoid expression and genotype comparisons | — | not stated |
-
Group differences were summarized as the posterior 'confidence in the difference' (fraction of the difference distribution above zero) with 95% compatibility intervals.↳ Could also: A frequentist mixed-effects model (e.g., generalized linear mixed model) with reported p-values and confidence intervals could also be used. — A frequentist summary would provide the p-values and effect estimates many readers and meta-analyses expect, complementing the Bayesian credible intervals already reported.
-
Counts of marker-positive cells were modeled with a binomial model and synapse counts with a zero-inflated Poisson model.↳ Could also: A negative-binomial or zero-inflated negative-binomial formulation could also be applied to count outcomes. — A negative-binomial component would accommodate overdispersion beyond the Poisson mean–variance assumption, which can be useful when count variability is large.
-
Software and package details for the Bayesian models were not stated in the provided text.↳ Could also: Reporting the specific tools and versions (e.g., R with brms/Stan, priors, and convergence diagnostics) could also be included. — Naming software, priors, and diagnostics (R-hat, effective sample size) aids exact reproducibility and lets readers assess model fit.
-
Inference relied on posterior distributions without an explicit multiplicity adjustment across the family of comparisons.↳ Could also: For the microarray/GO analyses, a Benjamini-Hochberg FDR (or hierarchical shrinkage) approach could also be reported. — An explicit FDR or hierarchical/partial-pooling scheme would document control of false discoveries across the many genes and contrasts tested.
-
Sample sizes were described as numbers of organoids, mice, and observations per analysis.↳ Could also: A brief statement of how sample size was determined (or an a priori power/precision rationale) could also be provided. — Describing the basis for n helps readers gauge the precision the design was intended to achieve, especially for the smaller per-genotype mouse groups.
-
Central estimates were reported as the posterior mode with 95% compatibility intervals.↳ Could also: The posterior mean or median with the interval could also be reported. — Median/mean summaries are less sensitive to the shape of skewed posteriors and are a common convention, making cross-study comparison straightforward.
Result convergence & founder nodes
Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.
-
Female rd1 host mice form more de novo synapses per host rod bipolar cell with retinal grafts than male hosts, independent of graft genotype.imaging female rd1 mouse retina up 2021×1papers★ This paper is the founder (earliest)
-
Nrl-GFP reporter fluorescence time courses and peak intensities are similar across wild-type, Bhlhb4-KO, and Islet1-KO retinal organoids, indicating unimpaired photoreceptor differentiation in vitro.imaging mouse retinal organoid none 2021×1papers★ This paper is the founder (earliest)
-
SCGN+ cone bipolar cells are reduced in Islet1-KO but not Bhlhb4-KO retinal grafts after transplantation into rd1 mice.imaging rd1 mouse retina down 2021×1papers★ This paper is the founder (earliest)
-
The probability of a host rod bipolar cell forming any synapse with a retinal graft overlaps broadly across wild-type, Bhlhb4-KO, and Islet1-KO genotypes, with only modest improvement in Bhlhb4-KO.imaging rd1 mouse retina mixed 2021×1papers★ This paper is the founder (earliest)
-
De novo host-graft synapses per host rod bipolar cell are increased in Bhlhb4-KO and Islet1-KO grafts compared to wild-type grafts in rd1 mice.imaging rd1 mouse retina up 2021×1papers★ This paper is the founder (earliest)
-
Photoreceptor content of retinal grafts does not differ significantly between wild-type, Bhlhb4-KO, and Islet1-KO lines after transplantation into rd1 mice.imaging rd1 mouse retina none 2021×1papers★ This paper is the founder (earliest)
-
PKCα+ rod bipolar cells are markedly reduced in Bhlhb4-KO and Islet1-KO retinal grafts versus wild-type after subretinal transplantation into rd1 mice.imaging rd1 mouse retina down 2021×1papers★ This paper is the founder (earliest)
-
VSX2 (Chx10)+ inner cells are reduced in Bhlhb4-KO grafts and show diminished immunosignal in Islet1-KO grafts after transplantation into rd1 mice.imaging rd1 mouse retina down 2021×1papers★ This paper is the founder (earliest)
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-34409267 (KO-graft / Matsuyama et al. 2021, iScience)
Paper: "Genetically engineered stem cell-derived retinal grafts for improved retinal reconstruction after transplantation." Repo: https://github.com/matsutakehoyo/KO-graft @ commit f48d71a2206910b9585b2184e6baccdcd3b39180 (2021-07-29). Public data: GEO GSE178653 ("Microarray analysis of mouse retinal organoids", Agilent SurePrint G3 Mouse GE 8x60K v2.0 = GPL21810, 9 samples GSM5395271–GSM5395279 = 3 lines {wt, Bhlhb4−/−, Islet1−/−} × 3 days {DD10,DD16,DD23}).
How the paper splits into modules (from mouse/README.md)
| Module | Figures | Pipeline | Input data | In scope? |
|---|---|---|---|---|
| organoid fluorescence | Fig1, SFig1 | R + Stan growth model | image-derived fluorescence table (NOT shipped) | OUT (no data) |
| organoid microarray | Fig1, SFig1–2 | R (tidyverse) load→filter→log2FC→GO | GSE178653 (PUBLIC) | IN |
| cell type | Fig2, SFig3 | Stan binomial | cell-count table (NOT shipped) | OUT (no data) |
| synapse | Fig3, SFig4 | Stan ZIP hierarchical | synapse counts (NOT shipped) | OUT (no data) |
| mERG | Fig4, SFig6–7 | R + Stan (MED64 MEA) | MEA recordings (NOT shipped) | OUT (no data) |
| two photon | SFig8 | Stan gamma-hurdle | Ca-imaging table (NOT shipped) | OUT (no data) |
| RGC_1s | Fig5–6, SFig10–11 | Stan clogit/lognormal | MEA recordings (NOT shipped) | OUT (no data) |
| SAS (behaviour) | Fig7, SFig12 | Stan binomial | behaviour table (NOT shipped) | OUT (no data) |
| human organoids | Fig4,6,7 | Stan (ordered_probit etc.) | IHC/MEA tables (NOT shipped) | OUT (no data) |
Only the microarray module has public input data (GSE178653). All other modules are Bayesian (Stan) re-analyses of wet-lab measurement tables that are not deposited in the repo or GEO, so they cannot be reproduced (would require the authors' private measurement spreadsheets). This is an honest scope limit, not a quality judgement: the heavy-stats core of the paper is wet-lab-data-bound.
In-scope reproduction target (microarray, Fig 1 / SFig 1–2)
The repo scripts load micro array data.R + gene comparison.R implement:
load 9-sample Agilent normalized signal → drop control probes + probes below
background → average duplicate probes → compute log2(KO/wt) per DD →
flag genes with >4-fold change (|log2FC|>2) → GO enrichment.
Concrete, falsifiable paper claims to reproduce (from Results/Fig 1 text):
- C1 Platform: 56,605 probes covering 37,074 genes (sanity check vs GPL21810 / data).
- C2 Threshold = fourfold (|log2FC|>2) — matches code
filter(log2_exp > 2 / < -2). - C3 Bhlhb4−/− shows lower expression of photoreceptor markers Cnga1, Cnga2, Cnga3, Gucy2d, Rho (vs wt).
- C4 Islet1−/− shows higher expression of Cabp4, Cnga1, Cngb1, Gnat2, Gngt1, Rho, Rcvrn, Sag (vs wt).
- C5 Overall: expression pattern of retinal genes "largely unaltered" between genotypes (few genes pass the 4-fold cut) — quantify by per-line up/down counts.
Out of the hard last 20%: exact GO-enrichment level1/level2 bar charts (SFig2)
need un-shipped helper files (go.obo, take.R, gen-log.R, genes.xlsx,
Agilent AllAnnotations_*.csv); not attempted — direction-of-marker-genes (C3/C4)
is the auditable core and is attempted instead.
Compute: all on «our HPC» SLURM (GEOquery + tidyverse conda env). Data on «infra».
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
Only the mouse retinal-organoid microarray module (Fig 1, SFig 1-2) has public input data (GEO GSE178653), and on that module the reproduction is faithful: Bhlhb4-/- lower markers (5/5), Islet1-/- higher markers (8/8 at mature DD23), and the 'largely unaltered' pattern all reproduce in the reported direction with no fabrication signal. The only deviations are on our/technical side — C1's Agilent design-count nomenclature (56,605 vs 27,306 symbols) and C4's wash-out when naively averaging across differentiation days, both explainable. The dominant limitation is data availability: 8 of 9 modules are wet-lab Stan re-analyses with no deposited data, so they are out of scope, not quality failures. Overall a solid reproduction with explainable deviations → yellow.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.