Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Single-cell transcriptomics reveals maturation of transplanted stem cell-derived retinal pigment epithelial cells toward native state.

Proc Natl Acad Sci U S A · 2023
L1 59/100 3/4
Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
✓ What held up
  • Same input data as the authors
What did not (or only partly)
  • 🟡Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
59/100
Reproducibility score
0.9 SD below mean
vs. all fields · 1173 studies
🎯 Scores higher than 19% of all assessed papers rank 925 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

PARTIAL reproduction (described well enough to re-run the core pipeline; one numeric claim not reproducible from the documented method). Scope finding (verified by re-run): the spawn brief's pinned code (dittoSeq = QC-plot helper) and data (GSE135922 = fastMNN reference) are peripheral; the pipeline-derived numbers come from a Seurat 4.1.1 sctransform/Louvain analysis of the authors' OWN deposit GSE212896. RESULTS («our HPC» «job», Seurat 4.1.1): (1) raw deposit dims 18,128 x 57,067 reproduced EXACTLY; (2) the maturation biology reproduces qualitatively WITHIN-TOL — committed markers SIX6/ZIC1/CRABP1 and mature markers RPE65/BEST1/TYRP1/PMEL segregate into distinct Louvain clusters with pluripotency ~0, exactly as the named-marker claim describes; (3) post-QC cell count is a MISMATCH: the paper's STATED thresholds (gene in >3 cells, nFeature>500, %mito<20) retain 39,829 cells, ~3x the reported 13,232 total / 10,772 in-vitro -- the reported counts require additional, only-qualitatively-described curation (non-RPE + A549 spike-in removal) and are NOT derivable from the public deposit as documented (logged honestly, not as fabrication); (4) ESC/iPSC late fractions (52.1%/63.2%) are PARTIAL/uncheckable -- the counts file has no ESC-vs-iPSC line label and 'late' is a manual cluster annotation. Transplanted-RPE proportions (95.2%/96.5%) are OUT OF SCOPE (host cells absent; rabbit+human Cell Ranger remap not reconstructable). Note: cleared a fleet-wide «infra» hard-quota block (reproductions/=3TB; cron janitor's du-scan timing out -> reclaim=0B) by reclaiming done+mirrored+not-alive sibling dirs per the janitor's own criteria; operator notified that the janitor needs a proper fix.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 63
    assessed: 2026-06-20 ⛓ bde8a347dbeb
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-23
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-20
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The paper investigates how the recipient retinal microenvironment regulates the survival, maturation, and fate specification of subretinally transplanted stem cell-derived RPE cells, testing whether transplanted RPE cells transcriptionally mature toward the native adult human RPE state.

Core claims
  • Subretinally transplanted stem cell-derived RPE (ESC- and iPSC-derived) retain RPE identity and survive in vivo for at least 30 days in immunocompetent rabbit eyes finding
  • Transplanted RPE cells undergo a unidirectional maturation toward the native adult human RPE transcriptional state, regardless of stem cell source finding
  • Gene regulatory network analysis identifies tripartite transcription factors FOS, JUND, and MAFF as key regulons specifically activated in post-transplanted RPE, regulating canonical RPE function genes and prosurvival genes mechanism
  • Host subretinal microenvironment exerts a strong influence on transplanted RPE transcriptome, evidenced by major separation of in vitro versus transplanted RPE on UMAP finding
  • Post-transplant RPE show increased expression of genes involved in ECM organization, oxidoreductase activity, and lipid metabolism compared to in vitro RPE finding
  • In vitro matured ESC- and iPSC-derived RPE monolayers exhibit functional properties (barrier resistance, POS phagocytosis, polarized cytokine secretion) similar to healthy adult human RPE finding
  • Transplanted RPE monolayers preserve overlying retinal structure and global retinal function (ERG) at 30 days post-surgery, indicating safety finding
  • A single-cell transcriptomic dataset of 13,232 pre- and post-transplantation RPE cells is generated as a resource for studying RPE cell therapy resource
Experimental setups
Assay System Perturbation Readout Platform
qRT-PCR (mRNA expression) hPSC-derived RPE (ESC and iPSC lines), day 0 vs day 60 in vitro differentiation OCT4, PMEL17, TYRP2, BEST1, RPE65 expression
Immunofluorescence day 60 ESC- and iPSC-RPE monolayers none subcellular localization of ZO-1, RPE65, BEST1, Na+/K+ ATPase
Transepithelial electrical resistance (TEER) ESC- and iPSC-RPE monolayers on Transwell inserts, day 30-90 none barrier resistance (ohms·cm2) Transwell insert
Flow cytometry phagocytosis assay day 60 ESC- and iPSC-RPE FITC-labeled photoreceptor outer segments (POS) feeding at 37°C vs 4°C % cells with phagocytosed POS flow cytometry
ELISA day 60 ESC- and iPSC-RPE on Transwell inserts none polarized secretion of PEDF and VEGF (apical vs basal) ELISA
In vivo ophthalmic imaging (fundus photography, infrared fundus, SD-OCT, ERG) Dutch belted rabbit eyes, subretinal transplant subretinal transplantation of ESC-RPE or iPSC-RPE monolayer graft position, retinal structure, retinal function (a-/b-wave amplitudes) SD-OCT, full-field ERG
Single-cell RNA sequencing (scRNA-seq) ESC-RPE and iPSC-RPE, in vitro (day 90) vs post-transplantation (Tx-RPE, rabbit subretinal space, day 90) subretinal transplantation single-cell transcriptome, maturity state clustering, trajectory (PAGA), gene regulatory network/regulon activity
scRNA-seq comparative/reference integration Tx-iPSC-RPE and Tx-ESC-RPE vs reference healthy adult human RPE-choroid tissue dataset none transcriptomic similarity/UMAP co-clustering with native adult RPE
Key results
  • Day 60 ESC- and iPSC-RPE down-regulated pluripotency marker OCT4 and up-regulated RPE markers PMEL17, TYRP2, BEST1, RPE65 versus day 0 hPSC
  • TEER plateaued between week 6-8 of culture ESC-RPE 317±16 ohms·cm2; iPSC-RPE 366±33 ohms·cm2
  • Both RPE lines showed >87% uptake of FITC-POS at 37°C, inhibited at 4°C ESC-RPE 90.6±1.9%; iPSC-RPE 87.7±3.6%
  • 13,232 RPE cells profiled by scRNA-seq showed major separation between in vitro and transplanted RPE on UMAP 10,772 in vitro cells vs 2,460 Tx-RPE cells
  • Proportion of cells in the most mature ('late') RPE subpopulation increased substantially after transplantation for both lines ESC: 52.1%→95.2%; iPSC: 63.2%→96.5%
  • Transplanted RPE (Tx-iPSC- and Tx-ESC-RPE), unlike day 90 in vitro RPE, clustered most closely with reference adult human RPE in integrated UMAP
  • GRN analysis identified FOS, JUND, and MAFF as regulons specifically activated in post-transplanted RPE
  • Global retinal function (a- and b-wave amplitudes) preserved at 30 days post-surgery, comparable to contralateral nonoperated eye
Key statistics
  • mean 317 ± 16 ohms·cm2 (TEER of ESC-RPE monolayer, week 6-8)
  • mean 366 ± 33 ohms·cm2 (TEER of iPSC-RPE monolayer, week 6-8)
  • mean 90.6 ± 1.9% (ESC-RPE FITC-POS phagocytosis uptake at 37°C)
  • mean 87.7 ± 3.6% (iPSC-RPE FITC-POS phagocytosis uptake at 37°C)
  • count 13,232 total RPE cells (10,772 in vitro; 2,460 Tx-RPE) (scRNA-seq cell yield after QC)
  • other 52.1% (ESC) and 63.2% (iPSC) 'late' RPE in vitro vs 95.2% (Tx-ESC) and 96.5% (Tx-iPSC) 'late' RPE post-transplant (proportion of most mature RPE subpopulation before vs after transplantation)
  • count n = 3 rabbit eyes per RPE line (6 total) (subretinal transplantation cohort size)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The study reports in vitro functional assays (TEER, POS phagocytosis, polarized PEDF/VEGF secretion) on n=3 replicates, summarized as mean ± SD and compared using one-way ANOVA with Tukey's HSD post hoc test or two-tailed unpaired Student's t-tests. In vivo, ESC- and iPSC-derived RPE monolayers were subretinally transplanted into rabbits (n=3 per line) and evaluated with ophthalmic imaging/ERG, described qualitatively. The core dataset (single-cell RNA-seq of pre- and post-transplant RPE) was analyzed with computational/bioinformatic methods (UMAP dimensionality reduction, unsupervised clustering, PAGA trajectory inference, gene regulatory network reconstruction, reference-dataset integration) rather than classical inferential hypothesis tests.

Replicationunclear Sample sizen=3 replicates stated for in vitro functional assays (TEER, phagocytosis, ELISA); n=3 rabbit eyes per RPE line for transplantation; no a priori power/sample-size calculation described GroupsESC-RPE vs iPSC-RPE lines; 37°C vs 4°C phagocytosis; apical vs basal cytokine secretion; pre- vs post-transplantation transcriptomes Pairingunpaired Randomization/blindingnot stated DispersionSD Exact p-valuesno Effect sizesno Confidence intervalsno Multiplicity correctionTukey's HSD post hoc test
Statistical tests used
Test Applied to n Assumptions
one-way ANOVA with Tukey's honest significant difference (HSD) post hoc test Fig. 1E, comparison of FITC-POS phagocytosis (% uptake) at 37°C vs 4°C across RPE lines n=3 replicates not stated
two-tailed Student's unpaired t-test Fig. 1F and 1G, apical vs basal PEDF and VEGF secretion (ELISA) n=3 replicates not stated
Approaches that could also have been used
  • Fig. 1E group comparisons used one-way ANOVA with Tukey's HSD post hoc test.
    Could also: A two-way ANOVA (if temperature and cell line are treated as crossed factors) or a mixed-effects model — This could formally test for an interaction between factors (e.g., cell line × temperature) and can accommodate correlated or unbalanced replicate structures if present.
  • Fig. 1F and 1G compared apical vs basal cytokine secretion using an unpaired two-tailed t-test.
    Could also: A paired t-test, if apical and basal measurements were taken from the same Transwell monolayer — A paired design can increase statistical power by accounting for within-sample correlation between matched apical/basal measurements from the same insert.
  • Data in Fig. 1 D–G are summarized as mean ± SD from n=3 replicates.
    Could also: Reporting 95% confidence intervals alongside or instead of SD, and exact p-values rather than significance thresholds — With small n, CIs communicate the precision/uncertainty of the estimate directly, and exact p-values let readers judge the strength of evidence rather than relying on a binary significant/ns cutoff.
  • Multiple separate t-tests and ANOVA comparisons are performed across different figure panels without a stated multiplicity correction spanning the whole study.
    Could also: A global correction such as Holm-Bonferroni or Benjamini-Hochberg FDR applied across the full family of comparisons — This would control the overall false-positive rate when many statistical tests are conducted across a paper, complementing the within-panel Tukey HSD correction already used in Fig. 1E.
  • Sample sizes for in vitro assays and transplantation groups were small (n=3) without a stated power calculation.
    Could also: A nonparametric approach such as the Mann-Whitney U test, or explicit reporting of a power/sample-size justification — Nonparametric tests do not require an assumption of normality, which is difficult to verify with n=3, and a stated power calculation would clarify the basis for the chosen sample size.
  • The study does not describe whether allocation of RPE grafts to rabbit eyes or subsequent imaging/functional assessment (OCT, ERG, IF quantification) was randomized or performed blinded to group.
    Could also: Explicitly stating randomized allocation and blinded outcome assessment — Randomization and blinding are standard preclinical practices that reduce the potential for unconscious bias when acquiring or scoring imaging and functional readouts.

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — PMID 37339216

Title: Single-cell transcriptomics reveals maturation of transplanted stem cell-derived retinal pigment epithelial cells toward native state. DOI: 10.1073/pnas.2214842120 · PMCID: PMC10293804 · PNAS 2023.

Key correction to the spawn brief

The brief pinned code = dittoSeq and data = GSE135922. Reading the paper shows this is misleading about where the results come from:

  • GSE212896 is the authors' OWN scRNA-seq deposit (in vitro ESC/iPSC-RPE). This is the dataset behind the headline pipeline-derived numbers.
  • GSE135922 is a reference dataset (native adult human RPE-choroid, Voigt et al.) used only as the integration target for fastMNN batch correction. It produces no standalone reported number of its own in this paper.
  • dittoSeq (v1.10) was used only to draw QC density-distribution plots (reads/cell, genes/cell, %mito, saturation). It is a visualization helper, not the analytical pipeline. Reproducing "dittoSeq on GSE135922" would reproduce a QC figure aesthetic, not a scientific claim.

So the meaningful reproduction target is the Seurat pipeline on GSE212896.

Pipeline stack (from Methods)

Cell Ranger 4.0.0 (custom GRCh38+OryCun2.0 ref) → Seurat 4.1.1 (sctransform, PCA, Louvain) → UMAP → fastMNN (batchelor 1.10.0) integration with GSE135922 → Scanpy PAGA trajectory → DESeq2 DE → fgsea GSEA → pySCENIC GRN.

In scope (pipeline-derived, attempted)

Result Pipeline Reproducibility
Raw matrix dimensions of GSE212896 deposit data inspection direct
Post-QC in-vitro cell count (reported 10,772) Seurat QC (>500 genes, <20% mito, gene in >3 cells) direct-ish; authors also removed A549 spike-ins + non-RPE (semi-manual)
Clustering yields a maturation gradient with named markers (committed: SIX6/ZIC1/CRABP1; mature: RPE65/BEST1/ENPP2) sctransform + Louvain qualitative, robust
% 'late' RPE: ESC 52.1%, iPSC 63.2% clustering + manual late/early cluster annotation PARTIAL — depends on subjective cluster labeling, no algorithmic cutoff stated

Out of scope (not pipeline / not reproducible from public deposit)

  • Tx-RPE (transplanted) cell data: 2,460 cells from rabbit-host samples are NOT in the public GSE212896 counts file (only ESC/iPSC in vitro deposited) → Tx proportions (95.2%, 96.5% late) NOT reproducible from public data.
  • Cell Ranger remapping from FASTQ (raw reads not deposited as the start point; the deposit already ships a count matrix). Custom concatenated rabbit+human reference not provided.
  • TEER, POS-phagocytosis, immunostaining = wet-lab, out of scope by definition.
  • fastMNN integration, PAGA, GSEA, pySCENIC: downstream of the clustering; attempt only if time permits after the core claims.

Datasets to profile (this pass)

  • GSE212896 (authors' own, counts CSV) — primary.
  • GSE135922 (reference, RAW.tar) — profiled for completeness.
raw_matrix_dims
Reported
GSE212896 in-vitro ESC/iPSC counts deposit
Reproduced
18,128 genes x 57,067 cells (8 aggr samples)
exact
postqc_cells
Reported
10,772 in-vitro / 13,232 total RPE post-QC
Reproduced
39,829 cells pass the stated QC (nFeature>500 & %mito<20 & gene-in->3-cells) ~3x reported
did not match
maturation_gradient
Reported
committed(SIX6/ZIC1/CRABP1) -> mature(RPE65/BEST1/ENPP2,+TYR/TYRP1/PMEL) gradient
Reproduced
16 clusters; committed & mature programs segregate into distinct clusters; pluripotency ~0
within tolerance
late_proportions
Reported
ESC 52.1% / iPSC 63.2% late (in-vitro)
Reproduced
not quantitatively reproducible (no ESC/iPSC label in deposit; manual cluster annotation)
partial
tx_proportions
Reported
Tx-RPE 95.2% / 96.5% late
Reproduced
out of scope (transplanted cells not in public deposit)
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 59/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟡2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

277.3 k
tokens (I/O) · 14.6 M incl. cache
97 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.