Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

In vivo prime editing rescues alternating hemiplegia of childhood in mice.

Cell · 2025
L1 84/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +3
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
  • Reported values are derivable from the shared data
  • The central claim held under reproduction
What did not (or only partly)
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡The deviation was non-trivial in magnitude
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
84/100
Reproducibility score
0.6 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 63% of all assessed papers rank 392 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough to reproduce: YES. Ran CRISPResso2 2.3.4 on the paper's own amplicon-seq (SRA/ENA PRJNA1211588, 42 runs) on «our HPC» SLURM (final «job», COMPLETED 1h31m) to reproduce the central Figure 1J-N result: pathogenic-allele correction + indel rates for the four prime-edited ATP1A3 AHC mutations (D801N, E815K, L839P, G947R) +/-MLH1dn and the ABE-corrected G947R-A, in patient-derived iPSCs. Sample->figure->guide mapping from Table S3; amplicon refs by in-silico PCR of ST3B primers vs ATP1A3 NG_008015.1; pathogenic codons via NM_152296.5; correction = (WT%_edited - WT%_noedit)/(100 - WT%_noedit)x100 using matched No-edit controls (all 5 controls ~50% baseline WT = heterozygous lines). 1:1 OUTCOME: correction reproduces within tolerance in 7/9 conditions (6 exact incl. BOTH D801N now exact -- 43.8 vs 43, 18.0 vs 18; L839P 69.6/71.4 vs 70/71; G947R-C 75.7 vs 74; ABE 90.0 vs 90; 1 within-tol G947R-C 49.6 vs 47); indels match in 7/9 (5 exact). The headline editing efficiencies reproduce essentially exactly => genuine, data-derivable, NO fabrication indicated. DIFFERENT: E815K (Fig 1K) reproduces ~7-11 pts HIGH on correction and LOW on indels in both conditions -- a consistent methodological offset (E815K's 19-nt-spacer silent-edit design makes the paper's precise-haplotype definition stricter than our single-position WT-base readout; narrow E815K QWC undercounts indels), NOT fabrication. NOT ATTEMPTED: in-vivo mouse phenotyping (wet-lab); HEK293T screens Fig S2-S5 (same pipeline, not needed for headline); Fig 3 off-target quantification (Table S2 corrupt at PMC). All compute on «our HPC» SLURM compute nodes; env built in-job; FASTQs on «infra»; «host» holds only small results. Grades provisional; a human reviewer decides (see AUDIT.md).

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 83
    assessed: 2026-06-16 ⛓ 2c213e998654
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-22
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-18
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The paper tests whether prime editing (PE) and base editing (BE) can precisely correct pathogenic ATP1A3/Atp1a3 mutations in human cells and in vivo in mouse models to rescue the neurological phenotypes of alternating hemiplegia of childhood (AHC).

Core claims
  • PE and BE strategies efficiently correct five prevalent ATP1A3 mutations (D801N, E815K, L839P, G947R-A, G947R-C) in HEK293T cells and AHC patient-derived iPSCs, with 43%-90% correction in iPSCs. finding
  • Bystander editing at adjacent bases limits base editor correction of the D801N, E815K, and L839P mutations, necessitating use of prime editing instead. finding
  • AAV9-mediated in vivo prime editing via P0 ICV injection corrects Atp1a3 D801N and E815K mutations in the CNS of two AHC mouse models, achieving up to 48% DNA correction and 73% mRNA correction in bulk brain cortex. finding
  • In vivo PE rescues clinically relevant AHC phenotypes in mice, including hippocampal Atp1a3 ATPase activity, paroxysmal spells, motor defects, cognition deficits, and greatly extends lifespan. finding
  • PE correction strategies show substantially lower and fewer off-target editing events than ABE8e in patient-derived iPSCs, as measured by CIRCLE-seq and rhAmpSeq. finding
  • Adding 'dead' sgRNAs (dsgRNAs) that disrupt local chromatin accessibility enhances PE correction efficiency at multiple target sites. method
  • Clustering silent edits near the corrected mutation and MLH1dn supplementation increase mismatch-repair evasion, improving PE correction purity. mechanism
  • Sequence divergence between human ATP1A3 and mouse Atp1a3 (90.8% coding identity) required redesign of guide RNAs and PAMs for mouse-specific PE strategies. method
Experimental setups
Assay System Perturbation Readout Platform
prime editing / base editing (transfection screen) HEK293T cell lines engineered with pathogenic ATP1A3 variants PE/CBE/ABE with various sgRNA/epegRNA/ngRNA designs pathogenic allele correction % and indel rate
prime editing / ABE base editing AHC patient-derived induced pluripotent stem cells (iPSCs) PE with epegRNA/ngRNA/dsgRNA (PE6 variants, MLH1dn) or ABE8e pathogenic allele correction % and indel rate
CIRCLE-seq AHC patient-derived iPSCs (genomic DNA in vitro) epegRNAs, ngRNAs, dsgRNAs, ABE sgRNA nominated genome-wide off-target cleavage sites
rhAmpSeq (RNaseH-dependent amplification sequencing) AHC patient-derived iPSCs treated with PE/ABE PE or ABE editing vs donor-matched mock indels and substitutions (log2 fold-change) at top CIRCLE-seq nominated off-target sites
prime editing (cell line generation) mouse N2a cell line (monoclonal, D801N) PE correction strategy redesigned for mouse Atp1a3 sequence editing/correction efficiency
AAV9-mediated dual-vector in vivo prime editing (ICV injection at P0) B6C3.Atp1a3 D801N/+ and B6C3.Atp1a3 E815K/+ AHC mouse models in vivo PE (epegRNA/ngRNA delivered via dual AAV9) DNA correction % and mRNA correction % in bulk brain cortex AAV9 viral vector
ATPase activity assay hippocampal tissue from D801N/E815K mice in vivo PE treatment vs untreated Atp1a3 Na+/K+ ATPase activity restoration
behavioral phenotyping (paroxysmal spell induction, motor tests, cognition tests, lifespan monitoring) D801N and E815K AHC mouse models in vivo PE treatment vs untreated mutant vs WT paroxysmal spell frequency, motor performance, cognitive performance, survival/lifespan
Key results
  • G947R-A ABE editing achieved efficient, precise correction with no observed bystander editing up to 40% correction
  • PE correction of D801N, E815K, L839P, and G947R-C in patient iPSCs with MLH1dn supplementation 43%, 63%, 70%, 74% pathogenic allele correction respectively
  • ABE correction of G947R-A in patient-derived iPSCs 90% pathogenic allele correction, 0.3% indels
  • dsgRNA non-PAM-44 improved D801N pathogenic allele correction 1.4-fold improvement, 55% correction
  • In vivo AAV9 PE corrected Atp1a3 mutations in mouse brain cortex up to 48% DNA correction and 73% mRNA correction
  • PE strategies showed fewer confirmed off-target loci than ABE8e among assayed sites 8 of 457 sites (PE) vs 16 of 457 sites (ABE8e)
  • Alternative deaminase ABE7.10 eliminated confirmed off-target sites compared to ABE8e for G947R-A correction 0/95 sites (ABE7.10) vs 16/95 sites (ABE8e)
  • In vivo PE treatment ameliorated paroxysmal spells, motor and cognition deficits, and extended lifespan in D801N mice
Key statistics
  • count 1 in 1,000,000 (estimated incidence of AHC)
  • fold_change 70% (proportion of AHC cases caused by ATP1A3 mutations)
  • fold_change 40%, 20%, 10% (approximate prevalence of D801N, E815K, and G947R variants among ATP1A3-associated AHC cases)
  • other 90.8% coding sequence identity (sequence divergence between human ATP1A3 and mouse Atp1a3)
  • fold_change 43%/8%, 63%/5.2%, 70%/1%, 74%/4% (correction/indels) (iPSC pathogenic allele correction and indel rates for D801N, E815K, L839P, G947R-C with MLH1dn)
  • fold_change up to 48% DNA correction, 73% mRNA correction (in vivo AAV9 PE correction in bulk mouse brain cortex)
  • count 24 of 457 assayed sites confirmed off-target (off-target validation across PE and ABE guide RNAs in iPSCs via rhAmpSeq)
  • fold_change 1.8-fold, 1.3-fold, 1.1-fold, 1.6-fold (ngRNA-driven improvements in pathogenic allele correction for D801N, E815K, L839P, G947R-C respectively)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This paper reports prime editing (PE) and base editing (BE) correction of ATP1A3/Atp1a3 mutations in HEK293T cells, patient-derived iPSCs, and two AHC mouse models. Editing outcomes are primarily quantified as pathogenic allele correction rates and indel rates expressed as percentages. In vivo efficacy is assessed via ATPase activity rescue, behavioral phenotyping, body weight, and survival in treated versus untreated AHC mice. Off-target editing is evaluated by CIRCLE-seq nomination followed by rhAmpSeq quantification at nominated sites, with results expressed as log2 fold-change versus donor-matched mock controls. The provided text excerpt is truncated before the full in vivo results and methods sections, so formal statistical tests for behavioral and survival endpoints are not visible in the supplied text.

Replicationunclear GroupsWT vs D801N+/- vs E815K+/- mice (treated and untreated); HEK293T and iPSC lines with vs without MLH1dn; multiple editor variants and guide RNA configurations Pairingunclear Randomization/blindingnot stated Dispersionnone Confidence intervalsno
Statistical tests used
Test Applied to n Assumptions
log2 fold-change relative to donor-matched mock controls; no named formal test visible in excerpt off-target editing quantification at 457 rhAmpSeq-amplifiable CIRCLE-seq nominated sites (Figures 2A–2C) 457 amplifiable sites of 480 nominated not stated
fold-change ratio comparison between optimization conditions (no named statistical test) epegRNA, ngRNA, and dsgRNA optimization across HEK293T and iPSC conditions (Figures S2–S5) not stated
not described in available text (likely survival/log-rank; mentioned as outcome only) lifespan comparison between treated and untreated D801N and E815K mice not stated
not described in available text behavioral outcomes: motor deficits, cognitive deficits, paroxysmal spell frequency not stated
not described in available text hippocampal Atp1a3 ATPase activity across WT, untreated, and PE-treated mice not stated
Approaches that could also have been used
  • Pathogenic allele correction rates are reported as single percentage values without any measure of variability across replicates
    Could also: Report mean ± SD or 95% CI across biological replicates of the editing experiment — Dispersion measures allow readers to assess reproducibility and distinguish robust from variable editing outcomes; for small n typical in iPSC work, SD conveys spread and CI conveys precision of the estimate
  • Off-target editing at 457 sites is characterized as log2 fold-change versus mock without a formal statistical threshold or multiplicity adjustment
    Could also: Apply a negative binomial or Poisson model to per-site read counts comparing edited versus mock, followed by Benjamini-Hochberg FDR correction across all tested sites — A statistically principled per-site test with FDR control would provide a consistent, reproducible decision threshold for calling confirmed off-targets and would reduce the influence of sequencing noise on site classification
  • Gains from individual optimization steps (dsgRNA, ngRNA, PE6 variants) are each reported as isolated fold-changes without joint modeling of all factors
    Could also: Factorial ANOVA or a linear mixed model treating editor variant, guide type, and their interactions as fixed factors — A factorial design analysis would partition variance attributable to each component and their interactions, enabling more principled conclusions about which factors independently contribute to editing gains
  • Lifespan is described as an outcome (extension in D801N mice) but the statistical approach is not stated in the available text
    Could also: Kaplan-Meier survival curves compared with the log-rank (Mantel-Cox) test; Cox proportional hazards model if covariates such as dose or sex are present — The log-rank test is the standard nonparametric approach for comparing survival distributions between groups and correctly accounts for censoring, which is typical in animal survival studies
  • Multiple behavioral endpoints (motor, cognitive, paroxysmal spells, body weight) are assessed in the same animals across multiple groups
    Could also: Mixed-effects ANOVA or linear mixed model with repeated measures per animal, combined with Holm or Bonferroni correction across behavioral outcome families — Multiple correlated outcomes measured in the same animals introduce within-subject correlation and inflated family-wise error; a mixed model accounts for correlation structure and multiplicity correction controls false-positive rates across endpoints
  • ATPase activity rescue is listed as a biochemical outcome comparing hippocampal tissue across genotype and treatment groups
    Could also: One-way ANOVA followed by Dunnett's test (each group vs. WT reference) or Tukey HSD (all pairwise) — When several mutant or treated groups are each compared to a single WT reference, Dunnett's test is optimized for that comparison structure and is more powerful than general all-pairs post-hoc tests
Software: CIRCLE-seq (in vitro cleavage off-target nomination method) · rhAmpSeq (RNaseH-dependent amplification and sequencing)

Result convergence & founder nodes

Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
29
Impact: medium
Foundation confidence
Built on 1 assessed reference(s) · mean reproducibility 81/100
stands on reproducible work
Topics

Assessed papers, coloured by verdict. Click a node to open it.

Built on (assessed references) (1)
Cited by (assessed papers) (0)
  • No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

#132777 RRID in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet
#174820 RRID in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet
#174824 RRID in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet
#207853 RRID in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet
6VPC PDBe in Figure (http://semanticscience.org/resource/SIO_000080)
no other assessed paper uses this yet
8WUT PDBe in Figure (http://semanticscience.org/resource/SIO_000080)
no other assessed paper uses this yet
A63881 ENA in Table (http://semanticscience.org/resource/SIO_000419)
no other assessed paper uses this yet
Addgene_132775 RRID in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet
Addgene_174038 RRID in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet
Addgene_178113 RRID in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet
Addgene_207852 RRID in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet
Addgene_207854 RRID in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet
B23318 ENA in Table (http://semanticscience.org/resource/SIO_000419)
no other assessed paper uses this yet
CCL-131 RRID in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet
CRL-3216 RRID in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet
MGI:5654213 RRID in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet
MGI:7461667 RRID in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet
MMRRC #071287 RRID in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet
MMRRC #071376 RRID in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet
PRJNA1211588 BioProject in Table (http://semanticscience.org/resource/SIO_000419)
no other assessed paper uses this yet
RRID:AB_10918613 RRID in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet
RRID:AB_1642205 RRID in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Figures / tables: Figure 1JFigure 1KFigure 1LFigure 1MFigure 1N
Figure 1J (+) MLH1dn D801N correction
Reported
43%
Reproduced
43.8%
exact
Figure 1J (+) MLH1dn D801N indel
Reported
8%
Reproduced
7.66%
exact
Figure 1J (-) MLH1dn D801N correction
Reported
18%
Reproduced
18.0%
exact
Figure 1J (-) MLH1dn D801N indel
Reported
15%
Reproduced
14.78%
exact
Figure 1K (+) MLH1dn E815K correction
Reported
63%
Reproduced
70.0%
partial
Figure 1K (+) MLH1dn E815K indel
Reported
5.2%
Reproduced
1.35%
partial
Figure 1K (-) MLH1dn E815K correction
Reported
36%
Reproduced
47.3%
partial
Figure 1K (-) MLH1dn E815K indel
Reported
9.3%
Reproduced
2.34%
did not match
Figure 1L (+) MLH1dn L839P correction
Reported
70%
Reproduced
69.6%
exact
Figure 1L (+) MLH1dn L839P indel
Reported
1%
Reproduced
1.38%
exact
Figure 1L (-) MLH1dn L839P correction
Reported
71%
Reproduced
71.4%
exact
Figure 1L (-) MLH1dn L839P indel
Reported
1.0%
Reproduced
0.92%
exact
Figure 1M (+) MLH1dn G947R-C correction
Reported
74%
Reproduced
75.7%
exact
Figure 1M (+) MLH1dn G947R-C indel
Reported
4%
Reproduced
3.47%
within tolerance
Figure 1M (-) MLH1dn G947R-C correction
Reported
47%
Reproduced
49.6%
within tolerance
Figure 1M (-) MLH1dn G947R-C indel
Reported
11%
Reproduced
9.18%
within tolerance
Figure 1N ABE G947R-A correction
Reported
90%
Reproduced
90.0%
exact
Figure 1N ABE G947R-A indel
Reported
0.3%
Reproduced
0.28%
exact

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 84/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟢5. Derivability / plausibility
🟡6. Severity of the deviation
🟢7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +3

This is a strong, data-derivable reproduction: 7 of 9 conditions reproduce the reported Fig 1J–N correction values within ~2.6 points (6 exact), including the headline ABE 90% and D801N 43%, using the authors' own SRA data (PRJNA1211588) — no fabrication indicated and the central claim of efficient pathogenic-allele correction holds. The one deviation, E815K (Fig 1K), runs 7–11 points high on correction and low on indels in a consistent direction, attributable to our side — an ambiguous 'specified genotype'/QWC definition for its silent-edit epegRNA design, not an authors' or data defect. Severity is moderate and confined to a single variant; overall a solid reproduction with an explainable methodological deviation.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

718.3 k
tokens (I/O) · 109.2 M incl. cache
263 min
runtime · 0.28 CPU-h
2.6 GB
peak RAM
3 (2 failed)
HPC jobs
hummel
machine