In vivo prime editing rescues alternating hemiplegia of childhood in mice.
The main results reproduced: recomputed values matched the published ones within tolerance.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Same input data as the authors
- ✓Reported values were directly comparable
- ✓Reported values are derivable from the shared data
- ✓The central claim held under reproduction
- 🟡A deviation arose in the data or preprocessing
- 🟡A deviation was attributed to the published material
- 🟡The deviation was non-trivial in magnitude
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Described well enough to reproduce: YES. Ran CRISPResso2 2.3.4 on the paper's own amplicon-seq (SRA/ENA PRJNA1211588, 42 runs) on «our HPC» SLURM (final «job», COMPLETED 1h31m) to reproduce the central Figure 1J-N result: pathogenic-allele correction + indel rates for the four prime-edited ATP1A3 AHC mutations (D801N, E815K, L839P, G947R) +/-MLH1dn and the ABE-corrected G947R-A, in patient-derived iPSCs. Sample->figure->guide mapping from Table S3; amplicon refs by in-silico PCR of ST3B primers vs ATP1A3 NG_008015.1; pathogenic codons via NM_152296.5; correction = (WT%_edited - WT%_noedit)/(100 - WT%_noedit)x100 using matched No-edit controls (all 5 controls ~50% baseline WT = heterozygous lines). 1:1 OUTCOME: correction reproduces within tolerance in 7/9 conditions (6 exact incl. BOTH D801N now exact -- 43.8 vs 43, 18.0 vs 18; L839P 69.6/71.4 vs 70/71; G947R-C 75.7 vs 74; ABE 90.0 vs 90; 1 within-tol G947R-C 49.6 vs 47); indels match in 7/9 (5 exact). The headline editing efficiencies reproduce essentially exactly => genuine, data-derivable, NO fabrication indicated. DIFFERENT: E815K (Fig 1K) reproduces ~7-11 pts HIGH on correction and LOW on indels in both conditions -- a consistent methodological offset (E815K's 19-nt-spacer silent-edit design makes the paper's precise-haplotype definition stricter than our single-position WT-base readout; narrow E815K QWC undercounts indels), NOT fabrication. NOT ATTEMPTED: in-vivo mouse phenotyping (wet-lab); HEK293T screens Fig S2-S5 (same pipeline, not needed for headline); Fig 3 off-target quantification (Table S2 corrupt at PMC). All compute on «our HPC» SLURM compute nodes; env built in-job; FASTQs on «infra»; «host» holds only small results. Grades provisional; a human reviewer decides (see AUDIT.md).
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 83assessed: 2026-06-16 ⛓ 2c213e998654
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-22
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-18no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetThe paper tests whether prime editing (PE) and base editing (BE) can precisely correct pathogenic ATP1A3/Atp1a3 mutations in human cells and in vivo in mouse models to rescue the neurological phenotypes of alternating hemiplegia of childhood (AHC).
- ★ PE and BE strategies efficiently correct five prevalent ATP1A3 mutations (D801N, E815K, L839P, G947R-A, G947R-C) in HEK293T cells and AHC patient-derived iPSCs, with 43%-90% correction in iPSCs. finding
- ★ Bystander editing at adjacent bases limits base editor correction of the D801N, E815K, and L839P mutations, necessitating use of prime editing instead. finding
- ★ AAV9-mediated in vivo prime editing via P0 ICV injection corrects Atp1a3 D801N and E815K mutations in the CNS of two AHC mouse models, achieving up to 48% DNA correction and 73% mRNA correction in bulk brain cortex. finding
- ★ In vivo PE rescues clinically relevant AHC phenotypes in mice, including hippocampal Atp1a3 ATPase activity, paroxysmal spells, motor defects, cognition deficits, and greatly extends lifespan. finding
- ★ PE correction strategies show substantially lower and fewer off-target editing events than ABE8e in patient-derived iPSCs, as measured by CIRCLE-seq and rhAmpSeq. finding
- Adding 'dead' sgRNAs (dsgRNAs) that disrupt local chromatin accessibility enhances PE correction efficiency at multiple target sites. method
- Clustering silent edits near the corrected mutation and MLH1dn supplementation increase mismatch-repair evasion, improving PE correction purity. mechanism
- Sequence divergence between human ATP1A3 and mouse Atp1a3 (90.8% coding identity) required redesign of guide RNAs and PAMs for mouse-specific PE strategies. method
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| prime editing / base editing (transfection screen) | HEK293T cell lines engineered with pathogenic ATP1A3 variants | PE/CBE/ABE with various sgRNA/epegRNA/ngRNA designs | pathogenic allele correction % and indel rate | — |
| prime editing / ABE base editing | AHC patient-derived induced pluripotent stem cells (iPSCs) | PE with epegRNA/ngRNA/dsgRNA (PE6 variants, MLH1dn) or ABE8e | pathogenic allele correction % and indel rate | — |
| CIRCLE-seq | AHC patient-derived iPSCs (genomic DNA in vitro) | epegRNAs, ngRNAs, dsgRNAs, ABE sgRNA | nominated genome-wide off-target cleavage sites | — |
| rhAmpSeq (RNaseH-dependent amplification sequencing) | AHC patient-derived iPSCs treated with PE/ABE | PE or ABE editing vs donor-matched mock | indels and substitutions (log2 fold-change) at top CIRCLE-seq nominated off-target sites | — |
| prime editing (cell line generation) | mouse N2a cell line (monoclonal, D801N) | PE correction strategy redesigned for mouse Atp1a3 sequence | editing/correction efficiency | — |
| AAV9-mediated dual-vector in vivo prime editing (ICV injection at P0) | B6C3.Atp1a3 D801N/+ and B6C3.Atp1a3 E815K/+ AHC mouse models | in vivo PE (epegRNA/ngRNA delivered via dual AAV9) | DNA correction % and mRNA correction % in bulk brain cortex | AAV9 viral vector |
| ATPase activity assay | hippocampal tissue from D801N/E815K mice | in vivo PE treatment vs untreated | Atp1a3 Na+/K+ ATPase activity restoration | — |
| behavioral phenotyping (paroxysmal spell induction, motor tests, cognition tests, lifespan monitoring) | D801N and E815K AHC mouse models | in vivo PE treatment vs untreated mutant vs WT | paroxysmal spell frequency, motor performance, cognitive performance, survival/lifespan | — |
- ▲ G947R-A ABE editing achieved efficient, precise correction with no observed bystander editing up to 40% correction
- ▲ PE correction of D801N, E815K, L839P, and G947R-C in patient iPSCs with MLH1dn supplementation 43%, 63%, 70%, 74% pathogenic allele correction respectively
- ▲ ABE correction of G947R-A in patient-derived iPSCs 90% pathogenic allele correction, 0.3% indels
- ▲ dsgRNA non-PAM-44 improved D801N pathogenic allele correction 1.4-fold improvement, 55% correction
- ▲ In vivo AAV9 PE corrected Atp1a3 mutations in mouse brain cortex up to 48% DNA correction and 73% mRNA correction
- ▼ PE strategies showed fewer confirmed off-target loci than ABE8e among assayed sites 8 of 457 sites (PE) vs 16 of 457 sites (ABE8e)
- ▼ Alternative deaminase ABE7.10 eliminated confirmed off-target sites compared to ABE8e for G947R-A correction 0/95 sites (ABE7.10) vs 16/95 sites (ABE8e)
- ▲ In vivo PE treatment ameliorated paroxysmal spells, motor and cognition deficits, and extended lifespan in D801N mice
- count 1 in 1,000,000 (estimated incidence of AHC)
- fold_change 70% (proportion of AHC cases caused by ATP1A3 mutations)
- fold_change 40%, 20%, 10% (approximate prevalence of D801N, E815K, and G947R variants among ATP1A3-associated AHC cases)
- other 90.8% coding sequence identity (sequence divergence between human ATP1A3 and mouse Atp1a3)
- fold_change 43%/8%, 63%/5.2%, 70%/1%, 74%/4% (correction/indels) (iPSC pathogenic allele correction and indel rates for D801N, E815K, L839P, G947R-C with MLH1dn)
- fold_change up to 48% DNA correction, 73% mRNA correction (in vivo AAV9 PE correction in bulk mouse brain cortex)
- count 24 of 457 assayed sites confirmed off-target (off-target validation across PE and ABE guide RNAs in iPSCs via rhAmpSeq)
- fold_change 1.8-fold, 1.3-fold, 1.1-fold, 1.6-fold (ngRNA-driven improvements in pathogenic allele correction for D801N, E815K, L839P, G947R-C respectively)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This paper reports prime editing (PE) and base editing (BE) correction of ATP1A3/Atp1a3 mutations in HEK293T cells, patient-derived iPSCs, and two AHC mouse models. Editing outcomes are primarily quantified as pathogenic allele correction rates and indel rates expressed as percentages. In vivo efficacy is assessed via ATPase activity rescue, behavioral phenotyping, body weight, and survival in treated versus untreated AHC mice. Off-target editing is evaluated by CIRCLE-seq nomination followed by rhAmpSeq quantification at nominated sites, with results expressed as log2 fold-change versus donor-matched mock controls. The provided text excerpt is truncated before the full in vivo results and methods sections, so formal statistical tests for behavioral and survival endpoints are not visible in the supplied text.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| log2 fold-change relative to donor-matched mock controls; no named formal test visible in excerpt | off-target editing quantification at 457 rhAmpSeq-amplifiable CIRCLE-seq nominated sites (Figures 2A–2C) | 457 amplifiable sites of 480 nominated | not stated |
| fold-change ratio comparison between optimization conditions (no named statistical test) | epegRNA, ngRNA, and dsgRNA optimization across HEK293T and iPSC conditions (Figures S2–S5) | — | not stated |
| not described in available text (likely survival/log-rank; mentioned as outcome only) | lifespan comparison between treated and untreated D801N and E815K mice | — | not stated |
| not described in available text | behavioral outcomes: motor deficits, cognitive deficits, paroxysmal spell frequency | — | not stated |
| not described in available text | hippocampal Atp1a3 ATPase activity across WT, untreated, and PE-treated mice | — | not stated |
-
Pathogenic allele correction rates are reported as single percentage values without any measure of variability across replicates↳ Could also: Report mean ± SD or 95% CI across biological replicates of the editing experiment — Dispersion measures allow readers to assess reproducibility and distinguish robust from variable editing outcomes; for small n typical in iPSC work, SD conveys spread and CI conveys precision of the estimate
-
Off-target editing at 457 sites is characterized as log2 fold-change versus mock without a formal statistical threshold or multiplicity adjustment↳ Could also: Apply a negative binomial or Poisson model to per-site read counts comparing edited versus mock, followed by Benjamini-Hochberg FDR correction across all tested sites — A statistically principled per-site test with FDR control would provide a consistent, reproducible decision threshold for calling confirmed off-targets and would reduce the influence of sequencing noise on site classification
-
Gains from individual optimization steps (dsgRNA, ngRNA, PE6 variants) are each reported as isolated fold-changes without joint modeling of all factors↳ Could also: Factorial ANOVA or a linear mixed model treating editor variant, guide type, and their interactions as fixed factors — A factorial design analysis would partition variance attributable to each component and their interactions, enabling more principled conclusions about which factors independently contribute to editing gains
-
Lifespan is described as an outcome (extension in D801N mice) but the statistical approach is not stated in the available text↳ Could also: Kaplan-Meier survival curves compared with the log-rank (Mantel-Cox) test; Cox proportional hazards model if covariates such as dose or sex are present — The log-rank test is the standard nonparametric approach for comparing survival distributions between groups and correctly accounts for censoring, which is typical in animal survival studies
-
Multiple behavioral endpoints (motor, cognitive, paroxysmal spells, body weight) are assessed in the same animals across multiple groups↳ Could also: Mixed-effects ANOVA or linear mixed model with repeated measures per animal, combined with Holm or Bonferroni correction across behavioral outcome families — Multiple correlated outcomes measured in the same animals introduce within-subject correlation and inflated family-wise error; a mixed model accounts for correlation structure and multiplicity correction controls false-positive rates across endpoints
-
ATPase activity rescue is listed as a biochemical outcome comparing hippocampal tissue across genotype and treatment groups↳ Could also: One-way ANOVA followed by Dunnett's test (each group vs. WT reference) or Tukey HSD (all pairwise) — When several mutant or treated groups are each compared to a single WT reference, Dunnett's test is optimized for that comparison structure and is more powerful than general all-pairs post-hoc tests
Result convergence & founder nodes
Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.
-
Dual spacer guide RNA at non-PAM -44 position improves prime editing correction of ATP1A3 D801N by 1.4-fold to 55% in HEK293T cells.other hek293t up 2025×1papers★ This paper is the founder (earliest)
-
Prime editing and adenine base editing correct five AHC-causing ATP1A3 pathogenic variants at 43-90% efficiency in patient-derived iPSCs.other human ipsc up 2025×1papers★ This paper is the founder (earliest)
-
Prime editing generates fewer off-target edits than adenine base editing in AHC patient-derived iPSCs, with only 8 of 457 nominated sites showing any elevation, none exceeding 0.5%.other human ipsc down 2025×1papers★ This paper is the founder (earliest)
-
AAV9-delivered prime editing corrects Atp1a3 D801N and E815K mutations in AHC mouse brain cortex with up to 48% DNA and 73% mRNA correction efficiency in vivo.other mouse brain cortex up 2025×1papers★ This paper is the founder (earliest)
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
Assessed papers, coloured by verdict. Click a node to open it.
- No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
This is a strong, data-derivable reproduction: 7 of 9 conditions reproduce the reported Fig 1J–N correction values within ~2.6 points (6 exact), including the headline ABE 90% and D801N 43%, using the authors' own SRA data (PRJNA1211588) — no fabrication indicated and the central claim of efficient pathogenic-allele correction holds. The one deviation, E815K (Fig 1K), runs 7–11 points high on correction and low on indels in a consistent direction, attributable to our side — an ambiguous 'specified genotype'/QWC definition for its silent-edit epegRNA design, not an authors' or data defect. Severity is moderate and confined to a single variant; overall a solid reproduction with an explainable methodological deviation.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.