Prime editing in mice reveals the essentiality of a single base in driving tissue-specific gene expression.
The main results reproduced, with only marginal, non-material deviations.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Same input data as the authors
- ✓Reported values were directly comparable
- ✓No authors-side cause for any deviation
- ✓Reported values are derivable from the shared data
- ✓The central claim held under reproduction
- 🟡A deviation arose in the data or preprocessing
- 🟡The deviation was non-trivial in magnitude
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Described well enough to reproduce the CENTRAL pipeline result 1:1. NOTE a registry link mismatch: code_url=changeseq (CHANGE-seq off-target tool) does NOT pair with data_accession=GSE158388 (RNA-seq aorta transcriptomes); the CHANGE-seq reads behind the paper's 105/188 off-target claim are not in this accession and have no resolvable public accession. So per P16 we reproduced the pipeline whose data IS public: the paper's terminal RNA-seq stage, DESeq2 differential expression, on the deposited GSE158388 count matrices, run on «our HPC» («job», R 4.5.3 / DESeq2 1.50.2 vs paper's 1.22.1; data on «infra»). RESULTS: the paper's headline — disrupting a single CArG base virtually abolishes Tspan2 in aorta — reproduces cleanly in BOTH experiments: Tspan2 log2FC -3.06 (-88%, sg/HDR) and -3.45 (-91%, pe/PE2), both padj<1e-24, matching the reported ~90% reduction (C1, C3 within-tol). The deposited NormCounts are an EXACT DESeq2 normalization of the deposited raw counts (Pearson=1.000, machine-precision; C4 exact) — a positive internal-consistency/anti-fabrication signal that the processed data is bona-fide DESeq2 output. ONE honest discrepancy (C2): the paper lists Tspan2os among the only significantly reduced target genes, but in the deposited counts Tspan2os is barely expressed (baseMean <2.5) and is NOT significant (padj ~0.97-0.99), though its direction is down; likely a DESeq2-version / independent-filtering / annotation difference vs Suppl Table S2 (human should check S2) — not flagged as fabrication. NOT ATTEMPTED (honest): (a) CHANGE-seq off-target counts (105/188) — CHANGE-seq raw reads unavailable under this accession (data_unavailable for that sub-result); (b) upstream STAR/featureCounts re-alignment from SRA SRP285017 — the hard 20%, skipped because the deposited count matrix already pins the pipeline output the DE claim needs; (c) wet-lab/phenotype results (non-pipeline). Overall: status=partial because the secondary Tspan2os significance claim does not reproduce and CHANGE-seq is out of scope, but the paper's PRIMARY quantitative result reproduced 1:1. All grades are provisional and must be confirmed by a human against reproduction/outputs/deseq2_results.json and the «infra» data.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 66assessed: 2026-06-15 ⛓ 3a91049831d9
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-15
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusIs a single transcription factor binding site (a CArG box in the Tspan2 promoter) necessary for tissue-specific Tspan2 expression in mice, and can two-component prime editing (PE2) edit this single base in vivo with higher fidelity than three-component CRISPR-mediated HDR?
- ★ A PE2-mediated single-base (C>G) substitution in the Tspan2 CArG box causes cell-specific loss of Tspan2 mRNA in aorta and bladder but not heart or brain, mirroring HDR-mediated 3-bp substitution. finding
- ★ PE2 directs high-fidelity editing with no on-target indels above background and no detected off-target mutations, unlike HDR which produced indels in all founders and off-target mutations. finding
- ★ A single base within the CArG box co-regulates an mRNA/long noncoding RNA antisense gene pair (Tspan2 and Tspan2os), with Tspan2os nearly abolished in aorta and bladder of edited mice. mechanism
- ★ Prime editing (PE2) can be applied in mice bred through the germline to model noncoding SNVs in transcription factor binding sites. method
- The Tspan2 CArG box is an SRF-binding site located 539 bp upstream of the human TSPAN2 transcription start site and is conserved across mammals including mouse. finding
- Immuno-RNA FISH (LMOD1 + Tspan2) confirms loss of Tspan2 mRNA in vascular smooth muscle cells of CArG box mutant mice. method
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| CRISPR HDR genome editing (Cas9 protein + sgRNA + ssODN, 3 bp substitution) | mouse zygotes / C57 mice (germline) | KO of CArG box via 3 bp substitution (HDR) | editing/genotype, germline transmission | CRISPOR-designed sgRNA |
| Prime editing (PE2: Cas9 nickase-reverse transcriptase mRNA + pegRNA, single C>G transversion) | mouse zygotes / mice (germline) | single-base C>G substitution in CArG box (PE2) | editing/genotype, germline transmission | in vitro-transcribed pCMV-PE2 mRNA, synthetic pegRNA |
| Quantitative RT-PCR (qRT-PCR) | mouse aorta, bladder, heart, brain | Tspan2 CArG box mutant (sg/sg and peg/peg) | Tspan2 and Tspan2os mRNA relative expression / Ct values | — |
| Immunofluorescence + RNA fluorescence in situ hybridization (immuno-RNA FISH) | mouse aorta, heart coronary vessels, brain (vascular smooth muscle cells) | CArG box mutant (sg/sg and peg/peg) | spatial localization of Tspan2 mRNA and LMOD1 protein | — |
| Targeted amplicon sequencing (on-target) and rhAmpSeq (off-target) | genomic DNA from spleen of HDR and PE2 founder mice | HDR vs PE2 editing | % correct editing, % indels, off-target mutations | CRISPResso analysis; rhAmpSeq |
| Bulk RNA-seq | mouse aortae (Tspan2 +/+ vs sg/sg or peg/peg) | CArG box mutant | genome-wide differential gene expression | — |
| CHANGE-seq (genome-wide off-target prediction) | in vitro with wild type Cas9 nuclease + sgRNA or pegRNA | none (nuclease complexed with guide) | predicted genome-wide off-target sites | CHANGE-seq |
- ▼ Tspan2 mRNA sharply attenuated in aorta and bladder of HDR Tspan2 sg/sg mice, with little change in heart or brain; intermediate in heterozygotes
- ▼ Tspan2 mRNA virtually abolished in aorta and bladder of PE2 Tspan2 peg/peg mice with little change in brain and heart ~90% decrease
- ▼ Tspan2os lncRNA nearly abrogated in aorta and bladder of C>G transversion homozygous mice ~90% decrease
- – All HDR founders showed undesired on-target indels, whereas no PE2 founders displayed indels above background
- ▼ Off-target analysis revealed mutations in many HDR founders but none in PE2 founders
- ▼ Only Tspan2 and Tspan2os were significantly reduced in both HDR and PE2 mutant aorta by bulk RNA-seq; no distal off-target effects ~90% decrease
- mean 55.65% (range 1.67–95.56%) correct on-target editing reads (HDR founder mice)
- mean 20.74% (range 2.66–50.94%) correct on-target editing reads (PE2 founder mice)
- mean 40.11% indels (range 0.91–93.91%) (undesired indels in HDR founders)
- count 105 and 188 predicted off-targets for Cas9 with pegRNA or sgRNA respectively (CHANGE-seq off-target prediction)
- count 20/37 (54%) founders correctly edited; 4/20 (20%) with indels (HDR founder genotyping)
- count 12/47 (26%) founders correctly edited (PE2 founder genotyping)
- other specificity scores 80 (MIT) and 93 (CFD); 0,1,0,12,68 predicted off-targets with 0–4 mismatches (CRISPOR protospacer analysis)
- count Mendelian: 12 +/+, 24 peg/+, 10 peg/peg (PE2 F1 intercross genotypes)
Statistical methods review
Model: opusA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
The study is a comparative experimental design contrasting HDR- and PE2-mediated editing of a single TFBS in mice, with molecular readouts (qRT-PCR of Tspan2/Tspan2os across tissues and genotypes, RNA FISH, targeted/whole-genome sequencing, and bulk RNA-seq). Quantitative gene-expression results are shown as relative mean values with standard deviation across small numbers of mice per genotype, and significance is indicated by asterisks denoting p < 0.05; on/off-target editing is summarized as percentages and ranges across founders. The specific statistical test(s) generating the p-values, multiplicity handling, and analysis software for the qRT-PCR and RNA-seq comparisons are not explicitly named in the available text.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| unspecified significance test (asterisks indicate p < 0.05) | qRT-PCR of Tspan2 mRNA across tissues/genotypes (Fig. 1c,d) and related expression comparisons | n = 5–7 mice/genotype (Fig. 1c); n = 6–7 mice/tissue (Fig. 1d); n = 4 mice/genotype (Fig. 3c); n = 4 aortae/genotype (Fig. 5b) | not stated |
| differential expression analysis for bulk RNA-seq (method not named in available text) | aorta RNA-seq, Tspan2+/+ vs Tspan2 sg/sg or Tspan2 peg/peg (Fig. 7) | n = 4 aortae for each genotype | not stated |
-
Quantitative qRT-PCR results were summarized as relative mean ± standard deviation (STD).↳ Could also: Showing the same data with a 95% confidence interval, or overlaying individual data points alongside the mean, would also be possible. — With small n per genotype, plotting individual points and a CI can additionally convey the actual spread and overlap between groups, which many readers find informative.
-
Group differences were reported categorically with asterisks denoting p < 0.05.↳ Could also: Reporting exact p-values together with an effect size (e.g., fold-change with its interval) would also be an option. — Exact values and effect sizes give readers the magnitude of differences and let them apply their own significance thresholds, complementing the categorical marker.
-
Multiple genotype-by-tissue comparisons of expression were each evaluated for significance.↳ Could also: A single ANOVA framework (e.g., two-way ANOVA across genotype and tissue) with a post-hoc multiple-comparison correction such as Tukey HSD could also be applied. — An omnibus model with post-hoc correction would also control the family-wise error rate across the set of related comparisons in one analysis.
-
The specific test used to derive the p < 0.05 markers for qRT-PCR is not named in the available text.↳ Could also: Stating the exact test used (e.g., two-tailed Student's t-test or Mann-Whitney U for small samples) along with its assumptions could also be done. — Naming the test and whether normality/variance assumptions were checked would let readers reproduce the analysis exactly; for very small n a rank-based test is one common choice.
-
Bulk RNA-seq differential expression was reported via scatter plots with 'significantly regulated' genes, with the analysis method not specified in the available text.↳ Could also: A documented pipeline such as DESeq2 or edgeR/limma-voom with Benjamini-Hochberg FDR control could also be used and named. — Specifying the differential-expression tool and a stated FDR threshold would also make the genome-wide multiplicity handling explicit and reproducible, which is helpful given n = 4 per genotype.
-
Editing fidelity (on/off-target) was summarized descriptively as means and ranges across founders.↳ Could also: A formal between-platform comparison (e.g., Mann-Whitney U on per-founder editing/indel percentages for HDR vs PE2) with reported effect size could also be presented. — An explicit test would also quantify the HDR-vs-PE2 difference in indel/off-target rates with an associated uncertainty rather than relying on descriptive ranges alone.
Result convergence & founder nodes
Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.
-
Off-target sequencing reveals mutations in many HDR founder mice but none in PE2 founders, demonstrating that in vivo prime editing eliminates off-target mutagenesis.other mouse spleen down 2021×1papers★ This paper is the founder (earliest)
-
All HDR-edited founder mice carry undesired on-target indels whereas no PE2 founders display indels above background, demonstrating superior on-target precision for prime editing.other mouse spleen mixed 2021×1papers★ This paper is the founder (earliest)
-
TSPAN2 mRNA is ~90% reduced in aorta and bladder but not heart or brain in prime-edited (PE2) CArG-box C>G mutant homozygous mice.qPCR mouse aorta down 2021×1papers★ This paper is the founder (earliest)
-
TSPAN2OS lncRNA is ~90% reduced in aorta and bladder of CArG-box C>G transversion homozygous mice.qPCR mouse aorta down 2021×1papers★ This paper is the founder (earliest)
-
Bulk RNA-seq of CArG-box mutant aorta shows TSPAN2 and TSPAN2OS are the only significantly reduced transcripts, with no distal off-target transcriptional effects.RNA-seq mouse aorta down 2021×1papers★ This paper is the founder (earliest)
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
Assessed papers, coloured by verdict. Click a node to open it.
- No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-33722289
Paper: Gao P, Lyu Q, Ghanam AR, … Tsai SQ, Long X, Miano JM. "Prime editing in mice reveals the essentiality of a single base in driving tissue-specific gene expression." Genome Biol 2021;22:83. PMID 33722289 · PMCID PMC7962346 · DOI 10.1186/s13059-021-02304-3.
Registry links (note the mismatch):
- code_url = https://github.com/tsailabSJ/changeseq (Tsai-lab CHANGE-seq tool)
- data_accession = GSE158388 (RNA-seq, aorta transcriptomes — NOT CHANGE-seq)
Key finding about the registry links
GSE158388 is RNA-seq ("Expression profiling by high throughput sequencing":
aorta transcriptomes of sgCArG-WT vs -mutant @12 wk, and peCArG-WT vs -mutant
@8 wk). The linked code repo changeseq is the CHANGE-seq genome-wide
off-target-detection pipeline, which operates on tagmentation/CHANGE-seq
sequencing reads — it cannot be run on this RNA-seq dataset. The code↔data
pairing in the registry is a text-mining artifact: the paper uses CHANGE-seq
(Fig 8a: 105 / 188 predicted off-targets for pegRNA / sgRNA Cas9), but the
CHANGE-seq raw reads are not deposited under GSE158388 and no separate
public accession for them is given. So changeseq is not reproducible from the
data this RU points to.
Per brief rule P16 (a third-party tool on the paper's own data is equally valid), the reproducible pipeline here is the one whose input data is actually public: the RNA-seq differential-expression pipeline behind the paper's central transcriptome claim.
In scope (attempted)
The paper's RNA-seq pipeline (Methods): bcl2fastq 2.19.1 → FastP 0.20.0 → STAR 2.7.0f (GRCm38 / GENCODE-M22) → featureCounts (subread 1.6.4) → DESeq2 1.22.1, p-value threshold 0.05.
GEO ships the processed outputs of this pipeline as supplementary files:
GSE158388_experiment1_deSeq2_counts.txt.gz(raw gene counts, sg experiment)GSE158388_experiment1_deSeq2_NormCounts.txt.gz(DESeq2-normalized counts, sg)GSE158388_experiment2_deSeq2_counts.txt.gz(raw gene counts, pe experiment)GSE158388_experiment2_deSeq2_NormCounts.txt.gz(DESeq2-normalized counts, pe)
Sample groups (from GEO sample titles):
- experiment1 (sg, HDR/Cas9): 5 WT vs 5 sgCArG-mutant
- experiment2 (pe, PE2): 4 WT vs 4 peCArG-mutant
Reproduction = re-run the terminal stage (DESeq2) on the deposited raw count matrices and check two things a human can verify 1:1:
- C1 / C2 — Tspan2 (and Tspan2os) differential expression. Paper: "~90% reduction" / "virtually abolished" in mutant aortas; "the only target genes significantly reduced … were Tspan2 and Tspan2os." We compute log2FC + adjusted-p for Tspan2/Tspan2os in both experiments and check sign, magnitude (~90% down ⇒ log2FC ≈ −3.3) and significance (padj < 0.05).
- C3 — normalization self-consistency. DESeq2 median-of-ratios size
factors are deterministic, so our DESeq2-normalized counts must match the
deposited
*_NormCounts.txtmatrices ~1:1. This is a clean internal check that the deposited processed data is genuinely DESeq2 output (anti-fabrication control).
Out of scope (not attempted) — with reasons
- CHANGE-seq off-target counts (105 / 188, Fig 8a). Requires CHANGE-seq raw
reads, which are NOT in GSE158388 and have no resolvable public accession.
This is the part the linked
changeseqrepo would address, but the data is unavailable → cannot reproduce. (data_unavailable for that sub-result.) - Upstream RNA-seq stages (STAR alignment, featureCounts). Would require the raw FASTQs (SRA SRP285017). Re-aligning is the "hard 20%": large download + long compute, and the deposited count matrix already pins the pipeline output we need for the DE claim. Skipped per 80/20; the count matrix is the authors' own featureCounts output, so starting DESeq2 from it is faithful.
- Wet-lab / phenotype results (editing efficiencies, founder genotyping, qPCR, histology, smooth-muscle phenotypes) — not pipeline-deri
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
The paper's primary result — single CArG-base disruption virtually abolishes Tspan2 in aorta — reproduces 1:1 from the deposited GSE158388 counts in both sg/HDR (-88%, padj 6.3e-25) and pe/PE2 (-91%, padj 1.3e-27) experiments, and C4 confirms the deposited NormCounts are genuine DESeq2 output to machine precision (a positive internal-consistency signal). The one genuine deviation is secondary: Tspan2os, reported as significantly reduced, is barely expressed (baseMean <2.5) and not significant on reproduction — most plausibly a DESeq2-version/independent-filtering/annotation difference on our side rather than an authors' defect. The CHANGE-seq off-target counts (105/188) are simply not in this accession, so they are out of scope, not a discrepancy. Overall: solid reproduction with one explainable secondary deviation — central conclusion holds.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.