Corpus 1,286 assessed · 1,187 scored · 648 reproduced ≥75 · 174 flagged ·∅ 73.9/100
← New search

Prime editing in mice reveals the essentiality of a single base in driving tissue-specific gene expression.

Genome Biol · 2021
L1 66/100 PQI 86
Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +2
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
  • No authors-side cause for any deviation
  • Reported values are derivable from the shared data
  • The central claim held under reproduction
What did not (or only partly)
  • 🟡A deviation arose in the data or preprocessing
  • 🟡The deviation was non-trivial in magnitude
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
66/100
Reproducibility score
0.4 SD below mean
vs. all fields · 1187 studies
🎯 Scores higher than 29% of all assessed papers rank 836 of 1187 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough to reproduce the CENTRAL pipeline result 1:1. NOTE a registry link mismatch: code_url=changeseq (CHANGE-seq off-target tool) does NOT pair with data_accession=GSE158388 (RNA-seq aorta transcriptomes); the CHANGE-seq reads behind the paper's 105/188 off-target claim are not in this accession and have no resolvable public accession. So per P16 we reproduced the pipeline whose data IS public: the paper's terminal RNA-seq stage, DESeq2 differential expression, on the deposited GSE158388 count matrices, run on «our HPC» («job», R 4.5.3 / DESeq2 1.50.2 vs paper's 1.22.1; data on «infra»). RESULTS: the paper's headline — disrupting a single CArG base virtually abolishes Tspan2 in aorta — reproduces cleanly in BOTH experiments: Tspan2 log2FC -3.06 (-88%, sg/HDR) and -3.45 (-91%, pe/PE2), both padj<1e-24, matching the reported ~90% reduction (C1, C3 within-tol). The deposited NormCounts are an EXACT DESeq2 normalization of the deposited raw counts (Pearson=1.000, machine-precision; C4 exact) — a positive internal-consistency/anti-fabrication signal that the processed data is bona-fide DESeq2 output. ONE honest discrepancy (C2): the paper lists Tspan2os among the only significantly reduced target genes, but in the deposited counts Tspan2os is barely expressed (baseMean <2.5) and is NOT significant (padj ~0.97-0.99), though its direction is down; likely a DESeq2-version / independent-filtering / annotation difference vs Suppl Table S2 (human should check S2) — not flagged as fabrication. NOT ATTEMPTED (honest): (a) CHANGE-seq off-target counts (105/188) — CHANGE-seq raw reads unavailable under this accession (data_unavailable for that sub-result); (b) upstream STAR/featureCounts re-alignment from SRA SRP285017 — the hard 20%, skipped because the deposited count matrix already pins the pipeline output the DE claim needs; (c) wet-lab/phenotype results (non-pipeline). Overall: status=partial because the secondary Tspan2os significance claim does not reproduce and CHANGE-seq is out of scope, but the paper's PRIMARY quantitative result reproduced 1:1. All grades are provisional and must be confirmed by a human against reproduction/outputs/deseq2_results.json and the «infra» data.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 66
    assessed: 2026-06-15 ⛓ 3a91049831d9
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-15
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-09-19

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The study tests whether a single transcription factor binding site (a CArG box in the Tspan2 promoter) is necessary for tissue-specific Tspan2 expression in mice, and compares the fidelity/efficiency of CRISPR-HDR versus prime editing (PE2) for introducing precise edits at this site in vivo.

Core claims
  • A three-base-pair (HDR) or single-base (PE2) substitution in the Tspan2 CArG box causes cell-specific loss of Tspan2 mRNA in aorta and bladder, but not heart or brain finding
  • PE2 prime editing achieves precise on-target single-base substitution with no detectable on-target or off-target indels, unlike HDR which produced indels in all founders and off-targets in many founders finding
  • Loss of the CArG box nearly abolishes expression of the overlapping antisense lncRNA Tspan2os in aorta and bladder finding
  • Bulk RNA-seq showed no distal off-target effects on gene expression in CArG-mutant aortae from either HDR or PE2 mice finding
  • In vitro transcription of PE2 mRNA combined with a synthetic pegRNA can be injected into mouse zygotes to achieve germline-transmitted precision editing method
  • CHANGE-seq combined with rhAmpSeq targeted sequencing can be used to comprehensively evaluate genome-wide off-target editing in founder mice method
  • Immuno-RNA FISH validated cell-specific (vascular smooth muscle) loss of Tspan2 mRNA in both HDR and PE2 CArG mutant mice finding
Experimental setups
Assay System Perturbation Readout Platform
qRT-PCR mouse tissues (aorta, bladder, heart, brain), Tspan2 sg/sg (HDR) and Tspan2 peg/peg (PE2) mice HDR or PE2-mediated CArG box substitution Tspan2 and Tspan2os mRNA levels
Allele-specific PCR and Sanger sequencing mouse zygotes/founder mice Cas9/sgRNA/ssODN (HDR) injection genotype confirmation of CArG box edit
Restriction digestion PCR and Sanger sequencing mouse zygotes/founder mice PE2 mRNA + synthetic pegRNA injection genotype confirmation of C>G transversion
Immunofluorescence (LMOD1) combined with RNA FISH mouse aorta, heart, brain (vascular smooth muscle cells) HDR or PE2 CArG box mutation spatial localization of Tspan2 mRNA and smooth muscle marker protein
Targeted sequencing (CRISPResso analysis) spleen genomic DNA from HDR and PE2 founder mice HDR (n=11) or PE2 (n=12) editing on-target editing frequency and indel frequency CRISPResso
Bulk RNA-seq mouse aorta, Tspan2+/+ vs Tspan2 sg/sg or Tspan2 peg/peg HDR or PE2 CArG box mutation genome-wide differential gene expression
CHANGE-seq genomic DNA, Cas9 complexed with sgRNA or pegRNA none (in vitro nuclease off-target profiling) genome-wide predicted off-target cleavage sites
rhAmpSeq targeted sequencing HDR and PE2 founder mice at CHANGE-seq/CasOFFinder predicted sites (244 + 13 sites) HDR or PE2 editing off-target mutation frequency rhAmpSeq
Key results
  • Tspan2 mRNA sharply attenuated in aorta and bladder of Tspan2 sg/sg mice; little change in heart or brain ~90% decrease vs Tspan2+/+
  • Tspan2 mRNA virtually abolished in aorta and bladder of Tspan2 peg/peg mice ~90% decrease vs Tspan2+/+
  • Tspan2os RNA near abrogated in aorta and bladder of PE2 CArG mutant mice
  • On-target editing frequency was higher in HDR founders than PE2 founders, but PE2 showed no on-target indels while all HDR founders had indels HDR mean 55.65% (range 1.67–95.56%) vs PE2 mean 20.74% (range 2.66–50.94%); HDR indels mean 40.11% (range 0.91–93.91%)
  • CHANGE-seq revealed fewer predicted off-targets for pegRNA than sgRNA 105 (pegRNA) vs 188 (sgRNA)
  • No distal effects on gene expression detected in CArG mutant aortae by bulk RNA-seq; no overlap in significantly changed genes between HDR and PE2 datasets
  • Off-target CRISPOR predicted sites did not show reduced expression of adjacent genes
Key statistics
  • other mean on-target editing 55.65% (range 1.67–95.56%) (HDR founder mice on-target editing frequency)
  • other mean on-target editing 20.74% (range 2.66–50.94%) (PE2 founder mice on-target editing frequency)
  • other mean indel frequency 40.11% (range 0.91–93.91%) (HDR founder mice on-target indels)
  • count 105 vs 188 predicted off-targets (CHANGE-seq off-target sites for pegRNA vs sgRNA respectively)
  • other MIT score 80, CFD score 93 (CRISPOR specificity scores for the protospacer)
  • fold_change ~90% decrease (Tspan2 and Tspan2os expression in mutant vs wild-type aorta)
  • count 20/37 (54%) HDR founders with correct editing; 12/47 (26%) PE2 founders with correct editing (founder genotyping success rates)
  • mean n = 5–7 mice/genotype (HDR qRT-PCR); n = 4 mice/genotype (PE2 qRT-PCR) (sample sizes for tissue mRNA quantification)

Statistical methods review

Model: opus

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The study is a comparative experimental design contrasting HDR- and PE2-mediated editing of a single TFBS in mice, with molecular readouts (qRT-PCR of Tspan2/Tspan2os across tissues and genotypes, RNA FISH, targeted/whole-genome sequencing, and bulk RNA-seq). Quantitative gene-expression results are shown as relative mean values with standard deviation across small numbers of mice per genotype, and significance is indicated by asterisks denoting p < 0.05; on/off-target editing is summarized as percentages and ranges across founders. The specific statistical test(s) generating the p-values, multiplicity handling, and analysis software for the qRT-PCR and RNA-seq comparisons are not explicitly named in the available text.

Replicationbiological Sample sizereported as number of mice (or aortae) per genotype/tissue per figure (e.g., n = 4 to 7); no formal power/sample-size calculation described Groupswild type vs heterozygous vs homozygous CArG-box mutants (HDR and PE2), across tissues (aorta, bladder, heart, brain) Pairingunpaired Randomization/blindingnot stated DispersionSD Exact p-valuesno Effect sizesno Confidence intervalsno Multiplicity correctionnone stated
Statistical tests used
Test Applied to n Assumptions
unspecified significance test (asterisks indicate p < 0.05) qRT-PCR of Tspan2 mRNA across tissues/genotypes (Fig. 1c,d) and related expression comparisons n = 5–7 mice/genotype (Fig. 1c); n = 6–7 mice/tissue (Fig. 1d); n = 4 mice/genotype (Fig. 3c); n = 4 aortae/genotype (Fig. 5b) not stated
differential expression analysis for bulk RNA-seq (method not named in available text) aorta RNA-seq, Tspan2+/+ vs Tspan2 sg/sg or Tspan2 peg/peg (Fig. 7) n = 4 aortae for each genotype not stated
Approaches that could also have been used
  • Quantitative qRT-PCR results were summarized as relative mean ± standard deviation (STD).
    Could also: Showing the same data with a 95% confidence interval, or overlaying individual data points alongside the mean, would also be possible. — With small n per genotype, plotting individual points and a CI can additionally convey the actual spread and overlap between groups, which many readers find informative.
  • Group differences were reported categorically with asterisks denoting p < 0.05.
    Could also: Reporting exact p-values together with an effect size (e.g., fold-change with its interval) would also be an option. — Exact values and effect sizes give readers the magnitude of differences and let them apply their own significance thresholds, complementing the categorical marker.
  • Multiple genotype-by-tissue comparisons of expression were each evaluated for significance.
    Could also: A single ANOVA framework (e.g., two-way ANOVA across genotype and tissue) with a post-hoc multiple-comparison correction such as Tukey HSD could also be applied. — An omnibus model with post-hoc correction would also control the family-wise error rate across the set of related comparisons in one analysis.
  • The specific test used to derive the p < 0.05 markers for qRT-PCR is not named in the available text.
    Could also: Stating the exact test used (e.g., two-tailed Student's t-test or Mann-Whitney U for small samples) along with its assumptions could also be done. — Naming the test and whether normality/variance assumptions were checked would let readers reproduce the analysis exactly; for very small n a rank-based test is one common choice.
  • Bulk RNA-seq differential expression was reported via scatter plots with 'significantly regulated' genes, with the analysis method not specified in the available text.
    Could also: A documented pipeline such as DESeq2 or edgeR/limma-voom with Benjamini-Hochberg FDR control could also be used and named. — Specifying the differential-expression tool and a stated FDR threshold would also make the genome-wide multiplicity handling explicit and reproducible, which is helpful given n = 4 per genotype.
  • Editing fidelity (on/off-target) was summarized descriptively as means and ranges across founders.
    Could also: A formal between-platform comparison (e.g., Mann-Whitney U on per-founder editing/indel percentages for HDR vs PE2) with reported effect size could also be presented. — An explicit test would also quantify the HDR-vs-PE2 difference in indel/off-target rates with an associated uncertainty rather than relying on descriptive ranges alone.
Software: CRISPOR (guide design / off-target prediction) · CRISPResso (indel/editing quantification) · CasOFFinder (off-target site prediction) · CHANGE-seq (genome-wide off-target detection) · rhAmpSeq (targeted sequencing of predicted sites)

Result convergence & founder nodes

Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
98
Impact: high
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

Assessed papers, coloured by verdict. Click a node to open it.

Built on (assessed references) (0)
  • No assessed neighbours yet — the network grows as more papers are assessed.
Cited by (assessed papers) (1)

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

GSE158388 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-33722289

Paper: Gao P, Lyu Q, Ghanam AR, … Tsai SQ, Long X, Miano JM. "Prime editing in mice reveals the essentiality of a single base in driving tissue-specific gene expression." Genome Biol 2021;22:83. PMID 33722289 · PMCID PMC7962346 · DOI 10.1186/s13059-021-02304-3.

Registry links (note the mismatch):

Key finding about the registry links

GSE158388 is RNA-seq ("Expression profiling by high throughput sequencing": aorta transcriptomes of sgCArG-WT vs -mutant @12 wk, and peCArG-WT vs -mutant @8 wk). The linked code repo changeseq is the CHANGE-seq genome-wide off-target-detection pipeline, which operates on tagmentation/CHANGE-seq sequencing reads — it cannot be run on this RNA-seq dataset. The code↔data pairing in the registry is a text-mining artifact: the paper uses CHANGE-seq (Fig 8a: 105 / 188 predicted off-targets for pegRNA / sgRNA Cas9), but the CHANGE-seq raw reads are not deposited under GSE158388 and no separate public accession for them is given. So changeseq is not reproducible from the data this RU points to.

Per brief rule P16 (a third-party tool on the paper's own data is equally valid), the reproducible pipeline here is the one whose input data is actually public: the RNA-seq differential-expression pipeline behind the paper's central transcriptome claim.

In scope (attempted)

The paper's RNA-seq pipeline (Methods): bcl2fastq 2.19.1 → FastP 0.20.0 → STAR 2.7.0f (GRCm38 / GENCODE-M22) → featureCounts (subread 1.6.4) → DESeq2 1.22.1, p-value threshold 0.05.

GEO ships the processed outputs of this pipeline as supplementary files:

  • GSE158388_experiment1_deSeq2_counts.txt.gz (raw gene counts, sg experiment)
  • GSE158388_experiment1_deSeq2_NormCounts.txt.gz (DESeq2-normalized counts, sg)
  • GSE158388_experiment2_deSeq2_counts.txt.gz (raw gene counts, pe experiment)
  • GSE158388_experiment2_deSeq2_NormCounts.txt.gz (DESeq2-normalized counts, pe)

Sample groups (from GEO sample titles):

  • experiment1 (sg, HDR/Cas9): 5 WT vs 5 sgCArG-mutant
  • experiment2 (pe, PE2): 4 WT vs 4 peCArG-mutant

Reproduction = re-run the terminal stage (DESeq2) on the deposited raw count matrices and check two things a human can verify 1:1:

  1. C1 / C2 — Tspan2 (and Tspan2os) differential expression. Paper: "~90% reduction" / "virtually abolished" in mutant aortas; "the only target genes significantly reduced … were Tspan2 and Tspan2os." We compute log2FC + adjusted-p for Tspan2/Tspan2os in both experiments and check sign, magnitude (~90% down ⇒ log2FC ≈ −3.3) and significance (padj < 0.05).
  2. C3 — normalization self-consistency. DESeq2 median-of-ratios size factors are deterministic, so our DESeq2-normalized counts must match the deposited *_NormCounts.txt matrices ~1:1. This is a clean internal check that the deposited processed data is genuinely DESeq2 output (anti-fabrication control).

Out of scope (not attempted) — with reasons

  • CHANGE-seq off-target counts (105 / 188, Fig 8a). Requires CHANGE-seq raw reads, which are NOT in GSE158388 and have no resolvable public accession. This is the part the linked changeseq repo would address, but the data is unavailable → cannot reproduce. (data_unavailable for that sub-result.)
  • Upstream RNA-seq stages (STAR alignment, featureCounts). Would require the raw FASTQs (SRA SRP285017). Re-aligning is the "hard 20%": large download + long compute, and the deposited count matrix already pins the pipeline output we need for the DE claim. Skipped per 80/20; the count matrix is the authors' own featureCounts output, so starting DESeq2 from it is faithful.
  • Wet-lab / phenotype results (editing efficiencies, founder genotyping, qPCR, histology, smooth-muscle phenotypes) — not pipeline-deri
Figures / tables: Fig 6TableFig 7Fig 8a
C1
Reported
Tspan2 ~90% reduced / virtually abolished in sgCArG (HDR) mutant aortas
Reproduced
DESeq2 log2FC=-3.06 (-88.0%), padj=6.3e-25 (5 WT vs 5 mut)
within tolerance
C3
Reported
Tspan2 virtually abolished in peCArG (prime-edited PE2) mutant aortas
Reproduced
DESeq2 log2FC=-3.45 (-90.9%), padj=1.3e-27 (4 WT vs 4 mut)
within tolerance
C4
Reported
Deposited NormCounts are DESeq2-1.22.1 normalized output of the deposited raw counts
Reproduced
Our DESeq2 median-of-ratios normalization == deposited *_NormCounts: Pearson=1.000, max rel diff ~5e-15 across 55,453 genes x all samples
exact
C2
Reported
Tspan2os among the only target genes significantly reduced (with Tspan2)
Reproduced
Tspan2os down (sg -74%, pe -82%) but NOT significant (padj 0.99/0.97; baseMean 0.65/2.46 — barely expressed)
did not match
COOS1+COOS2
Reported
CHANGE-seq predicted 105 (pegRNA) / 188 (sgRNA) off-targets (Fig 8a)
Reproduced
NOT ATTEMPTED — out of scope
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 66/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟡3. Location of the main deviation
🟢4. Cause of the deviation
🟢5. Derivability / plausibility
🟡6. Severity of the deviation
🟢7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +2

The paper's primary result — single CArG-base disruption virtually abolishes Tspan2 in aorta — reproduces 1:1 from the deposited GSE158388 counts in both sg/HDR (-88%, padj 6.3e-25) and pe/PE2 (-91%, padj 1.3e-27) experiments, and C4 confirms the deposited NormCounts are genuine DESeq2 output to machine precision (a positive internal-consistency signal). The one genuine deviation is secondary: Tspan2os, reported as significantly reduced, is barely expressed (baseMean <2.5) and not significant on reproduction — most plausibly a DESeq2-version/independent-filtering/annotation difference on our side rather than an authors' defect. The CHANGE-seq off-target counts (105/188) are simply not in this accession, so they are out of scope, not a discrepancy. Overall: solid reproduction with one explainable secondary deviation — central conclusion holds.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at [email protected].

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

121.3 k
tokens (I/O) · 6.5 M incl. cache
12 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.