Massively parallel genomic perturbations with multi-target CRISPR interrogates Cas9 activity and DNA repair at endogenous sites.
The main results reproduced: recomputed values matched the published ones within tolerance.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Same input data as the authors
- ✓Reported values were directly comparable
- ✓No authors-side cause for any deviation
- ✓Reported values are derivable from the shared data
- ✓Any deviation was negligible
- ✓The central claim held under reproduction
- 🟡A deviation arose in the data or preprocessing
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
REPRODUCED. Repo = rogerzou/multitargetCRISPR @ e06c27e (NOT chipseq_pcRNA -- that is a different paper; corrected). Paper is described well enough to reproduce. (A) The paper's headline on-target counts -- CT 145 / GG 126 / TA 117 -- reproduce EXACTLY by an independent re-derivation from the raw hg38 genome (exact 20bp protospacer + NGG PAM scan, both strands); all 388 cut-site coordinates match the shipped mgRNA_target_sites.xlsx (not a tautology -- derived from the reference, not the shipped file). (B) The 5,236 genome-wide GG dCas9 binding sites (Fig 4a) reproduce to within ~1.6%: aligning dCas9 3h ChIP (kGG-cas9-3h, SRR19744254/53) vs a no-Cas9 control (WT-cas9-rep1, SRR19744223) with the authors' bowtie2 --local hg38 + MAPQ>=25/markdup pipeline, calling peaks with macs2 -f BAMPE -g hs, and applying the authors' own custom target-assignment step (src/mtss.macs_gen: keep FE-peaks with an identifiable GG protospacer+GG-PAM within 200bp) yields 5,321 sites at fold-enrichment>=8 (the threshold used in script_3) and 5,975 at fold-enrichment>=4 (Methods). The reported 5,236 sits squarely in this range. Note: the raw macs2 FE>=4 peak count alone is ~11,000 -- the target-assignment filter is essential and the Methods' description maps cleanly onto the shipped code. NOT attempted (out of quick scope): the >40,000 putative mgRNA universe (A5) and amplicon editing kinetics (C1). No fabrication indicators: all reproduced numbers are derivable from public data + shipped code.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 98assessed: 2026-06-22 ⛓ b60f6f05a35f
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-22
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-22no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetWhether a single multi-target guide RNA (mgRNA) directing Cas9 to hundreds of well-mapped endogenous genomic sites simultaneously can enable high-throughput, generalizable characterization of Cas9 binding/cleavage dynamics and the ensuing cellular DNA damage response, while controlling for guide RNA sequence identity.
- ★ Multi-target gRNAs (mgRNAs) can direct Cas9 to over a hundred well-mapped endogenous genomic sites simultaneously, enabling massively parallel, high-throughput interrogation of Cas9 activity via short-read sequencing method
- ★ Cas9 departs rapidly from genomic DNA after cleavage, preferentially releasing the PAM-proximal DNA end, which facilitates MRE11 loading on that side mechanism
- ★ Cas9 binding to genomic DNA is enhanced at chromatin-accessible regions finding
- ★ Cleavage efficiency by bound Cas9 is higher near transcribed regions finding
- ★ DNA damage-induced chromatin decompaction is reversible, spatially limited (under 2 kb) and occurs with kinetics of approximately 1 hour finding
- ★ MRE11 enrichment is restricted to target sites with two or fewer mismatches, whereas dCas9 binding tolerates over eight mismatches, with both enhanced when mismatches are PAM-distal finding
- High heterogeneity in dCas9 and MRE11 enrichment occurs even between identical on-target sequences, likely due to epigenetic factors finding
- mgRNAs generate mutation signatures (predominantly 1-bp insertions) consistent with known Cas9-induced DSB repair profiles finding
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| In silico genome-wide target site search / bowtie2 alignment simulation | human (hg38), mouse, zebrafish genomes | none | number of on-target sites per gRNA; proportion of ambiguous sequencing reads | bowtie2 |
| High-throughput amplicon sequencing (indel/mutation analysis) | HeLa cells with Dox-inducible Cas9 | mgRNA-directed Cas9 (10-target and other mgRNAs), Dox induction over 0/2/6/10 days | indel rate, mutation type (deletions, SNVs, insertions) | PCR amplification + high-throughput sequencing |
| ChIP-seq for Cas9 and MRE11 | HEK293T cells | electroporation of Cas9 protein pre-assembled with mgRNA ('CT', 'GG', 'TA'), 3 h | Cas9/MRE11 binding enrichment, read span vs. abut classification at cut sites | ChIP-seq; peak calling with MACS2 |
| ChIP-seq for DNA damage markers 53BP1 and γH2AX | HEK293T cells | Cas9/mgRNA delivery | enrichment correlation with MRE11/Cas9 signal | ChIP-seq |
| ChIP-seq for Cas9 and MRE11 | induced pluripotent stem cells (WTC-11 iPSCs) | Cas9/mgRNA delivery | enrichment at target sites, correlation with HEK293T and between replicates | ChIP-seq |
| ChIP-seq for dCas9 (nuclease-dead) | HEK293T cells | dCas9/mgRNA ('GG') delivery | number and mismatch tolerance of binding sites | ChIP-seq |
- ▲ 10-target mgRNA in HeLa cells produced detectable indels at 8 of 10 target sites over 10 days 8/10 sites
- – Mutation type distributions were highly reproducible between biological replicates r=0.99
- – MRE11 ChIP-seq reads predominantly abutted cut sites while Cas9 reads predominantly spanned cut sites, indicating rapid post-cleavage Cas9 departure
- ▲ PAM-distal read species (dist+4) was significantly more enriched than PAM-proximal species, supporting preferential PAM-proximal Cas9 release P = 3×10^-80, 3×10^-57, 3×10^-16 (CT, TA, GG)
- ▲ 'GG' mgRNA showed significantly stronger Cas9 association with the PAM-proximal side compared to 'CT' and 'TA' P = 0.35, 2.6×10^-22, 6.6×10^-19
- – 5,236 dCas9 binding sites detected for the 'GG' mgRNA, comparable to off-target numbers from single-targeting gRNAs 5,236 sites
- – MRE11 enrichment restricted to ≤2 mismatch sites; dCas9 binding detected at sites with >8 mismatches, especially PAM-distal mismatches >8 mismatches (dCas9) vs ≤2 (MRE11)
- – Median distance between adjacent Cas9 binding sites was much smaller than between adjacent on-target sites 265 kb vs 13.2 Mb
- correlation r = 0.99 (reproducibility of mutation-type percentages between biological replicates)
- pvalue 3×10^-80, 3×10^-57, 3×10^-16 (dist+4 vs. prox+4/prox-4 read enrichment for CT, TA, GG mgRNAs (Student's t-test))
- pvalue 0.35, 2.6×10^-22, 6.6×10^-19 (comparison of PAM-proximal read numbers between gRNA sequences (Student's t-test))
- count 5,236 (dCas9 binding sites detected for 'GG' mgRNA by ChIP-seq)
- count 8/10 (target sites with detectable indels from 10-target mgRNA in HeLa cells)
- other 265 kb (median) (distance between adjacent MACS2-detected Cas9 binding sites)
- other 13.2 Mb (median) (distance between adjacent on-target sites)
- count >40,000 sequences, 2 to >1,000 on-target sites each (in silico discovery of candidate mgRNA target sequences in hg38)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
The paper is a technical/methods report describing a multiplexed CRISPR system (mgRNAs) and validating it primarily through descriptive and comparative genomic/ChIP-seq analyses. Quantitative comparisons (e.g., of ChIP-seq read categories between PAM-proximal/distal sides, or between different guide RNA sequences) were assessed with two-sided, unadjusted Student's t-tests, and relationships between variables (e.g., replicate reproducibility, enrichment correlations) were assessed with Pearson correlation coefficients. Results were reported largely as violin plots, scatter plots with correlation coefficients, and exact p-values, with biological replicates (typically two) averaged for reproducibility figures.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| two-sided unadjusted Student's t-test | Fig. 3m–o: comparison of dist+4 read counts vs. sum of prox+4 and prox−4 read counts, per target sequence (CT, TA, GG) | — | not stated |
| two-sided unadjusted Student's t-test | Fig. 3p: comparison of PAM-proximal read counts (prox+4 + prox−4) between different gRNA sequences (CT, TA, GG) | — | not stated |
| Pearson correlation coefficient | Fig. 1n: reproducibility of mutation-type percentages across biological replicates (r = 0.99) | — | not stated |
| Pearson correlation (with P value) | Fig. 3h: MRE11 PAM-proximal bias vs. Cas9 PAM-distal bias across target sites | — | not stated |
| Pearson correlation coefficient | Fig. 2o,p: ChIP–seq enrichment correlation between HEK293T cells and iPSCs; Fig. 2f,q,r: correlation between biological replicates | — | not stated |
| Linear regression (slope estimation) | Fig. 3d–f: PAM-proximal vs. PAM-distal RPM for MRE11 and Cas9 at each on-target site | — | not stated |
-
Multiple pairwise two-sided Student's t-tests were used to compare PAM-proximal read counts across three gRNA sequences (CT, TA, GG) in Fig. 3p, each reported with its own unadjusted P value.↳ Could also: A one-way ANOVA followed by a post-hoc test (e.g., Tukey HSD) across the three gRNA groups — An omnibus ANOVA with post-hoc correction would also control the family-wise error rate when making all pairwise comparisons among the three groups simultaneously, which can be a useful complement to performing several separate two-group tests.
-
P values from the Student's t-tests are described as 'unadjusted' across multiple figure panels and comparisons (e.g., three P values in Fig. 3m–o and three in Fig. 3p).↳ Could also: A multiplicity correction such as Benjamini-Hochberg FDR or Bonferroni across the full set of comparisons performed — Applying a correction across the family of related comparisons would also help control the overall false-positive rate when multiple statistical tests are performed on related data from the same experiment.
-
Reproducibility between biological replicates was assessed using Pearson correlation coefficients (e.g., r = 0.99 in Fig. 1n; correlations in Fig. 2f, q, r).↳ Could also: Bland-Altman analysis or an intraclass correlation coefficient (ICC) — These approaches would also assess replicate agreement, and can additionally reveal systematic bias or magnitude-dependent variability between replicates that a single correlation coefficient does not directly show.
-
Relationships between ChIP-seq enrichment signals (e.g., MRE11 vs. Cas9, HEK293T vs. iPSC enrichment) were assessed using Pearson correlation.↳ Could also: Spearman rank correlation — A rank-based correlation would also capture monotonic association between variables and can be less sensitive to outliers or non-normal distributions, which are common in ChIP-seq enrichment data.
-
Several key validation results (e.g., mutation rate and mutational signature analyses in Fig. 1l,m) are based on averages of two biological replicates without an accompanying formal variability or hypothesis test.↳ Could also: Reporting a dispersion measure (e.g., range or SD across replicates) or, where feasible, increasing replicate number for a formal statistical comparison — Adding a dispersion measure or additional replicates would also allow a more quantitative assessment of the consistency of the observed averages beyond visual inspection of two data points.
-
Significance for the t-tests is reported primarily via exact P values (e.g., P < 1 × 10−15) without an accompanying standardized effect size.↳ Could also: Reporting an effect size measure (e.g., Cohen's d or a fold-difference with confidence interval) alongside the P value — An effect size would also convey the magnitude of the difference between groups, which can be a useful complement to the P value, particularly for very large sample sizes where small differences can yield very small P values.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-36064968
Paper: Zou RS, Marin-Gonzalez A, Liu Y, Liu HB, Shen L, Dveirin RK, Luo JXJ, Kalhor R, Ha T. Massively parallel genomic perturbations with multi-target CRISPR interrogates Cas9 activity and DNA repair at endogenous sites. Nat Cell Biol 2022. PMID 36064968 / PMC9481459 / DOI 10.1038/s41556-022-00975-z.
Code (verbatim Code-availability): "Analysis code is available on GitHub (https://github.com/rogerzou/multitargetCRISPR)."
NOTE: registry/scaffold listed chipseq_pcRNA — that is a DIFFERENT paper (photocleavable gRNA, SRA PRJNA622564). Corrected in code/code.json.
Data (verbatim Data-availability): "Deep-sequencing data generated for this study have been deposited in Sequence Read Archive under BioProject accession PRJNA733683." Plus public epigenetic tracks (ATAC SRR6418075; DNase ENCFF120XFB; H3K4me1/3, H3K9me3, H3K27ac, H3K36me3 ENCODE; MNase ERR2403161; RNA-seq SRR5627161). Genome hg38 (GCF_000001405.26) / hg19.
Concept
Multi-target gRNAs (mgRNAs) are 20-bp protospacers that perfectly match a repetitive
motif in Alu elements (CCTGTAGTCCCAGCTAC + a 3-nt variable end), so a single gRNA
directs Cas9/dCas9 to MANY endogenous loci at once. Three mgRNAs are named by their
PAM-proximal end: GG, CT, TA. This gives massively parallel, internally-controlled
measurement of Cas9 binding/cutting and DNA-repair kinetics across the genome.
Sequencing assays in PRJNA733683 (229 ENA runs)
- ChIP-seq: 115 runs — Cas9 / dCas9 / MRE11 / γH2AX / 53BP1, mgRNA variants (GG,CT,TA,kGG,iGG,pcGG,cgGG), timepoints + reps.
- ATAC-seq: 60 runs — chromatin accessibility at cut sites over time (GG/dCas9/dNick/pcGG/cgGG).
- AMPLICON: 54 runs — 38 ct10/ct20 multi-target editing timeseries (T0/T2/T6/T10 days, R1/R2, Ind/NI) + 16 BLISS DSB libraries.
Pipeline-derived results (IN SCOPE)
| # | Result | Reported value (paper loc) | Pipeline | Tractability |
|---|---|---|---|---|
| A | mgRNA on-target site counts | CT 145, GG 126, TA 117 (Results, "three mgRNAs ('CT','GG','TA') with 145, 126 and 117 on-target sites") + mgRNA_target_sites.xlsx | script_1_putative.py: scan hg38 for the 20-mer protospacer perfect matches + NGG PAM | LOW compute (genome scan); PRIMARY / quick-minimum |
| A' | Putative mgRNA universe | "Over 40,000 20-bp sequences, each targeting 2 to >1,000 putative sites" (Results) | script_1_putative.py | LOW-MED compute |
| B | GG dCas9 genome-wide binding sites | 5,236 dCas9 binding sites for GG mgRNA (Fig. 4a); MACS2 fold-enrichment ≥4 | process_reads.sh (bowtie2 hg38, MAPQ≥25, dedup) → process_macs2.sh (MACS2) | MED compute (download+align dCas9 ChIP + control) |
| C | Multi-target editing kinetics | indels at 8/10 sites over 10 days, 10-target; ED Fig.1m,n for 20-target (Fig. 1l) | amplicon indel caller (script_3_kinetics.py / lib) on ct10/ct20 runs | LOW-MED compute (amplicon = small) |
OUT OF SCOPE (wet-lab / manual / external / not pipeline-derived here)
- Microscopy / flow / cloning / mgRNA design wet-lab steps.
- Random-forest enrichment predictors (Methods) — downstream ML, attempt only if A/B/C land and time allows.
- Hi-C / insulation / iPSC lineage scripts (script_4/8/9/10/11) — auxiliary; not core claims; deprioritised.
- Public epigenetic tracks are inputs (feature annotation), not reproduced.
Plan
- (light) Result A: re-derive CT/GG/TA perfect-match site counts from hg38 → compare to 145/126/117 + xlsx. Quick minimum.
- (med) Result B: align GG dCas9 ChIP-seq + control on «our HPC», MACS2 FE≥4 → compare peak count to 5,236.
- (opt) Result C: amplicon indel rates on ct10/ct20 → compare to "8/10 sites". All heavy compute on «our HPC» SLURM; data on «infra» only.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
Result A reproduces exactly and independently: the reported 145/126/117 (388 total) on-target sites were re-derived from the raw hg38 genome, with all 388 coordinates matching the shipped supplementary file — a genuine reproduction, not a tautology. Result B (5,236 GG dCas9 binding sites, Fig. 4a) reproduces within ~1.6% (5,321 at FE>=8) using the authors' full bowtie2+macs2+custom target-assignment pipeline; the small gap is on our side (macs2 version drift, merged-vs-single-replicate choice), not the authors'. No fabrication indicators — every value is derivable from public data and shipped code. Overall a strong reproduction graded yellow only because B1 carries a small, fully explainable deviation that required reconstructing the authors' custom peak-filtering step.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.