Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Protocol for assessing regulatory elements in murine heart using an AAV9-based massively parallel reporter assay.

STAR Protoc · 2025
L1 93/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Every question reproduced
-1 pts
From: “every question reproduced”
Total score -7
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
  • No relevant deviation in data/preprocessing
  • No authors-side cause for any deviation
  • Reported values are derivable from the shared data
  • Any deviation was negligible
  • The central claim held under reproduction
  • Overall, the reproduction was clean
What did not (or only partly)
  • Every checked point held up.
How its reproducibility compares
93/100
Reproducibility score
1.1 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 85% of all assessed papers rank 154 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough -> clean 1:1 reproduction. This is a STAR Protocols paper; its only computational component is the authors' aavMPRA read-counting pipeline (github mengm5/aavMPRA @ d30aee4, MIT). The repo ships a complete worked example: test FASTQ (100 read-pairs/sample), a pre-built bowtie index, parameters, and the expected output of EVERY pipeline step under test/. I re-ran src/main.py in mutagenesis mode on «our HPC» (SLURM 2175751, ~2.5s pipeline) inside a conda env with the repo's pinned tool versions (python 3.12.2, cutadapt 4.9, bowtie 1.3.1, samtools 1.20). Comparing my 3_readCounts/* to the shipped test/3_readCounts/*: 15/15 files content-identical, and the PRIMARY output mergedSamplesCount.txt is BYTE-FOR-BYTE identical (sha256 a2ffe0fb...). The 6 non-byte-identical files differ ONLY in line order (sort/uniq intermediates with non-fixed tie-break); every value/count matches. NOT attempted: all wet-lab steps (AAV9 packaging, mouse work, tissue, library prep, sequencing - non-computational); the 'common' bowtie2 mode (no shipped expected output to grade against); and downstream enhancer-activity/hit-calling biology (not in this repo - it stops at read counts). Data note: the Zenodo record is the GitHub release zip, not separate raw sequencing data, so reproduction is bounded to the shipped test data - not a fabrication concern, but it limits scope. No fabrication signal: all expected values are fully regenerable from shipped code+data.

💻 Code ↗ 🗄 Data: 10.5281/zenodo.15099106

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 93
    assessed: 2026-06-14 ⛓ 27df6423b9e0
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-14
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The paper addresses whether an AAV9-based massively parallel reporter assay (MPRA), using a self-transcribing STARR-seq-style design, can quantify the in vivo activity of thousands of candidate regulatory elements/enhancers specifically in murine heart tissue, overcoming the lack of physiological context in cell-culture-based (lenti-)MPRA methods.

Core claims
  • An AAV9-based in vivo MPRA (AAV-MPRA) workflow, combined with a companion informatics platform, can dissect and quantify enhancer activity in mouse heart. method
  • AAV serotype-dependent tissue tropism enables the assay to restrict enhancer activity quantification to a specific tissue or cell type. mechanism
  • A self-annealing PCR strategy can double the maximum commercial oligo pool synthesis length (200-300 nt) to a 400 bp full-length enhancer library by splitting each region into two overlapping (20 bp) half-length oligos that anneal and extend. method
  • Enhancer candidates cloned into the 3' UTR of a reporter driven by a weak mini hsp68 promoter allow enhancers to drive their own transcription (STARR-seq-based design), with normalized transcript counts from amplicon sequencing reflecting enhancer activity. mechanism
  • Lentivirus-mediated MPRA (lenti-MPRA) performs well in vitro via 'in-genome' readouts, especially in hard-to-transfect cells, but cultured cells lack the physiological environment needed to reflect true in vivo enhancer activity. finding
  • The gold standard for enhancer activity measurement, enhancer-reporter knock-in mice, is too costly and time-consuming to apply to parallel screening of many enhancer candidates. finding
  • The protocol has been successfully validated for quantifying cardiac enhancers in vivo. finding
  • Addition of UMIs during library preparation removes PCR duplicates prior to sequencing-based quantification. method
Experimental setups
Assay System Perturbation Readout Platform
Self-annealing touchdown PCR and amplification (oligo pool library construction) Agilent SurePrint 5-7.5K oligo pool (cell-free DNA) none double-stranded 400 bp full-length enhancer DNA fragment yield/purity Thermal cycler; NanoDrop2000
Restriction digestion, ligation, and ligation-efficiency screening (Sanger sequencing, colony PCR) Stbl3 competent E. coli (Weidi Bio) NotI-HF/AscI digestion and T4 ligase cloning of insert into STARR vector proportion of correctly ligated colonies (target >90%), band size at 456 bp Agarose gel electrophoresis; Sanger sequencing
Plasmid library amplification via electroporation and maxi-prep Agilent SURE electrocompetent E. coli electroporation with ligated plasmid library colony count (target ~100x initial oligo number), plasmid DNA concentration/purity Bio-Rad Gene Pulser Xcell electroporator; NanoDrop
AAV9 packaging and production HEK293T cells (ATCC CRL-3216) PEI-mediated co-transfection with Rep-Cap, pHGT-1 helper plasmids and library plasmid cell morphology/transfection efficiency at 60-72 h; purified rAAV9 viral particle yield Microscope (Mshot MF52-N); iodixanol gradient ultracentrifugation (VTi70 rotor)
AAV9 injection-based in vivo MPRA delivery (preparation stage described) Wild-type neonatal (P0) C57BL/6N mice, male and female AAV9 delivery of enhancer reporter library not detailed in provided text (protocol setup only)
Key results
  • Full-length amplified oligo library yields expected DNA concentration of 130-170 ng/uL with A260/A280 = 1.8-2.0 and A260/A230 > 2.0
  • Digested insert fragment runs at approximately 456 bp on gel, distinguishable from ~5000 bp vector band 456 bp
  • Ligation efficiency screening requires the proportion of correct ligation to exceed 90% before proceeding >90%
  • Electroporated bacterial colony number must reach approximately 100x the initial oligo number to ensure library complexity 100-fold
  • Final plasmid library concentration exceeds 800 ng/uL with A260/A280 = 1.9-2.1 and A260/A230 > 2.0 >800 ng/uL
Key statistics
  • other A260/A280 = 1.8-2.0, A260/A230 > 2.0 (quality control for amplified full-length oligo library DNA)
  • other 130-170 ng/uL (expected concentration of purified full-length oligo library)
  • other 10-15 ng/uL (total yield 100-300 ng) (expected concentration of gel-purified insert fragments)
  • other 8:1 molar ratio (insert to vector ratio used in ligation reaction)
  • other >90% (required proportion of correctly ligated colonies before proceeding to plasmid library preparation)
  • count 100x initial oligo number (target total bacterial colony number after electroporation to ensure library representation)
  • other >800 ng/uL, A260/A280 = 1.9-2.1, A260/A230 > 2.0 (expected concentration/purity of final maxi-prepped plasmid library)
  • count 9-12 single colonies picked per ligation test (number of colonies sampled for Sanger sequencing to check ligation mutation rate)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This is a methods protocol paper (STAR Protocols) describing an AAV9-based massively parallel reporter assay (MPRA) for quantifying cardiac enhancer activity in vivo in mice. The analytic strategy centers on amplicon sequencing of reporter transcripts, UMI-based PCR-duplicate removal, Bowtie2 read mapping, and normalization of RNA transcript counts against DNA input counts using positive and negative control sequences as benchmarks. Formal inferential statistical tests and effect-size reporting are not described in the provided text, which is characteristic of a protocol-format publication focused on procedural reproducibility.

Replicationunclear GroupsEnhancer candidates vs. negative controls (random non-coding or validated inactive sequences) vs. positive controls (validated high-activity sequences) Pairingna Randomization/blindingnot stated Dispersionnone
Approaches that could also have been used
  • Enhancer activity is quantified by normalizing RNA transcript counts to DNA input counts, with positive and negative controls used to set thresholds and reduce batch effects
    Could also: Formal count-based differential expression frameworks such as DESeq2 or edgeR could model RNA/DNA ratios across replicate animals, providing shrinkage-estimated log2 fold-changes, standard errors, and FDR-adjusted p-values for each candidate element — Likelihood-ratio or Wald tests within a negative-binomial model account for overdispersion typical in sequencing count data and yield statistically calibrated activity calls, enabling direct comparison across studies
  • UMI sequences are added to remove PCR duplicates before counting
    Could also: Probabilistic UMI deduplication tools such as UMI-tools or fastp could be applied, with explicit reporting of the duplication rate per library as a quality metric — Reporting the fraction of reads collapsed by UMI deduplication allows readers to assess PCR amplification uniformity and the effective sequencing depth for each replicate
  • Positive and negative control sequences are included to define an activity threshold
    Could also: A z-score or robust z-score (median ± MAD) calculated from the empirical null distribution of negative controls could be used to assign a quantitative significance threshold, or a mixture model could formally separate active from inactive elements — Data-driven thresholding anchored to the negative-control distribution provides a statistically explicit and reproducible activity cutoff that is less sensitive to subjective visual inspection of control ranges
  • The protocol describes injection of a single pooled AAV library into neonatal mice without specifying the number of biological replicates or a power calculation
    Could also: Pre-specifying the number of independent animals (biological replicates) per condition based on a power calculation — informed by pilot variance estimates for RNA/DNA ratios — could be included in the protocol design step — Explicit replication and power planning allow readers to assess reproducibility and to determine whether observed activity differences are likely to be detectable at a desired effect size
  • Library coverage is verified by requiring at least 100× colony representation relative to the initial oligo number, but downstream sequencing depth targets are not formally stated
    Could also: A rarefaction analysis or explicit minimum-reads-per-element threshold (e.g., ≥100 reads per barcode in both RNA and DNA libraries) could be stated as a quality-control criterion before proceeding to activity calls — Declaring a minimum-coverage threshold makes the pipeline more reproducible across laboratories and helps identify under-represented elements that should be excluded from downstream comparisons
  • The self-annealing PCR strategy and cloning efficiency are assessed qualitatively (Sanger sequencing of 9–12 colonies, >90% correct ligation required)
    Could also: Next-generation sequencing of the input plasmid library — reporting the coefficient of variation of element representation and the fraction of elements within a defined fold-coverage window — could supplement qualitative checks — Quantitative library uniformity metrics provide an objective indicator of how evenly all candidate elements are represented before AAV packaging, which directly affects the statistical power to detect activity differences for low-representation elements
Software: Bowtie2 · ImageJ · Custom MPRA analysis pipeline (aavMPRA, GitHub / Zenodo DOI 10.5281/zenodo.15099106)

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
0
Impact: low
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-40252222

Title: Protocol for assessing regulatory elements in murine heart using an AAV9-based massively parallel reporter assay. STAR Protocols 2025. Wang K, Meng M, et al. DOI 10.1016/j.xpro.2025.103782.

Code: https://github.com/mengm5/aavMPRA (MIT, default branch main, pinned commit d30aee4ab366bf1cf2053a8bd39fea3e4181c35f, 2024-12-03). Data: zenodo 10.5281/zenodo.15099106 — this record is the GitHub release archive mengm5/aavMPRA-v1.0.0.zip (61.6 MB ≈ repo size). It is not a separate raw-sequencing dataset; it is a snapshot of the same repo. The only analysis data shipped is the small test dataset inside the repo.

Paper type

This is a STAR Protocols methods paper. The bulk of the article is a wet-lab protocol (AAV9 library design/cloning, AAV9 packaging, mouse tail-vein injection, heart harvest, DNA/RNA extraction, library prep, sequencing). The only computational component is the aavMPRA read-counting pipeline that converts paired-end FASTQ into per-oligo UMI read counts.

In scope (pipeline-derived, attempted)

The aavMPRA pipeline (src/main.py) run on the shipped test data (data/test{1,2}_{1,2}.fastq.gz, 100 read-pairs/sample) with the shipped pre-built bowtie index (index/mutagenesis_index/mutagenesis), shipped parameter.txt, mutagenesis mode (bowtie 1.3.1). The repo ships the full worked-example expected outputs at every step under test/, so the reproduction target is well-defined and 1:1 checkable:

step tool expected-output file(s) (shipped under test/)
1 correctReads cutadapt 4.9 1_correctReads/Test{1,2}_R{1,2}_rm*.fastq.gz, logs
2 mapReads bowtie 1.3.1 2_mapReads/Test{1,2}.sam
3 readCounts python+samtools 3_readCounts/Test{1,2}_readCounts.txt, …_UMIs.txt, …_uniqueUMIs.txt, …_matchedPairs.txt, mergedSamplesCount.txt

Primary claim reproduced: the final per-oligo merged read-count table test/3_readCounts/mergedSamplesCount.txt (and the per-sample readCounts). Deterministic given pinned tool versions; whole pipeline ran in ~2.4 s in the authors' worked example (test/time.txt).

Out of scope (not attempted, why)

  • All wet-lab steps — AAV9 packaging, animal work, tissue, library prep, sequencing. Not computational; cannot be reproduced in silico.
  • common mode (bowtie2) — the shipped worked example / expected outputs were generated in mutagenesis mode (the SAM in test/2_mapReads is bowtie, parameter.txt mapReads rows are bowtie -m 1 -n 2). common mode has no shipped expected output to compare against, so it is not graded.
  • Downstream biology (enhancer activity = RNA/DNA ratio, hit calling, figures in the parent research paper) — not part of this protocol repo; the repo stops at read counts.

Reproduction strategy

Run python src/main.py -o OUT -f fastq_info.txt -p data/parameter.txt -m mutagenesis -i index/mutagenesis_index/mutagenesis --gz on «our HPC» inside a conda env built from the repo's pinned aavMPRA_environment.yml, then compare OUT/3_readCounts/* byte/line-wise against the shipped test/3_readCounts/*.

C1
Reported
mergedSamplesCount.txt (44 oligo rows + header), SHA256 a2ffe0fb2eb687b3be8e59751c4679a95f9f5281f232ee1b8a2fa5b91029fdfd
Reproduced
SHA256 a2ffe0fb2eb687b3be8e59751c4679a95f9f5281f232ee1b8a2fa5b91029fdfd (byte-identical; diff empty)
exact
C2
Reported
Test2_readCounts.txt (22 rows), SHA256 1309ae13224f720acf7ea64cd6df322986b485a9574ea8083d86b73b68c87716
Reproduced
SHA256 1309ae13224f720acf7ea64cd6df322986b485a9574ea8083d86b73b68c87716 (byte-identical)
exact
C3
Reported
Test1_readCounts.txt (22 oligo->count rows)
Reproduced
Identical counts for every oligo; differs only in row order (sort|uniq -c tie-break)
within tolerance
C4
Reported
All 15 readCounts artifacts in test/3_readCounts (UMIs, uniqueUMIs, matchedPairs, readsName, readCounts, merged)
Reproduced
15/15 content-identical (sorted-line match); 9/15 byte-identical, 6/15 differ only in line order of sort/uniq intermediates; 0 missing
within tolerance

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 93/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟢3. Location of the main deviation
🟢4. Cause of the deviation
🟢5. Derivability / plausibility
🟢6. Severity of the deviation
🟢7. Core claim
🟢8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Every question reproduced
-1 pts
From: “every question reproduced”
Total score -7

This STAR Protocols paper's only computational component — the aavMPRA read-counting pipeline — reproduces 1:1 on the authors' shipped worked example: mergedSamplesCount.txt is byte-for-byte identical (sha256 a2ffe0fb…), C2 byte-identical, and all 15 readCounts files content-identical. The only deviations (C3/C4) are line ordering in sort/uniq intermediates from a non-fixed tie-break — every count matches, so this is a technical/expected non-difference on our side, not an authors' or data defect. The one honest caveat is scope: the Zenodo deposit is just the GitHub release zip, so reproduction is bounded to the shipped test data rather than the real murine-heart sequencing run — a coverage limit (q1/q2 context), not a derivability or correctness concern. No fabrication signal; overall a clean green.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

83.5 k
tokens (I/O) · 7.3 M incl. cache
10 min
runtime · 0.01 CPU-h
2 GB
peak RAM
1
HPC jobs
hummel
machine