Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Distributed biotin-streptavidin transcription roadblocks for mapping cotranscriptional RNA folding.

Nucleic Acids Res · 2017
L1 77/100 PQI 84
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +3
✓ What held up
  • Same input data as the authors
  • No relevant deviation in data/preprocessing
  • No authors-side cause for any deviation
  • Reported values are derivable from the shared data
  • Any deviation was negligible
What did not (or only partly)
  • 🟡Reported values were only indirectly comparable
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
77/100
Reproducibility score
at the mean
vs. all fields · 1173 studies
🎯 Scores higher than 50% of all assessed papers rank 572 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough -> 1:1 reproduction of the in-scope pipeline output. The paper's reported computational results are SHAPE-Seq reactivity spectra from the Spats read-mapping/reactivity pipeline (Cotrans_SHAPE-Seq_Tools builds targets; spats_shape_seq computes reactivities; Python 2.7). The spats repo (commit 9a141e8) ships the paper's own fluoride-riboswitch data as a regression fixture: test/cotrans/ds.spats = 25000 paired reads (23673 unique) of the 132 nt Fluoride_wt construct PLUS a deposited test_validation result set, with the exact run parameters embedded. On «our HPC» («job», conda py2.7 env on «infra») we reprocessed the raw reads with Spats and reproduced the deposited result set with ZERO differences across two independent algorithms (find_partial, lookup) -> exact, deterministic. NOT attempted (last-20%, documented): full SRA re-alignment of PRJNA374354 (heavy, would not change verdict), RMDB numeric matrix cross-check, the native C++ engine (sqlite3.h header missing), and the wet-lab gel/biochemical numbers (roadblocking 80-87%/30%, backtracking 3-4 nt; non-pipeline, out of scope). Honest caveat: the C1/C2 comparison target is the repo-deposited expected output regenerated from the repo-deposited reads, not a number transcribed from the paper PDF; this is the cleanest fully-pinned 1:1 available and is independently re-runnable.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 77
    assessed: 2026-06-16 ⛓ 9aedfb0fc855
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-16
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-16
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

Can a sequence-independent biotin–streptavidin transcription roadblocking strategy replace the EcoRI E111Q (Gln111) roadblock to distribute stalled transcription elongation complexes across all transcript lengths for cotranscriptional SHAPE-Seq, enabling broadly applicable nucleotide-resolution mapping of cotranscriptional RNA folding?

Core claims
  • A sequence-independent biotin–streptavidin (SAv) roadblocking strategy using randomly biotinylated DNA templates can stall TECs across all template positions for cotranscriptional SHAPE-Seq, simplifying template preparation and reducing cost. method
  • Randomly distributed biotin–SAv roadblocks identify the same RNA structural transitions in the B. cereus crcB fluoride riboswitch decision-making process previously identified using EcoRI E111Q. finding
  • EcoRI E111Q (Gln111) maps nascent RNA structure to specific transcript lengths more precisely than biotin–SAv roadblocks. finding
  • Biotin–SAv roadblocking efficiency is DNA strand dependent: template strand biotin roadblocks are far more efficient than nontemplate strand roadblocks. finding
  • Biotin–SAv roadblocking uses entirely commercially available reagents, reducing experimental costs and simplifying materials preparation, increasing accessibility of cotranscriptional SHAPE-Seq. resource
  • The complementary strengths of biotin–SAv and Gln111 roadblocks can be leveraged through proposed experimental guidelines. method
Experimental setups
Assay System Perturbation Readout Platform
In vitro radiolabeled transcription (single-round, [α-32P]-UTP) E. coli RNAP holoenzyme with J23119/SRP biotinylated DNA templates single internal biotin–SAv roadblock (template strand +33/+42 or nontemplate strand +33) roadblocking efficiency and TEC stall position via denaturing PAGE band quantification Amersham Biosciences Typhoon 9400 Variable Mode Imager; ImageQuant
Cotranscriptional SHAPE-Seq (in vitro transcription + SHAPE probing + paired-end sequencing) E. coli RNAP holoenzyme with randomly biotinylated B. cereus crcB fluoride riboswitch DNA template biotin–SAv roadblock at 1×/2×/4× biotin density; ±10 mM NaF (ligand) per-nucleotide SHAPE reactivity matrices across transcript lengths (BzCN + vs DMSO −) Illumina HiSeq2500 (2×36 bp) and NextSeq500 (2×37 bp); Spats v1.0.1
Exonuclease III footprinting E. coli holo-RNAP halted A26 TEC on λ PR (pIA226) radiolabeled linear DNA templates (DNA-RI vs DNA-B) Gln111 (on EcoRI site) vs SAv (on biotin-dT) roadblock at matched position RNAP boundary/position via ExoIII digestion on denaturing PAGE
GreB cleavage assay E. coli holo-RNAP halted A26 TEC on λ PR linear DNA templates Gln111 or SAv roadblock; ±100 nM GreB transcript cleavage products on 15% polyacrylamide–7 M urea gels
Electrophoretic mobility shift assay (EMSA) Randomly biotinylated DNA templates (unmodified reverse primer) ±SAv (DNA:SAv binding) DNA mobility shift on agarose gel ChemiDoc (Bio-Rad), GelRed/UV transillumination
Pellet/supernatant separation in vitro transcription E. coli RNAP with biotinylated templates bound to streptavidin paramagnetic particles biotin–SAv roadblock with magnetic pull-down partitioning of roadblocked vs run-off RNA between pellet and supernatant Promega streptavidin magnesphere paramagnetic particles
Key results
  • Template strand biotin–SAv roadblocks stall TECs as a cluster of stops 7–13 nt upstream of the biotinylation site 80–87% efficiency
  • Nontemplate strand biotin–SAv roadblocks stall TECs as a more defined single stop but with low efficiency 30% efficiency
  • Randomly distributed biotin–SAv roadblocks reproduce the same fluoride riboswitch structural transitions previously identified with EcoRI E111Q
  • EcoRI E111Q maps nascent RNA structure to specific transcript lengths more precisely than biotin–SAv
Key statistics
  • other 80–87% (template strand roadblocking efficiency) (Template strand biotin–SAv roadblock stall efficiency)
  • other 30% (nontemplate strand roadblocking efficiency) (Nontemplate strand biotin–SAv roadblock stall efficiency)
  • count 7–13 nucleotides upstream of biotinylation site (Cluster of TEC stop positions for template strand roadblock)
  • count 4–6 M reads (HiSeq2500); 9–16 M reads (NextSeq500) (Reads mapped per sample to calculate reactivity matrices)
  • other BzCN t1/2 of 250 ms (SHAPE reagent benzoyl cyanide reaction half-life)
  • count 30 s (TEC stall time before SHAPE modification)

Statistical methods review

Model: opus

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This is a molecular methods paper that develops and benchmarks a biotin–streptavidin transcription roadblocking strategy for cotranscriptional SHAPE-Seq. Findings are reported primarily as quantitative measurements from gel/band quantification (e.g. roadblocking efficiency as a percentage of roadblocked versus total products) and as per-nucleotide SHAPE reactivity matrices computed from paired-end sequencing read counts using the Spats pipeline. The work relies on replication and direct quantitative comparison rather than on formal hypothesis-testing statistics; no inferential significance tests are explicitly described in the methods/results provided.

Replicationmixed Sample size1× biotin fluoride riboswitch was performed as replicate 1 (HiSeq2500) plus replicates 2–4 done in a different laboratory with different reagents (NextSeq500); 2× and 4× biotin libraries also sequenced; no power/sample-size calculation stated Groupsbiotin–SAv vs. EcoRI E111Q (Gln111) roadblocking; template vs. nontemplate strand biotinylation; +/- SHAPE reagent (BzCN vs DMSO); +/- NaF riboswitch ligand Pairingna Randomization/blindingnot stated Dispersionnone Effect sizesno Confidence intervalsno Multiplicity correctionnone stated
Approaches that could also have been used
  • Roadblocking efficiencies and other quantities are reported as single percentage values (e.g. 80–87%, 30%).
    Could also: Reporting each value as a mean across replicates accompanied by a dispersion measure such as SD, an IQR, or a 95% confidence interval. — Adding an explicit measure of spread alongside the central value conveys measurement variability and is often preferred, particularly with small numbers of replicates.
  • Replicate experiments (e.g. 1× biotin replicates 1–4, including runs in a separate laboratory) are used to demonstrate reproducibility, compared descriptively.
    Could also: Quantifying agreement between replicates with a correlation coefficient (e.g. Pearson/Spearman) on per-nucleotide reactivities, or a concordance metric across transcript lengths. — A numerical reproducibility statistic would summarize replicate agreement in a single value and make cross-condition or cross-laboratory consistency directly comparable.
  • Biotin–SAv and Gln111 roadblocking strategies (and template vs. nontemplate strand) are compared by direct inspection of measured values and reactivity profiles.
    Could also: A formal group comparison such as a t-test or Mann–Whitney U on the replicate-level efficiency measurements between conditions. — An inferential test would attach a quantitative measure of confidence to the observed differences between roadblocking approaches, complementing the descriptive comparison.
  • Per-nucleotide SHAPE reactivities are presented as reactivity matrices to identify structural transitions.
    Could also: Pairing reactivity values with replicate-derived uncertainty estimates (e.g. standard error per nucleotide) or a thresholding/statistical model for calling significant reactivity changes between conditions. — Propagating per-nucleotide uncertainty would help distinguish reactivity changes that exceed measurement noise from those within it when interpreting structural transitions.
  • Differences across +/- ligand (NaF) conditions are interpreted from the reactivity profiles.
    Could also: Applying a multiplicity-aware procedure (e.g. Benjamini–Hochberg FDR) when many nucleotide positions are compared between conditions. — When many positions are evaluated simultaneously, a false-discovery-rate control would manage the family of comparisons while flagging condition-dependent positions.
Software: Spats (cotranscriptional SHAPE-Seq read mapping/reactivity calculation) v1.0.1 · Cotrans_SHAPE-Seq_Tools (Cotrans_targets.py; LucksLab GitHub) · ImageQuant (gel band quantification)

Result convergence & founder nodes

Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
47
Impact: medium
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

PRJNA374354 BioProject in Data Availability (http://purl.obolibrary.org/obo/IAO_0000611)
no other assessed paper uses this yet

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-28398514

Paper: Strobel EJ, Watters KE, Nedialkov Y, Artsimovitch I, Lucks JB. Distributed biotin-streptavidin transcription roadblocks for mapping cotranscriptional RNA folding. Nucleic Acids Res 2017. PMID 28398514 / PMC5499547 / doi:10.1093/nar/gkx233.

Code: https://github.com/LucksLab/Cotrans_SHAPE-Seq_Tools (+ the analysis engine https://github.com/LucksLab/spats , pip spats_shape_seq, Python 2.7). Data: SRA BioProject PRJNA374354 (raw reads); reactivity spectra in RMDB (Supp Table S4).

Pipeline map (Methods + repo)

The reported computational results are SHAPE-Seq reactivity spectra for cotranscriptional RNA folding intermediates. Per Methods: "Reads were mapped and processed for Spats v1.0.1 as described previously"; Cotrans_targets.py (from Cotrans_SHAPE-Seq_Tools) builds the per-length target FASTA, then Spats aligns the paired-end reads, assigns each read to the (+) RRRY / (−) YYYR channel and a transcript length × stop site, and computes nucleotide-resolution reactivities (rho/theta) per transcript length.

In scope (pipeline-derived → attempted)

  • read-processing → reactivity result set for the fluoride riboswitch (B. cereus crcB, the paper's primary cotranscriptional model RNA). The Spats repository ships, as a regression fixture, the paper's own data: test/cotrans/ds.spats = a SQLite PairDB of 25,000 paired-end reads (23,673 unique) for the 132-nt Fluoride_wt construct, plus a deposited test_validation reactivity/assignment result set (25,000 result rows). The exact run parameters are embedded (run_data: cotrans=True, masks RRRY/YYYR, linker CTGACTCGGGCACCAAGGAC, adapters, cotrans_minimum_length=20, algorithm find_partial). Reproduction = reprocess the raw reads with Spats and reproduce the deposited result set, across the independent find_partial, lookup, and native algorithms.

Out of scope / not attempted (the hard last ~20%, with reason)

  • Full SRA re-alignment of all constructs (PRJNA374354, 4–16 M reads/sample across many BioSamples): the read-processing→reactivity step is already exactly exercised by the bundled 25k-read fixture; downloading/aligning the entire BioProject would not change the 1:1 verdict and is heavy compute. Skipped.
  • RMDB numeric cross-check of published reactivity values: the repo's bundled test_validation set is the deposited expected output for this pipeline; matching it is the cleaner, fully-pinned comparison. RMDB matrix-by-matrix comparison left out (last-20%).
  • Wet-lab quantities — roadblocking efficiency (80–87% template / 30% nontemplate), backtracking 3–4 nt protection: gel/biochemical measurements, NOT pipeline-derived → out of scope.
  • MATLAB plotting (Cotrans_matrix_rhos_processing_*.m): GUI/figure cosmetics, not a quantitative result.
C1
Reported
test_validation result set (25000 rows) = deterministic Spats cotrans output for the fluoride riboswitch
Reproduced
0 differing rows / 25000 (find_partial)
exact
C2
Reported
same deposited result set, independent algorithm
Reproduced
0 differing rows / 25000 (lookup)
exact
C3
Reported
132 nt B. cereus crcB fluoride riboswitch, two-channel (RRRY+/YYYR-) cotranscriptional SHAPE-Seq, 25000 read pairs
Reproduced
132 nt; 25000 pairs/23673 unique; RRRY+=365 YYYR-=134 assigned; lengths 23-130 (98 distinct in fixture)
within tolerance
X1
Reported
roadblocking efficiency 80-87% template / 30% nontemplate
Reproduced
not attempted (wet-lab, non-pipeline)
partial
X2
Reported
biotin-SAv backtracking 3-4 nt extra protection
Reproduced
not attempted (wet-lab, non-pipeline)
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 77/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟡2. Endpoint comparability
🟢3. Location of the main deviation
🟢4. Cause of the deviation
🟢5. Derivability / plausibility
🟢6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +3

The in-scope computational claim is reproduced exactly and deterministically — 0 differing rows / 25000 across two independent Spats engines on the authors' own deposited fluoride-riboswitch fixture, every input and expected output checksummed and re-runnable. The honest limitation, on our side (not the authors'), is that the comparison target is the repository's deposited test_validation set regenerated from the repository's deposited reads, not a value transcribed from the paper PDF or the RMDB reactivity matrices, so the paper's headline figure values were never numerically cross-checked. The paper's central wet-lab/methodological numbers (roadblocking 80-87%/30%, backtracking 3-4 nt) are non-pipeline and were out of scope. Verdict: a clean, high-quality pipeline reproduction with explainable, well-documented coverage limits — yellow overall, no authors'-side defect or fabrication signal.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

101.6 k
tokens (I/O) · 6.2 M incl. cache
14 min
runtime · 0.01 CPU-h
1.9 GB
peak RAM
1
HPC jobs
hummel
machine