Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Sequential structure probing of cotranscriptional RNA folding intermediates.

Nat Commun · 2025
L1 77/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Every question reproduced
-1 pts
From: “every question reproduced”
Total score -7
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
  • No relevant deviation in data/preprocessing
  • No authors-side cause for any deviation
  • Reported values are derivable from the shared data
  • Any deviation was negligible
  • The central claim held under reproduction
  • Overall, the reproduction was clean
What did not (or only partly)
  • Every checked point held up.
How its reproducibility compares
77/100
Reproducibility score
at the mean
vs. all fields · 1173 studies
🎯 Scores higher than 50% of all assessed papers rank 572 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

REPRODUCED 1:1. Ran authors' own tool TECtools v1.2.0 (cotrans_preprocessor, commit 47d9d59) on the paper's own SRA reads (PRJNA992462) on «our HPC» SLURM. The headline transcript-length / 3'-end read distribution (Fig 2b / Supp Fig 1a) matches in both replicates: +127 NPOM-dT arrest = 74.6%/75.3% (paper ~75%), +126 = 4.66%/4.63% (~5%), upstream of +128 = 98.6%/98.6% (~98%). Mapping rate 95.2%, 100% sequence match. Tool transcript-length = paper +N + 66 (leader offset), over-determined by 3 landmarks. NOT attempted (hard 20%): per-nt ShapeMapper2 reactivity, normalization, draw_intermediates structures, other RNAs in project, all wet-lab steps.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 77
    assessed: 2026-06-20 ⛓ bc8916465400
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-20
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-20
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The paper tests whether chemical probing can directly measure the cotranscriptional rearrangement of nascent RNA structures — i.e., whether an early RNA folding intermediate can be shown to rearrange into a later structure — rather than only inferring such rearrangements from separate end-point (static TEC) structure probing measurements.

Core claims
  • TECprobe-LM (linked-multipoint Transcription Elongation Complex RNA structure probing) directly measures cotranscriptional rearrangement of RNA structures by sequentially arresting RNAP at two or more template positions and chemically probing nascent RNA at each point method
  • TECprobe-LM detects rearrangement of a non-native intermediate hairpin into the native E. coli SRP RNA structure upon transcription from +127 to +161 finding
  • TECprobe-LM visualizes folding of the C. beijerinckii pfl ZTP riboswitch aptamer (pseudoknot and P3 formation) and its expression platform (terminator hairpin) finding
  • TECprobe-LM visualizes folding of the B. cereus crcB fluoride riboswitch aptamer and expression platform finding
  • Pseudoknot folding in the pfl ZTP aptamer becomes possible over a single nucleotide addition cycle, from +102 to +103 finding
  • Given the ~14 nt footprint of E. coli RNAP, the RNA exit channel can accommodate at least 3 pseudoknot base pairs of the folding ZTP aptamer mechanism
  • Washing roadblocked transcription elongation complexes to remove NTPs does not perturb nascent RNA structure finding
  • TECprobe-LM reactivity profiles agree with end-point TECprobe-VL (variable length) profiles in all but a small number of noted cases finding
Experimental setups
Assay System Perturbation Readout Platform
TECprobe-LM (SHAPE-MaP-based cotranscriptional RNA chemical probing) E. coli SRP RNA, in vitro transcription with E. coli RNAP RNAP positioned at +127 (NPOM-caged-dT stall), then UV-uncaged and chased to +161 (biotin-streptavidin roadblock) SHAPE reactivity (BzCN probing) of pre-wash, pre-chase, and post-chase RNA populations; transcript length distribution
TECprobe-LM C. beijerinckii pfl ZTP riboswitch aptamer, in vitro transcription with E. coli RNAP, ± ZMP RNAP positioned at +102, chased to +120; also 0 mM vs 1 mM ZMP ligand SHAPE reactivity of pre-wash/pre-UV, pre-chase, and post-chase samples; transcript length distribution
TECprobe-LM (three-point format) C. beijerinckii pfl ZTP riboswitch aptamer, in vitro transcription Increased NTP concentration and reduced wash volume to favor translocation to ≥+103; RNAP chased from NPOM-caged-dT stall to +103-106, then to +116-122 roadblock SHAPE reactivity of pre-UV, pre-chase, post-chase samples; transcript length distribution
TECprobe-LM C. beijerinckii pfl ZTP riboswitch expression platform (terminator hairpin) RNAP positioned at +111, chased to +143 (downstream of termination site) SHAPE reactivity reflecting terminator hairpin nucleation and ZMP-dependent aptamer folding
Denaturing PAGE SRP RNA and pfl ZTP riboswitch in vitro transcription reactions none (validation of transcript populations) Transcript length/abundance distribution
TECprobe-VL (variable-length, systematic end-point cotranscriptional structure probing) SRP RNA and pfl ZTP riboswitch none (comparator dataset from prior/parallel experiments) Reactivity profiles compared against TECprobe-LM profiles
Key results
  • Upon transcription from +127 to +161, nucleotides in the SRP RNA intermediate hairpin loop (33–41) became non-reactive while native-structure flexible nucleotides remained reactive, indicating refolding of the non-native intermediate into native SRP RNA structure
  • Transcription from +102 to +103/+104 caused decreased reactivity at nucleotides 24–26, A88, C89, and G91 in the pfl ZTP aptamer, indicating pseudoknot folding occurs within a single nucleotide addition cycle
  • Chasing RNAP to +120 caused decreased reactivity at nucleotides 84–86 (P3 folding) and, in the presence of ZMP, additional decreased reactivity in P1, at A34/A38 (L2), and at U94/G95 (L3) consistent with ligand binding
  • Washing roadblocked TECs to remove NTPs did not perturb RNA structure (pre-wash vs pre-chase reactivity profiles matched)
  • In the SRP RNA experiment, fraction of aligned reads mapping beyond +127 increased from ~2% to ~57% after chase, with ~40% mapping to the biotin-streptavidin roadblock enrichment sites (+158–162) ~2% to ~57%
  • Increasing NTP concentration and reducing wash volume caused ~86% of TECs to transcribe to +103–106 upon NPOM cage release, resolving the mixed population seen in the standard two-point format ~86%
  • TECprobe-LM reactivity profiles agreed with TECprobe-VL profiles except for two noted discrepancies attributed to RNAP backtracking during TECprobe-VL
Key statistics
  • count ~98% of aligned reads mapped upstream of +128; ~75% mapped to +127; ~5% mapped to +126 (SRP RNA pre-wash/pre-chase transcript length distribution)
  • count fraction of reads beyond +127 increased from ~2% to ~57%; ~40% at +158–162 (SRP RNA post-chase transcript distribution)
  • count ~91% of reads mapped upstream of +103; ~60% mapped to +102; ~4% mapped to +101 (pfl ZTP riboswitch pre-wash sample)
  • count fraction beyond +103 increased from ~2% to >80–90%; ~80–84% mapped to +116–122 (pfl ZTP riboswitch post-chase sample)
  • count ~86% of TECs transcribed to +103 to +106 upon NPOM cage release (three-point TECprobe-LM with increased NTP concentration)
  • count ~95% of TECs at +103 in pre-chase samples transcribed downstream upon NTP addition (assessing persistence of hyper-translocation at +103)
  • other ~14 nt (footprint of E. coli RNAP on RNA, cited from prior literature)
  • count ~6% of aligned reads mapped to +103 in pre-wash sample (indicates RNAP nucleotide insertion opposite NPOM-caged-dT in poly-dT tract)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The paper describes a method-development study (TECprobe-LM) for cotranscriptional RNA structure probing, reporting chemical probing reactivity profiles for defined transcript positions/conditions (e.g., pre-wash, pre-chase, post-chase; ±ZMP). Results are presented as reactivity profiles averaged from n = 2 replicates, with individual replicate values plotted as points alongside the mean trace, and comparisons between conditions (e.g., pre-chase vs. post-chase, TECprobe-LM vs. TECprobe-VL) are described qualitatively based on patterns of increased/decreased reactivity at specific nucleotides. The available text does not describe a formal inferential statistical test (e.g., t-test, ANOVA) applied to these comparisons.

Replicationunclear Sample sizeReactivity profile traces are stated to be 'the average of n = 2 replicates,' with individual replicate values shown as points on the plots; no further description of sample size determination or power was found in the provided text. GroupsSequential transcription-arrest samples (pre-wash/pre-UV, pre-chase, post-chase) at different RNAP positions, and ± ligand (ZMP) conditions, compared by reactivity profile. Pairingunclear Randomization/blindingnot stated Dispersionunclear Exact p-valuesno Effect sizesno Confidence intervalsno
Approaches that could also have been used
  • Reactivity profiles are summarized as the mean of n = 2 replicates, with individual replicate points shown on the plot.
    Could also: Reporting an explicit dispersion measure (e.g., SD, range, or a bootstrap-based interval) alongside the mean — With only two replicates, showing both raw points and a simple dispersion statistic can help readers gauge the consistency of the measurement in addition to seeing the individual values.
  • Differences in reactivity between sequential samples (e.g., pre-chase vs. post-chase, or +ZMP vs. −ZMP) are described qualitatively as increases or decreases at specific nucleotides.
    Could also: A formal per-nucleotide statistical comparison (e.g., paired t-test or a nonparametric equivalent across replicates) with a multiple-testing correction such as Benjamini-Hochberg FDR — Because reactivity is assessed at many nucleotide positions per profile, a formal per-position test with FDR control could complement the descriptive comparison by quantifying which reactivity changes are more or less likely to reflect a consistent shift versus replicate-to-replicate variability.
  • Each condition is based on n = 2 replicates.
    Could also: Increasing the number of replicates (e.g., n ≥ 3) in future experiments of this type — A larger replicate number would support more standard variance estimation and enable conventional inferential statistics, if hypothesis testing of specific reactivity differences were desired.
  • TECprobe-LM profiles are compared to TECprobe-VL profiles descriptively (e.g., noting where profiles 'agreed' or showed exceptions).
    Could also: A quantitative similarity metric between profiles (e.g., Pearson/Spearman correlation or root-mean-square deviation across nucleotide positions) — A quantitative concordance metric could provide a numerical summary of agreement between the two methods' reactivity profiles alongside the descriptive comparison already presented.

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-40450030

Paper: Szyjka CE, Kelly SL, Strobel EJ. Sequential structure probing of cotranscriptional RNA folding intermediates. Nat Commun 2025. DOI 10.1038/s41467-025-60425-w · PMCID PMC12126503.

Code: https://github.com/e-strobel-lab/TECtools (release tag v1.2.0) + https://github.com/e-strobel-lab/TECprobe_visualization (tag v1.0.0). Data: SRA BioProject PRJNA992462 (76 paired-end runs; ENA mirrors FASTQ).

The technique / pipeline (from Methods)

TECprobe-LM / -ML / -VL = cotranscriptional RNA structure probing. The bioinformatic pipeline, per Methods:

  1. cotrans_preprocessor MAKE_3pEND_TARGETS mode → build 3′-end + intermediate transcript target references from the template sequence.
  2. cotrans_preprocessor PROCESS_MULTI mode → adapter trimming via fastp, then demultiplex reads by 3′-end identity (transcript length) and channel (modified/untreated). Produces per-3′-end read sets + a transcript-length (3′-end position) read-count distribution.
  3. ShapeMapper2 run per intermediate transcript → alignment + per-nucleotide mutation/reactivity.
  4. process_TECprobeVL_profiles → whole-dataset normalization (min depth 50,000 reads/nt, max background mutation rate 0.05, top 10% reactivities for the normalization factor; mask + leftmost position excluded), replicate merging.
  5. mkmtrx → reactivity matrix; draw_intermediates → secondary structure via RNAstructure Fold.

In scope (pipeline-derived, reproducible)

  • [PRIMARY / 80] 3′-end (transcript-length) read distribution from cotrans_preprocessor PROCESS_MULTI demultiplexing on the paper's own raw reads. This is the cleanest, lowest-compute 1:1 target. The paper reports concrete percentages of aligned reads per 3′-end position:
    • Fig 2b (SRP RNA): ~98% of reads map upstream of NPOM-caged-dT at +128; ~75% to +127; ~5% to +126.
    • Fig 3b (pfl ZTP aptamer): ~86% of TECs arrested at NPOM site transcribed to +103–+106.
    • Fig 5c (crcB fluoride): >80% of aligned reads map to roadblock sites +67–+73. These are demultiplexing-stage outputs — reproducible without the full ShapeMapper/folding stack.
  • [SECONDARY / if time] per-nucleotide reactivity profile for one intermediate via ShapeMapper2 + process_TECprobeVL_profiles, qualitative agreement with a published trace.

Out of scope (not attempted, with reason)

  • Wet-lab steps (transcription, NPOM caging, library prep) — not computational.
  • draw_intermediates secondary-structure figures — depend on RNAstructure Fold + the full normalized matrix; deep into the "hard 20%".
  • Full 76-run reprocessing — compute-heavy and not needed to test the pinnable claims; we reproduce a representative subset (one experiment, n=2).

Reproduction plan («our HPC» / «infra»)

  • Clone TECtools v1.2.0 on «infra», build via build_TECprobe.sh (C, no heavy deps).
  • Fetch a representative SRP-RNA run pair (e.g. SRR25193903 rep1 / SRR25193898 rep2) from ENA into «infra» inside the compute job.
  • Run MAKE_3pEND_TARGETS + PROCESS_MULTI; tabulate the 3′-end read-count distribution; compute the % at +126/+127/+128 and compare to Fig 2b.
  • Grade in agreement.json; keep large intermediates on «infra», copy only the small distribution table + percentages to «host».
Figures / tables: Fig 2bFig 1a
C1
Reported
~75% of aligned SRP-RNA reads at +127 (NPOM-dT arrest, Fig 2b)
Reproduced
rep1 74.61%, rep2 75.26%
exact
C2
Reported
~5% at +126 (Fig 2b)
Reproduced
rep1 4.66%, rep2 4.63%
within tolerance
C3
Reported
~98% of aligned reads upstream of NPOM-dT at +128 (Fig 2b)
Reproduced
rep1 98.60%, rep2 98.58%
exact
C4
Reported
~40% at biotin-streptavidin roadblock +158-162 (chased sample)
Reproduced
processing samples 2/3
partial
C5
Reported
arrest/enrichment at +142 (chased sample)
Reproduced
processing samples 2/3
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 77/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟢3. Location of the main deviation
🟢4. Cause of the deviation
🟢5. Derivability / plausibility
🟢6. Severity of the deviation
🟢7. Core claim
🟢8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Every question reproduced
-1 pts
From: “every question reproduced”
Total score -7

This is a near-textbook 1:1 reproduction: the authors' own tool (TECtools v1.2.0, commit 47d9d59) was run on the authors' own deposited SRA reads (PRJNA992462), and the headline transcript-length distribution (Fig 2b) reproduces in both replicates to <1% absolute — +127 NPOM-dT arrest 74.61%/75.26% vs ~75%, +126 4.66%/4.63% vs ~5%, ≤127 upstream 98.60%/98.58% vs ~98%. Mapping QC is strong (95.18% 3'-end mapped, 100% expected-sequence match). The only gap is that the secondary chased-sample claims C4/C5 were still finalizing («job») and not graded, but these do not bear on the central conclusion. No fabrication concern — every checked value is directly derivable from the shared data.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

457.4 k
tokens (I/O) · 45.1 M incl. cache
92 min
runtime · 0.68 CPU-h
1.2 GB
peak RAM
2
HPC jobs
hummel
machine