Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

R2DT is a framework for predicting and visualising RNA secondary structure using templates.

Nat Commun · 2021
L1 87/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +6
✓ What held up
  • Reported values were directly comparable
  • No authors-side cause for any deviation
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🟡A deviation arose in the data or preprocessing
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
87/100
Reproducibility score
0.7 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 72% of all assessed papers rank 301 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

REPRODUCED (P16: ran the authors' own published tool R2DT on its own deposited example data, using the authors' own validation suite). R2DT's central computational claim - predicting & visualising RNA 2D structure from sequence via templates - reproduces cleanly on «our HPC» (apptainer, rnacentral/r2dt:latest = v2.2): the authors' unittest suite passes 123/125, with EVERY core structure-drawing test producing colored SVGs BYTE-IDENTICAL to the shipped references across all RNA classes (CRW & RiboVision rRNA, GtRNAdb tRNA, RNase P, Rfam, tmRNA, auto-detect, force-template, template-free); 10 example sequences were independently drawn to valid SVG (samples saved). The v1.1 paper-deposit's CM-library composition also reproduces (all 5 CM-database tests pass: CRW 884 / LSU 21 / SSU 9 / RNaseP 20). HONEST CAVEATS: no v1.1 container image exists anywhere (old version tags purged; v1.1 Dockerfile pulls now-moved source URLs), so we used the maintained v2.2 image and byte-identity is vs v2.2's own references; running v1.1 code under v2.2 binaries fails the DRAW tests on runtime-version grounds (ribotyper/tRNAscan/Infernal-1.1.5), not structure grounds. Table 1 exact total drifts with release version (3521 deposit / 3647 paper / 4680 current) though several per-category counts match exactly (RiboVision SSU 8, tRNA 74, RNase P 19, RiboVision LSU 21). NOT ATTEMPTED: Table 2 RNAcentral-scale application (>13M/16M sequences) - infeasible at our scale. No fabrication signals: all reported template categories are physically present and the tool reproduces its documented behaviour; count drift is transparent versioning.

💻 Code ↗ 🗄 Data: 10.5281/zenodo.4700588

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 68
    assessed: 2026-06-18 ⛓ e039fd56037f
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-24
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

Existing RNA 2D structure visualisation tools fail to produce consistent, recognisable, standardised layouts at scale, so the authors ask whether a template-based framework (R2DT) can automatically predict and visualise RNA secondary structure in community-accepted layouts across the full diversity of known structured RNAs.

Core claims
  • R2DT is a template-based computational framework/pipeline that predicts and visualises RNA 2D structure in standardised, community-accepted layouts method
  • R2DT is built on a library of 3,647 templates representing the majority of known structured RNAs resource
  • R2DT applied to RNAcentral produced >13 million diagrams, creating the world's largest RNA 2D structure dataset finding
  • Template selection uses Infernal/Ribovore covariance models (with tRNAscan-SE 2.0 for tRNAs) to classify input sequences to the best-matching template method
  • 2D structure diagrams are generated by folding sequences with Infernal cmalign against the selected covariance model and rendering with the Traveler software method
  • 3D structural data was used to revise covariation-based CRW rRNA templates, correcting base pairs and modelling species-specific expansion segments mechanism
  • Isotype- and domain-specific tRNA templates were built to capture distinct identity elements across bacterial, archaeal, and eukaryotic tRNAs resource
  • XRNA-GT, a modified version of XRNA, allows manual refinement of R2DT-generated templates for community contribution resource
Experimental setups
Assay System Perturbation Readout Platform
covariance model-based sequence classification (template selection) RNAcentral ncRNA sequences (>13 million) none top-scoring template/covariance model assignment Infernal cmsearch via ribotyper.pl (Ribovore v0.40)
tRNA-specific sequence classification query ncRNA sequences not matched by Ribovore none isotype- and domain-specific tRNA model matches tRNAscan-SE 2.0
structure-guided sequence alignment (folding) input RNA sequences none predicted 2D structure in dot-bracket notation compatible with template Infernal cmalign
2D diagram rendering aligned RNA sequence + template coordinates none SVG 2D structure diagram with nucleotide-level annotation (identical/inserted/modified/repositioned) Traveler software
template accuracy/taxonomic benchmarking RefSeq rRNA sequences <10,000 nt (23,843 sequences, July 2020) none taxonomic rank agreement between sequence and selected template; per-nucleotide positional match rate R2DT pipeline
template curation from 3D structural models 16S and 23S (SSU/LSU) rRNA structures none corrected/added base pairs, modelled expansion segments RiboVision-derived 3D structural data
Key results
  • R2DT produced more than 13 million 2D structure diagrams from RNAcentral, the largest RNA 2D structure dataset created >13 million diagrams
  • At least 94% of nucleotides in generated 2D diagrams were positioned identically to the template across all taxonomic ranks tested ≥94%
  • Selected rRNA templates matched input sequences' taxonomy predominantly at kingdom level, less often at more specific ranks 55.5% kingdom, 20.0% phylum, 16.1% class
  • Template library comprises 3,647 templates, of which 103 were manually curated specifically for this project 3647 total; 103 manual
  • 68 isotype-specific and 6 domain-specific tRNA templates were generated to capture tRNA structural diversity 68 + 6 templates
  • 3D structure-based revision of LSU/SSU rRNA templates enabled fully automatic single-page standard-orientation LSU diagrams, not previously possible
Key statistics
  • count >13 million (2D diagrams generated from RNAcentral sequences)
  • count 3647 (total templates in the R2DT template library)
  • count 103 (manually curated templates created for this project)
  • other ≥94% (nucleotides matching template position, across all taxonomic ranks)
  • other 55.5% kingdom, 20.0% phylum, 16.1% class (taxonomic rank at which RefSeq rRNA sequences matched selected templates)
  • count 23,843 (RefSeq rRNA sequences (<10,000 nt) used for template selection benchmark, July 2020)
  • count 68 (isotype-specific tRNA templates (plus 6 domain-specific templates))
  • count 56 (templates in rPredictorDB, the only comparable alternative method, as of July 2020)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

R2DT is a computational methods paper describing a template-based RNA secondary structure visualisation pipeline; its validation section relies entirely on descriptive statistics rather than inferential testing. Accuracy is assessed by comparing taxonomic concordance between 23,843 RefSeq rRNA input sequences and the templates selected by the pipeline, and by measuring per-nucleotide positional agreement between input sequences and templates. All results are reported as proportions or percentage thresholds; no formal hypothesis tests, confidence intervals, or effect-size measures are described in the text provided (the paper text is truncated before the full validation section).

Replicationunclear Sample size23,843 RefSeq rRNA sequences (shorter than 10,000 nt, as of July 2020); no formal power analysis or sample-size justification described GroupsInput rRNA sequences vs. selected templates, stratified by taxonomic rank Pairingna Randomization/blindingnot stated Dispersionnone Confidence intervalsno
Statistical tests used
Test Applied to n Assumptions
Descriptive proportion — taxonomic rank concordance Template-selection validation: proportion of 23,843 RefSeq rRNA sequences whose selected template shared each taxonomic rank (kingdom 55.5%, phylum 20.0%, class 16.1%, etc.) 23,843 rRNA sequences from RefSeq shorter than 10,000 nt, as of July 2020 not stated
Descriptive proportion — per-nucleotide positional agreement Quality of 2D diagrams: proportion of nucleotides placed in the same position as the corresponding template nucleotide, assessed per taxonomic rank not stated
Approaches that could also have been used
  • Per-nucleotide positional agreement was reported as a single lower-bound threshold ('at least 94%') with no distributional information across sequences
    Could also: The mean and standard deviation (or median and IQR) of per-sequence nucleotide agreement rates, or a cumulative distribution plot, could also characterise variability across the full benchmark set — A single threshold reveals the minimum but not the spread; distributional summaries would show whether most sequences cluster near 100% or whether a meaningful tail of sequences performs near the threshold
  • Template-selection accuracy was assessed only for rRNA, where species-level taxonomic annotation is available for comparison
    Could also: A held-out cross-validation scheme (e.g., leave-one-out or k-fold across templates) applied to additional RNA classes could also estimate generalisation to sequences not well represented in the template library — Cross-validation is a standard approach for evaluating bioinformatics classifiers and would quantify performance for divergent or novel sequences across all RNA families, not only rRNA
  • Taxonomic concordance was summarised as independent proportions at each taxonomic rank without accounting for the ordinal hierarchy of ranks
    Could also: A weighted concordance metric (e.g., a normalised taxonomic distance score or an ordinal kappa) could also summarise agreement across the full taxonomic hierarchy in a single interpretable number — Such metrics treat kingdom-level agreement and phylum-level agreement as qualitatively different outcomes, which can make overall template-selection performance easier to compare across RNA families or pipeline versions
  • The benchmark dataset comprised all available RefSeq rRNA sequences at one time point, which may over-represent well-sequenced taxa
    Could also: A stratified random sample balanced across taxonomic groups, RNA types, or sequence-length bins could also serve as a benchmark set — Stratified sampling would reduce the influence of dominant clades on aggregate statistics and make coverage gaps or biases in the template library more visible
  • Comparison with the only related tool (rPredictorDB) was limited to a qualitative count of templates (56 vs. 3,647)
    Could also: A quantitative head-to-head benchmark on the subset of sequences supported by both tools — comparing nucleotide positional agreement or an expert-rated visual quality score — could also be conducted — A common-set quantitative comparison would allow readers to assess relative layout quality more objectively than template count alone, which reflects scope rather than accuracy
  • The validation used a single snapshot of RefSeq sequences (July 2020) without reporting uncertainty around the reported proportions
    Could also: Bootstrap confidence intervals around the reported percentages (e.g., 55.5% at kingdom level) could also be computed to convey sampling uncertainty — Even for large n, reporting a CI makes it clear how stable the estimate is and facilitates comparison with future evaluations on updated datasets
Software: Infernal (cmsearch, cmalign) · Ribovore (ribotyper.pl) 0.40 · tRNAscan-SE 2.0 · Traveler · XRNA-GT

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Figures / tables: Fig 1Table
core-method-draw
Reported
R2DT predicts & visualises RNA 2D structure via templates (all RNA classes; Title, Fig 1-5)
Reproduced
Authors' own unittest suite 123/125 pass; ALL core structure-drawing byte-comparison tests produce colored SVGs byte-identical to shipped references (CRW/RiboVision LSU+SSU/RNaseP/GtRNAdb/Rfam/tmRNA/single-entry/force/template-free); 10 example seqs independently drawn to valid SVG
exact
cm-library-v1.1
Reported
v1.1 deposit CM library: CRW 884, RiboVision LSU 21, SSU 9, RNaseP 20
Reproduced
all 5 TestCovarianceModelDatabase tests pass on the zenodo v1.1 deposit
exact
T1-total
Reported
3647 templates total (Table 1)
Reproduced
3521 (v1.1 deposit models.json) / 4680 (v2.2 current)
partial
T1-ribovision-ssu
Reported
8 (Table 1)
Reproduced
8
exact
T1-trna
Reported
74 (Table 1)
Reproduced
74
exact
T1-rnasep
Reported
19 (Table 1)
Reproduced
19
exact
T1-ribovision-lsu
Reported
21 (Table 1)
Reproduced
21
exact
T1-rfam
Reported
2671 (Table 1)
Reproduced
~2659 / 3007 accessions
within tolerance
T1-crw
Reported
~654 SSU + 200 5S (Table 1)
Reproduced
884 CRW CMs (v1.1) / 661 (v2.2)
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 87/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟢2. Endpoint comparability
🟡3. Location of the main deviation
🟢4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +6

This is a self-reproduction of the authors' own tool (R2DT) by counting its deposited template library against Table 1. Several categories reproduce exactly (RiboVision SSU 8, tRNA 74, RNase P 19) and others within tolerance (Rfam 2659/2671, LSU 20/21), while CRW SSU (576/654) and 5S (171/200) and the total (3521/3647, 96.5%) fall short. The deviation sits on the input/version side — the Feb-2021 v1.1 Zenodo snapshot predates the June-2021 published counts and multiple representations exist — not on the authors' computation. Severity is moderate and direction-preserving; the central framework claim holds at the library level, but full confirmation is limited because the pipeline run was still in progress and the Table 2 large-scale validation was not attempted.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

312.1 k
tokens (I/O) · 21.3 M incl. cache
121 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.