R2DT is a framework for predicting and visualising RNA secondary structure using templates.
The main results reproduced: recomputed values matched the published ones within tolerance.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Reported values were directly comparable
- ✓No authors-side cause for any deviation
- 🟡Could not use the authors’ exact input data
- 🟡A deviation arose in the data or preprocessing
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The deviation was non-trivial in magnitude
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
REPRODUCED (P16: ran the authors' own published tool R2DT on its own deposited example data, using the authors' own validation suite). R2DT's central computational claim - predicting & visualising RNA 2D structure from sequence via templates - reproduces cleanly on «our HPC» (apptainer, rnacentral/r2dt:latest = v2.2): the authors' unittest suite passes 123/125, with EVERY core structure-drawing test producing colored SVGs BYTE-IDENTICAL to the shipped references across all RNA classes (CRW & RiboVision rRNA, GtRNAdb tRNA, RNase P, Rfam, tmRNA, auto-detect, force-template, template-free); 10 example sequences were independently drawn to valid SVG (samples saved). The v1.1 paper-deposit's CM-library composition also reproduces (all 5 CM-database tests pass: CRW 884 / LSU 21 / SSU 9 / RNaseP 20). HONEST CAVEATS: no v1.1 container image exists anywhere (old version tags purged; v1.1 Dockerfile pulls now-moved source URLs), so we used the maintained v2.2 image and byte-identity is vs v2.2's own references; running v1.1 code under v2.2 binaries fails the DRAW tests on runtime-version grounds (ribotyper/tRNAscan/Infernal-1.1.5), not structure grounds. Table 1 exact total drifts with release version (3521 deposit / 3647 paper / 4680 current) though several per-category counts match exactly (RiboVision SSU 8, tRNA 74, RNase P 19, RiboVision LSU 21). NOT ATTEMPTED: Table 2 RNAcentral-scale application (>13M/16M sequences) - infeasible at our scale. No fabrication signals: all reported template categories are physically present and the tool reproduces its documented behaviour; count drift is transparent versioning.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 68assessed: 2026-06-18 ⛓ e039fd56037f
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-24
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetExisting RNA 2D structure visualisation tools fail to produce consistent, recognisable, standardised layouts at scale, so the authors ask whether a template-based framework (R2DT) can automatically predict and visualise RNA secondary structure in community-accepted layouts across the full diversity of known structured RNAs.
- ★ R2DT is a template-based computational framework/pipeline that predicts and visualises RNA 2D structure in standardised, community-accepted layouts method
- ★ R2DT is built on a library of 3,647 templates representing the majority of known structured RNAs resource
- ★ R2DT applied to RNAcentral produced >13 million diagrams, creating the world's largest RNA 2D structure dataset finding
- ★ Template selection uses Infernal/Ribovore covariance models (with tRNAscan-SE 2.0 for tRNAs) to classify input sequences to the best-matching template method
- ★ 2D structure diagrams are generated by folding sequences with Infernal cmalign against the selected covariance model and rendering with the Traveler software method
- ★ 3D structural data was used to revise covariation-based CRW rRNA templates, correcting base pairs and modelling species-specific expansion segments mechanism
- Isotype- and domain-specific tRNA templates were built to capture distinct identity elements across bacterial, archaeal, and eukaryotic tRNAs resource
- XRNA-GT, a modified version of XRNA, allows manual refinement of R2DT-generated templates for community contribution resource
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| covariance model-based sequence classification (template selection) | RNAcentral ncRNA sequences (>13 million) | none | top-scoring template/covariance model assignment | Infernal cmsearch via ribotyper.pl (Ribovore v0.40) |
| tRNA-specific sequence classification | query ncRNA sequences not matched by Ribovore | none | isotype- and domain-specific tRNA model matches | tRNAscan-SE 2.0 |
| structure-guided sequence alignment (folding) | input RNA sequences | none | predicted 2D structure in dot-bracket notation compatible with template | Infernal cmalign |
| 2D diagram rendering | aligned RNA sequence + template coordinates | none | SVG 2D structure diagram with nucleotide-level annotation (identical/inserted/modified/repositioned) | Traveler software |
| template accuracy/taxonomic benchmarking | RefSeq rRNA sequences <10,000 nt (23,843 sequences, July 2020) | none | taxonomic rank agreement between sequence and selected template; per-nucleotide positional match rate | R2DT pipeline |
| template curation from 3D structural models | 16S and 23S (SSU/LSU) rRNA structures | none | corrected/added base pairs, modelled expansion segments | RiboVision-derived 3D structural data |
- – R2DT produced more than 13 million 2D structure diagrams from RNAcentral, the largest RNA 2D structure dataset created >13 million diagrams
- – At least 94% of nucleotides in generated 2D diagrams were positioned identically to the template across all taxonomic ranks tested ≥94%
- – Selected rRNA templates matched input sequences' taxonomy predominantly at kingdom level, less often at more specific ranks 55.5% kingdom, 20.0% phylum, 16.1% class
- – Template library comprises 3,647 templates, of which 103 were manually curated specifically for this project 3647 total; 103 manual
- – 68 isotype-specific and 6 domain-specific tRNA templates were generated to capture tRNA structural diversity 68 + 6 templates
- – 3D structure-based revision of LSU/SSU rRNA templates enabled fully automatic single-page standard-orientation LSU diagrams, not previously possible
- count >13 million (2D diagrams generated from RNAcentral sequences)
- count 3647 (total templates in the R2DT template library)
- count 103 (manually curated templates created for this project)
- other ≥94% (nucleotides matching template position, across all taxonomic ranks)
- other 55.5% kingdom, 20.0% phylum, 16.1% class (taxonomic rank at which RefSeq rRNA sequences matched selected templates)
- count 23,843 (RefSeq rRNA sequences (<10,000 nt) used for template selection benchmark, July 2020)
- count 68 (isotype-specific tRNA templates (plus 6 domain-specific templates))
- count 56 (templates in rPredictorDB, the only comparable alternative method, as of July 2020)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
R2DT is a computational methods paper describing a template-based RNA secondary structure visualisation pipeline; its validation section relies entirely on descriptive statistics rather than inferential testing. Accuracy is assessed by comparing taxonomic concordance between 23,843 RefSeq rRNA input sequences and the templates selected by the pipeline, and by measuring per-nucleotide positional agreement between input sequences and templates. All results are reported as proportions or percentage thresholds; no formal hypothesis tests, confidence intervals, or effect-size measures are described in the text provided (the paper text is truncated before the full validation section).
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Descriptive proportion — taxonomic rank concordance | Template-selection validation: proportion of 23,843 RefSeq rRNA sequences whose selected template shared each taxonomic rank (kingdom 55.5%, phylum 20.0%, class 16.1%, etc.) | 23,843 rRNA sequences from RefSeq shorter than 10,000 nt, as of July 2020 | not stated |
| Descriptive proportion — per-nucleotide positional agreement | Quality of 2D diagrams: proportion of nucleotides placed in the same position as the corresponding template nucleotide, assessed per taxonomic rank | — | not stated |
-
Per-nucleotide positional agreement was reported as a single lower-bound threshold ('at least 94%') with no distributional information across sequences↳ Could also: The mean and standard deviation (or median and IQR) of per-sequence nucleotide agreement rates, or a cumulative distribution plot, could also characterise variability across the full benchmark set — A single threshold reveals the minimum but not the spread; distributional summaries would show whether most sequences cluster near 100% or whether a meaningful tail of sequences performs near the threshold
-
Template-selection accuracy was assessed only for rRNA, where species-level taxonomic annotation is available for comparison↳ Could also: A held-out cross-validation scheme (e.g., leave-one-out or k-fold across templates) applied to additional RNA classes could also estimate generalisation to sequences not well represented in the template library — Cross-validation is a standard approach for evaluating bioinformatics classifiers and would quantify performance for divergent or novel sequences across all RNA families, not only rRNA
-
Taxonomic concordance was summarised as independent proportions at each taxonomic rank without accounting for the ordinal hierarchy of ranks↳ Could also: A weighted concordance metric (e.g., a normalised taxonomic distance score or an ordinal kappa) could also summarise agreement across the full taxonomic hierarchy in a single interpretable number — Such metrics treat kingdom-level agreement and phylum-level agreement as qualitatively different outcomes, which can make overall template-selection performance easier to compare across RNA families or pipeline versions
-
The benchmark dataset comprised all available RefSeq rRNA sequences at one time point, which may over-represent well-sequenced taxa↳ Could also: A stratified random sample balanced across taxonomic groups, RNA types, or sequence-length bins could also serve as a benchmark set — Stratified sampling would reduce the influence of dominant clades on aggregate statistics and make coverage gaps or biases in the template library more visible
-
Comparison with the only related tool (rPredictorDB) was limited to a qualitative count of templates (56 vs. 3,647)↳ Could also: A quantitative head-to-head benchmark on the subset of sequences supported by both tools — comparing nucleotide positional agreement or an expert-rated visual quality score — could also be conducted — A common-set quantitative comparison would allow readers to assess relative layout quality more objectively than template count alone, which reflects scope rather than accuracy
-
The validation used a single snapshot of RefSeq sequences (July 2020) without reporting uncertainty around the reported proportions↳ Could also: Bootstrap confidence intervals around the reported percentages (e.g., 55.5% at kingdom level) could also be computed to convey sampling uncertainty — Even for large n, reporting a CI makes it clear how stable the estimate is and facilitates comparison with future evaluations on updated datasets
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
This is a self-reproduction of the authors' own tool (R2DT) by counting its deposited template library against Table 1. Several categories reproduce exactly (RiboVision SSU 8, tRNA 74, RNase P 19) and others within tolerance (Rfam 2659/2671, LSU 20/21), while CRW SSU (576/654) and 5S (171/200) and the total (3521/3647, 96.5%) fall short. The deviation sits on the input/version side — the Feb-2021 v1.1 Zenodo snapshot predates the June-2021 published counts and multiple representations exist — not on the authors' computation. Severity is moderate and direction-preserving; the central framework claim holds at the library level, but full confirmation is limited because the pipeline run was still in progress and the Table 2 large-scale validation was not attempted.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.