Rfam 15: RNA families database in 2025.
The main results reproduced, with only marginal, non-material deviations.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- Nothing in this column.
- 🟡Could not use the authors’ exact input data
- 🟡Reported values were only indirectly comparable
- 🟡A deviation arose in the data or preprocessing
- 🟡A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The deviation was non-trivial in magnitude
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Described well enough to reproduce the in-artifact, pipeline-derived parts 1:1. Paper = Rfam 15 NAR database update; the only self-contained runnable computational component is the rfam-3d-seed-alignments pipeline (P16: the authors' own published tool, commit a22cc82). REPRODUCED: (C4) the documented per-PDB tool fr3d_2d.py 2QUS_B against the README's expected output -> sequence byte-identical, secondary structure 97.1% identical with a single extra base pair (G1-C64) that the live RNA 3D Hub API now annotates as cWW but didn't when the example was authored -> deterministic logic reproduces exactly, delta is upstream-data drift not a code bug. (C1) The '143 families detected' headline reproduces as 136-140 from the pipeline's committed input mapping (gap = moving .preview snapshot). NOT REPRODUCIBLE FROM THE ARTIFACT: (C2/C3) the release-curation numbers 65 families / 298 chains / 42 PKs are Rfam production-DB incorporation counts (which family SEEDs were revised in release 15 since 14.5), a different and smaller population than the repo's cumulative output (129 families / 962 chains / 28 PK-families); reported transparently, graded partial, explicitly NOT flagged as fabrication. NOT ATTEMPTED (skipped 20%): a full Infernal cmalign re-run of a family, because add_3d.py re-pulls the latest SEED+mapping at run time, producing a fresh computation against moved inputs rather than a reproduction of any stated value; the deterministic fr3d_2d.py example already reproduces the core 3D-annotation step. No «our HPC»/SLURM used: the reproducible results are one RNA-3D-Hub REST call plus deterministic counts over small shipped text files; the genuinely heavy step has low reproduction value due to moving upstream inputs.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 68assessed: 2026-06-15 ⛓ 0574355e9d6f
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-15
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15no human curator yet
- Last updated
- 2026-09-19
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetThe paper describes whether systematically updating Rfam's underlying genome database, incorporating experimental 3D structure data, applying R-scape covariation analysis, and synchronizing with miRBase and ontology resources can improve the accuracy, completeness, and functional annotation of non-coding RNA family models in Rfam release 15.0.
- ★ Rfamseq was expanded to 26,106 genomes, a 76% increase over Rfam 14.0, incorporating UniProt 2024_03 reference proteomes and additional viral genomes. resource
- ★ An automated pipeline maps weekly-updated experimentally determined RNA 3D structures from PDB to Rfam SEED alignments using Infernal cmscan/cmalign and FR3D basepair annotation, enabling curator review. method
- ★ 65 Rfam families were revised using 3D structure information, improving consensus secondary structures including addition of pseudoknots and other base pairings. finding
- ★ R-scape CaCoFold covariation analysis was used to systematically review Rfam families and select 26 families for consensus secondary structure improvements. finding
- ★ GO and SO annotations were comprehensively reviewed and updated, increasing GO term coverage from 60% to 75% of families. finding
- ★ 14 new Hepatitis C Virus RNA families were added and 4 outdated families removed using a whole-genome viral alignment workflow. finding
- ★ MicroRNA families in Rfam were synchronized with miRBase, resulting in 1603 microRNA families built from miRBase-tracked sequences. finding
- Rfam is used as a training dataset for machine learning models such as AlphaFold 3 and for genomic ncRNA annotation in resources like Ensembl. resource
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Covariance model genome search (Infernal cmscan) | Rfamseq database (26,106 genomes, cellular organisms and viruses) | none | FULL hits per Rfam family / gathering threshold matches | Infernal |
| 3D structure-to-alignment mapping (cmalign) and basepair annotation | RNA chains from Protein Data Bank (PDB) | none | Consensus secondary structure updates, basepair annotations, pseudoknots | Infernal cmalign; FR3D |
| Covariation analysis (R-scape / CaCoFold) | Rfam SEED multiple sequence alignments | none | Statistically significant covarying basepairs, alternative consensus structures | R-scape |
| Ontology curation review | Rfam family annotations (GO and SO terms) | none | Number and specificity of GO/SO terms per family | — |
| Whole genome viral alignment workflow | Hepatitis C Virus, Flaviviridae, Coronaviridae genomes | none | Identification of new viral ncRNA structural families | — |
| microRNA sequence extraction and CM building | miRBase microRNA sequences mapped to Rfam accessions | none | New/replaced SEED alignments, gathering thresholds, family counts | Infernal; RNAcentral identifiers |
- ▲ Total Rfamseq sequence length increased from 3.68e11 to 1.46e12 nucleotides and number of species from 14,774 to 26,106. 76% increase in species
- ▲ Number of Rfam hits increased from 2,897,296 to 10,736,534. ~3.7-fold
- – Of 3302 families with hits in both releases, average family size grew by 166%, with 2335 gaining hits, 547 losing hits, and 26 families showing >10x growth (mostly microRNA families matching plant genomes). 166% average growth
- ▲ 65 families updated using 298 PDB structure chains; pseudoknots added in 42 of 65 families (six with two PKs, one with three). 42/65 families
- ▲ 26 families improved with R-scape CaCoFold, e.g. HEARO family (RF02033) gained 24 additional covarying basepairs. up to 24 additional basepairs
- ▲ GO term coverage rose from 2084/3431 (60%) in release 14.3 to 3157 families (75%) in 15.0; total GO terms increased from 3752 to 4446. 60% to 75%
- ▲ MicroRNA families synchronized with miRBase, reaching 1603 microRNA families, up from a baseline where only 28% of 1983 miRBase families matched the prior 529 Rfam microRNA families. 1603 families
- count 26,106 genomes (Rfamseq v15) (Rfamseq species count vs 14,774 in v14)
- fold_change 76% increase in genomes (Rfamseq v14 to v15 growth)
- count 10,736,534 Rfam hits (v15) vs 2,897,296 (v14) (Total FULL hits comparison)
- mean 166% average family growth (Families with hits in both 14.3 and 15.0)
- count 65 families updated, 298 structure chains (3D structure-based family improvements since release 14.5)
- count 26 families improved via R-scape (Families selected for CaCoFold-based updates)
- other GO coverage 75% (3157 families) vs 60% in 14.3 (GO term annotation coverage across common families)
- count 1603 microRNA families; 14 new HCV families added, 4 removed (microRNA and viral family counts in release 15.0)
Statistical methods review
Model: opusA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This is a database resource update paper describing changes between Rfam releases 14.x and 15.0; it relies on descriptive counting and comparison rather than inferential hypothesis testing. Reported quantities are tallies and proportions (e.g. numbers of genomes, families, FULL hits, GO/SO terms, percentage growth) and a domain-specific covariation analysis (R-scape CaCoFold) is used to identify base pairs with statistically significant covariation when refining consensus secondary structures. No experimental groups, sample-size justification, or classical significance tests are reported.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| R-scape covariation analysis (CaCoFold) identifying base pairs with statistically significant covariation | refinement of consensus secondary structures; 26 families updated (Table 2) and per-family base-pair support (e.g. FMN riboswitch RF00050, IS605-orfB-I RF03065 in Figure 2) | — | not stated |
-
Release-to-release changes are summarized with point counts and percentage changes (e.g. average family grew by 166%, 75% of families have a GO term).↳ Could also: Accompanying these point estimates with distributional summaries (median and IQR or a 95% confidence/credible interval for proportions) would also be possible. — Distributional summaries would additionally convey the spread and uncertainty around aggregate figures, which can be informative when averages are influenced by a few large families (e.g. the >10× growth microRNA families noted in the text).
-
The average family growth is reported as a mean percentage across families.↳ Could also: A median (or log-scale geometric mean) could also be reported alongside the mean. — Because the per-family growth distribution appears right-skewed (some families grew >10×), a median or geometric mean would additionally describe the typical family and complement the mean.
-
Per-family covariation support is described via counts of additional covarying base pairs identified by R-scape (Table 2).↳ Could also: The associated R-scape covariation E-values/statistics and any thresholds could also be tabulated. — Reporting the underlying significance statistics and cutoffs would add transparency about the strength of evidence behind each added base pair and aid reproducibility.
-
Comparisons between releases are presented descriptively without an explicit multiple-comparison framework across the thousands of families.↳ Could also: When framing any family-level change as 'significant', an explicit multiplicity strategy (e.g. Benjamini–Hochberg FDR) could also be applied. — A stated correction scheme would help control the expected number of false discoveries when many families are screened simultaneously, should the analysis move from descriptive to inferential.
Result convergence & founder nodes
Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.
-
14 new HCV RNA families added and 4 outdated families removed in Rfam 15, expanding HCV genomic RNA structural coverageother hepatitis-c-virus up 2025×1papers★ This paper is the founder (earliest)
-
Rfam families annotated with at least one GO term increased from 60% (v14.3) to 75% (v15)other rfam-families up 2025×1papers★ This paper is the founder (earliest)
-
Total GO terms assigned across Rfam families increased from 3752 (v14.3) to 4446 (v15)other rfam-families up 2025×1papers★ This paper is the founder (earliest)
-
26 Rfam families improved by R-scape CaCoFold covariation analysis; HEARO (RF02033) gained 24 additional covarying basepairsother rfam-seed-alignments up 2025×1papers★ This paper is the founder (earliest)
-
65 Rfam families updated using 298 PDB RNA chains; pseudoknots annotated in 42 of the 65 updated familiesother rfam-seed-alignments up 2025×1papers★ This paper is the founder (earliest)
-
Average Rfam family hit count grew 166% in Rfam v15 vs v14, with 26 families showing >10x growth (primarily plant miRNA families)other rfamseq up 2025×1papers★ This paper is the founder (earliest)
-
Rfam family annotation hits increased ~3.7-fold (2.9M to 10.7M) across the Rfamseq genome collection in Rfam v15 vs v14other rfamseq up 2025×1papers★ This paper is the founder (earliest)
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-39526405 (Rfam 15: RNA families database in 2025)
Paper: Ontiveros-Palacios et al., Nucleic Acids Res 2025. PMID 39526405 / PMC11701678 / DOI 10.1093/nar/gkae1023.
Reproduction-relevant code: https://github.com/Rfam/rfam-3d-seed-alignments @ commit a22cc827b5b99a3e97935384cc0812b73bc49668
("A pipeline for adding RNA 3D structures to Rfam seed alignments", Apache-2.0, last push 2025-02-28).
The Zenodo DOI in the brief (10.5281/zenodo.13919037) is the general rfam-family-pipeline 15.0 code snapshot, not the 3D data; the 3D pipeline + its shipped outputs live in the GitHub repo above.
Nature of the paper
This is a database update paper (NAR Database Issue). Most of its content (release statistics, new families, R-scape curation, web interface, taxonomy) is database-management / wet-lab-curation narrative — out of scope for computational pipeline reproduction. The one self-contained, runnable, pipeline-derived component with a public code artifact is the 3D-seed-alignment pipeline in section "Using RNA 3D structures to revise Rfam secondary structures."
Pipeline (in scope)
rfam-3d-seed-alignments: for each Rfam family that matches ≥1 PDB chain (mapping pdb_full_region.txt),
download the FR3D secondary-structure annotation from the RNA 3D Hub REST API (fr3d_2d.py), then
iteratively add the PDB sequence + 2D structure to the Rfam SEED via Infernal cmbuild --hand /
cmalign --mapstr --mapali, and finally relabel PDB ids with RNAcentral ids. Output: one
data/output/RFxxxxx.sto per family.
In-scope reproduction targets (what we attempt)
| id | reported in paper | how reproducible from the artifact |
|---|---|---|
| C1 | "pipeline detects 143 families which match ≥1 3D structure" | count distinct families in committed pdb_full_region.txt (the pipeline's own input mapping) |
| C4 | fr3d_2d.py 2QUS_B documented example output (README) |
run the tool, byte-compare to the documented FASTA+dot-bracket |
| (char) | shape of data/output/ (the shipped pipeline products) |
count families / added chains / pseudoknot families |
Out of scope / not attempted (and why)
- C2 "65 families updated, 298 chains" / C3 "PKs in 42 of 65" — these are Rfam-release-curation
numbers (which families' SEEDs were incorporated/revised in release 15 since 14.5), derived from
Rfam's internal SVN/production database, not regenerable from this repo. The repo ships the
cumulative pipeline output (129 families, all PDB-matching families over time), a different
population. Reported here for transparency, graded
partial/out-of-scope, not mismatch — no evidence of fabrication, just a different counted set. - Full Infernal re-run of a family (
add_3d.py RFxxxxx --nocache) — the explicitly-skipped last ~20%. It re-pulls the latest SEED fromsvn.rfam.organd the latest.previewPDB mapping at run time, so its output is a fresh computation against moved inputs, not a reproduction of any stated value and not byte-comparable to the Feb-2025 committed.sto. Low marginal evidence value; skipped per the 80/20 rule. The deterministic, documentedfr3d_2d.pyexample (C4) reproduces the core 3D-annotation step of exactly this pipeline.
Compute note
No heavy SLURM job was needed: the cleanly-reproducible pipeline-derived results are a single RNA-3D-Hub
REST call (fr3d_2d.py, deterministic, tiny) plus deterministic counts over small shipped text files.
The genuinely heavy part (full cmalign re-run) has low reproduction value (moving upstream inputs),
so «our HPC» was not used. All work is small derived results, kept in reproduction/outputs/.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
The reproducible, in-artifact components of this Rfam 15 database-update paper reproduce well: C4 (fr3d_2d.py 2QUS_B) is sequence byte-identical (69/69) with 97.1% SS agreement (single G1-C64 pair now cWW upstream), and C1 reproduces as 136-140 vs the reported 143 (~95-98%) off a moving .preview snapshot. The deviations sit on the input/data-availability side, not in core computation — upstream API/mapping drift plus a population mismatch. C2/C3 (65 families/298 chains, 42 PKs) are release-curation numbers held in Rfam's production DB and are not derivable from the shipped repo (whose cumulative output is 129 families/962 chains/28 PK-families), so they remain unchecked but show no fabrication signal. Overall: a solid reproduction with fully explainable deviations; severity is low for the comparable claims and the central methodology holds, but the headline curation counts could not be confirmed.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at [email protected].
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.