Rfam 15: RNA families database in 2025.
The main results reproduced, with only marginal, non-material deviations.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- Nothing in this column.
- 🟡Could not use the authors’ exact input data
- 🟡Reported values were only indirectly comparable
- 🟡A deviation arose in the data or preprocessing
- 🟡A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The deviation was non-trivial in magnitude
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Described well enough to reproduce the in-artifact, pipeline-derived parts 1:1. Paper = Rfam 15 NAR database update; the only self-contained runnable computational component is the rfam-3d-seed-alignments pipeline (P16: the authors' own published tool, commit a22cc82). REPRODUCED: (C4) the documented per-PDB tool fr3d_2d.py 2QUS_B against the README's expected output -> sequence byte-identical, secondary structure 97.1% identical with a single extra base pair (G1-C64) that the live RNA 3D Hub API now annotates as cWW but didn't when the example was authored -> deterministic logic reproduces exactly, delta is upstream-data drift not a code bug. (C1) The '143 families detected' headline reproduces as 136-140 from the pipeline's committed input mapping (gap = moving .preview snapshot). NOT REPRODUCIBLE FROM THE ARTIFACT: (C2/C3) the release-curation numbers 65 families / 298 chains / 42 PKs are Rfam production-DB incorporation counts (which family SEEDs were revised in release 15 since 14.5), a different and smaller population than the repo's cumulative output (129 families / 962 chains / 28 PK-families); reported transparently, graded partial, explicitly NOT flagged as fabrication. NOT ATTEMPTED (skipped 20%): a full Infernal cmalign re-run of a family, because add_3d.py re-pulls the latest SEED+mapping at run time, producing a fresh computation against moved inputs rather than a reproduction of any stated value; the deterministic fr3d_2d.py example already reproduces the core 3D-annotation step. No «our HPC»/SLURM used: the reproducible results are one RNA-3D-Hub REST call plus deterministic counts over small shipped text files; the genuinely heavy step has low reproduction value due to moving upstream inputs.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 68assessed: 2026-06-15 ⛓ 0574355e9d6f
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-15
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusThis paper describes major updates to the Rfam non-coding RNA families database in release 15.0, aiming to expand genomic coverage, improve structural and functional annotation quality, and synchronize microRNA families, thereby reinforcing Rfam's role in RNA research, genome annotation and machine learning model development.
- ★ Rfamseq was expanded to 26 106 genomes, a 76% increase, by incorporating the latest UniProt reference proteomes and additional viral genomes resource
- ★ Sixty-five RNA families were improved using experimentally determined 3D structures, refining consensus secondary structures and annotations method
- ★ R-scape CaCoFold covariation analysis was used to refine structural predictions in 26 families method
- ★ GO term coverage was increased to 75% of families through comprehensive GO and SO annotation updates resource
- ★ 14 new Hepatitis C Virus RNA families were added and 4 outdated families removed via a viral whole-genome alignment workflow resource
- ★ MicroRNA family synchronization with miRBase was completed, resulting in 1603 microRNA families resource
- An automated weekly pipeline maps PDB RNA 3D structures and FR3D basepairs to Rfam SEED alignments to aid curation method
- Rfam is the largest source of GO annotations, with >10 million sequences having an Rfam-based GO annotation, and is used to train ML models including AlphaFold 3 resource
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Covariance model homology search / genome annotation (Infernal cmscan/cmsearch) | Rfamseq genome collection (26 106 genomes) | none | FULL hits per family above gathering threshold | Infernal |
| RNA 3D structure to secondary structure mapping pipeline (cmscan/cmalign) | Rfam SEED alignments vs PDB RNA chains | none | mapped sequences, annotated basepairs/pseudoknots, updated consensus secondary structure | Infernal, FR3D, PDB |
| Covariation analysis (R-scape CaCoFold) | Rfam family multiple sequence alignments | none | number of covarying basepairs / alternative consensus secondary structures | R-scape |
| GO and Sequence Ontology annotation curation | Rfam families (3431 common families) | none | number/coverage of GO and SO terms per family | — |
| Viral RNA family construction from whole genome alignments | Hepatitis C Virus genomic alignments | none | new RNA family structures in coding/non-coding regions | — |
| MicroRNA family synchronization (CM build and search) | miRBase microRNA alignments mapped to Rfamseq | none | synchronized SEED alignments, gathering thresholds, microRNA family count | Infernal, miRBase, RNAcentral |
- ▲ Rfamseq grew from 14 774 species (v14) to 26 106 species (v15), with total sequence length rising from 3.68e11 to 1.46e12 76% increase in genomes
- ▲ Number of Rfam hits increased from 2 897 296 (v14) to 10 736 534 (v15) ~3.7-fold
- ▲ Of 3302 families with hits in both releases, the average family grew by 166%, with 26 families showing >10x growth (primarily plant microRNA families) 166% average growth
- ▲ 65 families were updated using 298 chains from 3D structures, with pseudoknots added in 42 of the 65 families 65 families, 298 chains
- ▲ 26 families were selected and improved using R-scape CaCoFold structures, e.g. HEARO (RF02033) gained 24 additional covarying basepairs up to 24 additional covarying basepairs
- ▲ Families with at least one GO term increased coverage from 60% (2084 families in 14.3) to 75% (3157 families in 15.0) 60% to 75%
- ▲ Total number of GO terms increased from 3752 to 4446 3752 to 4446
- ▲ 14 new HCV families added and 4 outdated families removed +14 / -4 families
- count 26 106 (number of species/genomes in Rfamseq v15)
- count 10 736 534 (number of Rfam hits in v15 (vs 2 897 296 in v14))
- fold_change 166% (average family growth in hits for families present in both 14.3 and 15.0)
- count 65 (families updated using 3D structures (with 298 chains))
- count 26 (families improved using R-scape CaCoFold)
- count 75% (fraction of families with at least one GO term in 15.0 (up from 60%))
- count 1603 (total microRNA families after miRBase synchronization)
- count 2985 (additional viral genomes fetched from PIR at 75% protein sequence identity)
Statistical methods review
Model: opusA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This is a database resource update paper describing changes between Rfam releases 14.x and 15.0; it relies on descriptive counting and comparison rather than inferential hypothesis testing. Reported quantities are tallies and proportions (e.g. numbers of genomes, families, FULL hits, GO/SO terms, percentage growth) and a domain-specific covariation analysis (R-scape CaCoFold) is used to identify base pairs with statistically significant covariation when refining consensus secondary structures. No experimental groups, sample-size justification, or classical significance tests are reported.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| R-scape covariation analysis (CaCoFold) identifying base pairs with statistically significant covariation | refinement of consensus secondary structures; 26 families updated (Table 2) and per-family base-pair support (e.g. FMN riboswitch RF00050, IS605-orfB-I RF03065 in Figure 2) | — | not stated |
-
Release-to-release changes are summarized with point counts and percentage changes (e.g. average family grew by 166%, 75% of families have a GO term).↳ Could also: Accompanying these point estimates with distributional summaries (median and IQR or a 95% confidence/credible interval for proportions) would also be possible. — Distributional summaries would additionally convey the spread and uncertainty around aggregate figures, which can be informative when averages are influenced by a few large families (e.g. the >10× growth microRNA families noted in the text).
-
The average family growth is reported as a mean percentage across families.↳ Could also: A median (or log-scale geometric mean) could also be reported alongside the mean. — Because the per-family growth distribution appears right-skewed (some families grew >10×), a median or geometric mean would additionally describe the typical family and complement the mean.
-
Per-family covariation support is described via counts of additional covarying base pairs identified by R-scape (Table 2).↳ Could also: The associated R-scape covariation E-values/statistics and any thresholds could also be tabulated. — Reporting the underlying significance statistics and cutoffs would add transparency about the strength of evidence behind each added base pair and aid reproducibility.
-
Comparisons between releases are presented descriptively without an explicit multiple-comparison framework across the thousands of families.↳ Could also: When framing any family-level change as 'significant', an explicit multiplicity strategy (e.g. Benjamini–Hochberg FDR) could also be applied. — A stated correction scheme would help control the expected number of false discoveries when many families are screened simultaneously, should the analysis move from descriptive to inferential.
Result convergence & founder nodes
Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.
-
14 new HCV RNA families added and 4 outdated families removed in Rfam 15, expanding HCV genomic RNA structural coverageother hepatitis-c-virus up 2025×1papers★ This paper is the founder (earliest)
-
Rfam families annotated with at least one GO term increased from 60% (v14.3) to 75% (v15)other rfam-families up 2025×1papers★ This paper is the founder (earliest)
-
Total GO terms assigned across Rfam families increased from 3752 (v14.3) to 4446 (v15)other rfam-families up 2025×1papers★ This paper is the founder (earliest)
-
26 Rfam families improved by R-scape CaCoFold covariation analysis; HEARO (RF02033) gained 24 additional covarying basepairsother rfam-seed-alignments up 2025×1papers★ This paper is the founder (earliest)
-
65 Rfam families updated using 298 PDB RNA chains; pseudoknots annotated in 42 of the 65 updated familiesother rfam-seed-alignments up 2025×1papers★ This paper is the founder (earliest)
-
Average Rfam family hit count grew 166% in Rfam v15 vs v14, with 26 families showing >10x growth (primarily plant miRNA families)other rfamseq up 2025×1papers★ This paper is the founder (earliest)
-
Rfam family annotation hits increased ~3.7-fold (2.9M to 10.7M) across the Rfamseq genome collection in Rfam v15 vs v14other rfamseq up 2025×1papers★ This paper is the founder (earliest)
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-39526405 (Rfam 15: RNA families database in 2025)
Paper: Ontiveros-Palacios et al., Nucleic Acids Res 2025. PMID 39526405 / PMC11701678 / DOI 10.1093/nar/gkae1023.
Reproduction-relevant code: https://github.com/Rfam/rfam-3d-seed-alignments @ commit a22cc827b5b99a3e97935384cc0812b73bc49668
("A pipeline for adding RNA 3D structures to Rfam seed alignments", Apache-2.0, last push 2025-02-28).
The Zenodo DOI in the brief (10.5281/zenodo.13919037) is the general rfam-family-pipeline 15.0 code snapshot, not the 3D data; the 3D pipeline + its shipped outputs live in the GitHub repo above.
Nature of the paper
This is a database update paper (NAR Database Issue). Most of its content (release statistics, new families, R-scape curation, web interface, taxonomy) is database-management / wet-lab-curation narrative — out of scope for computational pipeline reproduction. The one self-contained, runnable, pipeline-derived component with a public code artifact is the 3D-seed-alignment pipeline in section "Using RNA 3D structures to revise Rfam secondary structures."
Pipeline (in scope)
rfam-3d-seed-alignments: for each Rfam family that matches ≥1 PDB chain (mapping pdb_full_region.txt),
download the FR3D secondary-structure annotation from the RNA 3D Hub REST API (fr3d_2d.py), then
iteratively add the PDB sequence + 2D structure to the Rfam SEED via Infernal cmbuild --hand /
cmalign --mapstr --mapali, and finally relabel PDB ids with RNAcentral ids. Output: one
data/output/RFxxxxx.sto per family.
In-scope reproduction targets (what we attempt)
| id | reported in paper | how reproducible from the artifact |
|---|---|---|
| C1 | "pipeline detects 143 families which match ≥1 3D structure" | count distinct families in committed pdb_full_region.txt (the pipeline's own input mapping) |
| C4 | fr3d_2d.py 2QUS_B documented example output (README) |
run the tool, byte-compare to the documented FASTA+dot-bracket |
| (char) | shape of data/output/ (the shipped pipeline products) |
count families / added chains / pseudoknot families |
Out of scope / not attempted (and why)
- C2 "65 families updated, 298 chains" / C3 "PKs in 42 of 65" — these are Rfam-release-curation
numbers (which families' SEEDs were incorporated/revised in release 15 since 14.5), derived from
Rfam's internal SVN/production database, not regenerable from this repo. The repo ships the
cumulative pipeline output (129 families, all PDB-matching families over time), a different
population. Reported here for transparency, graded
partial/out-of-scope, not mismatch — no evidence of fabrication, just a different counted set. - Full Infernal re-run of a family (
add_3d.py RFxxxxx --nocache) — the explicitly-skipped last ~20%. It re-pulls the latest SEED fromsvn.rfam.organd the latest.previewPDB mapping at run time, so its output is a fresh computation against moved inputs, not a reproduction of any stated value and not byte-comparable to the Feb-2025 committed.sto. Low marginal evidence value; skipped per the 80/20 rule. The deterministic, documentedfr3d_2d.pyexample (C4) reproduces the core 3D-annotation step of exactly this pipeline.
Compute note
No heavy SLURM job was needed: the cleanly-reproducible pipeline-derived results are a single RNA-3D-Hub
REST call (fr3d_2d.py, deterministic, tiny) plus deterministic counts over small shipped text files.
The genuinely heavy part (full cmalign re-run) has low reproduction value (moving upstream inputs),
so «our HPC» was not used. All work is small derived results, kept in reproduction/outputs/.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
The reproducible, in-artifact components of this Rfam 15 database-update paper reproduce well: C4 (fr3d_2d.py 2QUS_B) is sequence byte-identical (69/69) with 97.1% SS agreement (single G1-C64 pair now cWW upstream), and C1 reproduces as 136-140 vs the reported 143 (~95-98%) off a moving .preview snapshot. The deviations sit on the input/data-availability side, not in core computation — upstream API/mapping drift plus a population mismatch. C2/C3 (65 families/298 chains, 42 PKs) are release-curation numbers held in Rfam's production DB and are not derivable from the shipped repo (whose cumulative output is 129 families/962 chains/28 PK-families), so they remain unchecked but show no fabrication signal. Overall: a solid reproduction with fully explainable deviations; severity is low for the comparable claims and the central methodology holds, but the headline curation counts could not be confirmed.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.