Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Conservation and losses of non-coding RNAs in avian genomes.

PLoS One · 2015
L1 100/100 PQI 99
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Every question reproduced
-1 pts
From: “every question reproduced”
Total score -7
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
  • No relevant deviation in data/preprocessing
  • No authors-side cause for any deviation
  • Reported values are derivable from the shared data
  • Any deviation was negligible
  • The central claim held under reproduction
  • Overall, the reproduction was clean
What did not (or only partly)
  • Every checked point held up.
How its reproducibility compares
100/100
Reproducibility score
1.5 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 95% of all assessed papers rank 1 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough to reproduce 1:1 EXACT. Target: Table 1 (the paper's central ncRNA-count summary across human, chicken, and 48 birds). Ran the authors' own scripts/makeTable.pl at pinned commit 9b3ebd8 on the repo's shipped intermediate annotation+expression data on «our HPC» («job», ~13 s, pure-Perl, deterministic). All 68 numeric cells reproduced exactly, including the Total row (7340 / 1080.0 / 1194 / 865 = 72.4%) and the RNA-seq false-positive-rate note (865/123/1194, FPR 10.3%). No fabrication concern: every published value is deterministically derivable from the deposited data via the deposited script. NOT attempted (the hard ~20%, by design): re-running the upstream Infernal/Rfam-11.0 + tRNAscan-SE + miRBase scan of 48 genome assemblies and the compete_clans merge — those heavy steps' outputs are deposited in the repo as data/{rfam,trnascan,mirbase,merged-annotations}/*.gff, so the downstream summary is reproducible without re-deriving them; and the heatmaps.R figures (visual, not numeric). This is a P16-style reproduction using authors' own code + shipped intermediates with an exact pinned commit.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 100
    assessed: 2026-06-16 ⛓ 35a684bb4bc5
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-16
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-16
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

Can homology-based (covariance model) bioinformatic methods, anchored on curated RNA families, accurately annotate conserved non-coding RNA loci across 48 avian genomes and distinguish genuine ncRNA gene losses from sequence divergence and assembly-related missing data?

Core claims
  • 34 lncRNA-associated loci are conserved between birds and mammals, and 12 of these were validated in chicken by RNA-seq. finding
  • Several human-characterized lncRNAs (e.g., HOXA11-AS1, HOTAIRM1, HOTTIP, PART1, PCA3, RMST, SOX2OT, ST7-OT3, NBR2, DLEU2) are syntenically conserved in birds despite unknown/non-conserved function. finding
  • Covariance-model homology search (Rfam CMs via INFERNAL, tRNAscan-SE, snoStrip, miRBase models) is a state-of-the-art approach for annotating conserved ncRNAs in novel vertebrate genomes. method
  • Apparent ncRNA 'losses' in birds fall into three categories: genuine gene loss, sequence/structural divergence beyond CM detection, and data missing due to microchromosome assembly difficulties. finding
  • Genuine losses in the avian lineage include the mir-106b/mir-93/mir-25 cluster (cluster II) and specific let-7 clusters (cluster A and cluster F). finding
  • The Y5 RNA paralog family is absent from all bird genomes but present in alligator and turtle, with a conserved Y4-Y3-Y1 cluster retained. finding
  • 66,879 ncRNA-similar loci conserved in >10% of avian genomes were classified into 626 families, mostly miRNAs and snoRNAs. resource
  • SNORD93 is unusually expanded with 92 copies in the tinamou genome versus 1–2 copies in all other vertebrate genomes. finding
Experimental setups
Assay System Perturbation Readout Platform
Covariance-model homology search (ncRNA annotation) 48 avian genomes none presence/absence and copy-number of conserved ncRNA families INFERNAL 1.1 cmsearch with Rfam v11.0 CMs
tRNA annotation 48 avian genomes none tRNA gene predictions, isoacceptor type, functional vs pseudogene classification tRNAscan-SE v1.3.1
miRNA homology search 48 bird genomes plus American alligator and green turtle out-groups none conserved miRNA hits (seed sequence + hairpin) INFERNAL v1.1rc3 CMs built from miRBase v19 (999 families)
snoRNA homology search avian genomes none snoRNA family annotations snoStrip (queries from human, platypus, chicken)
Small RNA-seq validation chicken (14 tissues, 27 samples) none expression evidence for predicted ncRNAs Illumina HiSeq2000 (Bioproject PRJNA204941); mapping with SEGEMEHL 0.1.9 to galGal4
Strand-specific RNA-seq validation whole chicken embryo (7 stages) none expression evidence for predicted ncRNAs Illumina HiSeq, dUTP strand-specific protocol (SRA SRP041863)
Key results
  • 34 chicken lncRNA loci conserved with mammals; 12 validated by RNA-seq 12/34 (35.3%)
  • Total ncRNA loci identified across 48 avian genomes conserved in >10% of genomes 66,879 loci in 626 families
  • SNORD93 expansion in tinamou genome relative to other vertebrates 92 copies vs 1–2 copies
  • tRNA copy-number reduction in birds versus human/turtle/alligator ~900 to ~280 copies (tRNA); ~580 to ~100 (pseudogenes)
  • Chicken transfer RNAs show high RNA-seq expression confirmation 278/300 (92.7%)
  • mir-17/mir-92 cluster II (mir-106b/mir-93/mir-25) absent in turtles, crocodiles and birds
  • Y5 RNA paralog absent from all bird genomes but present in alligator and turtle
  • Total chicken ncRNAs annotated and fraction confirmed by RNA-seq 865/1194 (72.4%)
Key statistics
  • count 66,879 loci (ncRNA-similar loci conserved in >10% of avian genomes)
  • count 626 families (families classifying the identified avian ncRNA loci)
  • count 92 copies (SNORD93 copies in tinamou genome (vs 1–2 in other vertebrates))
  • count 12 (35.3%) (chicken lncRNAs confirmed with RNA-seq out of 34)
  • count 1194 total chicken ncRNAs; 865 (72.4%) confirmed (total chicken ncRNA genes and RNA-seq confirmation)
  • other E-value < 5 x 10^-4 (E-value cutoff for retaining Rfam cmsearch hits above GA threshold)
  • count 971 million reads / 27 samples / 14 tissues (small RNA-seq validation dataset PRJNA204941)
  • count 1.46 billion reads / 7 stages (strand-specific embryo RNA-seq dataset SRP041863)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This is a large-scale comparative bioinformatics study annotating ncRNA loci across 48 avian genomes using covariance model (CM)-based homology search (INFERNAL 1.1 with Rfam v11.0 CMs, tRNAscan-SE, and miRBase-derived CMs). Results are reported descriptively as counts, medians, presence/absence heatmaps, and validation percentages; no formal inferential statistical tests (e.g., t-tests, regression) are applied. Significance is operationalized algorithmically via E-value thresholds and Rfam gathering (GA) thresholds rather than through classical hypothesis testing. RNA-seq read mapping in chicken is used to validate predicted loci, with expression thresholds set at a declared false-positive rate of less than 10%.

Replicationunclear Sample size48 avian genomes enumerated by name; no formal power analysis described; two outgroup genomes (American alligator, green turtle) used for comparison GroupsncRNA family presence/copy-number in birds vs. mammals/other vertebrates; chicken ncRNA predictions vs. RNA-seq expression evidence Pairingna Randomization/blindingnot stated Dispersionnone Effect sizesno Confidence intervalsno Multiplicity correctionnone stated; algorithm-level E-value and GA thresholds, plus a 10% cross-genome conservation filter, serve as implicit specificity controls
Statistical tests used
Test Applied to n Assumptions
E-value threshold (≤ 5×10⁻⁴) applied to cmsearch hits above Rfam GA threshold All ncRNA family annotations across 48 avian genomes 48 avian genomes; 66,879 loci retained not stated
10% prevalence conservation filter (family retained only if found in ≥10% of avian genomes) Post-annotation filtering of Rfam, miRNA, and snoRNA hits 48 avian genomes not stated
False-positive rate threshold (<10%) for RNA-seq expression confirmation Chicken ncRNA expression validation using two RNA-seq datasets 971 million reads (small RNA-seq, 27 samples, 14 tissues); 1.46 billion reads (strand-specific, 7 embryonic stages) not stated
Clan competition (best-hit retention for overlapping Rfam family hits within a clan) Resolution of overlapping annotations across all avian genomes na
Approaches that could also have been used
  • Conservation of ncRNA families was assessed by a fixed 10% prevalence threshold across 48 genomes
    Could also: Phylogenetic ancestral-state reconstruction methods (e.g., Dollo parsimony, maximum-likelihood models of gene gain/loss such as those in CAFE or BayesTraits) could also have been applied to the presence/absence matrix — Model-based approaches would also provide probabilistic estimates of gain and loss rates along specific lineages, allowing one to distinguish stochastic absence from lineage-specific loss with an associated confidence measure, rather than relying on a single prevalence cutoff
  • Copy-number summary across 48 birds was reported as the median (Table 1) with no measure of spread
    Could also: The interquartile range (IQR) or a bootstrapped 95% confidence interval around the median could also have been reported alongside the median — Adding a dispersion statistic would allow readers to assess how variable copy-numbers are across the avian clade, distinguishing families with uniformly low counts from those with a few extreme outliers like SNORD93 in tinamou
  • E-value and GA thresholds from INFERNAL/Rfam were used as the sole criteria for calling a hit 'significant'
    Could also: A precision-recall analysis or receiver-operating-characteristic (ROC) evaluation against a curated gold-standard set could also have been used to empirically benchmark threshold choice for this particular dataset — An empirical threshold evaluation would also quantify the sensitivity/specificity trade-off for these specific avian genomes, since model-specific GA thresholds were derived from diverse training data that may differ from avian sequence composition
  • RNA-seq expression validation used a single FPR cutoff (<10%) without a formal statistical model for read counts
    Could also: Negative-binomial count models (e.g., DESeq2 or edgeR) applied to per-locus read counts across the 27 RNA-seq samples could also have provided per-locus statistical evidence for expression above background — Model-based differential expression testing would also yield adjusted p-values and fold-changes over background regions, enabling a more granular ranking of validated loci and an explicit correction for the large number of loci tested simultaneously
  • Syntenic conservation of lncRNA domains was described qualitatively by visual inspection of genome browsers and gene-order diagrams
    Could also: A synteny-block scoring method (e.g., MCScanX or i-ADHoRe) could also have been applied to quantify the probability of observed gene-order conservation under a null model of random chromosomal rearrangement — A quantitative synteny score would also provide a statistical basis for distinguishing coincidental co-localization from genuinely conserved chromosomal neighborhoods, particularly useful when the number of co-occurring RNA domains is small
  • Losses and absences were partitioned into three qualitative categories (genuine loss, divergence, missing data) based on expert interpretation of search results
    Could also: A genome-completeness correction approach (e.g., using BUSCO scores or regression of detection rate on assembly N50/coverage) could also have been applied to estimate the expected detection rate given assembly quality, separating assembly-driven absences from biological losses more formally — Correcting for assembly completeness would also allow quantitative partitioning of absences into assembly-attributable versus biology-attributable fractions, reducing reliance on qualitative expert judgment for the 'missing data' category
Software: INFERNAL (cmsearch) 1.1 / 1.1rc3 · Rfam database v11.0 · tRNAscan-SE 1.3.1 · miRBase v19 · snoStrip · SEGEMEHL 0.1.9 · RALEE

Result convergence & founder nodes

Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
20
Impact: medium
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

PRJNA204941 BioProject in Methods (http://purl.org/orb/Methods)
also used by 1 paper:
GCA_000247815.2 GCA in Introduction (http://purl.org/orb/Introduction)
no other assessed paper uses this yet
RF01976 Rfam in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
SRP041863 ENA in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-25822729

Paper: Gardner PP et al. (2015) "Conservation and losses of non-coding RNAs in avian genomes." PLoS One 10(3):e0121797. PMCID PMC4378963.

Code: https://github.com/ppgardne/bird-genomes (authors' own; commit 9b3ebd87f678c3f908aab17758c0e568e3a94c1d, 2014-08-07; default branch master). Data: SRA PRJNA204941 (RNA-seq) + the Avian Phylogenomics genome assemblies.

The pipeline (what produced the paper's numbers)

  1. Annotation of 48 bird genomes (+ human + 3 reptile outgroups) for ncRNAs:
    • Rfam 11.0 covariance models scanned with Infernal (cmsearch, E ≤ 0.0005) → data/rfam/*.gff
    • tRNAscan-SE → data/trnascan/*.gff
    • miRBase homology → data/mirbase/*.gff
    • Stadler-group tools → data/stadler-annotations/*.gff
  2. Merge / overlap resolution across methods+clans (scripts/compete_clans.pl using data/clan_info.txt) → data/merged-annotations/*.gff, and the

    10%-conserved subset → data/conserved-merged-annotations/*.gff.

  3. Count-matrix build (scripts/gffs2heatmaps.pl) → data/R/*.dat (e.g. allRNA.dat: per-family copy number in every species).
  4. Summary table (scripts/makeTable.pl) → Table 1 of the manuscript: per-RNA-type ncRNA counts in human / median-of-48-birds / chicken, plus the number of chicken ncRNAs confirmed expressed by RNA-seq (max RNA_i > 13.0, thresholds {5,12} on McCarthy+Ulitsky tissue data, with shipped randomized negative controls giving the false-positive rate).
  5. Figures (scripts/heatmaps.R) → heatmaps/plots from data/R/*.dat.

In scope (attempted — clearly specified, low-compute, deterministic)

  • Table 1 in full — regenerate via scripts/makeTable.pl from the shipped data/R/allRNA.dat, data/rfam2type.txt, data/RNA-seq/*.dat, and data/conserved-merged-annotations/Gallus_gallus.gff. Pure-Perl, no external modules, fully deterministic (negative controls are precomputed and shipped). This is step 4 above and reproduces the paper's central quantitative summary.

Out of scope / not attempted (the hard ~20%, by design — 80/20 rule)

  • Steps 1–2 (Infernal/tRNAscan/miRBase scan + merge of 48 genomes). The heavy upstream compute. Its outputs are shipped in the repo (data/rfam/*.gff, data/trnascan/*.gff, …), so the downstream summary is reproducible without re-running it. Re-running cmsearch over 48 genome assemblies with the exact Rfam 11.0 / Infernal version is the expensive last 20%; not attempted.
  • Figures (heatmaps.R) — visual, not numeric claims; skipped (Table 1 is the clearer, machine-checkable target).
  • Wet-lab / external — none material to the pipeline numbers.

Reproduction strategy: authors' own code + authors' shipped intermediate data, exact pinned commit, run on «our HPC». Equivalent to verifying that the published Table 1 is faithfully and deterministically derivable from the deposited data.

Figures / tables: Table
C1
Reported
7340
Reproduced
7340
exact
C2
Reported
1080.0
Reproduced
1080.0
exact
C3
Reported
1194
Reproduced
1194
exact
C4
Reported
865 (72.4%)
Reproduced
865 (72.4%)
exact
C5
Reported
microRNA 356/499.5/427/280 (65.6%)
Reproduced
356/499.5/427/280 (65.6%)
exact
C6
Reported
Transfer RNA 1084/173.5/300/278 (92.7%)
Reproduced
1084/173.5/300/278 (92.7%)
exact
C7
Reported
C/D box snoRNA 281/120.0/106/90 (84.9%)
Reproduced
281/120.0/106/90 (84.9%)
exact
C8
Reported
Major spliceosomal 1754/48.5/71/32 (45.1%)
Reproduced
1754/48.5/71/32 (45.1%)
exact
C9
Reported
RNA-seq FPR 10.3% (123/1194)
Reproduced
FPR[10.3] negCtrl[123] genes[1194]
exact
C10
Reported
Full Table 1, 68 numeric cells + Total + FPR
Reproduced
all 68 cells + Total + FPR identical
exact

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 100/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟢3. Location of the main deviation
🟢4. Cause of the deviation
🟢5. Derivability / plausibility
🟢6. Severity of the deviation
🟢7. Core claim
🟢8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Every question reproduced
-1 pts
From: “every question reproduced”
Total score -7

Exact, clean reproduction. All 68 numeric cells of Table 1 plus the Total row (7340/1080.0/1194/865 = 72.4%) and the FPR note (123/1194 = 10.3%) match the published values bit-for-bit, produced by the authors' own makeTable.pl at pinned commit 9b3ebd8 on the repo's deposited data. Every published value is deterministically derivable from the deposited inputs — no fabrication concern. The only nuance, not a defect, is that this verifies the downstream summary from shipped intermediates (the upstream genome scan was not re-run), so it is a code-reuse/downstream reproduction rather than a full independent pipeline rebuild.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

87.1 k
tokens (I/O) · 6.9 M incl. cache
10 min
runtime · 0 CPU-h
0.1 GB
peak RAM
1
HPC jobs
hummel
machine