Identification and Characterization of Small Noncoding RNAs in Genome Sequences of the Edible Fungus Pleurotus ostreatus.
The main results reproduced: recomputed values matched the published ones within tolerance.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Same input data as the authors
- ✓Reported values were directly comparable
- ✓No relevant deviation in data/preprocessing
- ✓No authors-side cause for any deviation
- ✓Reported values are derivable from the shared data
- ✓Any deviation was negligible
- ✓The central claim held under reproduction
- ✓Overall, the reproduction was clean
- Every checked point held up.
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Described well enough -> 1:1 reproduction of the deterministic, pipeline-derived outputs. This is a genome-assembly + sncRNA-annotation paper (WGS PRJNA327267, not sRNA-seq). Reproduced against the deposited 2016 assembly GCA_001956935.1 (ASM195693v1, strain CCMSSC00389, md5 e70c2fe1f45c7f91413fc95dba2ab88f). Four descriptive assembly stats recomputed from the FASTA: scaffolds 2529 (EXACT), N50 394787 bp (EXACT), GC 49.54% (EXACT over total length), size 34.86 Mb called bases ~= reported 34.9 Mb (total span 35.82 Mb incl gap-Ns). tRNA annotation re-run with the paper's exact tool tRNAscan-SE v1.3.1 (default Eukaryotic mode): 184 standard-AA tRNAs vs reported 185 (off by 1, 0.5%; 194 incl 10 pseudogenes), and length range 71-144 nt reproduces EXACTLY. No fabrication signal: every value derivable from shipped data, reproduced exactly or off-by-<=1. NOT attempted (harder 20%): miRNA=46 (BLASTn Rfam RF00003), snoRNA=7 / snRNA=4 (Infernal 1.0.3) - all depend on an unspecified Rfam release whose decade of content drift makes a 1:1 count unreliable; and the SeqPrep/Sickle read-cleaning step (no pinnable reported count). All heavy compute on «our HPC» SLURM (node n094); data kept on «infra», only small results on «host».
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 95assessed: 2026-06-16 ⛓ 6da8b3cdd32d
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-16
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-16no human curator yet
- Last updated
- 2026-09-19
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetThe study asks whether small noncoding RNAs (snRNAs, snoRNAs, tRNAs, miRNAs) can be systematically identified genome-wide in the basidiomycete edible fungus Pleurotus ostreatus, and how conserved these sncRNAs are within the species and across Agaricomycotina fungi.
- ★ Genome-scale identification detected 254 small noncoding RNAs (snRNAs, snoRNAs, tRNAs, miRNAs, and other Rfam-classified sncRNAs) in the P. ostreatus CCEF00389 genome assembly finding
- ★ snRNA U1 could not be identified in the CCEF00389 genome or in other basidiomycetous genomes, implying that if it exists it has a sequence highly divergent from other organisms finding
- sncRNAs were detected using Rfam/Infernal and BLASTn for snRNAs/snoRNAs, tRNAscan-SE for tRNAs, and BLASTn against Rfam miRNA family RF00003 for miRNAs method
- ★ snRNAs and most tRNAs are located in pseudo-UTR regions near gene boundaries, while miRNAs are commonly found within introns finding
- ★ Most sncRNAs (77.56%) are highly conserved within P. ostreatus (relative to strain PC15), but only a minority are conserved across other Agaricomycotina fungi and very few in the basal basidiomycete Ustilago maydis finding
- A 34.9-Mb draft genome assembly of P. ostreatus CCEF00389 with 13,438 predicted gene models was produced, the first released draft genome of a Chinese P. ostreatus strain resource
- The conserved microRNAs miR2673 and miR-4968-3p have been reported elsewhere to have many predicted target genes across species finding
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| whole-genome de novo sequencing (paired-end and mate-pair) | Pleurotus ostreatus monokaryon strain CCEF00389 | none | genome assembly (scaffolds) | Illumina HiSeq 2500 |
| RNA-seq (transcriptome sequencing) | P. ostreatus CCEF00389 mycelia | heat stress (37°C for 0, 0.5, 1, 1.5 h) | assembled transcriptome (TRINITY) | Illumina HiSeq 2500 |
| Rfam/Infernal and BLASTn homology search | CCEF00389 genome assembly | none | snRNA and snoRNA loci and classes | Infernal v1.0.3 / BLAST+ v2.2.31 |
| tRNAscan-SE prediction | CCEF00389 genome assembly | none | tRNA loci, length, and anticodons | tRNAscan-SE v1.3.1 |
| BLASTn against Rfam miRNA family (RF00003) | CCEF00389 genome assembly | none | mature miRNA sequences and counts | BLASTn (e-value 1e-3, word size 19) |
| gene prediction and functional annotation (BLAST+, GO, InterProScan) | CCEF00389 genome assembly | none | predicted gene models, GO terms, protein domains | BRAKER1 / BLAST+ / InterProScan |
| comparative genome alignment (BLASTn) | sncRNAs of CCEF00389 vs. genomes of 6 Agaricomycotina fungi and Ustilago maydis | none | presence/absence and sequence identity of homologous sncRNAs | — |
| hierarchical clustering (Spearman correlation of sequence identities) | sncRNA sequence identity matrix across fungal species | none | clustering of fungal species by sncRNA similarity | — |
- – 254 sncRNAs (snRNAs, snoRNAs, tRNAs, miRNAs, and other Rfam-classified RNAs) identified in the CCEF00389 genome assembly 254 total
- – snRNA U1 not found in CCEF00389 or in eight other basidiomycetous genome assemblies searched
- – Most tRNAs located within 500 bp of a gene boundary (pseudo-UTR region) 136/167, 81.44%
- – Majority of miRNAs located within introns of host genes 31/46, 67%
- – High proportion of CCEF00389 sncRNAs had homologous sncRNAs in P. ostreatus strain PC15 with high sequence identity 197/254, 77.56%; identities above 81.65%
- ▼ Very few sncRNAs had homologues in the non-Agaricomycotina basidiomycete Ustilago maydis 15/254, 5.9%
- – 185 tRNAs (71-144 nt) and 46 mature miRNAs (19-23 nt) identified in the genome assembly 185 tRNAs; 46 miRNAs
- – Genome assembly yielded 34.9-Mb size with 13,438 predicted gene models, GO annotation for 48.9% and domain annotation for 73.9% of genes 34.9 Mb; 13,438 genes
- count 254 sncRNAs (total sncRNAs identified in CCEF00389 genome assembly)
- count 185 tRNAs (tRNAs identified by tRNAscan-SE)
- count 46 mature miRNAs (miRNAs identified by BLASTn against Rfam RF00003)
- other 77.56% (197/254) (sncRNAs with homologues in P. ostreatus PC15, sequence identity above 81.65%)
- other 5.9% (15/254) (sncRNAs with homologues in Ustilago maydis)
- other 81.44% (136/167) (tRNAs located within 500 bp of gene boundary)
- other 67% (31/46) (miRNAs located within introns)
- count 13,438 gene models; 34.9-Mb genome assembly (predicted genes and genome size from ~81 million Illumina reads at ~300x coverage)
Statistical methods review
Model: opusA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This is a computational genomics study that assembled and annotated a Pleurotus ostreatus monokaryon genome and identified 254 small noncoding RNAs using sequence-homology and covariance-model tools (Rfam/Infernal/BLAST, tRNAscan-SE). Findings were reported descriptively as counts, percentages, sequence-identity thresholds, and genomic-locus distances rather than through inferential hypothesis testing. Cross-species conservation was summarized with sequence identities and visualized by hierarchical clustering using a Spearman-correlation-based dissimilarity; no formal significance tests, replication-based comparisons, or p-values were reported.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Hierarchical clustering using Spearman correlation coefficient of sequence identities as the dissimilarity measure | Figure 3, clustering of P. ostreatus CCEF00389 against six other Agaricomycotina fungi and Ustilago maydis | 254 sncRNAs (identity set to zero where no match found) | not stated |
-
Cross-species conservation was assessed with a fixed sequence-identity threshold (e.g., matches above ~80–81.65% counted as conserved) using BLASTn homology searches.↳ Could also: An e-value- or bit-score-based criterion, or covariance-model/structure-aware searches (e.g., Infernal) for all RNA classes, could also be used to call homologues. — Structure-aware or score-based criteria can capture noncoding RNAs whose sequence diverges while structure is retained, which complements identity-only cutoffs for highly variable ncRNAs.
-
Relationships among species were summarized by hierarchical clustering with a Spearman-correlation dissimilarity over sequence identities, with missing matches set to zero.↳ Could also: A model-based phylogenetic reconstruction (e.g., maximum-likelihood or Bayesian trees) or bootstrap resampling on the clustering could also be applied. — Adding resampling-based support values or a phylogenetic model would convey how robust the groupings are and provide an explicit measure of confidence in the inferred relationships.
-
Conservation and distribution results were reported as raw counts and percentages (e.g., 77.56%, 67% of miRNAs in introns).↳ Could also: Accompanying these proportions with confidence intervals (e.g., binomial/Wilson intervals) would also be an option. — Interval estimates around proportions convey the precision of each percentage given the finite number of loci, which complements the point estimates.
-
Genomic locus enrichment (e.g., sncRNAs near gene boundaries / in introns vs. other regions) was described by tabulating distances and proportions.↳ Could also: A formal enrichment test against a null/background distribution (e.g., permutation of loci, chi-square, or Fisher's exact test) could also be used. — A background-model comparison would quantify whether the observed positional preferences exceed what is expected by chance, adding an inferential layer to the descriptive counts.
-
The miRNA detection outcome was shown to depend on the BLASTn 'word size' parameter (46 vs. 10 matches at different settings).↳ Could also: A sensitivity analysis across a range of parameters, or specialized miRNA-prediction pipelines (e.g., miRDeep-style tools using sequencing read support), could also be reported. — Systematically reporting how results vary with key parameters, or incorporating expression evidence, would help readers gauge the stability of the detected set.
Result convergence & founder nodes
Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.
-
77.56% of sncRNAs from P. ostreatus CCEF00389 have homologues in strain PC15 with >81.65% sequence identityother pleurotus ostreatus ccef00389 vs pc15 2016×1papers★ This paper is the founder (earliest)
-
46 mature miRNAs (19–23 nt) identified in P. ostreatus CCEF00389, with 67% located within intronsother pleurotus ostreatus ccef00389 2016×1papers★ This paper is the founder (earliest)
-
254 small noncoding RNAs identified in P. ostreatus CCEF00389, comprising 0.054% of the genomeother pleurotus ostreatus ccef00389 2016×1papers★ This paper is the founder (earliest)
-
7 snoRNAs (including SNORD24, SNORD46), 1 RNase MRP, and 5 Hammerhead ribozymes identified in P. ostreatus CCEF00389other pleurotus ostreatus ccef00389 2016×1papers★ This paper is the founder (earliest)
-
185 tRNAs (71–144 nt) identified in P. ostreatus CCEF00389, with 81% located within 500 bp of gene boundariesother pleurotus ostreatus ccef00389 2016×1papers★ This paper is the founder (earliest)
-
U1 snRNA is absent from P. ostreatus CCEF00389 and 8 other basidiomycete genomes; U2, U4, U5, and U6 spliceosomal snRNAs are presentother pleurotus ostreatus ccef00389 2016×1papers★ This paper is the founder (earliest)
-
51 sncRNAs are conserved across all 6 surveyed Agaricomycotina genomes; only 5.9% have homologues in Ustilago maydisother pleurotus ostreatus vs agaricomycotina 2016×1papers★ This paper is the founder (earliest)
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-27703969
Paper: Qu et al. 2016, Identification and Characterization of Small Noncoding RNAs in Genome Sequences of the Edible Fungus Pleurotus ostreatus, BioMed Res Int. DOI 10.1155/2016/2503023.
Nature of the work
This is a whole-genome shotgun assembly + sncRNA annotation paper, NOT a small-RNA-seq study. Illumina genomic reads (~81 M reads, ~300× coverage) were quality-trimmed (SeqPrep + Sickle), assembled (PLATANUS + L_RNA_Scaffolder), gene-modelled (BRAKER1), then sncRNAs were annotated on the assembly with several standard tools. Deposited as BioProject PRJNA327267, WGS MAYC00000000.1, assembly GCA_001956935.1 (ASM195693v1), strain CCMSSC00389.
The brief lists the "code" as SeqPrep (jstjohn/SeqPrep) — a read-merging / adapter-trimmer used only in the read-cleaning step. Per brief rule P16, applying the paper's named third-party tools to the paper's own deposited data is a valid reproduction. The deposited artefact that downstream sncRNA results derive from is the assembly FASTA, so we reproduce against that.
IN SCOPE (clearly-specified, deterministic, low-hanging — the 80%)
| # | Reported result | Paper location | Pipeline / tool | How we reproduce |
|---|---|---|---|---|
| A | Genome size 34.9 Mb | Abstract / Results | PLATANUS assembly (deposited) | recompute total length from GCA_001956935.1 FASTA |
| B | 2,529 scaffolds | Results / assembly | (deposited) | count records in FASTA |
| C | GC content 49.54% | Results | (deposited) | recompute GC% from FASTA |
| D | Scaffold N50 394,787 bp | Results | (deposited) | recompute N50 from FASTA |
| E | 185 tRNAs, length 71–144 nt | Results / Table | tRNAscan-SE v1.3.1 | run tRNAscan-SE 1.3.1 on the assembly, count + length range |
A–D are descriptive stats recomputed from the deposited assembly (cross-check that the deposited data matches the paper's described assembly). E is a genuine re-execution of a named annotation tool on the deposited genome.
OUT OF SCOPE / harder 20% (not attempted, or attempted only if time)
- miRNA: 46 mature via BLASTn of Rfam (e-value 1e-3, word size 19). Needs an Rfam release pin (paper does not state version); Rfam content drift over a decade makes a 1:1 count brittle. Low-confidence target → deferred.
- snoRNA (7), snRNA (4) via Infernal 1.0.3 + Rfam. Same Rfam-version uncertainty; Infernal 1.0.3 is a 2009 release. Deferred.
- Read cleaning with SeqPrep/Sickle → "cleaned reads": the paper gives no exact cleaned-read count to compare against (only the rule "Q<20 or N>10% or len<25 removed"); raw-read SRA is large. No pinnable target → not a useful 1:1.
- All wet-lab / assembly-engineering steps (PLATANUS run, BRAKER1 gene models) are out of scope (not re-run; we only verify deposited descriptive stats).
Reference genome
GCA_001956935.1_ASM195693v1_genomic.fna.gz (NCBI FTP). NOTE: a later GCA_001956935.2 (CCMSSC00389_v2, chromosome-level, 136 scaffolds, GC 51%, 35.09 Mb) exists — that is the updated assembly and must NOT be used; the 2016 paper's numbers match v1.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
Reproduced 1:1 against the authors' own deposited 2016 assembly GCA_001956935.1: scaffolds (2,529), N50 (394,787 bp), GC (49.54%) and tRNA length range (71-144 nt) are exact, genome size matches on rounding (34.86→34.9 Mb), and the tRNA count is off by a single borderline call (184 vs 185, 0.5%). No deviation sits in core computation in a substantive way and every value is derivable from shared data — no fabrication signal. The harder Rfam-dependent counts (miRNA/snoRNA/snRNA) were responsibly deferred rather than asserted, which does not bear on the reproduced claims. Overall a clean, high-quality reproduction.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at [email protected].
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.