Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Identification and Characterization of Small Noncoding RNAs in Genome Sequences of the Edible Fungus Pleurotus ostreatus.

Biomed Res Int · 2016
L1 95/100 PQI 93
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Every question reproduced
-1 pts
From: “every question reproduced”
Total score -7
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
  • No relevant deviation in data/preprocessing
  • No authors-side cause for any deviation
  • Reported values are derivable from the shared data
  • Any deviation was negligible
  • The central claim held under reproduction
  • Overall, the reproduction was clean
What did not (or only partly)
  • Every checked point held up.
How its reproducibility compares
95/100
Reproducibility score
1.2 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 89% of all assessed papers rank 105 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough -> 1:1 reproduction of the deterministic, pipeline-derived outputs. This is a genome-assembly + sncRNA-annotation paper (WGS PRJNA327267, not sRNA-seq). Reproduced against the deposited 2016 assembly GCA_001956935.1 (ASM195693v1, strain CCMSSC00389, md5 e70c2fe1f45c7f91413fc95dba2ab88f). Four descriptive assembly stats recomputed from the FASTA: scaffolds 2529 (EXACT), N50 394787 bp (EXACT), GC 49.54% (EXACT over total length), size 34.86 Mb called bases ~= reported 34.9 Mb (total span 35.82 Mb incl gap-Ns). tRNA annotation re-run with the paper's exact tool tRNAscan-SE v1.3.1 (default Eukaryotic mode): 184 standard-AA tRNAs vs reported 185 (off by 1, 0.5%; 194 incl 10 pseudogenes), and length range 71-144 nt reproduces EXACTLY. No fabrication signal: every value derivable from shipped data, reproduced exactly or off-by-<=1. NOT attempted (harder 20%): miRNA=46 (BLASTn Rfam RF00003), snoRNA=7 / snRNA=4 (Infernal 1.0.3) - all depend on an unspecified Rfam release whose decade of content drift makes a 1:1 count unreliable; and the SeqPrep/Sickle read-cleaning step (no pinnable reported count). All heavy compute on «our HPC» SLURM (node n094); data kept on «infra», only small results on «host».

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 95
    assessed: 2026-06-16 ⛓ 6da8b3cdd32d
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-16
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-16
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

The paper sets out to perform the first genome-scale identification and characterization of small noncoding RNAs (sncRNAs) in a basidiomycete, the edible fungus Pleurotus ostreatus, and to analyze their genomic distribution and evolutionary conservation across Agaricomycotina fungi.

Core claims
  • 254 small noncoding RNAs (snRNAs, snoRNAs, tRNAs, miRNAs) were detected in the P. ostreatus CCEF00389 genome assembly, the first genome-scale identification of sncRNAs for a basidiomycete. finding
  • snRNA U1 was not found in CCEF00389 or other basidiomycete genomes, implying that if it exists in basidiomycetes its sequence varies significantly from other organisms. finding
  • snRNAs and most tRNAs (88.6%) are located in pseudo-UTR regions, whereas miRNAs are commonly found in introns. finding
  • Most sncRNAs (77.56%) are highly conserved within P. ostreatus, but only ~20% are conserved across Agaricomycotina fungi, indicating most P. ostreatus sncRNAs are not broadly conserved. finding
  • A 34.9-Mb draft genome of P. ostreatus strain CCEF00389 was assembled with 13,438 predicted gene models, the first released draft genome of a P. ostreatus strain in China. resource
  • sncRNAs were identified computationally by aligning Rfam sequences with BLAST+/Infernal, predicting tRNAs with tRNAscan-SE, and detecting miRNAs via BLASTn of Rfam miRNA sequences. method
  • BLASTn word size strongly affects miRNA detection; word size 19 found 46 miRNAs while word size 20 found only 10. finding
  • Only 10 sncRNAs (all miRNAs) were conserved across all selected basidiomycetes, and only 15 of 254 (5.9%) had homologues in Ustilago maydis. finding
Experimental setups
Assay System Perturbation Readout Platform
Whole genome de novo sequencing and assembly P. ostreatus monokaryon strain CCEF00389 (derived from dikaryon CCMSSC00389) none genome assembly, gene models, sncRNA loci Illumina HiSeq 2500; PLATANUS, L_RNA_Scaffolder, BRAKER1
Bulk RNA-seq / transcriptome sequencing P. ostreatus mycelia under heat stress (37°C for 0, 0.5, 1, 1.5 h) heat stress (37°C) transcriptome to guide genome assembly/annotation Illumina HiSeq 2500; TRINITY de novo assembler
Computational sncRNA detection (snRNA/snoRNA, Rfam homology search) CCEF00389 genome assembly none identified snRNAs, snoRNAs, RNase MRP, ribozymes BLAST+ and Infernal version 1.0.3
tRNA prediction CCEF00389 genome assembly none tRNA loci, anticodons, lengths tRNAscan-SE version 1.3.1
miRNA detection by sequence alignment CCEF00389 genome assembly none mature miRNAs identified BLASTn (Rfam RF00003, e-value 1e-3, word size 19)
Evolutionary conservation / homology alignment and hierarchical clustering CCEF00389 sncRNAs vs genomes of 6 Agaricomycotina fungi and Ustilago maydis none sequence identities, conservation, clustering by Spearman correlation
Gene functional annotation CCEF00389 predicted proteins none GO annotations and protein domains BLAST+ v2.2.31 (NR, Refseq), InterProScan
Key results
  • 34.9-Mb genome assembled from ~81 million Illumina reads, generating 13,438 gene models 34.9 Mb; ~300x coverage; 13,438 genes
  • 254 sncRNAs detected, accounting for 0.054% of the CCEF00389 genome 254 sncRNAs; 0.054% of genome
  • 185 tRNAs identified (length 71–144 nt); most located within 500 bp of gene boundary 185 tRNAs; 136/167 (81.44%) within 500 bp
  • 46 mature miRNAs identified (length 19–23 nt); 67% located in introns 46 miRNAs; 31/46 (67%) in introns
  • snRNA U1 absent from CCEF00389 and 8 other basidiomycete genomes; U2,U4,U5,U6 found 4 of 5 spliceosomal snRNAs found
  • 197 of 254 sncRNAs (77.56%) had homologues in P. ostreatus PC15 with identity above 81.65% 77.56% (197/254); identity >81.65%
  • Only 15 of 254 (5.9%) sncRNAs had homologues in Ustilago maydis; 51 found in all 6 Agaricomycotina genomes, 74 in at least five 5.9% (15/254); 51 and 74 sncRNAs
  • 7 snoRNAs identified (3 snoZ13_snr52, plus snosnR60_Z15, SNORD24, Afu_455, SNORD46); plus 1 RNase MRP and 5 Hammerhead ribozymes 7 snoRNAs; 6 other sncRNAs
Key statistics
  • count 254 sncRNAs (total sncRNAs detected in CCEF00389 genome)
  • other 0.054% (fraction of CCEF00389 genome occupied by sncRNA sequence)
  • count 34.9-Mb genome; 13,438 gene models (genome assembly size and predicted genes)
  • count 88.6% (snRNAs and most tRNAs located in pseudo-UTR regions (abstract))
  • other 77.56% (197 out of 254) (sncRNAs with homologues in P. ostreatus PC15, identity >81.65%)
  • other 5.9% (15 out of 254) (sncRNAs with homologues in Ustilago maydis)
  • count 67% (31 out of 46) (miRNAs located in introns)
  • count 81.44% (136 out of 167) (tRNAs located within 500 bp of gene boundary)

Statistical methods review

Model: opus

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This is a computational genomics study that assembled and annotated a Pleurotus ostreatus monokaryon genome and identified 254 small noncoding RNAs using sequence-homology and covariance-model tools (Rfam/Infernal/BLAST, tRNAscan-SE). Findings were reported descriptively as counts, percentages, sequence-identity thresholds, and genomic-locus distances rather than through inferential hypothesis testing. Cross-species conservation was summarized with sequence identities and visualized by hierarchical clustering using a Spearman-correlation-based dissimilarity; no formal significance tests, replication-based comparisons, or p-values were reported.

Replicationunclear Sample sizeDescribed as a single genome assembly (one monokaryon, CCEF00389) of ~34.9 Mb at ~300x coverage; 254 sncRNAs detected; no replicate genomes or power analysis described GroupsCCEF00389 sncRNAs vs. genome assemblies of other Agaricomycotina fungi and U. maydis Pairingna Randomization/blindingna Dispersionnone Exact p-valuesno Effect sizesno Confidence intervalsno Multiplicity correctionnone stated
Statistical tests used
Test Applied to n Assumptions
Hierarchical clustering using Spearman correlation coefficient of sequence identities as the dissimilarity measure Figure 3, clustering of P. ostreatus CCEF00389 against six other Agaricomycotina fungi and Ustilago maydis 254 sncRNAs (identity set to zero where no match found) not stated
Approaches that could also have been used
  • Cross-species conservation was assessed with a fixed sequence-identity threshold (e.g., matches above ~80–81.65% counted as conserved) using BLASTn homology searches.
    Could also: An e-value- or bit-score-based criterion, or covariance-model/structure-aware searches (e.g., Infernal) for all RNA classes, could also be used to call homologues. — Structure-aware or score-based criteria can capture noncoding RNAs whose sequence diverges while structure is retained, which complements identity-only cutoffs for highly variable ncRNAs.
  • Relationships among species were summarized by hierarchical clustering with a Spearman-correlation dissimilarity over sequence identities, with missing matches set to zero.
    Could also: A model-based phylogenetic reconstruction (e.g., maximum-likelihood or Bayesian trees) or bootstrap resampling on the clustering could also be applied. — Adding resampling-based support values or a phylogenetic model would convey how robust the groupings are and provide an explicit measure of confidence in the inferred relationships.
  • Conservation and distribution results were reported as raw counts and percentages (e.g., 77.56%, 67% of miRNAs in introns).
    Could also: Accompanying these proportions with confidence intervals (e.g., binomial/Wilson intervals) would also be an option. — Interval estimates around proportions convey the precision of each percentage given the finite number of loci, which complements the point estimates.
  • Genomic locus enrichment (e.g., sncRNAs near gene boundaries / in introns vs. other regions) was described by tabulating distances and proportions.
    Could also: A formal enrichment test against a null/background distribution (e.g., permutation of loci, chi-square, or Fisher's exact test) could also be used. — A background-model comparison would quantify whether the observed positional preferences exceed what is expected by chance, adding an inferential layer to the descriptive counts.
  • The miRNA detection outcome was shown to depend on the BLASTn 'word size' parameter (46 vs. 10 matches at different settings).
    Could also: A sensitivity analysis across a range of parameters, or specialized miRNA-prediction pipelines (e.g., miRDeep-style tools using sequencing read support), could also be reported. — Systematically reporting how results vary with key parameters, or incorporating expression evidence, would help readers gauge the stability of the detected set.
Software: BLAST+ (BLASTn) 2.2.31 · Infernal 1.0.3 · tRNAscan-SE 1.3.1 · Rfam 11 · PLATANUS · L_RNA_Scaffolder · TRINITY · BRAKER1 · SeqPrep · Sickle · InterProScan

Result convergence & founder nodes

Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
33
Impact: medium
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

RF00003 Rfam in Methods (http://purl.org/orb/Methods)
also used by 1 paper:
PRJNA327267 BioProject in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-27703969

Paper: Qu et al. 2016, Identification and Characterization of Small Noncoding RNAs in Genome Sequences of the Edible Fungus Pleurotus ostreatus, BioMed Res Int. DOI 10.1155/2016/2503023.

Nature of the work

This is a whole-genome shotgun assembly + sncRNA annotation paper, NOT a small-RNA-seq study. Illumina genomic reads (~81 M reads, ~300× coverage) were quality-trimmed (SeqPrep + Sickle), assembled (PLATANUS + L_RNA_Scaffolder), gene-modelled (BRAKER1), then sncRNAs were annotated on the assembly with several standard tools. Deposited as BioProject PRJNA327267, WGS MAYC00000000.1, assembly GCA_001956935.1 (ASM195693v1), strain CCMSSC00389.

The brief lists the "code" as SeqPrep (jstjohn/SeqPrep) — a read-merging / adapter-trimmer used only in the read-cleaning step. Per brief rule P16, applying the paper's named third-party tools to the paper's own deposited data is a valid reproduction. The deposited artefact that downstream sncRNA results derive from is the assembly FASTA, so we reproduce against that.

IN SCOPE (clearly-specified, deterministic, low-hanging — the 80%)

# Reported result Paper location Pipeline / tool How we reproduce
A Genome size 34.9 Mb Abstract / Results PLATANUS assembly (deposited) recompute total length from GCA_001956935.1 FASTA
B 2,529 scaffolds Results / assembly (deposited) count records in FASTA
C GC content 49.54% Results (deposited) recompute GC% from FASTA
D Scaffold N50 394,787 bp Results (deposited) recompute N50 from FASTA
E 185 tRNAs, length 71–144 nt Results / Table tRNAscan-SE v1.3.1 run tRNAscan-SE 1.3.1 on the assembly, count + length range

A–D are descriptive stats recomputed from the deposited assembly (cross-check that the deposited data matches the paper's described assembly). E is a genuine re-execution of a named annotation tool on the deposited genome.

OUT OF SCOPE / harder 20% (not attempted, or attempted only if time)

  • miRNA: 46 mature via BLASTn of Rfam (e-value 1e-3, word size 19). Needs an Rfam release pin (paper does not state version); Rfam content drift over a decade makes a 1:1 count brittle. Low-confidence target → deferred.
  • snoRNA (7), snRNA (4) via Infernal 1.0.3 + Rfam. Same Rfam-version uncertainty; Infernal 1.0.3 is a 2009 release. Deferred.
  • Read cleaning with SeqPrep/Sickle → "cleaned reads": the paper gives no exact cleaned-read count to compare against (only the rule "Q<20 or N>10% or len<25 removed"); raw-read SRA is large. No pinnable target → not a useful 1:1.
  • All wet-lab / assembly-engineering steps (PLATANUS run, BRAKER1 gene models) are out of scope (not re-run; we only verify deposited descriptive stats).

Reference genome

GCA_001956935.1_ASM195693v1_genomic.fna.gz (NCBI FTP). NOTE: a later GCA_001956935.2 (CCMSSC00389_v2, chromosome-level, 136 scaffolds, GC 51%, 35.09 Mb) exists — that is the updated assembly and must NOT be used; the 2016 paper's numbers match v1.

A
Reported
34.9 Mb (genome size)
Reproduced
34.86 Mb called bases / 35.82 Mb total span
within tolerance
B
Reported
2,529 scaffolds
Reproduced
2529
exact
C
Reported
GC 49.54%
Reproduced
49.5378% (GC/total)
exact
D
Reported
scaffold N50 394,787 bp
Reproduced
394787
exact
E
Reported
185 tRNAs
Reproduced
184 standard-AA (194 incl 10 pseudogenes)
within tolerance
F
Reported
tRNA length 71-144 nt
Reproduced
71-144 nt
exact

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 95/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟢3. Location of the main deviation
🟢4. Cause of the deviation
🟢5. Derivability / plausibility
🟢6. Severity of the deviation
🟢7. Core claim
🟢8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Every question reproduced
-1 pts
From: “every question reproduced”
Total score -7

Reproduced 1:1 against the authors' own deposited 2016 assembly GCA_001956935.1: scaffolds (2,529), N50 (394,787 bp), GC (49.54%) and tRNA length range (71-144 nt) are exact, genome size matches on rounding (34.86→34.9 Mb), and the tRNA count is off by a single borderline call (184 vs 185, 0.5%). No deviation sits in core computation in a substantive way and every value is derivable from shared data — no fabrication signal. The harder Rfam-dependent counts (miRNA/snoRNA/snRNA) were responsibly deferred rather than asserted, which does not bear on the reproduced claims. Overall a clean, high-quality reproduction.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

101.1 k
tokens (I/O) · 10.1 M incl. cache
12 min
runtime · 0.02 CPU-h
1.9 GB
peak RAM
2
HPC jobs
hummel
machine