sRNAbench and sRNAtoolbox 2019: intuitive fast small RNA profiling and differential expression.
Part of the results reproduced; minor but material deviations remained.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- Nothing in this column.
- 🟡Could not use the authors’ exact input data
- 🔴Reported values were only indirectly comparable
- 🟡A deviation arose in the data or preprocessing
- 🔴A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The deviation was non-trivial in magnitude
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
▸Reproduction agent’s raw note
Described well enough as software, but NOT a pipeline-reproducible result: this is a NAR Web Server paper for sRNAbench/sRNAtoolbox. The listed code (github.com/miRTop/mirtop, commit 18238f46, live/MIT) and data (SRR2105509: Homo sapiens ncRNA-Seq, 4.29M reads, ~143MB; study PRJNA290097) both resolve, so this is not repo_gone/no_code/data_unavailable. The blocker is no_expected_result: the paper reports no pipeline-derived number tied to the listed repo+accession. SRR2105509 occurs in the paper ONLY as a command-line syntax example ('SRR2105509:SRR2105510 would merge both SRA runs'), never analysed for a reported value; mirtop is named only as an interoperable mirGFF3 export target with no reported output. The single pinnable reported number (GEO study counts 280->764) is a non-pipeline scientometric figure with an undisclosed query that does not reproduce under any natural Entrez query (closest 367->626, 1.71x). Figure 1 re-analysis numbers come from an external comparison dataset (not SRR2105509) with no shipped runnable workflow/thresholds (out of scope). NOT attempted: a full sRNAbench->mirtop run on SRR2105509, because its output would have no paper value to be graded against (would demonstrate 'tool runs', not 'reported result reproduces') -- skipped per 80/20 as low-value compute. Honest outcome: artifacts resolvable, but no auditable expected result -> drop.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessmentassessed: 2026-06-16 ⛓ 41e0d97019ee
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-16
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-16no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opus- ★ sRNAtoolbox 2019 adds all major small RNA library preparation protocols (including UMI-based) to sRNAbench with automatic protocol-specific preprocessing. resource
- ★ sRNAbench now supports a batch mode allowing an unlimited number of read files to be profiled at once with the same parameters, improving reproducibility and user time efficiency. method
- ★ sRNAde computes consensus differential expression across five methods (edgeR, DESeq, DESeq2, NOISeq, Student's t-test). method
- ★ The scope was expanded with 90 genome assemblies from Ensembl, virus/bacteria collections from NCBI, and microRNA reference sequences from miRBase and MirGeneDB. resource
- isomiR classification can be performed either hierarchically (one category per read) or fuzzy (multiple categories per read). method
- Single nucleotide variants are detected at the precursor level from reported mismatches across all input formats. method
- Genome-mapped reads are visualized via UCSC/Ensembl track hubs and downloadable bedGraph, bigWig and bed files instead of the previous jBrowse instance. method
- The number of microRNA sequencing studies deposited in GEO nearly tripled from 2014 to 2018, motivating fast easy-to-use miRNA-seq tools. finding
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| small RNA-seq (miRNA-seq) expression profiling via sRNAbench batch mode | human (12 biological samples from SRA study SRP046046) | none | microRNA/small RNA expression values, read length distribution, RNA type distribution, adapter-dimer fraction | Illumina library preparation protocol |
| consensus differential expression analysis via sRNAde | human SRP046046 samples grouped by condition | other (two-group comparison) | differentially expressed microRNAs (over/underexpressed), fold-changes, consensus across methods | edgeR, DESeq, DESeq2, NOISeq, Student's t-test |
| genome mapping of small RNA reads | Ensembl reference genome assemblies | none | genome-mapped read counts, isomiRs, highly redundant reads | bowtie1 (seed length 20 nt) |
- ▲ One sample (BJAB exosomes, SRR1563017) showed nearly 60% adapter-dimer reads while samples were generally below 20%, indicating possible library preparation issues. ~60% vs <20%
- – Overlap of microRNAs with log2 fold-change >1 or <-1 across the five DE methods was high, suggesting normalization methods have moderate impact on fold-changes. 34 out of 49
- ▲ Only one microRNA showed statistically significant overexpression across all five methods, as Student's t-test and NOISeq are much stricter. 1 out of 32
- – DESeq, DESeq2 and edgeR showed the highest mutual overlap of differentially expressed microRNAs. 11 out of 32
- – Percentage of microRNAs among RNA types varied across samples in the study. 10% to 70%
- count 280 (2014) to 764 (2018) miRNA-seq GEO studies (GEO miRNA-seq study deposits nearly tripled 2014-2018)
- count 34 out of 49 (microRNAs with |log2 fold-change|>1 overlapping across five DE methods)
- count 1 out of 32 (microRNAs significantly overexpressed in all five DE methods)
- count 11 out of 32 (overlap of DE microRNAs among DESeq, DESeq2 and edgeR)
- count 12 (biological samples in SRP046046 study, one run per sample)
- count 90 (genome assemblies contained in current sRNAtoolbox)
- count ~60% (adapter-dimer reads in BJAB exosomes sample SRR1563017)
- other 21-22 nt (expected narrow read-length peak for microRNAs)
Statistical methods review
Model: opusA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This is a web-server/software paper describing the 2019 release of sRNAtoolbox for small RNA-seq profiling and differential expression rather than a hypothesis-testing study. Its statistical content is embodied in the differential expression module (sRNAde), which runs five methods—edgeR, DESeq, DESeq2, NOISeq and a Student's t-test—and reports a consensus of differentially expressed microRNAs. Results are illustrated with a public dataset (SRP046046, 12 biological samples) and presented descriptively via boxplots, heatmaps, volcano plots, UpSet intersection plots and log2 fold-change overlaps, without formal inference about the tool's performance.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Student's t-test | one of five differential expression methods in sRNAde, applied to two-group miRNA expression comparisons in the SRP046046 working example | — | not stated |
| edgeR (negative binomial GLM) | differential expression method in sRNAde consensus | — | not stated |
| DESeq (negative binomial) | differential expression method in sRNAde consensus | — | not stated |
| DESeq2 (negative binomial Wald) | differential expression method in sRNAde consensus | — | not stated |
| NOISeq | differential expression method in sRNAde consensus | — | not stated |
-
Differentially expressed microRNAs are identified by running five methods and taking their consensus/overlap (e.g. via UpSet plots).↳ Could also: One could pre-specify a single primary method (e.g. DESeq2 or edgeR) with a stated adjusted-significance threshold, optionally reporting the others as a sensitivity analysis. — A pre-specified primary analysis gives a single, clearly defined error-rate interpretation, while consensus plus sensitivity analyses also communicate robustness across methods.
-
A Student's t-test is offered as one of the count-based DE methods.↳ Could also: A non-parametric test (e.g. Mann-Whitney U) or a count-model approach (negative binomial as in DESeq2/edgeR) could also be applied to read-count data. — Count-model or rank-based approaches accommodate the discrete, often-skewed nature and small replicate numbers typical of miRNA-seq, and would add a complementary view of significance.
-
Multiple-testing adjustment is not explicitly described in the main text for the consensus comparison.↳ Could also: Reporting Benjamini–Hochberg FDR-adjusted P-values per method, with the FDR threshold stated, could also be done. — Explicitly stating the FDR method and cutoff makes the family of tests and the controlled error rate transparent across the many microRNAs tested.
-
Group differences and expression spread are summarized graphically (boxplots, heatmaps, volcano plots) in a worked example.↳ Could also: Accompanying numerical summaries such as medians with IQR, or means with SD/95% CI, could also be tabulated. — Numerical dispersion and interval estimates convey magnitude and uncertainty alongside the visualizations, which is especially informative for small sample sizes.
-
Effect is reported as log2 fold-change with a fixed threshold (|log2FC|>1), with +1 added to expression values to avoid division by zero.↳ Could also: Shrinkage-based effect estimates (e.g. DESeq2's apeglm/lfcShrink) or fold-changes paired with their confidence intervals could also be reported. — Shrunken estimates and CIs temper inflated fold-changes from low-count microRNAs and convey the precision of each effect, complementing the pseudocount approach already noted in the paper.
Result convergence & founder nodes
Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.
-
One sample (BJAB exosomes) showed ~60% adapter-dimer reads versus <20% in others, flagging a library-preparation issue.RNA-seq human bjab exosome up 2019×1papers★ This paper is the founder (earliest)
-
DESeq, DESeq2 and edgeR showed the highest mutual overlap (11/32) of differentially expressed microRNAs.RNA-seq human none 2019×1papers★ This paper is the founder (earliest)
-
Across five DE methods, microRNAs with |log2 fold-change|>1 overlapped highly (34/49), indicating normalization choice has only moderate impact on fold-changes.RNA-seq human none 2019×1papers★ This paper is the founder (earliest)
-
The fraction of microRNAs among RNA types varied widely across samples (10% to 70%).RNA-seq human mixed 2019×1papers★ This paper is the founder (earliest)
-
Only one microRNA was significantly overexpressed across all five DE methods, reflecting the stricter Student's t-test and NOISeq.RNA-seq human up 2019×1papers★ This paper is the founder (earliest)
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
Assessed papers, coloured by verdict. Click a node to open it.
- No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
Downstream reach in the literature
2 downstream papers · 1 datasetsHow widely the datasets deposited by this paper are reused across the whole literature (Europe PMC), beyond our assessed set. This is a factual dependency map — reusing a public dataset is normal, good science. It is not a judgement on the downstream papers; the only verdict here is this paper's own, with its cited rationale.
- Fetal Bovine Serum RNA Interferes with the Cell Cult... 2016 · 112 cites
- Using Small RNA Deep Sequencing Data to Detect Human... 2016 · 17 cites
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
scope.md — pmid-31114926
Paper: Aparicio-Puerta E et al. "sRNAbench and sRNAtoolbox 2019: intuitive fast small RNA profiling and differential expression." Nucleic Acids Res 2019. PMID 31114926 · PMCID PMC6602500 · DOI 10.1093/nar/gkz415.
Paper type: NAR Web Server issue article — i.e. a software/web-service description paper, not a primary data-analysis study. This is decisive for reproducibility scoping (see verdict).
Listed artifacts (from registry / harvester):
- Code: https://github.com/miRTop/mirtop → live, MIT, not archived,
default branch
master, latest commit18238f46ee6f29da8510237efacc42130571c01c(2025-04-07) as of 2026-06-16. (This is the miRTop community mirGFF3 standardisation tool, a third-party tool the paper interoperates with — not the authors' own sRNAbench source.) - Data:
sra:SRR2105509→ resolves. Homo sapiens, ncRNA-Seq (small RNA), Illumina, 4,287,674 reads / 218,671,374 bases, ~143 MB fastq.gz. Study PRJNA290097 "Plasma extracellular RNA profiles in healthy and cancer patients", sample SAMN03863548.
In-scope vs out-of-scope (pipeline-derived results)
Per the brief, only pipeline-derived computational results that are (a) reported as a specific value/figure/table and (b) regenerable from the listed code+data are in scope. Reading the full text (PMC6602500), the reported numbers are:
| Reported item (paper location) | Origin | In scope? | Why |
|---|---|---|---|
| "miRNA-seq studies in GEO nearly triplicated from 2014 (280) to 2018 (764)" (Intro) | GEO metadata count (database query) | No (scientometric, not a sequencing pipeline; query string not disclosed) | attempted anyway as a bonus auditable check — see below |
| Fig 1C: BJAB exosome sample (SRR1563017) "~60% adapter-dimers", others "<20%" | re-analysis of an external comparison dataset | No | not the listed accession; no shipped runnable workflow / thresholds |
| Fig 1E/F: "1 of 32 miRNAs sig. across all 5 DE methods"; "34 of 49 |log2FC|>1" | external DE comparison, 5 methods (edgeR/DESeq/DESeq2/NOISeq/t-test) | No | undisclosed processing + thresholds; no shipped script; data ≠ SRR2105509 |
| Tool/feature descriptions, reference DB versions (Ensembl 91, miRBase, MirGeneDB), bowtie1 seed=20, ≤10 genome hits | software documentation | No | not a reported numeric result to reproduce |
The listed code×data pair (mirtop × SRR2105509)
The harvester paired mirtop with SRR2105509. Critically, in the paper SRR2105509 appears only as a command-line syntax example in the working- example/help text:
"…
SRR2105509:SRR2105510would merge both SRA runs into a single job."
There is no reported result (no count, no figure, no table value) derived from running sRNAbench or mirtop on SRR2105509. SRR2105509 was never analysed in the paper for a reported value — it is a placeholder accession illustrating input syntax. Likewise mirtop is mentioned only as an interoperable export target ("sRNAbench output can be converted to the miRTop standardized format"), with no reported mirtop-derived number.
Therefore the implied reproduction — run mirtop on SRR2105509 and compare to the paper — has no expected result to compare against. One could run the sRNAbench→mirtop pipeline on SRR2105509 and obtain isomiR counts, but those numbers would be ungradeable (the paper reports none for this sample); that is "the tool runs", not "a reported result reproduces", so it is not attempted as a compute job (80/20: no auditable comparator ⇒ no value over the cost).
Bonus auditable check (out-of-scope, scientometric)
Reproduced the only pinnable reported number, the GEO triplication, via NCBI
Entrez (esearch db=gds … [PDAT]). Results are query-definition dependent and
do not reproduce 280→764 under any natural query:
| Query (db=gds, GSE series, by submission year) | 2014 | 2018 | ratio |
|---|---|---|---|
| reported in paper | 280 | 764 |
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
This is a NAR Web Server paper for sRNAbench/sRNAtoolbox: the listed code (mirtop, live/MIT @18238f46) and data (SRR2105509, resolves) both exist, but no pipeline-derived number is tied to that repo+accession — SRR2105509 is only a CLI syntax example and mirtop only an export target, so there is nothing to grade (no_expected_result). The lone pinnable figure (GEO study counts 2014=280→2018=764, 2.73x) is an out-of-scope scientometric count whose exact query is undisclosed; no natural Entrez query reproduces it (closest 367→626, 1.71x), though the increasing direction holds. The deviation is on the authors'/data-availability side (undisclosed, time-drifting query) but shows no fabrication signal — best explained by database reclassification — so this lands as solid-but-not-pipeline-reproducible (yellow), not a critical discrepancy.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.