Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

sRNAbench and sRNAtoolbox 2019: intuitive fast small RNA profiling and differential expression.

Nucleic Acids Res · 2019
not yet assessed 3/4
Why this verdict

Part of the results reproduced; minor but material deviations remained.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Input / endpoint not comparable 1:1
+1 pts
From: Q2 · Endpoint comparability 🔴
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +7
✓ What held up
  • Nothing in this column.
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🔴Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🔴A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
Reproduction agent’s raw note

Described well enough as software, but NOT a pipeline-reproducible result: this is a NAR Web Server paper for sRNAbench/sRNAtoolbox. The listed code (github.com/miRTop/mirtop, commit 18238f46, live/MIT) and data (SRR2105509: Homo sapiens ncRNA-Seq, 4.29M reads, ~143MB; study PRJNA290097) both resolve, so this is not repo_gone/no_code/data_unavailable. The blocker is no_expected_result: the paper reports no pipeline-derived number tied to the listed repo+accession. SRR2105509 occurs in the paper ONLY as a command-line syntax example ('SRR2105509:SRR2105510 would merge both SRA runs'), never analysed for a reported value; mirtop is named only as an interoperable mirGFF3 export target with no reported output. The single pinnable reported number (GEO study counts 280->764) is a non-pipeline scientometric figure with an undisclosed query that does not reproduce under any natural Entrez query (closest 367->626, 1.71x). Figure 1 re-analysis numbers come from an external comparison dataset (not SRR2105509) with no shipped runnable workflow/thresholds (out of scope). NOT attempted: a full sRNAbench->mirtop run on SRR2105509, because its output would have no paper value to be graded against (would demonstrate 'tool runs', not 'reported result reproduces') -- skipped per 80/20 as low-value compute. Honest outcome: artifacts resolvable, but no auditable expected result -> drop.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment
    assessed: 2026-06-16 ⛓ 41e0d97019ee
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-16
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-16
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Core claims
  • sRNAtoolbox 2019 adds all major small RNA library preparation protocols (including UMI-based) to sRNAbench with automatic protocol-specific preprocessing. resource
  • sRNAbench now supports a batch mode allowing an unlimited number of read files to be profiled at once with the same parameters, improving reproducibility and user time efficiency. method
  • sRNAde computes consensus differential expression across five methods (edgeR, DESeq, DESeq2, NOISeq, Student's t-test). method
  • The scope was expanded with 90 genome assemblies from Ensembl, virus/bacteria collections from NCBI, and microRNA reference sequences from miRBase and MirGeneDB. resource
  • isomiR classification can be performed either hierarchically (one category per read) or fuzzy (multiple categories per read). method
  • Single nucleotide variants are detected at the precursor level from reported mismatches across all input formats. method
  • Genome-mapped reads are visualized via UCSC/Ensembl track hubs and downloadable bedGraph, bigWig and bed files instead of the previous jBrowse instance. method
  • The number of microRNA sequencing studies deposited in GEO nearly tripled from 2014 to 2018, motivating fast easy-to-use miRNA-seq tools. finding
Experimental setups
Assay System Perturbation Readout Platform
small RNA-seq (miRNA-seq) expression profiling via sRNAbench batch mode human (12 biological samples from SRA study SRP046046) none microRNA/small RNA expression values, read length distribution, RNA type distribution, adapter-dimer fraction Illumina library preparation protocol
consensus differential expression analysis via sRNAde human SRP046046 samples grouped by condition other (two-group comparison) differentially expressed microRNAs (over/underexpressed), fold-changes, consensus across methods edgeR, DESeq, DESeq2, NOISeq, Student's t-test
genome mapping of small RNA reads Ensembl reference genome assemblies none genome-mapped read counts, isomiRs, highly redundant reads bowtie1 (seed length 20 nt)
Key results
  • One sample (BJAB exosomes, SRR1563017) showed nearly 60% adapter-dimer reads while samples were generally below 20%, indicating possible library preparation issues. ~60% vs <20%
  • Overlap of microRNAs with log2 fold-change >1 or <-1 across the five DE methods was high, suggesting normalization methods have moderate impact on fold-changes. 34 out of 49
  • Only one microRNA showed statistically significant overexpression across all five methods, as Student's t-test and NOISeq are much stricter. 1 out of 32
  • DESeq, DESeq2 and edgeR showed the highest mutual overlap of differentially expressed microRNAs. 11 out of 32
  • Percentage of microRNAs among RNA types varied across samples in the study. 10% to 70%
Key statistics
  • count 280 (2014) to 764 (2018) miRNA-seq GEO studies (GEO miRNA-seq study deposits nearly tripled 2014-2018)
  • count 34 out of 49 (microRNAs with |log2 fold-change|>1 overlapping across five DE methods)
  • count 1 out of 32 (microRNAs significantly overexpressed in all five DE methods)
  • count 11 out of 32 (overlap of DE microRNAs among DESeq, DESeq2 and edgeR)
  • count 12 (biological samples in SRP046046 study, one run per sample)
  • count 90 (genome assemblies contained in current sRNAtoolbox)
  • count ~60% (adapter-dimer reads in BJAB exosomes sample SRR1563017)
  • other 21-22 nt (expected narrow read-length peak for microRNAs)

Statistical methods review

Model: opus

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This is a web-server/software paper describing the 2019 release of sRNAtoolbox for small RNA-seq profiling and differential expression rather than a hypothesis-testing study. Its statistical content is embodied in the differential expression module (sRNAde), which runs five methods—edgeR, DESeq, DESeq2, NOISeq and a Student's t-test—and reports a consensus of differentially expressed microRNAs. Results are illustrated with a public dataset (SRP046046, 12 biological samples) and presented descriptively via boxplots, heatmaps, volcano plots, UpSet intersection plots and log2 fold-change overlaps, without formal inference about the tool's performance.

Replicationbiological Sample sizeThe working example used study SRP046046 with 12 biological samples and one run per sample; no power/sample-size calculation is described. Groupstwo groups (condition labels, e.g. healthy vs. cancer/treated) assigned by the user Pairingna Randomization/blindingna Dispersionunclear Exact p-valuesno Effect sizesyes Confidence intervalsno Multiplicity correctionnot explicitly stated in the text; the DE methods used (edgeR, DESeq, DESeq2, NOISeq) conventionally report adjusted values, and the paper refers to 'adjusted read counts' for multiple genomic mapping rather than multiple-testing adjustment
Statistical tests used
Test Applied to n Assumptions
Student's t-test one of five differential expression methods in sRNAde, applied to two-group miRNA expression comparisons in the SRP046046 working example not stated
edgeR (negative binomial GLM) differential expression method in sRNAde consensus not stated
DESeq (negative binomial) differential expression method in sRNAde consensus not stated
DESeq2 (negative binomial Wald) differential expression method in sRNAde consensus not stated
NOISeq differential expression method in sRNAde consensus not stated
Approaches that could also have been used
  • Differentially expressed microRNAs are identified by running five methods and taking their consensus/overlap (e.g. via UpSet plots).
    Could also: One could pre-specify a single primary method (e.g. DESeq2 or edgeR) with a stated adjusted-significance threshold, optionally reporting the others as a sensitivity analysis. — A pre-specified primary analysis gives a single, clearly defined error-rate interpretation, while consensus plus sensitivity analyses also communicate robustness across methods.
  • A Student's t-test is offered as one of the count-based DE methods.
    Could also: A non-parametric test (e.g. Mann-Whitney U) or a count-model approach (negative binomial as in DESeq2/edgeR) could also be applied to read-count data. — Count-model or rank-based approaches accommodate the discrete, often-skewed nature and small replicate numbers typical of miRNA-seq, and would add a complementary view of significance.
  • Multiple-testing adjustment is not explicitly described in the main text for the consensus comparison.
    Could also: Reporting Benjamini–Hochberg FDR-adjusted P-values per method, with the FDR threshold stated, could also be done. — Explicitly stating the FDR method and cutoff makes the family of tests and the controlled error rate transparent across the many microRNAs tested.
  • Group differences and expression spread are summarized graphically (boxplots, heatmaps, volcano plots) in a worked example.
    Could also: Accompanying numerical summaries such as medians with IQR, or means with SD/95% CI, could also be tabulated. — Numerical dispersion and interval estimates convey magnitude and uncertainty alongside the visualizations, which is especially informative for small sample sizes.
  • Effect is reported as log2 fold-change with a fixed threshold (|log2FC|>1), with +1 added to expression values to avoid division by zero.
    Could also: Shrinkage-based effect estimates (e.g. DESeq2's apeglm/lfcShrink) or fold-changes paired with their confidence intervals could also be reported. — Shrunken estimates and CIs temper inflated fold-changes from low-count microRNAs and convey the precision of each effect, complementing the pseudocount approach already noted in the paper.
Software: sRNAbench / sRNAtoolbox (sRNAde) 2019 release · edgeR · DESeq · DESeq2 · NOISeq · bowtie1 (read mapping)

Result convergence & founder nodes

Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
194
Impact: high
Foundation confidence
Built on 1 assessed reference(s) · mean reproducibility 94/100
stands on reproducible work
Topics

Assessed papers, coloured by verdict. Click a node to open it.

Built on (assessed references) (1)
Cited by (assessed papers) (0)
  • No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

SRP046046 ENA in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet

Downstream reach in the literature

2 downstream papers · 1 datasets

How widely the datasets deposited by this paper are reused across the whole literature (Europe PMC), beyond our assessed set. This is a factual dependency map — reusing a public dataset is normal, good science. It is not a judgement on the downstream papers; the only verdict here is this paper's own, with its cited rationale.

SRP046046 ENA reused by 3 papers in the literature
Most-cited downstream papers:

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

scope.md — pmid-31114926

Paper: Aparicio-Puerta E et al. "sRNAbench and sRNAtoolbox 2019: intuitive fast small RNA profiling and differential expression." Nucleic Acids Res 2019. PMID 31114926 · PMCID PMC6602500 · DOI 10.1093/nar/gkz415.

Paper type: NAR Web Server issue article — i.e. a software/web-service description paper, not a primary data-analysis study. This is decisive for reproducibility scoping (see verdict).

Listed artifacts (from registry / harvester):

  • Code: https://github.com/miRTop/mirtoplive, MIT, not archived, default branch master, latest commit 18238f46ee6f29da8510237efacc42130571c01c (2025-04-07) as of 2026-06-16. (This is the miRTop community mirGFF3 standardisation tool, a third-party tool the paper interoperates with — not the authors' own sRNAbench source.)
  • Data: sra:SRR2105509resolves. Homo sapiens, ncRNA-Seq (small RNA), Illumina, 4,287,674 reads / 218,671,374 bases, ~143 MB fastq.gz. Study PRJNA290097 "Plasma extracellular RNA profiles in healthy and cancer patients", sample SAMN03863548.

In-scope vs out-of-scope (pipeline-derived results)

Per the brief, only pipeline-derived computational results that are (a) reported as a specific value/figure/table and (b) regenerable from the listed code+data are in scope. Reading the full text (PMC6602500), the reported numbers are:

Reported item (paper location) Origin In scope? Why
"miRNA-seq studies in GEO nearly triplicated from 2014 (280) to 2018 (764)" (Intro) GEO metadata count (database query) No (scientometric, not a sequencing pipeline; query string not disclosed) attempted anyway as a bonus auditable check — see below
Fig 1C: BJAB exosome sample (SRR1563017) "~60% adapter-dimers", others "<20%" re-analysis of an external comparison dataset No not the listed accession; no shipped runnable workflow / thresholds
Fig 1E/F: "1 of 32 miRNAs sig. across all 5 DE methods"; "34 of 49 |log2FC|>1" external DE comparison, 5 methods (edgeR/DESeq/DESeq2/NOISeq/t-test) No undisclosed processing + thresholds; no shipped script; data ≠ SRR2105509
Tool/feature descriptions, reference DB versions (Ensembl 91, miRBase, MirGeneDB), bowtie1 seed=20, ≤10 genome hits software documentation No not a reported numeric result to reproduce

The listed code×data pair (mirtop × SRR2105509)

The harvester paired mirtop with SRR2105509. Critically, in the paper SRR2105509 appears only as a command-line syntax example in the working- example/help text:

"…SRR2105509:SRR2105510 would merge both SRA runs into a single job."

There is no reported result (no count, no figure, no table value) derived from running sRNAbench or mirtop on SRR2105509. SRR2105509 was never analysed in the paper for a reported value — it is a placeholder accession illustrating input syntax. Likewise mirtop is mentioned only as an interoperable export target ("sRNAbench output can be converted to the miRTop standardized format"), with no reported mirtop-derived number.

Therefore the implied reproduction — run mirtop on SRR2105509 and compare to the paperhas no expected result to compare against. One could run the sRNAbench→mirtop pipeline on SRR2105509 and obtain isomiR counts, but those numbers would be ungradeable (the paper reports none for this sample); that is "the tool runs", not "a reported result reproduces", so it is not attempted as a compute job (80/20: no auditable comparator ⇒ no value over the cost).


Bonus auditable check (out-of-scope, scientometric)

Reproduced the only pinnable reported number, the GEO triplication, via NCBI Entrez (esearch db=gds … [PDAT]). Results are query-definition dependent and do not reproduce 280→764 under any natural query:

Query (db=gds, GSE series, by submission year) 2014 2018 ratio
reported in paper 280 764
geo_triplication
Reported
GEO miRNA-seq studies 2014=280, 2018=764 (~2.73x)
Reproduced
best natural Entrez query 367->626 (1.71x); +microRNA 218->398 (1.83x); none reproduce 280/764
m.public.grade.no_expected_result
mirtop_on_srr2105509
Reported
NONE (SRR2105509 appears only as a CLI syntax example; no reported mirtop/sRNAbench output value)
Reproduced
not attempted (no comparator to grade against)
m.public.grade.no_expected_result

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 38/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🔴2. Endpoint comparability
🟡3. Location of the main deviation
🔴4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Input / endpoint not comparable 1:1
+1 pts
From: Q2 · Endpoint comparability 🔴
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +7

This is a NAR Web Server paper for sRNAbench/sRNAtoolbox: the listed code (mirtop, live/MIT @18238f46) and data (SRR2105509, resolves) both exist, but no pipeline-derived number is tied to that repo+accession — SRR2105509 is only a CLI syntax example and mirtop only an export target, so there is nothing to grade (no_expected_result). The lone pinnable figure (GEO study counts 2014=280→2018=764, 2.73x) is an out-of-scope scientometric count whose exact query is undisclosed; no natural Entrez query reproduces it (closest 367→626, 1.71x), though the increasing direction holds. The deviation is on the authors'/data-availability side (undisclosed, time-drifting query) but shows no fabrication signal — best explained by database reclassification — so this lands as solid-but-not-pipeline-reproducible (yellow), not a critical discrepancy.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

75.6 k
tokens (I/O) · 2.6 M incl. cache
19 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.