Systematic Assessment of Small RNA Profiling in Human Extracellular Vesicles.
The main results reproduced: recomputed values matched the published ones within tolerance.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓No relevant deviation in data/preprocessing
- ✓No authors-side cause for any deviation
- ✓Reported values are derivable from the shared data
- ✓Any deviation was negligible
- 🟡Could not use the authors’ exact input data
- 🟡Reported values were only indirectly comparable
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Described well enough to reproduce. Third-party-tool reproduction (P16): the paper's own tool findadapt (github.com/chc-code/findadapt, pinned commit 47bf1f1, 2024-01-04) was re-run on «our HPC» (SLURM «job»; python3.9 + pyahocorasick + cutadapt 5.2). findadapt's adapter-pattern detection is a deterministic function of the input FASTQ and is the cleanly-specified runnable pipeline output. (1) On the repo's three shipped demo datasets it auto-identified all 3 library kits correctly (QIAseq / NEXTflex / SMARTer). (2) For NEXTflex (GSE122068) the emitted .adapter.txt matched the repo's DOCUMENTED expected output essentially 1:1: 3' seed TGGAATTCTCGG and 4/4 random-base phase EXACT, ratios within 0.0014 (0.8659 vs 0.8667; 0.9725 vs 0.9711); read count 1162 vs 1177 (-1.3%) explained by tool-version drift (README example 2023-09 vs pinned 2024-01 commit). (3) On the PAPER'S OWN GSE100467 EV data (2 exosome + 1 serum run, 2M-read subsamples) findadapt recovered the expected Illumina TruSeq smallRNA 3' adapter seed TGGAATTCTCGG, phase 0, on all three; the lower detection ratio on the exosome runs (0.69-0.70) is consistent with the paper's own note that GSE100467 EV samples were low-quality (most excluded for <100k reads / <20% mapping). No fabrication signal: every value regenerated from public/shipped data. NOT attempted (80/20): the full downstream TIGER trimming/mapping/quantification pipeline and the paper's multi-dataset EV small-RNA profiling tables -- heavy, multi-dataset, and no single in-text scalar is cleanly pinnable (those numbers live in Supplementary).
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 90assessed: 2026-06-15 ⛓ 7134a4a170ec
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-15
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusThere is a lack of systematic assessment of the quality, technical biases, RNA composition, and RNA biotype enrichment in small RNA-seq profiling of extracellular vesicles (EVs); this study systematically evaluates these biases across cell types, biofluids, and conditions and characterizes the preferential loading of small RNA biotypes into EVs versus matched donor cells.
- ★ Different EV isolation methods vary in reproducibility for isolating small RNAs and have characteristic effects on small RNA composition, with differential ultracentrifugation showing the highest variability/lowest replicability. finding
- ★ rRNA fragments and tRNA fragments are relatively enriched in EVs, while miRNAs and snoRNA fragments are depleted in EVs compared to matched donor cells. finding
- ★ Only eight miRNAs are preferentially exported into EVs in a context-independent manner; selective release of most miRNAs into EVs is study-specific. finding
- ★ Reanalysis of 2756 publicly available human EV small RNA-seq samples from 83 studies provides a global picture of EV small RNA quality and preferential loading. resource
- ★ FindAdapt, a Python package, enables fast, accurate, automatic detection of adapter patterns without prior information, feeding the TIGER small RNA-seq pipeline. method
- Nine miRNAs (e.g., miR-451a, miR-486-5p, miR-122-5p) are highly abundant in EVs but not in donor cells, significantly enriched independent of isolation method. finding
- EVs from biofluids have significantly higher median small RNA proportions than EVs from cell lines and primary cell cultures. finding
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| small RNA-seq (reanalysis of public datasets) | human EVs from biofluids, cell lines, and primary cell cultures (2756 samples, 83 studies) | none (different EV isolation methods compared) | read counts after trimming, alignment rates to host/non-host genomes, proportional abundance of small RNA biotypes (miRNA, tRNA, rRNA, Y RNA, snRNA, snoRNA) | FindAdapt + TIGER pipeline; Cutadapt v2.10; FastQC v0.11.9; Bowtie1 v1.3.0; GENCODE GRCh37.p13, miRBase, GtRNAdb2 |
| small RNA-seq (EV vs matched donor cell comparison) | human EVs and matched donor cells (28 studies; 15 cell-line studies focused) | none (cellular vs EV compartment) | differential RNA biotype proportions and differential miRNA expression between EVs and cells | DESeq2; three-way ANOVA with EV isolation method as confounder |
| EV isolation/enrichment | human plasma and other biofluids/cells | Exo-Quick, ExoEasy, ExoRNeasy, qEV, total exosome isolation kit (Thermo), differential ultracentrifugation | small RNA proportion median and IQR (variability/reproducibility) | miRNeasy kits for RNA purification |
| EV validation (referenced in source studies) | human EVs | none | EV-enriched marker proteins (CD63, CD81, CD9), particle size/number | Western blotting; electron microscopy; nanoparticle tracking analysis; dynamic light scattering; tunable resistive pulse sensing; flow cytometry |
- ▼ miRNA proportion significantly lower in EVs than matched cellular levels p = 7.56 × 10^-64
- ▼ snoRNA fragments significantly lower in EVs than cells p = 4.19 × 10^-21
- ▲ rRNA fragments significantly higher in EVs than cells p = 1.58 × 10^-30
- ▲ tRNA fragments significantly higher in EVs than cells p = 4.67 × 10^-25
- ▲ Differential ultracentrifugation showed larger IQRs in small RNA proportions than other methods p = 0.006/0.002/0.03/0.008/0.26 vs Exo-Quick/ExoEasy/ExoRNeasy/qEV/total exosome kit
- ▲ EVs from biofluids had higher median small RNA proportions than cell lines/primary cultures p = 2.37 × 10^-5/0.02
- – miRNAs were the most abundant small RNA biotype in plasma EVs median proportion 39.6%
- ▲ Nine miRNAs highly enriched in EVs vs cells independent of isolation method log2FC > 3 and FDR < 0.01
- count 2756 samples from 83 studies (55 EVs only, 28 with EVs and matched donor cells) (datasets reanalyzed after filtering)
- pvalue 7.56 × 10^-64 (miRNA depletion in EVs vs cells, three-way ANOVA)
- pvalue 1.58 × 10^-30 (rRNA fragment enrichment in EVs vs cells)
- pvalue 4.67 × 10^-25 (tRNA fragment enrichment in EVs vs cells)
- pvalue 4.19 × 10^-21 (snoRNA fragment depletion in EVs vs cells)
- other 94.6% (datasets with >100,000 reads and >20% mapping to host+non-host genomes after trimming)
- other 39.6% (median miRNA proportion in plasma EVs; Y RNA median 15.9%, rRNA 10.5%, tRNA 3.2%)
- other 78.6% (datasets meeting both ERCC criteria (>100,000 host reads and >50% host genome reads))
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This is a large-scale meta-analysis reprocessing 2756 publicly available EV small RNA-seq samples from 83 studies through a uniform bioinformatics pipeline (FindAdapt + TIGER). Quality metrics were summarized as medians and IQRs and compared across isolation methods and donor source types using t-tests. RNA biotype enrichment in EVs versus matched donor cells was tested with three-way ANOVA (EV isolation method as a confounding factor). Per-study differential miRNA expression between EVs and matched cells was assessed with DESeq2 (Benjamini-Hochberg FDR, |FC|>1.5 threshold). Results are reported primarily as proportional abundances with exact p-values and log2 fold changes.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| t-test (direction unspecified) | Median small RNA proportions in EVs from biofluids vs. cell lines vs. primary cell cultures (Figure S1A); IQR of small RNA proportions in urine vs. other biofluids (Figure S1B) | — | not stated |
| t-test (direction unspecified) | Pairwise comparisons of IQR of small RNA proportions between differential ultracentrifugation and each of five other EV extraction methods (Figure 2A); p = 0.006, 0.002, 0.03, 0.008, 0.26 | — | not stated |
| three-way ANOVA with EV isolation method as confounding variable | Proportional enrichment/depletion of miRNA (p=7.56e-64), snoRNA (p=4.19e-21), rRNA (p=1.58e-30), and tRNA (p=4.67e-25) in EVs vs. matched donor cells (Figure 4A-D) | 15 cell-line studies (focused subset of 28 matched EV/cell studies) | not stated |
| DESeq2 negative binomial Wald test with EV isolation method as covariate | Differential miRNA expression between EVs and matched donor cells, per study and across studies (Figure 5B); FDR<0.05 and |log2FC|>log2(1.5) | 28 studies with both EVs and matched donor cells | not stated |
-
Five pairwise t-tests compared IQR of small RNA proportions between differential ultracentrifugation and each of five other isolation methods, without a stated family-wise correction↳ Could also: A one-way ANOVA or Kruskal-Wallis test across all six methods followed by a post-hoc procedure (e.g., Tukey HSD or Dunn's test with BH correction) could also have been applied — A unified omnibus test followed by a post-hoc correction controls the family-wise error rate across all pairwise comparisons simultaneously, which is a standard approach when comparing a continuous outcome across more than two groups
-
RNA biotype enrichment in EVs versus matched cells was assessed with three-way ANOVA, treating EV isolation method as a confounding fixed factor; samples are nested within studies across 15 heterogeneous experiments↳ Could also: A linear mixed-effects model with study as a random effect and EV isolation method (and cell type, if available) as fixed covariates could also have been used — Because samples are nested within studies, a mixed-effects model explicitly models the non-independence of within-study observations and the between-study variance, which is a common consideration in pooled reanalysis of multi-study data
-
Variability within studies was quantified as IQR and compared across isolation methods by applying a t-test to those IQR point estimates↳ Could also: A formal test of equality of variance (e.g., Levene's test or Brown-Forsythe test) applied to the raw proportion values within each method group could also have been used — Testing dispersion directly on the underlying observations is more sensitive to true differences in spread than comparing IQR summary statistics with a location test (t-test), which treats each study's IQR as a single observation
-
Differential miRNA expression between EVs and matched cells was assessed per study using DESeq2↳ Could also: edgeR (quasi-likelihood F-test) or limma-voom could also have been applied to the same count data, and a cross-study meta-analytic approach (e.g., combining per-study effect sizes) could have aggregated findings — DESeq2, edgeR, and limma-voom implement complementary distributional assumptions; comparing results across methods is a common sensitivity check, and formal meta-analysis frameworks (e.g., metafor) could pool per-study log-fold-change estimates with their standard errors to yield confidence intervals on the consensus effect
-
Proportional abundances and IQRs were reported without confidence intervals throughout↳ Could also: Bootstrap or exact binomial confidence intervals on proportions, and bootstrap confidence intervals on IQRs, could also have been reported — Confidence intervals convey both the magnitude of each estimate and its uncertainty, which is particularly informative when summarizing heterogeneous multi-study data where study-level sample sizes differ substantially
-
miRNA expression was normalized to the median value across all samples prior to identifying highly expressed miRNAs↳ Could also: TMM (trimmed mean of M-values), upper-quartile normalization, or quantile normalization could also have been applied at the count level before proportional summarization — Composition-aware normalization methods explicitly correct for differences in library composition — a known concern when aggregating libraries prepared with different kits across 83 studies — and are standard in multi-sample small RNA-seq workflows
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
This is a P16 third-party-tool reproduction: the paper's own deterministic tool findadapt was re-run and reproduced the repo's documented NEXTflex demo output essentially 1:1 (seed TGGAATTCTCGG and 4/4 random-base phase exact; ratios within 0.0014; read count 1162 vs 1177, a -1.3% diff cleanly attributable to README-vs-pinned-commit version drift), auto-called 3/3 demo kits correctly, and recovered the expected TruSeq adapter on the paper's own GSE100467 EV data. Every value is derivable from shared/public data — no fabrication signal (q5 green). The deviations are negligible and technical (q3/q4/q6 green). However, the comparison is to repo-documented output and known kit chemistries rather than to a paper-printed scalar, and the paper's actual headline EV small-RNA profiling tables were not attempted, so the central conclusion is only partially validated (q7 yellow) and the overall reproduction, while clean, is solid-with-caveats rather than a 1:1 of the published claims (q8 yellow).
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.