Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← Dataset search

SRR2174513

ENA first seen 2018

Provenance — who produced it, who reused it

Linked to 1 papers in the literature. Roles are inferred factual signals (who deposited the data vs who reused it), with counts — never a judgement about any author.

Reused by

1 further paper cites this accession but reuse could not be confirmed.

Deep data QC

83/100 · B

Standardized, field-standard QC computed by touching the data — every metric states how it was obtained · evidence: measured

What this means
claude:opus

This is a high-depth, miRNA-targeted small-RNA library (labeled bulk RNA-seq, miRNA-Seq on a HiSeq 2500) for which essentially all quality signals were directly measured, giving a confident reading (evidence_strength=1, so this is not provisional). The verdict is a solid B: base-call quality is the clear strength — 93.2% of bases at Q30 and a mean base quality of 36.7 mean the reads themselves are clean and reliably called, with zero adapter contamination and negligible N content, so trimming and mapping should be unproblematic. The one metric dragging the grade down is a 91.24% duplication rate, which scored 0; for most assays that would scream PCR/optical over-amplification, but for a small-RNA/miRNA library it is partly expected because the transcriptome is low-complexity (a small number of highly abundant mature miRNAs sequenced deeply at 51 bp). Practically, reuse the data with that caveat in mind: the per-base quality is trustworthy, but treat raw read counts cautiously and rely on UMI-style or post-collapse deduplication for any quantitative miRNA-abundance analysis, since the extreme duplication limits confidence in true molecular complexity.

Data type / assay
bulk-RNA-seq
Organism
Homo sapiens
Instrument
Illumina HiSeq 2500
Platform
ILLUMINA
N numbers (samples, groups)
1 runs
Metrics (value · how obtained)
checksum ok yes reported
total bases 775519566 reported
total reads 15206266 reported
n content pct 0.007 measured
pct q20 bases 96.8 measured
pct q30 bases 93.2 measured
gc content pct 46.6 measured
mean read length 51 measured
mean base quality 36.7 measured
adapter content pct 0 measured
duplication rate pct 91.24 measured
How this grade was computed
Weighted mean of 4 scored metric(s) → 83/100

The B grade is a transparent weighted average. Each metric below scored from 0–100% against the published bulk-RNA-seq thresholds, weighted by its importance; nothing is hidden or subjective.

pct q30 bases 93.2 measured ×1 100%
mean base quality 36.7 measured ×0.6 100%
adapter content pct 0 measured ×0.4 100%
duplication rate pct 91.24 measured ×0.4 0%
QC cost 20 s compute

measured = computed from the data · extrapolated/reported = derived or from the repository · dq-1.0

Scientific quality

Based on hands-on reproduction of the papers that use this dataset. A reproducible paper that stands on this data is positive evidence; a flagged one is a prompt to look closer — never a verdict on the dataset itself without the evidence.

1 studies use it 1 reproduced mean score 89