← Dataset search
ERR1437992
ENAProvenance — who produced it, who reused it
Linked to 0 papers in the literature. Roles are inferred factual signals (who deposited the data vs who reused it), with counts — never a judgement about any author.
No linked papers found in the corpus yet.
Deep data QC
80/100 · BStandardized, field-standard QC computed by touching the data — every metric states how it was obtained · evidence: measured
Data type / assay
WGS
Organism
Streptococcus pneumoniae
Instrument
Illumina HiSeq 2500
Platform
ILLUMINA
Read type
short-read
Files available
FASTQ (raw reads)
N numbers (samples, groups)
1 runs
Metrics (value · how obtained)
gc sd
6.51
measured
checksum ok
yes
reported
total bases
94611744
reported
total reads
488681
reported
n content pct
0
measured
pct q20 bases
99.9
measured
pct q30 bases
98
measured
pct reads q30
100
measured
sampled bases
47645800
measured
sampled reads
488681
measured
gc content pct
40
measured
polyg tail pct
0
measured
read length sd
8.3
measured
quality dropoff
3
measured
read length max
100
measured
read length min
50
measured
read length n50
100
measured
max base quality
40
measured
mean read length
97.5
measured
max n pct per pos
0.002
measured
mean base quality
37.3
measured
pct reads lt 100bp
16.04
measured
read length median
100
measured
adapter content pct
4.67
measured
median read quality
37.6
measured
duplication rate pct
25.03
measured
overrepresented top pct
0.02
measured
How this grade was computed
Weighted mean of 3 scored metric(s) → 80/100
The B grade is a transparent weighted average. Each metric below scored from 0–100% against the published WGS thresholds, weighted by its importance; nothing is hidden or subjective.
pct q30 bases
98
measured
×1
100%
duplication rate pct
25.03
measured
×0.5
47%
adapter content pct
4.67
measured
×0.4
74%
QC cost
11 s compute
measured = computed from the data · extrapolated/reported = derived or from the repository · dq-1.0