Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← Dataset search

SRX326768

SRA

Provenance — who produced it, who reused it

Linked to 0 papers in the literature. Roles are inferred factual signals (who deposited the data vs who reused it), with counts — never a judgement about any author.

No linked papers found in the corpus yet.

Deep data QC

97/100 · A

Standardized, field-standard QC computed by touching the data — every metric states how it was obtained · evidence: measured

What this means
claude:opus

This is whole-genome shotgun sequencing of the Asian longhorned beetle (Anoplophora glabripennis) on Illumina HiSeq 2000, and it earns a clean A (97/100): the reads are genuinely high quality, with 92.8% of bases at Q30 and a mean base quality of 36.4, meaning base calls are reliable enough to support confident variant calling and assembly. The metric that pulled the score down most is the 11.48% duplication rate, which is moderate and worth noting because PCR/optical duplicates inflate apparent coverage and can bias allele-frequency or coverage-based analyses unless flagged and removed; negligible adapter (0.05%) and N content (0.005%) mean little upfront trimming is needed. Critically, the headline counts — total_bases (~23.7 Gb) and total_reads — are reported rather than independently measured, but the QC-defining quality, duplication, adapter, and GC metrics were all actually measured, so the grade rests on real evidence rather than extrapolation. Overall this is a trustworthy dataset to reuse for ~100 bp paired WGS work; just deduplicate before any coverage- or frequency-sensitive step and treat the read/base totals as provider-stated until verified.

Data type / assay
WGS
Organism
Anoplophora glabripennis
Instrument
Illumina HiSeq 2000
Platform
ILLUMINA
N numbers (samples, groups)
1 runs
Metrics (value · how obtained)
checksum ok yes reported
total bases 23655093446 reported
total reads 117104423 reported
n content pct 0.005 measured
pct q20 bases 97.1 measured
pct q30 bases 92.8 measured
gc content pct 35.9 measured
mean read length 101 measured
mean base quality 36.4 measured
adapter content pct 0.05 measured
duplication rate pct 11.48 measured
How this grade was computed
Weighted mean of 3 scored metric(s) → 97/100

The A grade is a transparent weighted average. Each metric below scored from 0–100% against the published WGS thresholds, weighted by its importance; nothing is hidden or subjective.

pct q30 bases 92.8 measured ×1 100%
duplication rate pct 11.48 measured ×0.5 89%
adapter content pct 0.05 measured ×0.4 100%
QC cost 10 s compute

measured = computed from the data · extrapolated/reported = derived or from the repository · dq-1.0

Scientific quality

Based on hands-on reproduction of the papers that use this dataset. A reproducible paper that stands on this data is positive evidence; a flagged one is a prompt to look closer — never a verdict on the dataset itself without the evidence.

1 studies use it 1 reproduced mean score 84