Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← Dataset search

SRX326766

SRA

Provenance — who produced it, who reused it

Linked to 0 papers in the literature. Roles are inferred factual signals (who deposited the data vs who reused it), with counts — never a judgement about any author.

No linked papers found in the corpus yet.

Deep data QC

80/100 · B

Standardized, field-standard QC computed by touching the data — every metric states how it was obtained · evidence: measured

What this means
claude:opus

This is a deep whole-genome shotgun (WGS) Illumina HiSeq 2000 dataset for the Asian longhorned beetle (Anoplophora glabripennis), comprising ~206M 101-bp reads (~41.7 Gb), and it earns a solid grade B (80/100): broadly reusable but not pristine. The grade is held back chiefly by pct_q30_bases at 83.1% — meaning roughly one in six bases falls below Q30, which modestly erodes confidence at read ends and will translate into more base-calling noise during variant calling or assembly, though the Q20 rate of 91.9% and mean base quality of Q33.4 indicate the bulk of the data is reliable. On the positive side, near-zero adapter content (0.97%) and a moderate, manageable duplication rate (10.6%) mean little library or trimming cleanup is needed before use, and low N-content and a stable GC of 37.1% suggest no gross contamination or compositional artifacts. Note that the evidence_strength flag is low because the headline yield figures (total reads and bases, checksum) are reported rather than independently measured, so treat the depth/throughput claims as provisional pending a full measured pass, even though the core per-base quality metrics here were directly measured and can be trusted.

Data type / assay
WGS
Organism
Anoplophora glabripennis
Instrument
Illumina HiSeq 2000
Platform
ILLUMINA
N numbers (samples, groups)
1 runs
Metrics (value · how obtained)
checksum ok yes reported
total bases 41694717182 reported
total reads 206409491 reported
n content pct 0.01 measured
pct q20 bases 91.9 measured
pct q30 bases 83.1 measured
gc content pct 37.1 measured
mean read length 101 measured
mean base quality 33.4 measured
adapter content pct 0.97 measured
duplication rate pct 10.6 measured
How this grade was computed
Weighted mean of 3 scored metric(s) → 80/100

The B grade is a transparent weighted average. Each metric below scored from 0–100% against the published WGS thresholds, weighted by its importance; nothing is hidden or subjective.

pct q30 bases 83.1 measured ×1 66%
duplication rate pct 10.6 measured ×0.5 92%
adapter content pct 0.97 measured ×0.4 100%
QC cost 14 s compute

measured = computed from the data · extrapolated/reported = derived or from the repository · dq-1.0

Scientific quality

Based on hands-on reproduction of the papers that use this dataset. A reproducible paper that stands on this data is positive evidence; a flagged one is a prompt to look closer — never a verdict on the dataset itself without the evidence.

1 studies use it 1 reproduced mean score 84