Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← Dataset search

PRJEB57848

BioProject

Provenance — who produced it, who reused it

Linked to 0 papers in the literature. Roles are inferred factual signals (who deposited the data vs who reused it), with counts — never a judgement about any author.

No linked papers found in the corpus yet.

Deep data QC

59/100 · F

Standardized, field-standard QC computed by touching the data — every metric states how it was obtained · evidence: measured

What this means
claude:haiku

Amplicon sequencing from sorghum. Grade F (59/100) is driven by Q30 bases at 77.7% but more critically by catastrophic duplication at 75.68%, indicating severe PCR over-amplification or contamination. At this duplication level, biological signal is obscured; reuse is not viable without independent re-sequencing.

Data type / assay
amplicon
Organism
Sorghum bicolor
Instrument
Illumina MiSeq
Platform
ILLUMINA
N numbers (samples, groups)
120 runs
Metrics (value · how obtained)
checksum ok yes reported
total bases 2014474105 reported
total reads 6696294 reported
n content pct 0 measured
pct q20 bases 89.3 measured
pct q30 bases 77.7 measured
gc content pct 54.3 measured
mean read length 301 measured
mean base quality 32.8 measured
adapter content pct 0 measured
duplication rate pct 75.68 measured
How this grade was computed
Weighted mean of 2 scored metric(s) → 59/100

The F grade is a transparent weighted average. Each metric below scored from 0–100% against the published amplicon thresholds, weighted by its importance; nothing is hidden or subjective.

pct q30 bases 77.7 measured ×1 39%
adapter content pct 0 measured ×0.5 100%
QC cost 1 s compute

measured = computed from the data · extrapolated/reported = derived or from the repository · dq-1.0

Scientific quality

Based on hands-on reproduction of the papers that use this dataset. A reproducible paper that stands on this data is positive evidence; a flagged one is a prompt to look closer — never a verdict on the dataset itself without the evidence.

1 studies use it mean score 62