Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← Dataset search

PRJDB547

BioProject first seen 2016

Provenance — who produced it, who reused it

Linked to 1 papers in the literature. Roles are inferred factual signals (who deposited the data vs who reused it), with counts — never a judgement about any author.

Reused by

1 further paper cites this accession but reuse could not be confirmed.

Deep data QC

47/100 · F

Standardized, field-standard QC computed by touching the data — every metric states how it was obtained · evidence: measured

What this means
claude:opus

This is a whole-genome shotgun dataset for the cellulolytic anaerobe Acetivibrio straminisolvens, sequenced on the Ion Torrent PGM, and despite ample raw volume (~48.8 Gb across ~194.5M reads) it earns a failing grade driven overwhelmingly by base-call quality. The single decisive metric is pct_q30_bases at 49.2%, which scored 0/100 at the highest weight: fewer than half the bases meet the Q30 (≥99.9% accuracy) threshold, and a mean base quality of 27 with only 81.9% of bases reaching Q20 confirms an elevated per-base error rate that will inflate false-positive variant calls and fragment assemblies — a well-known Ion Torrent weakness in homopolymer-rich regions. The few things working in its favor are clean library prep (0% adapter content) and a low 7.07% duplication rate, but these are minor weights that cannot offset the quality deficit, and the 57.8% GC with 0% N-content suggests no gross compositional artifact. Note that the core quality metrics here are flagged as measured rather than extrapolated, so this read is reliable on its own terms; the grade is not provisional, and reuse should be limited to applications tolerant of high per-base error or paired with aggressive quality trimming.

Data type / assay
WGS
Organism
Acetivibrio straminisolvens JCM 21531
Instrument
Ion Torrent PGM
Platform
ION_TORRENT
N numbers (samples, groups)
360 runs
Metrics (value · how obtained)
checksum ok yes reported
total bases 48845782778 reported
total reads 194508305 reported
n content pct 0 measured
pct q20 bases 81.9 measured
pct q30 bases 49.2 measured
gc content pct 57.8 measured
mean read length 197.5 measured
mean base quality 27 measured
adapter content pct 0 measured
duplication rate pct 7.07 measured
How this grade was computed
Weighted mean of 3 scored metric(s) → 47/100

The F grade is a transparent weighted average. Each metric below scored from 0–100% against the published WGS thresholds, weighted by its importance; nothing is hidden or subjective.

pct q30 bases 49.2 measured ×1 0%
duplication rate pct 7.07 measured ×0.5 100%
adapter content pct 0 measured ×0.4 100%
QC cost 10 s compute

measured = computed from the data · extrapolated/reported = derived or from the repository · dq-1.0

Scientific quality

Based on hands-on reproduction of the papers that use this dataset. A reproducible paper that stands on this data is positive evidence; a flagged one is a prompt to look closer — never a verdict on the dataset itself without the evidence.

1 studies use it 1 reproduced mean score 89