Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← Dataset search

PRJNA222257

BioProject first seen 2016

Provenance — who produced it, who reused it

Linked to 3 papers in the literature. Roles are inferred factual signals (who deposited the data vs who reused it), with counts — never a judgement about any author.

Reused by

3 further papers cite this accession but reuse could not be confirmed.

Deep data QC

94/100 · A

Standardized, field-standard QC computed by touching the data — every metric states how it was obtained · evidence: measured

What this means
claude:opus

This is whole-genome shotgun sequencing of the bacterium Agrilactobacillus composti, generated on an Illumina HiSeq 2000 at short read length (101 bp). In QC terms it is a strong, reusable dataset: an A grade (94/100) is justified by excellent base-call accuracy — 91.3% of bases at Q30 and a mean base quality of 35.3 — together with negligible adapter (0.12%) and N content (0.001%), meaning reads can be confidently mapped or assembled with minimal trimming. The only real drag on the score is a duplication rate of 15.24%, which capped that metric at 77/100; this matters for reuse because PCR/optical duplicates inflate apparent coverage and can bias variant-allele fractions, so duplicate marking before any quantitative analysis is advisable. Note that the headline volume figures (total bases and reads) are reported rather than independently measured, but since every quality-determining metric here was actually measured (and most carry full weight), the grade rests on solid empirical footing rather than extrapolation.

Data type / assay
WGS
Organism
Agrilactobacillus composti DSM 18527 = JCM 14202
Instrument
Illumina HiSeq 2000
Platform
ILLUMINA
N numbers (samples, groups)
162 runs
Metrics (value · how obtained)
checksum ok yes reported
total bases 121475023354 reported
total reads 585536527 reported
n content pct 0.001 measured
pct q20 bases 96.9 measured
pct q30 bases 91.3 measured
gc content pct 52.4 measured
mean read length 101 measured
mean base quality 35.3 measured
adapter content pct 0.12 measured
duplication rate pct 15.24 measured
How this grade was computed
Weighted mean of 3 scored metric(s) → 94/100

The A grade is a transparent weighted average. Each metric below scored from 0–100% against the published WGS thresholds, weighted by its importance; nothing is hidden or subjective.

pct q30 bases 91.3 measured ×1 100%
duplication rate pct 15.24 measured ×0.5 77%
adapter content pct 0.12 measured ×0.4 100%
QC cost 14 s compute

measured = computed from the data · extrapolated/reported = derived or from the repository · dq-1.0

Scientific quality

Based on hands-on reproduction of the papers that use this dataset. A reproducible paper that stands on this data is positive evidence; a flagged one is a prompt to look closer — never a verdict on the dataset itself without the evidence.

1 studies use it 1 reproduced mean score 89