Provenance — who produced it, who reused it
Linked to 3 papers in the literature. Roles are inferred factual signals (who deposited the data vs who reused it), with counts — never a judgement about any author.
3 further papers cite this accession but reuse could not be confirmed.
Deep data QC
94/100 · AStandardized, field-standard QC computed by touching the data — every metric states how it was obtained · evidence: measured
This is whole-genome shotgun sequencing of the bacterium Agrilactobacillus composti, generated on an Illumina HiSeq 2000 at short read length (101 bp). In QC terms it is a strong, reusable dataset: an A grade (94/100) is justified by excellent base-call accuracy — 91.3% of bases at Q30 and a mean base quality of 35.3 — together with negligible adapter (0.12%) and N content (0.001%), meaning reads can be confidently mapped or assembled with minimal trimming. The only real drag on the score is a duplication rate of 15.24%, which capped that metric at 77/100; this matters for reuse because PCR/optical duplicates inflate apparent coverage and can bias variant-allele fractions, so duplicate marking before any quantitative analysis is advisable. Note that the headline volume figures (total bases and reads) are reported rather than independently measured, but since every quality-determining metric here was actually measured (and most carry full weight), the grade rests on solid empirical footing rather than extrapolation.
The A grade is a transparent weighted average. Each metric below scored from 0–100% against the published WGS thresholds, weighted by its importance; nothing is hidden or subjective.
measured = computed from the data · extrapolated/reported = derived or from the repository · dq-1.0
Scientific quality
Based on hands-on reproduction of the papers that use this dataset. A reproducible paper that stands on this data is positive evidence; a flagged one is a prompt to look closer — never a verdict on the dataset itself without the evidence.