Remapping the SRA: Drosophila melanogaster RNA-Seq data from the Sequence Read Archive
The sequence read archive (SRA) contains over 52 terabases or 482 billion reads from Drosophila melanogaster (as of June 2018). These data are massively underused by the community and include 14,423 RNA-Seq samples, that is roughly 7 times the size of modENCODE. Currently the major challenge is finding high quality datasets that are suitable for inclusion in new studies. To help the community overcome this hurdle, we re-processed all D. melanogaster RNA-Seq SRA experiments (SRXs) using an identi...
Provenance — who produced it, who reused it
Linked to 3 papers in the literature. Roles are inferred factual signals (who deposited the data vs who reused it), with counts — never a judgement about any author.
2 further papers cite this accession but reuse could not be confirmed.
Deep data QC
metadata only · no data-level QC for this typeStandardized, field-standard QC computed by touching the data — every metric states how it was obtained
No quantitative QC rubric exists for this data type yet, so it is deliberately left unscored — this is an honest "not applicable", not a poor rating.
measured = computed from the data · extrapolated/reported = derived or from the repository · dq-1.0 · provisional — verify independently