PRJNA477449
BioProjectProvenance — who produced it, who reused it
Linked to 0 papers in the literature. Roles are inferred factual signals (who deposited the data vs who reused it), with counts — never a judgement about any author.
No linked papers found in the corpus yet.
Deep data QC
100/100 · AStandardized, field-standard QC computed by touching the data — every metric states how it was obtained · evidence: measured
This is a human bulk RNA-seq dataset (Illumina NextSeq 500, ~811M reads / 61 Gb) that earns a clean A: every metric feeding the score was directly measured (evidence_strength = 1.0), so the grade is fully substantiated rather than provisional. The grade is driven chiefly by base-call quality — 93.5% of bases at Q30 and a mean base quality of 34.3 — which means individual base calls are highly reliable and should support sensitive, low-error alignment and expression quantification, while a measured 0% adapter content indicates the reads are clean and need no further trimming before reuse. The one metric worth a researcher's eye is the 22.47% duplication rate; this did not lower the grade and is unremarkable for RNA-seq, where highly expressed transcripts naturally generate duplicate fragments, but for any analysis sensitive to library complexity (e.g., low-input or allele-specific work) it is worth confirming whether duplicates are biological rather than PCR artifacts. Overall these are trustworthy, high-quality data suitable for standard differential-expression and transcript-quantification reuse.
The A grade is a transparent weighted average. Each metric below scored from 0–100% against the published bulk-RNA-seq thresholds, weighted by its importance; nothing is hidden or subjective.
measured = computed from the data · extrapolated/reported = derived or from the repository · dq-1.0
Scientific quality
Based on hands-on reproduction of the papers that use this dataset. A reproducible paper that stands on this data is positive evidence; a flagged one is a prompt to look closer — never a verdict on the dataset itself without the evidence.