Learning the regulatory grammar of yeast 5' untranslated regions from a large library of random sequences
Our ability to predict protein expression from DNA sequence alone remains poor, reflecting our limited understanding of cis-regulatory grammar and hampering the design of engineered genes for synthetic biology applications. Here, we generate a model that predicts the translational efficiency of the 5’ untranslated region (UTR) of mRNAs in the yeast Saccharomyces cerevisiae. We constructed a library of half a million 50-nucleotide- long random 5’ UTRs and assayed their activity in a massively par...
Provenance — who produced it, who reused it
Linked to 2 papers in the literature. Roles are inferred factual signals (who deposited the data vs who reused it), with counts — never a judgement about any author.
1 further paper cites this accession but reuse could not be confirmed.
Deep data QC
metadata only · no data-level QC for this typeStandardized, field-standard QC computed by touching the data — every metric states how it was obtained
No quantitative QC rubric exists for this data type yet, so it is deliberately left unscored — this is an honest "not applicable", not a poor rating.
measured = computed from the data · extrapolated/reported = derived or from the repository · dq-1.0 · provisional — verify independently