Corpus 1,273 assessed · 1,174 scored · 643 reproduced ≥75 · 169 flagged ·∅ 74.1/100
← Dataset search

GSE125218

GEO first seen 2019

Accurate annotation of human protein-coding small open reading frames

Organism
Homo sapiens
Samples
33
Type
Expression profiling by high...
Submitted
2019-01-16

Protein-coding small open reading frames (smORFs) are emerging as an important class of genes, however, the coding capacity of smORFs in the human genome is unclear. By integrating de novo transcriptome assembly and Ribo-Seq, we confidently annotate thousands of novel translated smORFs in three human cell lines. We find that smORF translation prediction is noisier than for annotated coding sequences, underscoring the importance of analyzing multiple experiments and footprinting conditions. These...

Provenance — who produced it, who reused it

Linked to 5 papers in the literature. Roles are inferred factual signals (who deposited the data vs who reused it), with counts — never a judgement about any author.

Deposited / produced by
Thomas F MartinezAlan SaghatelianMaxim N Shokhirev
Reused by

4 further papers cite this accession but reuse could not be confirmed.

Deep data QC

metadata only · no data-level QC for this type

Standardized, field-standard QC computed by touching the data — every metric states how it was obtained

Data type / assay
other
Organism
Homo sapiens

No quantitative QC rubric exists for this data type yet, so it is deliberately left unscored — this is an honest "not applicable", not a poor rating.

QC cost 20 s compute

measured = computed from the data · extrapolated/reported = derived or from the repository · dq-1.0 · provisional — verify independently