Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
EGASP: the human ENCODE Genome Annotation Assessment Project.
PMID 16925836 · PMC1810551 · Genome biology · 2006 · 8 claims · 6 setups
Best-performing computational gene prediction methods correctly predict at least one transcript for close to 70% of annotated genes in the ENCODE regions.
-
Full-text index only
Identification of common genetic variation that modulates alternative splicing.
PMID 17571926 · PMC1904363 · PLoS genetics · 2007 · 7 claims · 8 setups
Common SNPs located close to intron-exon boundaries are associated with and causally modulate alternative splicing patterns in human genes
-
Full-text index only
Brain progranulin expression in GRN-associated frontotemporal lobar degeneration.
PMID 19649643 · PMC3104467 · Acta neuropathologica · 2010 · 8 claims · 8 setups
GRN transcript haploinsufficiency, previously shown in blood-derived cells, does not hold in most brain regions of GRN mutation carriers
-
Full-text index only
Assessing the genomic evidence for conserved transcribed pseudogenes under selection.
PMID 19754956 · PMC2753554 · BMC genomics · 2009 · 8 claims · 8 setups
1750 transcribed pseudogene annotations (TPAs) were identified in the human genome, ~11.5% of all human pseudogene annotations.
-
Full-text index only
Large-scale analysis of Macaca fascicularis transcripts and inference of genetic divergence between M. fascicularis and M. mulatta.
PMID 18294402 · PMC2287170 · BMC genomics · 2008 · 8 claims · 6 setups
Constructed full-length-enriched cDNA libraries and determined 85,721 EST sequences and 9407 full-insert sequences from cynomolgus macaque brain (7 regions), testis, and liver
-
Full-text index only
The ENCODE Project at UC Santa Cruz.
PMID 17166863 · PMC1781110 · Nucleic acids research · 2007 · 8 claims · 4 setups
The UCSC ENCODE portal serves as the primary repository and access point for sequence-based ENCODE pilot phase data
-
Has reproduction · 85
Expansion of the SOS regulon of Vibrio cholerae through extensive transcriptome analysis and experimental validation.
PMID 29783948 · PMC5963079 · BMC genomics · 2018 · 8 claims · 8 setups
Whole transcriptome sequencing with extensive TSS mapping identified 3078 transcription start sites and 629 ncRNAs in V. cholerae N16961
-
Full-text index only
Pairagon+N-SCAN_EST: a model-based gene annotation pipeline.
PMID 16925839 · PMC1810554 · Genome biology · 2006 · 7 claims · 5 setups
Pairagon+N-SCAN_EST, using only native alignments, was as accurate as ENSEMBL and ExoGean in the EGASP mRNA/EST evidence assessment
-
Full-text index only
A third approach to gene prediction suggests thousands of additional human transcribed regions.
PMID 16543943 · PMC1391917 · PLoS computational biology · 2006 · 8 claims · 7 setups
A third basic concept for gene prediction exists, based on detecting strand-specific 'transcription footprints' (mutational and selectional biases) rather than gene structure or sequence similarity.
-
Full-text index only
Prediction-based approaches to characterize bidirectional promoters in the mammalian genome.
PMID 18366609 · PMC2386062 · BMC genomics · 2008 · 8 claims · 7 setups
The mapping algorithm identified 5,647 candidate bidirectional promoter regions in the mouse genome, similar in number to those previously found in human.
-
Full-text index only
Using several pair-wise informant sequences for de novo prediction of alternatively spliced transcripts.
PMID 16925842 · PMC1810557 · Genome biology · 2006 · 8 claims · 4 setups
MARS, an extension of the Twinscan algorithm, uses multiple pairwise informant genomes to predict human alternatively spliced transcripts de novo without expressed sequence information.
-
Full-text index only
Exogean: a framework for annotating protein-coding genes in eukaryotic genomic DNA.
PMID 16925841 · PMC1810556 · Genome biology · 2006 · 8 claims · 5 setups
Exogean is a framework using directed acyclic coloured multigraphs (DACMs) to represent biological objects (mRNA, ESTs, protein alignments, exons) and iteratively combine them into complex protein-coding transcript models.
-
Full-text index only
Systems biology of gene regulation fulfills its promise.
PMID 16719937 · PMC1779525 · Genome biology · 2006 · 8 claims · 8 setups
Suz12, a Polycomb Group complex component, has DNA targets identifiable by ChIP-chip and can silence large genomic regions in a cell-type-specific manner.
-
Full-text index only
SelenoDB 1.0 : a database of selenoprotein genes, proteins and SECIS elements.
PMID 18174224 · PMC2238826 · Nucleic acids research · 2008 · 6 claims · 5 setups
Standard genome annotation pipelines misannotate selenoprotein genes because they rely on UGA as a universal stop codon, failing to recognize its dual role as the selenocysteine-recoding codon.
-
Full-text index only
Novel CLCN1 mutations and clinical features of Korean patients with myotonia congenita.
PMID 19949657 · PMC2775849 · Journal of Korean medical science · 2009 · 7 claims · 8 setups
Sequencing of CLCN1 in 10 unrelated Korean MC patients identified nine different point mutations, six of which are novel (p.M128I, p.S189C, p.M373L, p.P480S, p.G523D, p.M609K).
-
Has reproduction · 50
Dynamics and regulation of mitotic chromatin accessibility bookmarking at single-cell resolution.
PMID 36696508 · PMC9876548 · Science advances · 2023 · 7 claims · 8 setups
Chromatin accessibility continually decreases from mitotic entry until metaphase, then gradually increases as chromosomes segregate.
-
Has reproduction · 83
MetaGT: A pipeline for de novo assembly of metatranscriptomes with the aid of metagenomic data.
PMID 36386613 · PMC9651917 · Frontiers in microbiology · 2022 · 7 claims · 4 setups
MetaGT is a pipeline that combines metatranscriptomic and metagenomic data from the same sample to assemble complete transcript sequences
-
Full-text index only
Satellog: a database for the identification and prioritization of satellite repeats in disease association studies.
PMID 15949044 · PMC1181805 · BMC bioinformatics · 2005 · 7 claims · 6 setups
Satellog is a database cataloging all pure 1-16 unit satellite repeats in the human genome with supplementary polymorphism, gene-location, and expression data for prioritizing repeats in disease-association studies.
-
Full-text index only
GENCODE: producing a reference annotation for ENCODE.
PMID 16925838 · PMC1810553 · Genome biology · 2006 · 8 claims · 8 setups
GENCODE annotation combines initial manual annotation by HAVANA, experimental validation, and refinement based on results to identify protein-coding genes in ENCODE regions
-
Full-text index only
Evola: Ortholog database of all human genes in H-InvDB with manual curation of phylogenetic trees.
PMID 17982176 · PMC2238928 · Nucleic acids research · 2008 · 6 claims · 7 setups
Evola combines genome synteny-based computational ortholog detection with manual curation of phylogenetic trees by experts to yield more reliable orthologs than automated pairwise methods