Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
Ensembl 2007.
PMID 17148474 · PMC1761443 · Nucleic acids research · 2007 · 8 claims · 7 setups
Ensembl added 18 new chordate genomes this year, increasing total genomes available from 15 to 33, the largest yearly increase to date.
-
Full-text index only
Systematic identification of pseudogenes through whole genome expression evidence profiling.
PMID 16945953 · PMC1636364 · Nucleic acids research · 2006 · 8 claims · 8 setups
Developed a novel bioinformatics method that identifies pseudogenes by profiling whole-genome transcript and protein expression evidence
-
Full-text index only
Ensembl 2008.
PMID 18000006 · PMC2238821 · Nucleic acids research · 2008 · 8 claims · 6 setups
The Ensembl regulatory build integrates multiple genome-wide functional genomics datasets to automatically annotate regulatory regions and assign putative functions across the genome.
-
Full-text index only
Ensembl 2009.
PMID 19033362 · PMC2686571 · Nucleic acids research · 2009 · 8 claims · 6 setups
Ensembl provides comprehensive, consistently annotated genome information for chordate genomes with automatically generated genesets and comparative genomics data
-
Full-text index only
Pairagon+N-SCAN_EST: a model-based gene annotation pipeline.
PMID 16925839 · PMC1810554 · Genome biology · 2006 · 7 claims · 5 setups
Pairagon+N-SCAN_EST, using only native alignments, was as accurate as ENSEMBL and ExoGean in the EGASP mRNA/EST evidence assessment
-
Full-text index only
Ensembl 2006.
PMID 16381931 · PMC1347495 · Nucleic acids research · 2006 · 8 claims · 5 setups
Ensembl now provides annotation for 19 genomes, up from 4 the previous year, including new mammalian (Rhesus macaque, Opossum), chordate (Ciona intestinalis), and yeast genomes.
-
Full-text index only
Vertebrate gene finding from multiple-species alignments using a two-level strategy.
PMID 16925840 · PMC1810555 · Genome biology · 2006 · 8 claims · 5 setups
DOGFISH cleanly separates a multi-species alignment classifier (RVM cascade) from an HMM-based structure predictor, avoiding tight coupling of alignment complexity with HMM formalism
-
Full-text index only
U7 snRNAs: a computational survey.
PMID 18267300 · PMC5054213 · Genomics, proteomics & bioinformatics · 2007 · 8 claims · 6 setups
A computational (BLAST-based) survey identified bona fide U7 snRNA genes with characteristic upstream promoter elements (PSE) across most vertebrate genomes examined, plus numerous pseudogenes.
-
Full-text index only
Ensembl's 10th year.
PMID 19906699 · PMC2808936 · Nucleic acids research · 2010 · 8 claims · 8 setups
Ensembl provides comprehensive gene annotation and integrated genomic resources (variation, regulation, comparative genomics) across a growing set of chordate genomes
-
Full-text index only
Ensembl 2005.
PMID 15608235 · PMC540092 · Nucleic acids research · 2005 · 8 claims · 4 setups
Ensembl's automatic gene build system can flexibly and reliably annotate a wide variety of genomes with limited species-specific evidence.
-
Full-text index only
Identification, characterization and comparative genomics of chimpanzee endogenous retroviruses.
PMID 16805923 · PMC1779541 · Genome biology · 2006 · 8 claims · 6 setups
The chimpanzee genome contains at least 42 separate families of endogenous retroviruses, 9 newly identified
-
Full-text index only
Predicting deleterious nsSNPs: an analysis of sequence and structural attributes.
PMID 16630345 · PMC1489951 · BMC bioinformatics · 2006 · 8 claims · 7 setups
Sequence conservation (PSIC score difference) at the nsSNP position is the single most useful attribute for predicting deleterious vs neutral status.
-
Full-text index only
Designating eukaryotic orthology via processed transcription units.
PMID 18445630 · PMC2425467 · Nucleic acids research · 2008 · 8 claims · 5 setups
Existing ortholog databases discard/ignore alternative splicing via all-against-all protein comparisons, causing ambiguous ortholog calls and misclassification of AS isoforms as in-paralogs
-
Full-text index only
The other side of comparative genomics: genes with no orthologs between the cow and other mammalian species.
PMID 20003425 · PMC2808326 · BMC genomics · 2009 · 7 claims · 4 setups
3,801 bovine genes have no orthologs in human, mouse and dog, and 1,010 human genes have no orthologs in cow despite having orthologs in mouse and dog
-
Full-text index only
SVC: structured visualization of evolutionary sequence conservation.
PMID 15991338 · PMC1160265 · Nucleic acids research · 2005 · 7 claims · 5 setups
SVC aligns protein-coding sequences of orthologous gene pairs and maps them back onto their encoding exons/introns to generate a scaffold of conserved gene structure.
-
Full-text index only
Evola: Ortholog database of all human genes in H-InvDB with manual curation of phylogenetic trees.
PMID 17982176 · PMC2238928 · Nucleic acids research · 2008 · 6 claims · 7 setups
Evola combines genome synteny-based computational ortholog detection with manual curation of phylogenetic trees by experts to yield more reliable orthologs than automated pairwise methods
-
Full-text index only
Sequence occurrence and structural uniqueness of a G-quadruplex in the human c-kit promoter.
PMID 17720713 · PMC2034477 · Nucleic acids research · 2007 · 8 claims · 4 setups
The native 22-nt c-kit87 sequence occurs only once in the entire human genome.
-
Full-text index only
A genome-wide survey demonstrates widespread non-linear mRNA in expressed sequences from multiple species.
PMID 16237125 · PMC1258171 · Nucleic acids research · 2005 · 8 claims · 6 setups
A genome-wide computational survey identifies 245 genes in mammals (264 across six species) that produce RREO events in expressed sequences
-
Full-text index only
Large-scale analysis of human alternative protein isoforms: pattern classification and correlation with subcellular localization signals.
PMID 15860772 · PMC1087780 · Nucleic acids research · 2005 · 8 claims · 8 setups
Constructed a large-scale dataset of 6876 human alternative protein isoforms from 2624 genes by combining H-Invitational full-length cDNA data and SwissProt VARSPLIC entries
-
Full-text index only
Inparanoid: a comprehensive database of eukaryotic orthologs.
PMID 15608241 · PMC540061 · Nucleic acids research · 2005 · 8 claims · 4 setups
The Inparanoid algorithm identifies true ortholog clusters by seeding on reciprocal best-matching pairs, gathering inparalogs (post-speciation duplicates) while excluding outparalogs (pre-speciation duplicates)