Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
DBD--taxonomically broad transcription factor predictions: new content and functionality.
PMID 18073188 · PMC2238844 · Nucleic acids research · 2008 · 8 claims · 3 setups
DBD is a database of predicted sequence-specific DNA-binding transcription factors covering over 700 publicly available proteomes, up from 150 in the initial version.
-
Full-text index only
Protein coding potential of retroviruses and other transposable elements in vertebrate genomes.
PMID 15716312 · PMC549403 · Nucleic acids research · 2005 · 8 claims · 5 setups
About 1000 genes across four vertebrate gene sets analyzed contain at least one RETRA marker protein domain
-
Full-text index only
What makes species unique? The contribution of proteins with obscure features.
PMID 16859532 · PMC1779552 · Genome biology · 2006 · 7 claims · 8 setups
POFs constitute 18-38% (average 26%) of a typical eukaryotic proteome
-
Full-text index only
Coiled-coil protein composition of 22 proteomes--differences and common themes in subcellular infrastructure and traffic control.
PMID 16288662 · PMC1322226 · BMC evolutionary biology · 2005 · 7 claims · 5 setups
Proteins with extended coiled-coil domains (>250 amino acids) are largely absent from bacterial genomes but present in archaea and eukaryotes.
-
Full-text index only
SysZNF: the C2H2 zinc finger gene database.
PMID 18974185 · PMC2686507 · Nucleic acids research · 2009 · 7 claims · 6 setups
SysZNF is a database that systematically catalogs C2H2-ZNF genes in human and mouse with physical location, gene models, expression probes, protein domains, homologs, and literature links
-
Full-text index only
Comparative analysis of plant genomes allows the definition of the "Phytolongins": a novel non-SNARE longin domain protein family.
PMID 19889231 · PMC2779197 · BMC genomics · 2009 · 8 claims · 6 setups
A novel, plant-specific family of longin-related proteins, the 'Phytolongins', was identified in land plant genomes.
-
Full-text index only
Reconstructing the evolution of the mitochondrial ribosomal proteome.
PMID 17604309 · PMC1950548 · Nucleic acids research · 2007 · 8 claims · 6 setups
The ancestral mitoribosome was of alpha-proteobacterial descent and more than doubled its protein content in most eukaryotic lineages.
-
Full-text index only
pTARGET: a web server for predicting protein subcellular localization.
PMID 16844995 · PMC1538910 · Nucleic acids research · 2006 · 7 claims · 3 setups
pTARGET web server predicts nine distinct subcellular localizations in eukaryotic non-plant proteins using an algorithm based on location-specific Pfam domain occurrence patterns and amino acid composition (AAC)
-
Full-text index only
Evolutionary cores of domain co-occurrence networks.
PMID 15788102 · PMC1079808 · BMC evolutionary biology · 2005 · 8 claims · 4 setups
The innermost (globally central) cores of protein domain co-occurrence networks gradually grow in size with increasing evolutionary/developmental complexity of the organism.
-
Full-text index only
Protein length in eukaryotic and prokaryotic proteomes.
PMID 15951512 · PMC1150220 · Nucleic acids research · 2005 · 7 claims · 5 setups
Eukaryotic proteins are significantly longer than prokaryotic proteins across virtually all functional categories and the majority of protein families
-
Full-text index only
InParanoid 7: new algorithms and tools for eukaryotic orthology analysis.
PMID 19892828 · PMC2808972 · Nucleic acids research · 2010 · 8 claims · 7 setups
InParanoid 7 expands the database by an order of magnitude to 100 species, 1.3 million proteins, and 42.7 million pairwise ortholog groups.
-
Full-text index only
Molecular phylogeny of the kelch-repeat superfamily reveals an expansion of BTB/kelch proteins in animals.
PMID 13678422 · PMC222960 · BMC bioinformatics · 2003 · 8 claims · 8 setups
The human genome encodes at least 71 kelch-repeat proteins
-
Full-text index only
Tracing the origin of functional and conserved domains in the human proteome: implications for protein evolution at the modular level.
PMID 17090320 · PMC1654190 · BMC evolutionary biology · 2006 · 8 claims · 5 setups
HHpred (HMM-HMM comparison) detects remote homologs in the human proteome with higher sensitivity than hmmpfam (HMMER), giving 10% more functional domain coverage and 20% higher residue coverage against Pfam-A families.
-
Has reproduction · 76
Bayesian prediction of microbial oxygen requirement.
PMID 26913185 · PMC4743139 · F1000Research · 2013 · 7 claims · 8 setups
A naive Bayesian classifier based on presence/absence of class-associated Pfam-A domains can distinguish three oxygen requirement classes (aerobe, anaerobe, facultative anaerobe) from genome sequence, unlike prior studies that only made pairwise distinctions.
-
Full-text index only
The genome of the simian and human malaria parasite Plasmodium knowlesi.
PMID 18843368 · PMC2656934 · Nature · 2008 · 8 claims · 7 setups
The P. knowlesi (H strain) nuclear genome was sequenced and assembled: 23.5 Mb across 14 chromosomes with 5,188 predicted protein-encoding genes.
-
Full-text index only
The EH1 motif in metazoan transcription factors.
PMID 16309560 · PMC1310626 · BMC genomics · 2005 · 8 claims · 5 setups
There is a statistically significant association between EH1hox motif HMM score and transcription factor function across human, Drosophila and C. elegans proteomes.
-
Full-text index only
The Bifidobacterium dentium Bd1 genome sequence reflects its genetic adaptation to the human oral cavity.
PMID 20041198 · PMC2788695 · PLoS genetics · 2009 · 8 claims · 8 setups
The B. dentium Bd1 genome was sequenced to completion, revealing a single circular 2,636,368 bp chromosome with 2,143 predicted ORFs
-
Full-text index only
Anopheles gambiae genome reannotation through synthesis of ab initio and comparative gene prediction algorithms.
PMID 16569258 · PMC1557760 · Genome biology · 2006 · 8 claims · 7 setups
An exon-gene-union (EGU) algorithm followed by an open-reading-frame-selection algorithm can synthesize ab initio (GENSCAN, GeneMark, SNAP) and comparative (Ensembl/Genewise) predictions into a single, more complete CDS set
-
Full-text index only
The protein-phosphatome of the human malaria parasite Plasmodium falciparum.
PMID 18793411 · PMC2559854 · BMC genomics · 2008 · 8 claims · 8 setups
P. falciparum possesses 27 putative protein phosphatase sequences across the four major PP families (PPP, PPM, PTP, NIF), plus 7 additional sequences predicted to dephosphorylate non-protein substrates, totaling 34.
-
Full-text index only
An emerging cyberinfrastructure for biodefense pathogen and pathogen-host data.
PMID 17984082 · PMC2239001 · Nucleic acids research · 2008 · 8 claims · 7 setups
The Biodefense Proteomics Resource Center (RC) is a public cyberinfrastructure that stores, integrates, and disseminates experimental data from seven Proteomics Research Centers (PRCs) on biodefense pathogens and host interactions