Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
The Universal Protein Resource (UniProt) in 2010.
PMID 19843607 · PMC2808944 · Nucleic acids research · 2010 · 8 claims · 5 setups
UniProt is a centralized, freely accessible, comprehensive knowledgebase of protein sequence and functional annotation maintained by the EBI, SIB and PIR consortium.
-
Full-text index only
MutDB: update on development of tools for the biochemical analysis of genetic variation.
PMID 17827212 · PMC2238958 · Nucleic acids research · 2008 · 7 claims · 5 setups
MutDB integrates dbSNP and Swiss-Prot genetic variation data with protein structural information, functional disruption prediction scores, and clinical phenotype links (OMIM, dbGAP)
-
Full-text index only
SNAP: predict effect of non-synonymous polymorphisms on function.
PMID 17526529 · PMC1920242 · Nucleic acids research · 2007 · 7 claims · 8 setups
SNAP, a neural network-based method using sequence-derived information, predicts whether a non-synonymous SNP is neutral or non-neutral for protein function
-
Full-text index only
Towards a comprehensive structural coverage of completed genomes: a structural genomics viewpoint.
PMID 17349043 · PMC1829165 · BMC bioinformatics · 2007 · 8 claims · 6 setups
A combined target-selection approach — pursuing both structurally uncharacterised domain families and additional targets from large structurally characterised superfamilies — is essential for comprehensive structural coverage of the genomes.
-
Full-text index only
Mapping proteins to disease terminologies: from UniProt to MeSH.
PMID 18460185 · PMC2367626 · BMC bioinformatics · 2008 · 8 claims · 7 setups
Developed a three-step procedure (disease name extraction, exact matching, partial/similarity-based matching) to map UniProtKB/Swiss-Prot disease names to MeSH terms
-
Full-text index only
Genome reannotation of Escherichia coli CFT073 with new insights into virulence.
PMID 19930606 · PMC2785843 · BMC genomics · 2009 · 8 claims · 7 setups
Reannotation excluded 608 CDSs from the original RefSeq annotation, mostly unfunctional 'hypothetical'/'putative' genes
-
Full-text index only
PA-GOSUB: a searchable database of model organism protein sequences with their predicted Gene Ontology molecular function and subcellular localization.
PMID 15608166 · PMC540074 · Nucleic acids research · 2005 · 7 claims · 4 setups
PA-GOSUB significantly extends the coverage of GO molecular function and subcellular localization annotations for 10 model organism proteomes compared with existing databases (GOA, Swiss-Prot).
-
Full-text index only
Columba: an integrated database of proteins, structures, and annotations.
PMID 15801979 · PMC1087474 · BMC bioinformatics · 2005 · 8 claims · 6 setups
COLUMBA physically integrates data from twelve protein structure-related databases (PDB, KEGG, Swiss-Prot, CATH, SCOP, Gene Ontology, ENZYME, etc.) into a single PostgreSQL data warehouse.
-
Full-text index only
Protein ranking by semi-supervised network propagation.
PMID 16723003 · PMC1810311 · BMC bioinformatics · 2006 · 8 claims · 5 setups
RankProp, a diffusion-based network propagation algorithm on a PSI-BLAST-derived protein similarity network, significantly outperforms local search methods (BLAST/PSI-BLAST) at detecting remote homologs.
-
Full-text index only
Predicting the phenotypic effects of non-synonymous single nucleotide polymorphisms based on support vector machines.
PMID 18005451 · PMC2216041 · BMC bioinformatics · 2007 · 8 claims · 5 setups
Parepro, an SVM-based method integrating three attribute sets (RD, MI, IE) derived from evolutionary and residue-property information, predicts whether an nsSNP is deleterious or neutral.
-
Full-text index only
Identification of deleterious non-synonymous single nucleotide polymorphisms using sequence-derived information.
PMID 18588693 · PMC2446391 · BMC bioinformatics · 2008 · 8 claims · 5 setups
A decision tree built on 10 selected sequence-derived features classifies SAPs as Disease or Polymorphism with 82.6% accuracy and 0.607 MCC in cross-validation.
-
Full-text index only
Phylogenetic analysis of RhoGAP domain-containing proteins.
PMID 17127216 · PMC5054073 · Genomics, proteomics & bioinformatics · 2006 · 7 claims · 6 setups
RhoGAP domain-containing proteins, sharing the conserved arginine residue, form a monophyletic group with a common ancestor.
-
Full-text index only
Inventory and analysis of the protein subunits of the ribonucleases P and MRP provides further evidence of homology between the yeast and human enzymes.
PMID 16998185 · PMC1636426 · Nucleic acids research · 2006 · 8 claims · 6 setups
Fungal Pop8 is evolutionarily related to the Rpp14/Pop5 protein family, suggesting Pop8 is the fungal orthologue of Rpp14
-
Full-text index only
Sequence variation in G-protein-coupled receptors: analysis of single nucleotide polymorphisms.
PMID 15784611 · PMC1069129 · Nucleic acids research · 2005 · 7 claims · 8 setups
Position-specific phylogenetic features describing evolutionary conservation at a site (e.g. SIFT score, normalized site entropy, residue frequency change) are the best individual discriminators of disease-causing versus neutral GPCR mutations.
-
Full-text index only
Expansion of the BioCyc collection of pathway/genome databases to 160 genomes.
PMID 16246909 · PMC1266070 · Nucleic acids research · 2005 · 8 claims · 6 setups
The BioCyc collection has been expanded to 160 pathway/genome databases (PGDBs) organized into three curation tiers.
-
Full-text index only
Searching for interpretable rules for disease mutations: a simulated annealing bump hunting strategy.
PMID 16984653 · PMC1618409 · BMC bioinformatics · 2006 · 8 claims · 6 setups
The proposed feature set outperforms existing published feature sets for predicting effects of amino acid substitutions
-
Full-text index only
Systematic identification of pseudogenes through whole genome expression evidence profiling.
PMID 16945953 · PMC1636364 · Nucleic acids research · 2006 · 8 claims · 8 setups
Developed a novel bioinformatics method that identifies pseudogenes by profiling whole-genome transcript and protein expression evidence
-
Full-text index only
Correlating novel variable and conserved motifs in the Hemagglutinin protein with significant biological functions.
PMID 18681973 · PMC2553082 · Virology journal · 2008 · 8 claims · 6 setups
14 MEME blocks were identified in the HA protein of H3N2 strains (1968-1999), with blocks 1, 2, 3, and 7 correlating with several biological functions
-
Full-text index only
Searching for new clues about the molecular cause of endomyocardial fibrosis by way of in silico proteomics and analytical chemistry.
PMID 19823676 · PMC2757908 · PloS one · 2009 · 8 claims · 4 setups
Cross-reactivity of antibodies against C-terminal sequences of ribosomal P proteins from several animals, plants and protozoa with heart tissue may mediate EMF similarly to how T. cruzi C-termini mediate Chaga's disease
-
Full-text index only
Improved mutation tagging with gene identifiers applied to membrane protein stability prediction.
PMID 19758467 · PMC2745585 · BMC bioinformatics · 2009 · 8 claims · 4 setups
MutationTagger achieves 87% F-measure for the mutation retrieval task on a benchmark dataset