Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
Automatic discovery of cross-family sequence features associated with protein function.
PMID 16409628 · PMC1395344 · BMC bioinformatics · 2006 · 8 claims · 6 setups
A self-supervised data mining approach can find relationships between sequence features and functional annotations without preconceived functional categories.
-
Full-text index only
Inparanoid: a comprehensive database of eukaryotic orthologs.
PMID 15608241 · PMC540061 · Nucleic acids research · 2005 · 8 claims · 4 setups
The Inparanoid algorithm identifies true ortholog clusters by seeding on reciprocal best-matching pairs, gathering inparalogs (post-speciation duplicates) while excluding outparalogs (pre-speciation duplicates)
-
Full-text index only
InParanoid 7: new algorithms and tools for eukaryotic orthology analysis.
PMID 19892828 · PMC2808972 · Nucleic acids research · 2010 · 8 claims · 7 setups
InParanoid 7 expands the database by an order of magnitude to 100 species, 1.3 million proteins, and 42.7 million pairwise ortholog groups.
-
Full-text index only
InParanoid 6: eukaryotic ortholog clusters with inparalogs.
PMID 18055500 · PMC2238924 · Nucleic acids research · 2008 · 8 claims · 3 setups
InParanoid 6 is an updated eukaryotic ortholog database covering 35 species (34 eukaryotes plus E. coli as outgroup), providing pairwise ortholog clusters with inparalogs for all species pairs.
-
Full-text index only
EPGD: a comprehensive web resource for integrating and displaying eukaryotic paralog/paralogon information.
PMID 17984073 · PMC2238967 · Nucleic acids research · 2008 · 8 claims · 8 setups
EPGD is a gene-centered, internet-accessible database integrating paralog family and paralogon information for 26 eukaryotic genomes.
-
Full-text index only
The Princeton Protein Orthology Database (P-POD): a comparative genomics analysis tool for biologists.
PMID 17712414 · PMC1942082 · PloS one · 2007 · 8 claims · 5 setups
P-POD is the first comparative genomics database to combine results from multiple computational ortholog/homolog prediction methods with manually curated literature-derived experimental evidence of functional conservation.
-
Full-text index only
Pegasys: software for executing and integrating analyses of biological sequences.
PMID 15096276 · PMC406494 · BMC bioinformatics · 2004 · 8 claims · 7 setups
Pegasys is a flexible, modular, customizable software system for executing and integrating heterogeneous biological sequence analysis tools
-
Full-text index only
The TIGR Gene Indices: clustering and assembling EST and known genes and integration with eukaryotic genomes.
PMID 15608288 · PMC540018 · Nucleic acids research · 2005 · 8 claims · 8 setups
The TIGR Gene Indices (TGI) are a collection of 77 species-specific databases that cluster and assemble EST and known gene sequences into tentative consensus (TC) sequences to identify and characterize expressed transcripts.
-
Full-text index only
What makes species unique? The contribution of proteins with obscure features.
PMID 16859532 · PMC1779552 · Genome biology · 2006 · 7 claims · 8 setups
POFs constitute 18-38% (average 26%) of a typical eukaryotic proteome
-
Full-text index only
The two faces of Alba: the evolutionary connection between proteins participating in chromatin structure and RNA metabolism.
PMID 14519199 · PMC328453 · Genome biology · 2003 · 8 claims · 5 setups
Sequence-profile (PSI-BLAST) searches unify archaeal Alba with eukaryotic RNase P/MRP subunits Rpp20/Pop7 and Rpp25, and with the ciliate macronuclear-development protein Mdp2, into a single Alba superfamily.
-
Full-text index only
MODBASE, a database of annotated comparative protein structure models and associated resources.
PMID 18948282 · PMC2686492 · Nucleic acids research · 2009 · 8 claims · 8 setups
MODBASE contains 5,152,695 reliable comparative protein structure models for 1,593,209 unique protein sequences.
-
Full-text index only
An SVD-based comparison of nine whole eukaryotic genomes supports a coelomate rather than ecdysozoan lineage.
PMID 15606920 · PMC544558 · BMC bioinformatics · 2004 · 8 claims · 7 setups
SVD-based analysis of tetrapeptide frequency vectors can compare whole eukaryotic proteomes without pre-defining orthologs or aligning homologous sites
-
Full-text index only
Inverse symmetry in complete genomes and whole-genome inverse duplication.
PMID 19898631 · PMC2771390 · PloS one · 2009 · 8 claims · 5 setups
Reverse and complement symmetries are essentially absent in genomic sequences at all scales.
-
Full-text index only
Metagenomic analysis of respiratory tract DNA viral communities in cystic fibrosis and non-cystic fibrosis individuals.
PMID 19816605 · PMC2756586 · PloS one · 2009 · 8 claims · 8 setups
CF phage communities are highly similar to each other, whereas Non-CF individuals have more distinct, variable phage communities reflecting transient environmental sampling
-
Full-text index only
Coverage of whole proteome by structural genomics observed through protein homology modeling database.
PMID 17146617 · PMC1769342 · Journal of structural and functional genomics · 2006 · 8 claims · 7 setups
FAMSBASE, a homology-modeling database of whole-genome ORFs, currently covers about 50% of predicted ORFs (368,724 of 734,193) across 276 genomes with modeled 3D structures.
-
Full-text index only
Database resources of the National Center for Biotechnology Information.
PMID 17170002 · PMC1781113 · Nucleic acids research · 2007 · 8 claims · 8 setups
NCBI maintains an integrated suite of database resources (Entrez, PubMed, RefSeq, dbSNP, BLAST, etc.) for molecular biology data retrieval and analysis