Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
The global landscape of sequence diversity.
PMID 17996061 · PMC2258180 · Genome biology · 2007 · 7 claims · 5 setups
Eukaryotic sequence datasets show substantially greater genetic diversity (higher sequence/gene family discovery rates) than bacterial datasets, likely related to differences in modes of genetic inheritance.
-
Full-text index only
Automatic discovery of cross-family sequence features associated with protein function.
PMID 16409628 · PMC1395344 · BMC bioinformatics · 2006 · 8 claims · 6 setups
A self-supervised data mining approach can find relationships between sequence features and functional annotations without preconceived functional categories.
-
Full-text index only
Comparative genomics of cyclin-dependent kinases suggest co-evolution of the RNAP II C-terminal domain and CTD-directed CDKs.
PMID 15380029 · PMC521075 · BMC genomics · 2004 · 8 claims · 6 setups
Cell-cycle related CDKs (orthologs of CDK1-6) are present in all sampled eukaryotic organisms, including the most ancestral protists.
-
Full-text index only
Pseudofam: the pseudogene families database.
PMID 18957444 · PMC2686518 · Nucleic acids research · 2009 · 8 claims · 7 setups
Pseudofam is an online database of pseudogene families built by mapping pseudogenes to Pfam protein families, providing query tools, statistics, and sequence alignments
-
Full-text index only
Polymorphix: a sequence polymorphism database.
PMID 15608242 · PMC540030 · Nucleic acids research · 2005 · 8 claims · 5 setups
Polymorphix is an ACNUC-structured database that organizes EMBL/GenBank sequences into within-species homologous sequence families using similarity and bibliographic criteria, with alignments, outgroups and phylogenetic trees provided.
-
Full-text index only
EPGD: a comprehensive web resource for integrating and displaying eukaryotic paralog/paralogon information.
PMID 17984073 · PMC2238967 · Nucleic acids research · 2008 · 8 claims · 8 setups
EPGD is a gene-centered, internet-accessible database integrating paralog family and paralogon information for 26 eukaryotic genomes.
-
Full-text index only
The Princeton Protein Orthology Database (P-POD): a comparative genomics analysis tool for biologists.
PMID 17712414 · PMC1942082 · PloS one · 2007 · 8 claims · 5 setups
P-POD is the first comparative genomics database to combine results from multiple computational ortholog/homolog prediction methods with manually curated literature-derived experimental evidence of functional conservation.
-
Full-text index only
An SVD-based comparison of nine whole eukaryotic genomes supports a coelomate rather than ecdysozoan lineage.
PMID 15606920 · PMC544558 · BMC bioinformatics · 2004 · 8 claims · 7 setups
SVD-based analysis of tetrapeptide frequency vectors can compare whole eukaryotic proteomes without pre-defining orthologs or aligning homologous sites
-
Full-text index only
Comprehensive genome analysis of 203 genomes provides structural genomics with new insights into protein family space.
PMID 16481312 · PMC1373602 · Nucleic acids research · 2006 · 8 claims · 7 setups
The number of protein families continues to expand steadily as more genomes are sequenced, showing no sign of saturation.
-
Full-text index only
How to find soluble proteins: a comprehensive analysis of alpha/beta hydrolases for recombinant expression in E. coli.
PMID 15804363 · PMC1079826 · BMC genomics · 2005 · 7 claims · 7 setups
Predicted solubility in E. coli (via CV-CV') depends on hydrolase size, phylogenetic origin, homologous family, and superfamily
-
Full-text index only
MODBASE, a database of annotated comparative protein structure models and associated resources.
PMID 18948282 · PMC2686492 · Nucleic acids research · 2009 · 8 claims · 8 setups
MODBASE contains 5,152,695 reliable comparative protein structure models for 1,593,209 unique protein sequences.
-
Full-text index only
DBD--taxonomically broad transcription factor predictions: new content and functionality.
PMID 18073188 · PMC2238844 · Nucleic acids research · 2008 · 8 claims · 3 setups
DBD is a database of predicted sequence-specific DNA-binding transcription factors covering over 700 publicly available proteomes, up from 150 in the initial version.
-
Full-text index only
Integrative multi-omics analysis of dietary fibre-induced modulations in the composition and function of chicken caecal microbiota.
PMID 41741451 · PMC13046836 · NPJ biofilms and microbiomes · 2026 · 6 claims · 6 setups
High inulin supplementation (4%) significantly altered caecal microbial composition and promoted broader microbial metabolic adaptations, indicating a strong fermentative response to soluble fibre