Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
A survey of integral alpha-helical membrane proteins.
PMID 19760129 · PMC2780624 · Journal of structural and functional genomics · 2009 · 8 claims · 8 setups
An automated annotation pipeline defines the integral membrane genome and family associations for 21,379 proteins from 34 genomes, most belonging to 598 Pfam-derived membrane protein families.
-
Full-text index only
Tracing the origin of functional and conserved domains in the human proteome: implications for protein evolution at the modular level.
PMID 17090320 · PMC1654190 · BMC evolutionary biology · 2006 · 8 claims · 5 setups
HHpred (HMM-HMM comparison) detects remote homologs in the human proteome with higher sensitivity than hmmpfam (HMMER), giving 10% more functional domain coverage and 20% higher residue coverage against Pfam-A families.
-
Full-text index only
Comprehensive genome analysis of 203 genomes provides structural genomics with new insights into protein family space.
PMID 16481312 · PMC1373602 · Nucleic acids research · 2006 · 8 claims · 7 setups
The number of protein families continues to expand steadily as more genomes are sequenced, showing no sign of saturation.
-
Full-text index only
Ratiocinative screen of eukaryotic integral membrane protein expression and solubilization for structure determination.
PMID 19031011 · PMC2756966 · Journal of structural and functional genomics · 2009 · 8 claims · 6 setups
A discovery-oriented pipeline using standardized single-condition methods (one expression system, one detergent, one SEC buffer) can efficiently triage large numbers of eukaryotic IMP targets to identify well-behaved candidates for crystallization
-
Full-text index only
Pseudofam: the pseudogene families database.
PMID 18957444 · PMC2686518 · Nucleic acids research · 2009 · 8 claims · 7 setups
Pseudofam is an online database of pseudogene families built by mapping pseudogenes to Pfam protein families, providing query tools, statistics, and sequence alignments
-
Full-text index only
CanPredict: a computational tool for predicting cancer-associated missense mutations.
PMID 17537827 · PMC1933186 · Nucleic acids research · 2007 · 8 claims · 7 setups
CanPredict is a web application providing public access to a random forest classifier that combines SIFT, LogR.E-value, and GOSS scores to predict whether a missense mutation is cancer-associated
-
Full-text index only
DBD--taxonomically broad transcription factor predictions: new content and functionality.
PMID 18073188 · PMC2238844 · Nucleic acids research · 2008 · 8 claims · 3 setups
DBD is a database of predicted sequence-specific DNA-binding transcription factors covering over 700 publicly available proteomes, up from 150 in the initial version.
-
Full-text index only
nf-core/proteinfamilies: a scalable pipeline for the generation of protein families.
PMID 41563008 · PMC12950615 · GigaScience · 2026 · 8 claims · 3 setups
nf-core/proteinfamilies is a scalable, parametrizable, open-source Nextflow pipeline that generates new protein families or assigns sequences to existing families using profile HMMs and MSAs
-
Full-text index only
Towards a comprehensive structural coverage of completed genomes: a structural genomics viewpoint.
PMID 17349043 · PMC1829165 · BMC bioinformatics · 2007 · 8 claims · 6 setups
A combined target-selection approach — pursuing both structurally uncharacterised domain families and additional targets from large structurally characterised superfamilies — is essential for comprehensive structural coverage of the genomes.
-
Full-text index only
Hybrid sequencing reveals incompleteness of the H37Rv reference genome and highlights lineage-specific genomic divergence in Mycobacterium tuberculosis.
PMID 42224013 · PMC13225438 · Microbial genomics · 2026 · 8 claims · 6 setups
The H37Rv_ref reference genome, sequenced in 1998 with early technology, is incomplete relative to modern hybrid-sequenced assemblies
-
Full-text index only
Comparison of complete nuclear receptor sets from the human, Caenorhabditis elegans and Drosophila genomes.
PMID 11532213 · PMC55326 · Genome biology · 2001 · 7 claims · 5 setups
The human genome contains fewer than 50 functional nuclear receptors, far fewer than C. elegans and about twice as many as Drosophila
-
Full-text index only
Genome wide survey of G protein-coupled receptors in Tetraodon nigroviridis.
PMID 16022726 · PMC1187884 · BMC evolutionary biology · 2005 · 8 claims · 8 setups
466 Tetraodon GPCRs (Tnig-GPCRs) were identified genome-wide, of which 457 had not been previously reported
-
Full-text index only
NovelFam3000--uncharacterized human protein domains conserved across model organisms.
PMID 16533400 · PMC1440326 · BMC genomics · 2006 · 8 claims · 7 setups
NovelFam3000 is an online data centre unifying bioinformatics resource links, news, comments, and user-submitted experimental data (including a Gene Characterization Index) for ~3000 uncharacterized Pfam-B/DUF domain families conserved across worm, fly, and human
-
Has reproduction · 95
Determining virus-host interactions and glycerol metabolism profiles in geographically diverse solar salterns with metagenomics.
PMID 28097058 · PMC5228507 · PeerJ · 2017 · 8 claims · 8 setups
Similar virus-host interactions and glycerol metabolism gene associations (notably dihydroxyacetone kinase with Haloquadratum/Halorubrum) exist across geographically diverse solar salterns
-
Full-text index only
MEROPS: the peptidase database.
PMID 19892822 · PMC2808883 · Nucleic acids research · 2010 · 8 claims · 5 setups
MEROPS is a manually curated hierarchical classification of peptidases and protein inhibitors organized into protein species, families, and clans based on sequence and structural homology.
-
Has reproduction · 92
Telomere-to-telomere reference genome for Panax ginseng highlights the evolution of saponin biosynthesis.
PMID 38883331 · PMC11179851 · Horticulture research · 2024 · 8 claims · 8 setups
A telomere-to-telomere reference genome of P. ginseng was assembled (3.45 Gb, 24 chromosomes, 77266 protein-coding genes)
-
Full-text index only
Local combinational variables: an approach used in DNA-binding helix-turn-helix motif prediction with sequence information.
PMID 19651875 · PMC2761287 · Nucleic acids research · 2009 · 8 claims · 7 setups
The LCV approach predicts HTH motifs with 93.29% accuracy, 93.93% sensitivity and 92.66% specificity using only primary sequence information
-
Full-text index only
SGCEdb: a flexible database and web interface integrating experimental results and analysis for structural genomics focusing on Caenorhabditis elegans.
PMID 16381914 · PMC1347399 · Nucleic acids research · 2006 · 8 claims · 8 setups
SGCEdb is a flexible, reusable database and web interface for reporting and analyzing structural genomics experiment results, focused on C. elegans
-
Has reproduction · 95
A whole genome duplication drives the genome evolution of Phytophthora betacei, a closely related species to Phytophthora infestans.
PMID 34740326 · PMC8571832 · BMC genomics · 2021 · 8 claims · 7 setups
P. betacei P8084 has the largest sequenced genome in the Phytophthora genus (270 Mb)