Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
Molecular phylogeny of the kelch-repeat superfamily reveals an expansion of BTB/kelch proteins in animals.
PMID 13678422 · PMC222960 · BMC bioinformatics · 2003 · 8 claims · 8 setups
The human genome encodes at least 71 kelch-repeat proteins
-
Full-text index only
Protein coding potential of retroviruses and other transposable elements in vertebrate genomes.
PMID 15716312 · PMC549403 · Nucleic acids research · 2005 · 8 claims · 5 setups
About 1000 genes across four vertebrate gene sets analyzed contain at least one RETRA marker protein domain
-
Full-text index only
What makes species unique? The contribution of proteins with obscure features.
PMID 16859532 · PMC1779552 · Genome biology · 2006 · 7 claims · 8 setups
POFs constitute 18-38% (average 26%) of a typical eukaryotic proteome
-
Full-text index only
Genetic diversity among five T4-like bacteriophages.
PMID 16716236 · PMC1524935 · Virology journal · 2006 · 8 claims · 8 setups
A core set of 82 conserved genes (T4-like genes) is present in all five genomes analyzed, clustered in large collinear blocks.
-
Full-text index only
CanPredict: a computational tool for predicting cancer-associated missense mutations.
PMID 17537827 · PMC1933186 · Nucleic acids research · 2007 · 8 claims · 7 setups
CanPredict is a web application providing public access to a random forest classifier that combines SIFT, LogR.E-value, and GOSS scores to predict whether a missense mutation is cancer-associated
-
Full-text index only
Comprehensive genome analysis of 203 genomes provides structural genomics with new insights into protein family space.
PMID 16481312 · PMC1373602 · Nucleic acids research · 2006 · 8 claims · 7 setups
The number of protein families continues to expand steadily as more genomes are sequenced, showing no sign of saturation.
-
Full-text index only
nf-core/proteinfamilies: a scalable pipeline for the generation of protein families.
PMID 41563008 · PMC12950615 · GigaScience · 2026 · 8 claims · 3 setups
nf-core/proteinfamilies is a scalable, parametrizable, open-source Nextflow pipeline that generates new protein families or assigns sequences to existing families using profile HMMs and MSAs
-
Full-text index only
The Vertebrate Genome Annotation (Vega) database.
PMID 15608237 · PMC540089 · Nucleic acids research · 2005 · 8 claims · 8 setups
Vega is a community database for browsing manual annotation of finished vertebrate genome sequences, based on an extended Ensembl-style schema.
-
Full-text index only
Bioinformatic mapping of AlkB homology domains in viruses.
PMID 15627404 · PMC544882 · BMC genomics · 2005 · 8 claims · 8 setups
AlkB-like domains are found in at least 22 different single-stranded RNA positive-strand plant viruses, mainly within a subgroup of the Flexiviridae family.
-
Full-text index only
Analysis of protein sequence and interaction data for candidate disease gene prediction.
PMID 17020920 · PMC1636487 · Nucleic acids research · 2006 · 8 claims · 7 setups
Combining CPS and CMP using known disease genes as input achieves sensitivity 0.52 and specificity 0.97, reducing candidate lists 13-fold
-
Full-text index only
SGCEdb: a flexible database and web interface integrating experimental results and analysis for structural genomics focusing on Caenorhabditis elegans.
PMID 16381914 · PMC1347399 · Nucleic acids research · 2006 · 8 claims · 8 setups
SGCEdb is a flexible, reusable database and web interface for reporting and analyzing structural genomics experiment results, focused on C. elegans
-
Has reproduction · 91
De Novo Assembly and Annotation of the Larval Transcriptome of Two Spadefoot Toads Widely Divergent in Developmental Rate.
PMID 31217263 · PMC6686947 · G3 (Bethesda, Md.) · 2019 · 8 claims · 8 setups
De novo transcriptome assemblies were generated for larval P. cultripes and S. couchii, providing new genomic resources for spadefoot toads
-
Has reproduction · 95
Determining virus-host interactions and glycerol metabolism profiles in geographically diverse solar salterns with metagenomics.
PMID 28097058 · PMC5228507 · PeerJ · 2017 · 8 claims · 8 setups
Similar virus-host interactions and glycerol metabolism gene associations (notably dihydroxyacetone kinase with Haloquadratum/Halorubrum) exist across geographically diverse solar salterns