Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Has reproduction · 95
Determining virus-host interactions and glycerol metabolism profiles in geographically diverse solar salterns with metagenomics.
PMID 28097058 · PMC5228507 · PeerJ · 2017 · 8 claims · 8 setups
Similar virus-host interactions and glycerol metabolism gene associations (notably dihydroxyacetone kinase with Haloquadratum/Halorubrum) exist across geographically diverse solar salterns
-
Full-text index only
Visualizing the genome: techniques for presenting human genome data and annotations.
PMID 12149135 · PMC119855 · BMC bioinformatics · 2002 · 8 claims · 4 setups
Web-based client-server genome browsers (e.g., LocusLink evidence viewer, UCSC genome browser) are limited by lack of true interactivity, requiring server round-trips for navigation
-
Full-text index only
SysZNF: the C2H2 zinc finger gene database.
PMID 18974185 · PMC2686507 · Nucleic acids research · 2009 · 7 claims · 6 setups
SysZNF is a database that systematically catalogs C2H2-ZNF genes in human and mouse with physical location, gene models, expression probes, protein domains, homologs, and literature links
-
Full-text index only
Protein coding potential of retroviruses and other transposable elements in vertebrate genomes.
PMID 15716312 · PMC549403 · Nucleic acids research · 2005 · 8 claims · 5 setups
About 1000 genes across four vertebrate gene sets analyzed contain at least one RETRA marker protein domain
-
Full-text index only
Analysis of the glutathione S-transferase (GST) gene family.
PMID 15607001 · PMC3500200 · Human genomics · 2004 · 8 claims · 4 setups
The complete human GST gene family comprises 16 genes in six subfamilies: alpha (GSTA), mu (GSTM), omega (GSTO), pi (GSTP), theta (GSTT) and zeta (GSTZ).
-
Full-text index only
Comprehensive genome analysis of 203 genomes provides structural genomics with new insights into protein family space.
PMID 16481312 · PMC1373602 · Nucleic acids research · 2006 · 8 claims · 7 setups
The number of protein families continues to expand steadily as more genomes are sequenced, showing no sign of saturation.
-
Full-text index only
The Vertebrate Genome Annotation (Vega) database.
PMID 15608237 · PMC540089 · Nucleic acids research · 2005 · 8 claims · 8 setups
Vega is a community database for browsing manual annotation of finished vertebrate genome sequences, based on an extended Ensembl-style schema.
-
Full-text index only
Filtering high-throughput protein-protein interaction data using a combination of genomic features.
PMID 15833142 · PMC1127019 · BMC bioinformatics · 2005 · 8 claims · 8 setups
A combination of three genomic features (interacting Pfam domains, GO annotations, sequence homology) using naive Bayesian networks predicts true protein-protein interactions with high sensitivity and good specificity.
-
Has reproduction · 76
Bayesian prediction of microbial oxygen requirement.
PMID 26913185 · PMC4743139 · F1000Research · 2013 · 7 claims · 8 setups
A naive Bayesian classifier based on presence/absence of class-associated Pfam-A domains can distinguish three oxygen requirement classes (aerobe, anaerobe, facultative anaerobe) from genome sequence, unlike prior studies that only made pairwise distinctions.
-
Full-text index only
Anopheles gambiae genome reannotation through synthesis of ab initio and comparative gene prediction algorithms.
PMID 16569258 · PMC1557760 · Genome biology · 2006 · 8 claims · 7 setups
An exon-gene-union (EGU) algorithm followed by an open-reading-frame-selection algorithm can synthesize ab initio (GENSCAN, GeneMark, SNAP) and comparative (Ensembl/Genewise) predictions into a single, more complete CDS set
-
Full-text index only
WebGestalt: an integrated system for exploring gene sets in various biological contexts.
PMID 15980575 · PMC1160236 · Nucleic acids research · 2005 · 8 claims · 6 setups
WebGestalt is an integrated web-based system composed of four modules: gene set management, information retrieval, organization/visualization, and statistics.
-
Full-text index only
Ratiocinative screen of eukaryotic integral membrane protein expression and solubilization for structure determination.
PMID 19031011 · PMC2756966 · Journal of structural and functional genomics · 2009 · 8 claims · 6 setups
A discovery-oriented pipeline using standardized single-condition methods (one expression system, one detergent, one SEC buffer) can efficiently triage large numbers of eukaryotic IMP targets to identify well-behaved candidates for crystallization
-
Full-text index only
MEROPS: the peptidase database.
PMID 19892822 · PMC2808883 · Nucleic acids research · 2010 · 8 claims · 5 setups
MEROPS is a manually curated hierarchical classification of peptidases and protein inhibitors organized into protein species, families, and clans based on sequence and structural homology.
-
Full-text index only
Analysis of protein sequence and interaction data for candidate disease gene prediction.
PMID 17020920 · PMC1636487 · Nucleic acids research · 2006 · 8 claims · 7 setups
Combining CPS and CMP using known disease genes as input achieves sensitivity 0.52 and specificity 0.97, reducing candidate lists 13-fold
-
Full-text index only
What makes species unique? The contribution of proteins with obscure features.
PMID 16859532 · PMC1779552 · Genome biology · 2006 · 7 claims · 8 setups
POFs constitute 18-38% (average 26%) of a typical eukaryotic proteome
-
Full-text index only
CanPredict: a computational tool for predicting cancer-associated missense mutations.
PMID 17537827 · PMC1933186 · Nucleic acids research · 2007 · 8 claims · 7 setups
CanPredict is a web application providing public access to a random forest classifier that combines SIFT, LogR.E-value, and GOSS scores to predict whether a missense mutation is cancer-associated
-
Full-text index only
Pseudofam: the pseudogene families database.
PMID 18957444 · PMC2686518 · Nucleic acids research · 2009 · 8 claims · 7 setups
Pseudofam is an online database of pseudogene families built by mapping pseudogenes to Pfam protein families, providing query tools, statistics, and sequence alignments
-
Full-text index only
Intrinsic structural disorder confers cellular viability on oncogenic fusion proteins.
PMID 19888473 · PMC2768585 · PLoS computational biology · 2009 · 8 claims · 5 setups
Translocation-related human proteins are significantly enriched in intrinsic structural disorder compared to all human proteins
-
Full-text index only
The genome of the simian and human malaria parasite Plasmodium knowlesi.
PMID 18843368 · PMC2656934 · Nature · 2008 · 8 claims · 7 setups
The P. knowlesi (H strain) nuclear genome was sequenced and assembled: 23.5 Mb across 14 chromosomes with 5,188 predicted protein-encoding genes.
-
Full-text index only
Genome wide survey of G protein-coupled receptors in Tetraodon nigroviridis.
PMID 16022726 · PMC1187884 · BMC evolutionary biology · 2005 · 8 claims · 8 setups
466 Tetraodon GPCRs (Tnig-GPCRs) were identified genome-wide, of which 457 had not been previously reported