Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
Molecular phylogeny of the kelch-repeat superfamily reveals an expansion of BTB/kelch proteins in animals.
PMID 13678422 · PMC222960 · BMC bioinformatics · 2003 · 8 claims · 8 setups
The human genome encodes at least 71 kelch-repeat proteins
-
Full-text index only
Interactome-transcriptome analysis reveals the high centrality of genes differentially expressed in lung cancer tissues.
PMID 16188928 · PMC4631381 · Bioinformatics (Oxford, England) · 2005 · 7 claims · 4 setups
Genes upregulated in squamous cell lung cancer are highly connected (well-connected) nodes in the protein interactome
-
Full-text index only
How to find soluble proteins: a comprehensive analysis of alpha/beta hydrolases for recombinant expression in E. coli.
PMID 15804363 · PMC1079826 · BMC genomics · 2005 · 7 claims · 7 setups
Predicted solubility in E. coli (via CV-CV') depends on hydrolase size, phylogenetic origin, homologous family, and superfamily
-
Full-text index only
Aberrant 5' splice sites in human disease genes: mutation pattern, nucleotide structure and comparison of computational tools that predict their utilization.
PMID 17576681 · PMC1934990 · Nucleic acids research · 2007 · 8 claims · 4 setups
Cryptic 5'ss are best predicted by computational algorithms that accommodate nucleotide dependencies (e.g., Markov model, maximum entropy, maximum dependence decomposition) rather than by weight-matrix models
-
Full-text index only
Efficient algorithms for probing the RNA mutation landscape.
PMID 18688270 · PMC2475669 · PLoS computational biology · 2008 · 8 claims · 4 setups
RNAmutants generalizes McCaskill's partition function algorithm to sum over the grand canonical ensemble of all secondary structures of all k-point mutants, simultaneously computing MFE(k) and Z(k) for each k
-
Full-text index only
A survey of integral alpha-helical membrane proteins.
PMID 19760129 · PMC2780624 · Journal of structural and functional genomics · 2009 · 8 claims · 8 setups
An automated annotation pipeline defines the integral membrane genome and family associations for 21,379 proteins from 34 genomes, most belonging to 598 Pfam-derived membrane protein families.
-
Full-text index only
The 10 sea urchin receptor for egg jelly proteins (SpREJ) are members of the polycystic kidney disease-1 (PKD1) family.
PMID 17629917 · PMC1934368 · BMC genomics · 2007 · 8 claims · 5 setups
Sea urchins possess 10 SpREJ (PKD1 family) genes, compared to five in humans, all defined by possession of a ~600 residue REJ domain
-
Full-text index only
Comparative genomics of fungal allergens and epitopes shows widespread distribution of closely related allergen and epitope orthologues.
PMID 17029625 · PMC1613252 · BMC genomics · 2006 · 8 claims · 3 setups
A database of 82 allergen sequences was compiled and used to search 22 fungal genomes for orthologues.
-
Full-text index only
pTARGET: a web server for predicting protein subcellular localization.
PMID 16844995 · PMC1538910 · Nucleic acids research · 2006 · 7 claims · 3 setups
pTARGET web server predicts nine distinct subcellular localizations in eukaryotic non-plant proteins using an algorithm based on location-specific Pfam domain occurrence patterns and amino acid composition (AAC)
-
Full-text index only
The Functional RNA Database 3.0: databases to support mining and annotation of functional RNAs.
PMID 18948287 · PMC2686472 · Nucleic acids research · 2009 · 8 claims · 5 setups
fRNAdb 3.0 is a completely rebuilt sequence database hosting a much larger collection of known/predicted non-coding RNA sequences with improved search functionality
-
Has reproduction · 71
Protein structure quality assessment based on the distance profiles of consecutive backbone Cα atoms.
PMID 24555103 · PMC3892923 · F1000Research · 2013 · 8 claims · 8 setups
The distance between consecutive backbone Cα atoms in high-quality structures is normally distributed with mean 3.8 Å and standard deviation 0.04 Å, justifying a reference state in which all consecutive Cα atoms are 3.8 Å apart.
-
Full-text index only
A computational screen for type I polyketide synthases in metagenomics shotgun data.
PMID 18953415 · PMC2568958 · PloS one · 2008 · 8 claims · 6 setups
Combining HMM domain searches with maximum-likelihood phylogenetic trees can discriminate true PKS I sequences from evolutionarily related but functionally different enzymes (e.g., FAS I) in metagenomic data.
-
Full-text index only
PA-GOSUB: a searchable database of model organism protein sequences with their predicted Gene Ontology molecular function and subcellular localization.
PMID 15608166 · PMC540074 · Nucleic acids research · 2005 · 7 claims · 4 setups
PA-GOSUB significantly extends the coverage of GO molecular function and subcellular localization annotations for 10 model organism proteomes compared with existing databases (GOA, Swiss-Prot).
-
Full-text index only
Widespread A-to-I RNA editing of Alu-containing mRNAs in the human transcriptome.
PMID 15534692 · PMC526178 · PLoS biology · 2004 · 8 claims · 6 setups
Intramolecular pairs of oppositely oriented Alu elements within the same pre-mRNA form dsRNA foldback structures that are major substrates for A-to-I RNA editing
-
Full-text index only
Molecular phylogeny of the antiangiogenic and neurotrophic serpin, pigment epithelium derived factor in vertebrates.
PMID 17020603 · PMC1609119 · BMC genomics · 2006 · 8 claims · 8 setups
A single PEDF gene is present in all examined vertebrate species but is absent from invertebrates (D. melanogaster, C. elegans, C. intestinalis)
-
Has reproduction · 76
Bayesian prediction of microbial oxygen requirement.
PMID 26913185 · PMC4743139 · F1000Research · 2013 · 7 claims · 8 setups
A naive Bayesian classifier based on presence/absence of class-associated Pfam-A domains can distinguish three oxygen requirement classes (aerobe, anaerobe, facultative anaerobe) from genome sequence, unlike prior studies that only made pairwise distinctions.
-
Full-text index only
QTL MatchMaker: a multi-species quantitative trait loci (QTL) database and query system for annotation of genes and QTL.
PMID 16381937 · PMC1347390 · Nucleic acids research · 2006 · 8 claims · 5 setups
QTL MatchMaker integrates QTL information with physical, genetic and cytogenetic maps across human, mouse and rat genomes
-
Full-text index only
Rapid identification of PAX2/5/8 direct downstream targets in the otic vesicle by combinatorial use of bioinformatics tools.
PMID 18828907 · PMC2760872 · Genome biology · 2008 · 8 claims · 8 setups
A combinatorial bioinformatics pipeline (evolutionary double filtering comparative genomics, GXD/ZFIN database queries, MEDLINE text mining) can rapidly and specifically identify PAX2/5/8 direct downstream targets in the otic vesicle
-
Full-text index only
Gene prediction in eukaryotes with a generalized hidden Markov model that uses hints from external sources.
PMID 16469098 · PMC1409804 · BMC bioinformatics · 2006 · 7 claims · 3 setups
AUGUSTUS+ extends the AUGUSTUS GHMM by combining intrinsic sequence information with extrinsic hints via an extended emission alphabet, so the GHMM jointly models the DNA sequence, gene structure, and hint collection.
-
Full-text index only
Bases and spaces: resources on the web for accessing the draft human genome.
PMID 11178254 · PMC138875 · Genome biology · 2000 · 8 claims · 8 setups
By combining currently available genomic databases and mapping resources (GenBank/Entrez, UniGene, RH maps, BAC fingerprint maps, Ensembl, NIX), it is possible to devise strategies that fully exploit the fragmentary draft human genome sequence.