Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
ABS: a database of Annotated regulatory Binding Sites from orthologous promoters.
PMID 16381947 · PMC1347478 · Nucleic acids research · 2006 · 7 claims · 6 setups
ABS is a public database of experimentally identified TF binding sites conserved in orthologous vertebrate gene promoters, manually curated from the literature.
-
Has reproduction · 95
The archives are half-empty: an assessment of the availability of microbial community sequencing data.
PMID 32859925 · PMC7455719 · Communications biology · 2020 · 8 claims · 6 setups
A large proportion of 16S rRNA amplicon sequencing studies contain data that is not available or not reusable despite being reported as deposited.
-
Has reproduction · 50
SMAC, a computational system to link literature, biomedical and expression data.
PMID 31324861 · PMC6642118 · Scientific reports · 2019 · 8 claims · 8 setups
SMAC is a tool that extracts, prioritises, integrates and analyses biomedical and molecular data according to user-defined terms
-
Full-text index only
PeroxisomeDB: a database for the peroxisomal proteome, functional genomics and disease.
PMID 17135190 · PMC1747181 · Nucleic acids research · 2007 · 8 claims · 6 setups
PeroxisomeDB integrates the complete peroxisomal proteome of Homo sapiens and Saccharomyces cerevisiae into interrelated 'Genes', 'Functions', 'Metabolic pathways' and 'Diseases' sections with links to NCBI, ENSEMBL and UCSC
-
Full-text index only
Gene Prospector: an evidence gateway for evaluating potential susceptibility genes and interacting risk factors for human diseases.
PMID 19063745 · PMC2613935 · BMC bioinformatics · 2008 · 8 claims · 5 setups
Gene Prospector is a Web-based application that selects and prioritizes potential disease-related genes using a curated, updated literature database of genetic association studies
-
Full-text index only
Filtering high-throughput protein-protein interaction data using a combination of genomic features.
PMID 15833142 · PMC1127019 · BMC bioinformatics · 2005 · 8 claims · 8 setups
A combination of three genomic features (interacting Pfam domains, GO annotations, sequence homology) using naive Bayesian networks predicts true protein-protein interactions with high sensitivity and good specificity.
-
Has reproduction · 76
Bayesian prediction of microbial oxygen requirement.
PMID 26913185 · PMC4743139 · F1000Research · 2013 · 7 claims · 8 setups
A naive Bayesian classifier based on presence/absence of class-associated Pfam-A domains can distinguish three oxygen requirement classes (aerobe, anaerobe, facultative anaerobe) from genome sequence, unlike prior studies that only made pairwise distinctions.
-
Full-text index only
Mutations associated with HNPCC predisposition -- Update of ICG-HNPCC/INSiGHT mutation database.
PMID 15528792 · PMC3839397 · Disease markers · 2004 · 8 claims · 4 setups
The ICG-HNPCC/INSiGHT mutation database has grown from 126 predisposing mutations (1997) to 448 mutations occurring in 748 families (2003 update)
-
Full-text index only
Protein coding potential of retroviruses and other transposable elements in vertebrate genomes.
PMID 15716312 · PMC549403 · Nucleic acids research · 2005 · 8 claims · 5 setups
About 1000 genes across four vertebrate gene sets analyzed contain at least one RETRA marker protein domain
-
Full-text index only
Policy implications of genetic information on regulation under the Clean Air Act: the case of particulate matter and asthmatics.
PMID 16507451 · PMC1392222 · Environmental health perspectives · 2006 · 8 claims · 4 setups
The Clean Air Act mandates protection of sensitive subpopulations, including asthmatics, from air pollution health effects, creating an opening for genetic susceptibility data in regulation.
-
Full-text index only
Dissecting microregulation of a master regulatory network.
PMID 18294391 · PMC2289817 · BMC genomics · 2008 · 8 claims · 6 setups
143 human miRNAs (termed p53-miRs) each contain at least one putative p53 binding site within 10 kb flanking sequence and are predicted to target at least one known gene
-
Full-text index only
MEROPS: the peptidase database.
PMID 19892822 · PMC2808883 · Nucleic acids research · 2010 · 8 claims · 5 setups
MEROPS is a manually curated hierarchical classification of peptidases and protein inhibitors organized into protein species, families, and clans based on sequence and structural homology.
-
Full-text index only
BioAfrica's HIV-1 proteomics resource: combining protein data with bioinformatics tools.
PMID 15757512 · PMC555852 · Retrovirology · 2005 · 8 claims · 3 setups
BioAfrica's HIV-1 Proteomics Resource integrates protein structure, gene expression, post-translational modification, functional activity and protein-macromolecule interaction data with bioinformatics tools in a single website.
-
Full-text index only
Phosphorylation states of cell cycle and DNA repair proteins can be altered by the nsSNPs.
PMID 16111488 · PMC1208866 · BMC cancer · 2005 · 8 claims · 4 setups
15 of 89 nsSNPs (16.9%) studied were predicted to abolish or create phosphorylation sites in 14 of 32 proteins (44.0%)
-
Full-text index only
Genome annotation errors in pathway databases due to semantic ambiguity in partial EC numbers.
PMID 16034025 · PMC1179732 · Nucleic acids research · 2005 · 7 claims · 4 setups
Partial EC numbers are semantically ambiguous, and databases that assign a gene to all reactions sharing the same partial EC number make a faulty inference, causing systematic misannotation.
-
Full-text index only
Predicting candidate genes for human deafness disorders: a bioinformatics approach.
PMID 16854223 · PMC1564145 · BMC genomics · 2006 · 8 claims · 4 setups
A bioinformatic approach combining expression databases and protein interaction data narrows ~2400 candidate genes across deafness loci to a manageable set of candidates.
-
Full-text index only
LOCATE: a mammalian protein subcellular localization database.
PMID 17986452 · PMC2238969 · Nucleic acids research · 2008 · 8 claims · 6 setups
LOCATE is a curated, web-accessible database housing membrane organization and subcellular localization data for mouse and human proteins.
-
Full-text index only
When does a protein become an allergen? Searching for a dynamic definition based on most advanced technology tools.
PMID 18477011 · PMC2607534 · Clinical and experimental allergy : journal of the British Society for Allergy and Clinical Immunology · 2008 · 7 claims · 8 setups
IgE-binding is not an intrinsic property of a protein but the result of an interaction between two molecules (antigen and IgE), so a molecule can only be classified relative to demonstrated IgE binding.
-
Full-text index only
Exome sequencing identifies the cause of a mendelian disorder.
PMID 19915526 · PMC2847889 · Nature genetics · 2010 · 8 claims · 7 setups
Exome sequencing of a small number of unrelated affected individuals, combined with filtering against public SNP databases and HapMap exomes, is sufficient to identify the causal gene for a monogenic disorder of unknown etiology.
-
Has reproduction · 69
Automatic discovery of 100-miRNA signature for cancer classification using ensemble feature selection.
PMID 31533612 · PMC6751684 · BMC bioinformatics · 2019 · 8 claims · 6 setups
An ensemble feature selection strategy using consensus of feature relevance across 8 classifier types identifies a 100-miRNA signature from a 1046-feature TCGA dataset