Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
Functional nsSNPs from carcinogenesis-related genes expressed in breast tissue: potential breast cancer risk alleles and their distribution across human populations.
PMID 16595073 · PMC3500178 · Human genomics · 2006 · 7 claims · 5 setups
A bioinformatics strategy cross-referencing carcinogenesis-related gene lists with breast-tissue expression data can identify candidate breast cancer risk nsSNPs.
-
Full-text index only
Gene Prospector: an evidence gateway for evaluating potential susceptibility genes and interacting risk factors for human diseases.
PMID 19063745 · PMC2613935 · BMC bioinformatics · 2008 · 8 claims · 5 setups
Gene Prospector is a Web-based application that selects and prioritizes potential disease-related genes using a curated, updated literature database of genetic association studies
-
Full-text index only
SuperCYP: a comprehensive database on Cytochrome P450 enzymes including a tool for analysis of CYP-drug interactions.
PMID 19934256 · PMC2808967 · Nucleic acids research · 2010 · 8 claims · 6 setups
SuperCYP is a comprehensive relational database aggregating CYP enzyme, drug metabolism, SNP/mutation, and structural information from literature and web resources.
-
Has reproduction · 95
Pathway-targeting gene matrix for Drosophila gene set enrichment analysis.
PMID 34710184 · PMC8553153 · PloS one · 2021 · 8 claims · 4 setups
Gene matrix files for GSEA are largely unavailable for Drosophila, limiting pathway-level enrichment analysis in this model organism
-
Full-text index only
Integration of text- and data-mining using ontologies successfully selects disease gene candidates.
PMID 15767279 · PMC1065256 · Nucleic acids research · 2005 · 7 claims · 6 setups
Integrating eVOC anatomical ontology-based text-mining of PubMed abstracts with data-mining of gene expression annotation successfully selects and prioritizes candidate disease genes
-
Full-text index only
Epidemiology of doublet/multiplet mutations in lung cancers: evidence that a subset arises by chronocoordinate events.
PMID 19005564 · PMC2579325 · PloS one · 2008 · 8 claims · 7 setups
Doublet mutations are significantly more frequent in EGFR (6.0%) and TP53 (2.3%) in human lung cancer than spontaneous doublets in mouse lacI (0.7%), about 8-fold and 3-fold higher respectively.
-
Full-text index only
Cataloging coding sequence variations in human genome databases.
PMID 18974781 · PMC2570488 · PloS one · 2008 · 8 claims · 7 setups
A significant proportion of CVs overlap between HGMD and dbSNP (4.36% of HGMD CVs registered in dbSNP; 8.11% of dbSNP CVs registered in HGMD), warranting caution when interpreting phenotypic relevance of concurrent CVs.
-
Full-text index only
PolySearch: a web-based text mining system for extracting relationships between human diseases, genes, mutations, drugs and metabolites.
PMID 18487273 · PMC2447794 · Nucleic acids research · 2008 · 8 claims · 7 setups
PolySearch supports more than 50 different classes of queries against nearly a dozen types of text, abstract, or bioinformatic databases
-
Has reproduction · 79
TSUNAMI: Translational Bioinformatics Tool Suite for Network Analysis and Mining.
PMID 33705981 · PMC9403021 · Genomics, proteomics & bioinformatics · 2021 · 8 claims · 6 setups
TSUNAMI is a freely accessible web-based tool suite that mines gene co-expression network (GCN) modules from public (GEO, TCGA) or user-uploaded numerical omics data and performs downstream gene set enrichment analysis.
-
Full-text index only
CGMIM: automated text-mining of Online Mendelian Inheritance in Man (OMIM) to identify genetically-associated cancers and candidate genes.
PMID 15796777 · PMC1274267 · BMC bioinformatics · 2005 · 8 claims · 2 setups
CGMIM is a Perl program that text-mines OMIM entries to identify cancer-gene associations and genetically-related cancer type pairs.
-
Full-text index only
Molecular phylogeny of the antiangiogenic and neurotrophic serpin, pigment epithelium derived factor in vertebrates.
PMID 17020603 · PMC1609119 · BMC genomics · 2006 · 8 claims · 8 setups
A single PEDF gene is present in all examined vertebrate species but is absent from invertebrates (D. melanogaster, C. elegans, C. intestinalis)
-
Full-text index only
Function2Gene: a gene selection tool to increase the power of genetic association studies by utilizing public databases and expert knowledge.
PMID 18631403 · PMC2500032 · BMC bioinformatics · 2008 · 6 claims · 5 setups
Function2Gene is a set of Perl programs that queries public databases (NCBI, GeneCards, Harvester, with Uniprot/Ensembl also supported) using expert-selected keywords to rank genes by prior probability of disease association.
-
Full-text index only
Improved mutation tagging with gene identifiers applied to membrane protein stability prediction.
PMID 19758467 · PMC2745585 · BMC bioinformatics · 2009 · 8 claims · 4 setups
MutationTagger achieves 87% F-measure for the mutation retrieval task on a benchmark dataset
-
Has reproduction · 75
An informatics research platform to make public gene expression time-course datasets reusable for more scientific discoveries.
PMID 33247935 · PMC7698665 · Database : the journal of biological databases and curation · 2020 · 8 claims · 6 setups
GETc enables discovery and visualization of time-course gene expression data and analytical results from GEO
-
Full-text index only
Evolutionary genomics reveals lineage-specific gene loss and rapid evolution of a sperm-specific ion channel complex: CatSpers and CatSperbeta.
PMID 18974790 · PMC2572835 · PloS one · 2008 · 8 claims · 6 setups
The CatSper channel complex (four CatSpers plus CatSperβ) originated as early as primitive metazoans such as the Cnidarian Nematostella vectensis
-
Full-text index only
Update of the G2D tool for prioritization of gene candidates to inherited diseases.
PMID 17478516 · PMC1933178 · Nucleic acids research · 2007 · 8 claims · 4 setups
G2D is a web server that prioritizes candidate genes for inherited diseases using three distinct algorithms based on different input information.
-
Full-text index only
Origin and diversification of the basic helix-loop-helix gene family in metazoans: insights from comparative genomics.
PMID 17335570 · PMC1828162 · BMC evolutionary biology · 2007 · 8 claims · 4 setups
An initial diversification of bHLHs occurred in the pre-Cambrian, prior to metazoan cladogenesis
-
Has reproduction · 85
Exploring microproteins from various model organisms using the mip-mining database.
PMID 37919660 · PMC10623795 · BMC genomics · 2023 · 5 claims · 4 setups
Mip-mining is a database of 336 curated RNA-seq datasets from 8626 samples across nine species, built specifically to explore microprotein functions under stress and disease conditions
-
Has reproduction · 94
Systematic assessment of pathway databases, based on a diverse collection of user-submitted experiments.
PMID 36088548 · PMC9487593 · Briefings in bioinformatics · 2022 · 8 claims · 6 setups
Well-established, hierarchically organized pathway annotation systems (e.g. GO, Reactome, KEGG) yield the best overall enrichment performance despite covering much of the human genome only in general terms.
-
Full-text index only
Mining expressed sequence tags identifies cancer markers of clinical interest.
PMID 17078886 · PMC1635568 · BMC bioinformatics · 2006 · 8 claims · 6 setups
An EST-mining approach (Fisher Exact Test on tumor vs. non-tumor library hit counts) identifies differentially expressed transcripts with an estimated false discovery rate below 22% when human and mouse screens are combined.