Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
A genome-wide survey of Major Histocompatibility Complex (MHC) genes and their paralogues in zebrafish.
PMID 16271140 · PMC1309616 · BMC genomics · 2005 · 8 claims · 4 setups
149 putative MHC gene loci and their paralogues were identified in the zebrafish genome using sequence similarity searches against the Zv4 draft assembly.
-
Full-text index only
Functional classification using phylogenomic inference.
PMID 16846248 · PMC1484587 · PLoS computational biology · 2006 · 8 claims · 1 setups
Functional annotation via top-hit database search transfer is used far more often in practice than phylogenomic inference, despite phylogenomic inference being more accurate.
-
Full-text index only
SuperCYP: a comprehensive database on Cytochrome P450 enzymes including a tool for analysis of CYP-drug interactions.
PMID 19934256 · PMC2808967 · Nucleic acids research · 2010 · 8 claims · 6 setups
SuperCYP is a comprehensive relational database aggregating CYP enzyme, drug metabolism, SNP/mutation, and structural information from literature and web resources.
-
Full-text index only
Efficient algorithms for probing the RNA mutation landscape.
PMID 18688270 · PMC2475669 · PLoS computational biology · 2008 · 8 claims · 4 setups
RNAmutants generalizes McCaskill's partition function algorithm to sum over the grand canonical ensemble of all secondary structures of all k-point mutants, simultaneously computing MFE(k) and Z(k) for each k
-
Has reproduction · 77
Representing and querying disease networks using graph databases.
PMID 27462371 · PMC4960687 · BioData mining · 2016 · 7 claims · 8 setups
Graph databases are well suited for representing biological information that is highly connected, semi-structured, and unpredictable.
-
Full-text index only
Functional coverage of the human genome by existing structures, structural genomics targets, and homology models.
PMID 16118666 · PMC1188274 · PLoS computational biology · 2005 · 8 claims · 5 setups
Existing PDB structures provide single-domain coverage for 37% of functional classes in the human genome and complete (whole-protein) structure coverage for 25%.
-
Full-text index only
TRED: a Transcriptional Regulatory Element Database and a platform for in silico gene regulation studies.
PMID 15608156 · PMC539958 · Nucleic acids research · 2005 · 8 claims · 5 setups
TRED is a database collecting both cis-regulatory elements (promoters) and trans-regulatory elements (transcription factor binding/regulation data) with linked access.
-
Full-text index only
Columba: an integrated database of proteins, structures, and annotations.
PMID 15801979 · PMC1087474 · BMC bioinformatics · 2005 · 8 claims · 6 setups
COLUMBA physically integrates data from twelve protein structure-related databases (PDB, KEGG, Swiss-Prot, CATH, SCOP, Gene Ontology, ENZYME, etc.) into a single PostgreSQL data warehouse.
-
Has reproduction · 83
Gene Expression Atlas update--a value-added database of microarray and sequencing-based functional genomics experiments.
PMID 22064864 · PMC3245177 · Nucleic acids research · 2012 · 8 claims · 5 setups
Gene Expression Atlas is an added-value database providing curated, re-annotated and statistically analysed gene expression data across cell types, organism parts, developmental stages, disease states and other biological/experimental conditions, derived from ArrayExpress Archive and the European Nucleotide Archive.
-
Full-text index only
Gene-disease relationship discovery based on model-driven data integration and database view definition.
PMID 19042916 · PMC2639000 · Bioinformatics (Oxford, England) · 2009 · 8 claims · 4 setups
Explicit gene–disease relationships can be formulated as candidate gene definitions (e.g., co-localization, dysregulation, functional similarity) that may include intermediary orthologous or interacting genes
-
Full-text index only
Using multiple alignments to improve seeded local alignment algorithms.
PMID 16100379 · PMC1185574 · Nucleic acids research · 2005 · 8 claims · 2 setups
Using information implicit in a multiple alignment to dynamically build a spaced-seed index weighted toward promising regions increases sensitivity of local alignment search compared to indexing a sequence alone
-
Has reproduction · 90
CONSULT: accurate contamination removal using locality-sensitive hashing.
PMID 34377979 · PMC8340999 · NAR genomics and bioinformatics · 2021 · 8 claims · 6 setups
CONSULT uses locality-sensitive hashing to test whether query k-mers fall within a user-defined Hamming distance of a reference k-mer database, allowing inexact matching against tens of thousands of microbial species.
-
Full-text index only
The RCSB PDB information portal for structural genomics.
PMID 16381872 · PMC1347482 · Nucleic acids research · 2006 · 7 claims · 5 setups
The RCSB PDB Structural Genomics Information Portal integrates three resources: Structural Genomics Initiatives, Targets (TargetDB/PepcDB), and Structures (functional coverage analysis).
-
Full-text index only
Mass spectrometry group has mass appeal.
PMID 15598607 · PMC1247668 · Environmental health perspectives · 2004 · 7 claims · 4 setups
Mass spectrometry can identify proteins and their post-translational modifications with high specificity and sensitivity, making it central to toxicoproteomics.
-
Full-text index only
Oligomeric protein structure networks: insights into protein-protein interactions.
PMID 16336694 · PMC1326230 · BMC bioinformatics · 2005 · 8 claims · 6 setups
Interface amino acid clusters identified at Imin=6% correlate well with residues losing accessible surface area (δASA) upon oligomerization
-
Full-text index only
Identification of candidate disease genes by integrating Gene Ontologies and protein-interaction networks: case study of primary immunodeficiencies.
PMID 19073697 · PMC2632920 · Nucleic acids research · 2009 · 8 claims · 5 setups
Combining high protein-interaction network scores with significant PID-related GO terms identifies novel PID candidate genes
-
Has reproduction · 86
RNASEQR--a streamlined and accurate RNA-seq sequence analysis program.
PMID 22199257 · PMC3315322 · Nucleic acids research · 2012 · 8 claims · 7 setups
RNASEQR is a new RNA-seq mapper/aligner that combines a BWT-based (Bowtie) transcriptomic/genomic alignment with hash-based BLAT local alignment in three sequential steps: transcriptome mapping, novel exon detection, and anchor-and-align novel splice junction identification.
-
Has reproduction · 67
Cyrface: An interface from Cytoscape to R that provides a user interface to R packages.
PMID 24715956 · PMC3962008 · F1000Research · 2013 · 8 claims · 6 setups
Cyrface is a Cytoscape app/Java library providing a general interface from Cytoscape (Java) to any R function or package.
-
Full-text index only
The human urinary proteome contains more than 1500 proteins, including a large proportion of membrane proteins.
PMID 16948836 · PMC1794545 · Genome biology · 2006 · 8 claims · 6 setups
Identified 1543 proteins in urine from ten healthy donors while essentially eliminating false-positive identifications
-
Full-text index only
Systems integration of biodefense omics data for analysis of pathogen-host interactions and identification of potential targets.
PMID 19779614 · PMC2745575 · PloS one · 2009 · 8 claims · 8 setups
A protein-centric data integration approach (Master Protein Directory) enables integration and mining of heterogeneous pathogen-host omics data across multiple research centers