Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
Columba: an integrated database of proteins, structures, and annotations.
PMID 15801979 · PMC1087474 · BMC bioinformatics · 2005 · 8 claims · 6 setups
COLUMBA physically integrates data from twelve protein structure-related databases (PDB, KEGG, Swiss-Prot, CATH, SCOP, Gene Ontology, ENZYME, etc.) into a single PostgreSQL data warehouse.
-
Full-text index only
An integrated database-pipeline system for studying single nucleotide polymorphisms and diseases.
PMID 19091018 · PMC2638159 · BMC bioinformatics · 2008 · 6 claims · 5 setups
Existing SNP/disease databases are fragmented; no combined resource widely supports gene-, SNP-, and disease-related information together
-
Full-text index only
SysPIMP: the web-based systematical platform for identifying human disease-related mutated sequences from mass spectrometry.
PMID 19036792 · PMC2686442 · Nucleic acids research · 2009 · 8 claims · 7 setups
SysPIMP is a web-based platform integrating disease mutation databases with X!Tandem and BLAST to identify disease-related mutated proteins from MS results
-
Has reproduction · 53
PulmonDB: a curated lung disease gene expression database.
PMID 31949184 · PMC6965635 · Scientific reports · 2020 · 6 claims · 6 setups
PulmonDB is a curated, web-based gene expression database and R package integrating microarray and RNA-seq data for COPD and IPF with manually curated controlled-vocabulary annotation.
-
Full-text index only
Functional coverage of the human genome by existing structures, structural genomics targets, and homology models.
PMID 16118666 · PMC1188274 · PLoS computational biology · 2005 · 8 claims · 5 setups
Existing PDB structures provide single-domain coverage for 37% of functional classes in the human genome and complete (whole-protein) structure coverage for 25%.
-
Full-text index only
DAVID Bioinformatics Resources: expanded annotation database and novel algorithms to better extract biology from large gene lists.
PMID 17576678 · PMC1933169 · Nucleic acids research · 2007 · 8 claims · 4 setups
The DAVID Gene Concept uses a single-linkage method to agglomerate tens of millions of gene/protein identifiers from NCBI, PIR, UniProt and other resources into unified DAVID genes.
-
Full-text index only
Sys-BodyFluid: a systematical database for human body fluid proteome research.
PMID 18978022 · PMC2686600 · Nucleic acids research · 2009 · 6 claims · 4 setups
Sys-BodyFluid is a web-based database integrating proteomic data from 11 human body fluids (plasma/serum, urine, cerebrospinal fluid, saliva, bronchoalveolar lavage fluid, synovial fluid, nipple aspirate fluid, tear fluid, seminal fluid, milk, amniotic fluid), containing over 10,000 proteins
-
Full-text index only
GeneTide--Terra Incognita Discovery Endeavor: a new transcriptome focused member of the GeneCards/GeneNote suite of databases.
PMID 15608261 · PMC540076 · Nucleic acids research · 2005 · 8 claims · 7 setups
GeneTide integrates UniGene, DoTS, AceView, BLAT/GeneLoc genomic alignment, and GeneAnnot probe-set data into a unified Consensus/Uniqueness/Score scheme to associate ESTs with GeneCards genes
-
Has reproduction · 77
Representing and querying disease networks using graph databases.
PMID 27462371 · PMC4960687 · BioData mining · 2016 · 7 claims · 8 setups
Graph databases are well suited for representing biological information that is highly connected, semi-structured, and unpredictable.
-
Full-text index only
TRED: a transcriptional regulatory element database, new entries and other development.
PMID 17202159 · PMC1899102 · Nucleic acids research · 2007 · 8 claims · 3 setups
TRED collects mammalian cis- and trans-regulatory elements together with experimental evidence, mapped onto assembled genomes
-
Full-text index only
ChimerDB--a knowledgebase for fusion sequences.
PMID 16381848 · PMC1347382 · Nucleic acids research · 2006 · 8 claims · 6 setups
ChimerDB integrates bioinformatics analysis of mRNA/EST sequences, manually collected literature data, and OMIM translocation data into a single fusion sequence knowledgebase
-
Full-text index only
Structure SNP (StSNP): a web server for mapping and modeling nsSNPs on protein structures with linkage to metabolic pathways.
PMID 17537826 · PMC1933130 · Nucleic acids research · 2007 · 7 claims · 5 setups
StSNP integrates dbSNP, PDB, KEGG, and NCBI Entrez data into a single web server for nsSNP analysis
-
Full-text index only
SpliceMiner: a high-throughput database implementation of the NCBI Evidence Viewer for microarray splice variant analysis.
PMID 17338820 · PMC1839109 · BMC bioinformatics · 2007 · 6 claims · 4 setups
EVDB is a comprehensive, non-redundant relational database of known human splice variants built from NCBI Entrez Gene and Evidence Viewer data
-
Full-text index only
Large-scale discovery of insertion hotspots and preferential integration sites of human transposed elements.
PMID 20008508 · PMC2836564 · Nucleic acids research · 2010 · 8 claims · 6 setups
Most TEs insert within specific 'hotspots' along the targeted TE rather than uniformly.
-
Has reproduction · 80
MCPmed: a call for Model Context Protocol-enabled bioinformatics web services for LLM-driven discovery.
PMID 41729821 · PMC12927880 · Briefings in bioinformatics · 2026 · 6 claims · 3 setups
Adapting MCP to bioinformatics web server backends provides a standardized, machine-actionable semantic layer linking API endpoints to scientific concepts and metadata.
-
Full-text index only
Gene-disease relationship discovery based on model-driven data integration and database view definition.
PMID 19042916 · PMC2639000 · Bioinformatics (Oxford, England) · 2009 · 8 claims · 4 setups
Explicit gene–disease relationships can be formulated as candidate gene definitions (e.g., co-localization, dysregulation, functional similarity) that may include intermediary orthologous or interacting genes
-
Full-text index only
Comparative Toxicogenomics Database: a knowledgebase and discovery tool for chemical-gene-disease networks.
PMID 18782832 · PMC2686584 · Nucleic acids research · 2009 · 8 claims · 5 setups
CTD is a manually curated knowledgebase that integrates chemical-gene interactions, chemical-disease relationships, and gene-disease relationships into a chemical-gene-disease triad
-
Full-text index only
An XML-based system for synthesis of data from disparate databases.
PMID 16501185 · PMC1513665 · Journal of the American Medical Informatics Association : JAMIA · 2006 · 8 claims · 2 setups
An XML-based data management framework (built on Mobius) supports integration of disparate data sources and large data sets for biomedical research applications.
-
Full-text index only
piRNABank: a web resource on classified and clustered Piwi-interacting RNAs.
PMID 17881367 · PMC2238943 · Nucleic acids research · 2008 · 6 claims · 4 setups
piRNABank is a web-accessible database storing empirically known piRNA sequences and annotations for human, mouse and rat.
-
Has reproduction · 84
Expression Atlas update--a database of gene and transcript expression from microarray- and sequencing-based functional genomics experiments.
PMID 24304889 · PMC3964963 · Nucleic acids research · 2014 · 8 claims · 6 setups
Expression Atlas is a value-added database providing gene, protein and splice variant expression across cell types, organism parts, developmental stages, diseases and other biological/experimental conditions, built from manually curated high-quality microarray and RNA-sequencing experiments from ArrayExpress.