Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
Gene Prospector: an evidence gateway for evaluating potential susceptibility genes and interacting risk factors for human diseases.
PMID 19063745 · PMC2613935 · BMC bioinformatics · 2008 · 8 claims · 5 setups
Gene Prospector is a Web-based application that selects and prioritizes potential disease-related genes using a curated, updated literature database of genetic association studies
-
Full-text index only
Web services and workflow management for biological resources.
PMID 16351751 · PMC1866383 · BMC bioinformatics · 2005 · 8 claims · 4 setups
Workflow management systems combined with Web Services are a promising ICT approach for automating access to and integration of biomedical data.
-
Full-text index only
Genome annotation errors in pathway databases due to semantic ambiguity in partial EC numbers.
PMID 16034025 · PMC1179732 · Nucleic acids research · 2005 · 7 claims · 4 setups
Partial EC numbers are semantically ambiguous, and databases that assign a gene to all reactions sharing the same partial EC number make a faulty inference, causing systematic misannotation.
-
Has reproduction · 74
Wide-Open: Accelerating public data release by automating detection of overdue datasets.
PMID 28594819 · PMC5464523 · PLoS biology · 2017 · 7 claims · 5 setups
Wide-Open is a general text-mining approach that automatically detects overdue datasets by scanning PubMed articles for dataset accession identifiers and querying repositories to determine if the datasets remain private.
-
Full-text index only
The Mammalian Phenotype Ontology as a tool for annotating, analyzing and comparing phenotypic information.
PMID 15642099 · PMC549068 · Genome biology · 2005 · 7 claims · 1 setups
The MP Ontology enables robust, standardized annotation of mammalian phenotypes for mutations, QTLs, and strains used as models of human biology and disease.
-
Full-text index only
Discovery of protein-protein interactions using a combination of linguistic, statistical and graphical information.
PMID 15941473 · PMC1164402 · BMC bioinformatics · 2005 · 8 claims · 5 setups
A combined linguistic+statistical+rule-based method achieves precision 0.61 and recall 0.97 (f=0.74) detecting yeast protein-protein interactions across 12,300 Medline abstracts.
-
Has reproduction · 83
Public Omics Explorer (POE): Enabling integrative semantic search across GEO omics datasets based on PubMed publications.
PMID 41282419 · PMC12636342 · Computational and structural biotechnology journal · 2025 · 6 claims · 4 setups
POE is a web platform that semantically links GEO datasets and ENA records through their associated PubMed publications for literature-informed dataset retrieval
-
Has reproduction · 43
StatsDB: platform-agnostic storage and understanding of next generation sequencing run metrics.
PMID 24627795 · PMC3938176 · F1000Research · 2013 · 8 claims · 6 setups
StatsDB is an open-source software package for storage and analysis of next generation sequencing run metrics, backed by an SQL (MySQL) database with Perl and Java APIs.
-
Full-text index only
An integrated database-pipeline system for studying single nucleotide polymorphisms and diseases.
PMID 19091018 · PMC2638159 · BMC bioinformatics · 2008 · 6 claims · 5 setups
Existing SNP/disease databases are fragmented; no combined resource widely supports gene-, SNP-, and disease-related information together
-
Has reproduction · 94
Systematic assessment of pathway databases, based on a diverse collection of user-submitted experiments.
PMID 36088548 · PMC9487593 · Briefings in bioinformatics · 2022 · 8 claims · 6 setups
Well-established, hierarchically organized pathway annotation systems (e.g. GO, Reactome, KEGG) yield the best overall enrichment performance despite covering much of the human genome only in general terms.
-
Full-text index only
Columba: an integrated database of proteins, structures, and annotations.
PMID 15801979 · PMC1087474 · BMC bioinformatics · 2005 · 8 claims · 6 setups
COLUMBA physically integrates data from twelve protein structure-related databases (PDB, KEGG, Swiss-Prot, CATH, SCOP, Gene Ontology, ENZYME, etc.) into a single PostgreSQL data warehouse.
-
Full-text index only
Systems integration of biodefense omics data for analysis of pathogen-host interactions and identification of potential targets.
PMID 19779614 · PMC2745575 · PloS one · 2009 · 8 claims · 8 setups
A protein-centric data integration approach (Master Protein Directory) enables integration and mining of heterogeneous pathogen-host omics data across multiple research centers