Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
Mapping proteins to disease terminologies: from UniProt to MeSH.
PMID 18460185 · PMC2367626 · BMC bioinformatics · 2008 · 8 claims · 7 setups
Developed a three-step procedure (disease name extraction, exact matching, partial/similarity-based matching) to map UniProtKB/Swiss-Prot disease names to MeSH terms
-
Has reproduction · 99
The systematic assessment of completeness of public metadata accompanying omics studies in the Gene Expression Omnibus data repository.
PMID 40926267 · PMC12421755 · Genome biology · 2025 · 8 claims · 3 setups
Over 25% of critical metadata are omitted, with only 74.8% of relevant phenotypes available in publications or public repositories.
-
Full-text index only
GENCODE: producing a reference annotation for ENCODE.
PMID 16925838 · PMC1810553 · Genome biology · 2006 · 8 claims · 8 setups
GENCODE annotation combines initial manual annotation by HAVANA, experimental validation, and refinement based on results to identify protein-coding genes in ENCODE regions
-
Has reproduction · 95
The archives are half-empty: an assessment of the availability of microbial community sequencing data.
PMID 32859925 · PMC7455719 · Communications biology · 2020 · 8 claims · 5 setups
More than half of surveyed amplicon sequencing studies were affected by lack of data deposition, improper file formatting, or inconsistent labeling that impede reuse.
-
Full-text index only
CGMIM: automated text-mining of Online Mendelian Inheritance in Man (OMIM) to identify genetically-associated cancers and candidate genes.
PMID 15796777 · PMC1274267 · BMC bioinformatics · 2005 · 8 claims · 2 setups
CGMIM is a Perl program that text-mines OMIM entries to identify cancer-gene associations and genetically-related cancer type pairs.
-
Full-text index only
BRENDA, AMENDA and FRENDA: the enzyme information system in 2007.
PMID 17202167 · PMC1899097 · Nucleic acids research · 2007 · 7 claims · 6 setups
BRENDA is the largest publicly available enzyme information system worldwide, manually curated from primary literature and covering all identified enzymes regardless of source.
-
Full-text index only
A new procedure for determining the genetic basis of a physiological process in a non-model species, illustrated by cold induced angiogenesis in the carp.
PMID 19852815 · PMC2771047 · BMC genomics · 2009 · 8 claims · 5 setups
The Conditional Stepped Reciprocal Best Hit (CSRBH) approach, combining direct RBH and zebrafish-stepped RBH (SRBH), outperformed other ortholog assignment methods and attained 8,726 carp-human functional homolog relationships for 16,650 carp contigs
-
Full-text index only
Extraction of human kinase mutations from literature, databases and genotyping studies.
PMID 19758464 · PMC2745582 · BMC bioinformatics · 2009 · 7 claims · 6 setups
A literature mining pipeline combining MutationFinder, false-positive filtering, and SVM-based classification can extract and disambiguate single-point mutation mentions from abstracts and full text