Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
Identification of "pathologs" (disease-related genes) from the RIKEN mouse cDNA dataset using human curation plus FACTS, a new biological information extraction system.
PMID 15115540 · PMC420239 · BMC genomics · 2004 · 6 claims · 3 setups
Bioinformatic sequence comparison of 60,770 RIKEN FANTOM2 mouse cDNA clones identified 2,578 sequences with 70-85% identity to known human disease genes/proteins
-
Full-text index only
The Vertebrate Genome Annotation (Vega) database.
PMID 15608237 · PMC540089 · Nucleic acids research · 2005 · 8 claims · 8 setups
Vega is a community database for browsing manual annotation of finished vertebrate genome sequences, based on an extended Ensembl-style schema.
-
Full-text index only
ABS: a database of Annotated regulatory Binding Sites from orthologous promoters.
PMID 16381947 · PMC1347478 · Nucleic acids research · 2006 · 7 claims · 6 setups
ABS is a public database of experimentally identified TF binding sites conserved in orthologous vertebrate gene promoters, manually curated from the literature.
-
Has reproduction · 89
Evolution of Highly Repetitive Silk Genes in the Luna Moth, Actias luna.
PMID 41738778 · PMC12962854 · Genome biology and evolution · 2026 · 8 claims · 6 setups
Eight sericin genes were identified in the A. luna genome, including two clusters of closely related paralogs (serB-D and serE-G)
-
Full-text index only
Polymorphix: a sequence polymorphism database.
PMID 15608242 · PMC540030 · Nucleic acids research · 2005 · 8 claims · 5 setups
Polymorphix is an ACNUC-structured database that organizes EMBL/GenBank sequences into within-species homologous sequence families using similarity and bibliographic criteria, with alignments, outgroups and phylogenetic trees provided.
-
Full-text index only
Sequence similarity network reveals common ancestry of multidomain proteins.
PMID 18475320 · PMC2377100 · PLoS computational biology · 2008 · 8 claims · 6 setups
Traditional homology definitions do not capture multidomain evolution; the authors extend the definition to include domain insertion via a common ancestral locus model.
-
Full-text index only
Mapping proteins to disease terminologies: from UniProt to MeSH.
PMID 18460185 · PMC2367626 · BMC bioinformatics · 2008 · 8 claims · 7 setups
Developed a three-step procedure (disease name extraction, exact matching, partial/similarity-based matching) to map UniProtKB/Swiss-Prot disease names to MeSH terms
-
Full-text index only
VectorBase: a home for invertebrate vectors of human pathogens.
PMID 17145709 · PMC1751530 · Nucleic acids research · 2007 · 8 claims · 5 setups
VectorBase is a web-accessible data repository for information about invertebrate vectors of human pathogens
-
Has reproduction · 77
Representing and querying disease networks using graph databases.
PMID 27462371 · PMC4960687 · BioData mining · 2016 · 7 claims · 8 setups
Graph databases are well suited for representing biological information because it is typically highly connected, semi-structured and unpredictable, unlike relational databases which require rigid schemas.
-
Full-text index only
GENCODE: producing a reference annotation for ENCODE.
PMID 16925838 · PMC1810553 · Genome biology · 2006 · 8 claims · 8 setups
GENCODE annotation combines initial manual annotation by HAVANA, experimental validation, and refinement based on results to identify protein-coding genes in ENCODE regions
-
Full-text index only
Evola: Ortholog database of all human genes in H-InvDB with manual curation of phylogenetic trees.
PMID 17982176 · PMC2238928 · Nucleic acids research · 2008 · 6 claims · 7 setups
Evola combines genome synteny-based computational ortholog detection with manual curation of phylogenetic trees by experts to yield more reliable orthologs than automated pairwise methods
-
Full-text index only
Identifying related L1 retrotransposons by analyzing 3' transduced sequences.
PMID 12734010 · PMC156586 · Genome biology · 2003 · 8 claims · 6 setups
L1 elements with transduction-derived 3' sequence (L1-TDs) can be computationally identified using RepeatMasker/TSDfinder and grouped into families sharing a common progenitor via BLAST comparison of downstream sequences.
-
Full-text index only
TreeFam: a curated database of phylogenetic trees of animal gene families.
PMID 16381935 · PMC1347480 · Nucleic acids research · 2006 · 7 claims · 6 setups
Tree-based inference of orthologs and paralogs is more robust than BLAST-based methods because evolutionary rates (and thus pairwise BLAST scores) vary across gene family members
-
Has reproduction · 84
Genome of the Asian longhorned beetle (Anoplophora glabripennis), a globally significant invasive species, reveals key functional and evolutionary innovations at the beetle-plant interface.
PMID 27832824 · PMC5105290 · Genome biology · 2016 · 8 claims · 7 setups
The A. glabripennis genome encodes a uniquely diverse arsenal of enzymes that degrade plant cell wall polysaccharide networks (cellulose, hemicellulose, pectin) and detoxify plant allelochemicals.
-
Full-text index only
GLIDA: GPCR--ligand database for chemical genomics drug discovery--database and tools update.
PMID 17986454 · PMC2238933 · Nucleic acids research · 2008 · 7 claims · 5 setups
GLIDA is a public relational database integrating biological information on GPCRs with chemical information on their ligands and their binding interactions.
-
Full-text index only
The gene guessing game.
PMID 11025532 · PMC2448377 · Yeast (Chichester, England) · 2000 · 8 claims · 6 setups
Published methods for estimating human gene number diverge widely, from ~30,000 to over 140,000 genes.
-
Full-text index only
SUPERFAMILY--sophisticated comparative genomics, data mining, visualization and phylogeny.
PMID 19036790 · PMC2686452 · Nucleic acids research · 2009 · 7 claims · 6 setups
SUPERFAMILY provides structural, functional and evolutionary annotation for proteins from all completely sequenced genomes using SCOP-based hidden Markov models
-
Has reproduction · 76
The genome and development-dependent transcriptomes of Pyronema confluens: a window into fungal evolution.
PMID 24068976 · PMC3778014 · PLoS genetics · 2013 · 8 claims · 8 setups
The 50 Mb P. confluens genome with 13,369 predicted protein-coding genes is more characteristic of higher filamentous ascomycetes than of the large, repeat-rich Tuber melanosporum genome, showing that the truffle's expanded genome is not typical of the Pezizales.