Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
Functional coverage of the human genome by existing structures, structural genomics targets, and homology models.
PMID 16118666 · PMC1188274 · PLoS computational biology · 2005 · 8 claims · 5 setups
Existing PDB structures provide single-domain coverage for 37% of functional classes in the human genome and complete (whole-protein) structure coverage for 25%.
-
Full-text index only
Automatic discovery of cross-family sequence features associated with protein function.
PMID 16409628 · PMC1395344 · BMC bioinformatics · 2006 · 8 claims · 6 setups
A self-supervised data mining approach can find relationships between sequence features and functional annotations without preconceived functional categories.
-
Has reproduction · 68
Mining the equine gut metagenome: poorly-characterized taxa associated with cardiovascular fitness in endurance athletes.
PMID 36192523 · PMC9529974 · Communications biology · 2022 · 8 claims · 8 setups
Built an integrated horse gut microbiome gene catalog (~25 million unique genes) and 372 metagenome-assembled genomes (MAGs) spanning 4179 genera and 95 phyla
-
Full-text index only
MODBASE: a database of annotated comparative protein structure models and associated resources.
PMID 16381869 · PMC1347422 · Nucleic acids research · 2006 · 8 claims · 7 setups
MODBASE is a database of automatically calculated comparative protein structure models covering all UniProt sequences matchable to a known structure
-
Full-text index only
Analysis of expressed sequence tags from Actinidia: applications of a cross species EST database for gene discovery in the areas of flavor, health, color and ripening.
PMID 18655731 · PMC2515324 · BMC genomics · 2008 · 7 claims · 6 setups
A collection of 132,577 ESTs from four Actinidia species was generated and clustered into 41,858 non-redundant clusters (18,070 TCs and 23,788 singletons)
-
Has reproduction · 24
MiGPC: a comprehensive catalog of enzybiotics from environmental metagenomes.
PMID 41888223 · PMC13172421 · Scientific reports · 2026 · 8 claims · 8 setups
MiGPC is the first genome-resolved metagenomic gene and protein catalog specifically targeted to enzybiotics
-
Full-text index only
The Universal Protein Resource (UniProt) in 2010.
PMID 19843607 · PMC2808944 · Nucleic acids research · 2010 · 8 claims · 5 setups
UniProt is a centralized, freely accessible, comprehensive knowledgebase of protein sequence and functional annotation maintained by the EBI, SIB and PIR consortium.
-
Full-text index only
ARED 3.0: the large and diverse AU-rich transcriptome.
PMID 16381826 · PMC1347415 · Nucleic acids research · 2006 · 7 claims · 6 setups
ARED 3.0 computationally mapped more than 4000 ARE-mRNAs to the human genome, representing 5-8% of human genes.
-
Full-text index only
An integrated database-pipeline system for studying single nucleotide polymorphisms and diseases.
PMID 19091018 · PMC2638159 · BMC bioinformatics · 2008 · 6 claims · 5 setups
Existing SNP/disease databases are fragmented; no combined resource widely supports gene-, SNP-, and disease-related information together
-
Has reproduction · 63
Comparative transcriptome analysis of tomato (Solanum lycopersicum) in response to exogenous abscisic acid.
PMID 24289302 · PMC4046761 · BMC genomics · 2013 · 8 claims · 7 setups
Exogenous ABA alters the expression of a majority (54.73%) of expressed tomato leaf transcripts, with 2,787 significantly differentially expressed genes, predominantly up-regulated.
-
Full-text index only
The global landscape of sequence diversity.
PMID 17996061 · PMC2258180 · Genome biology · 2007 · 7 claims · 5 setups
Eukaryotic sequence datasets show substantially greater genetic diversity (higher sequence/gene family discovery rates) than bacterial datasets, likely related to differences in modes of genetic inheritance.