Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
InParanoid 7: new algorithms and tools for eukaryotic orthology analysis.
PMID 19892828 · PMC2808972 · Nucleic acids research · 2010 · 8 claims · 7 setups
InParanoid 7 expands the database by an order of magnitude to 100 species, 1.3 million proteins, and 42.7 million pairwise ortholog groups.
-
Full-text index only
Evolutionary sequence analysis of complete eukaryote genomes.
PMID 15762985 · PMC1274250 · BMC bioinformatics · 2005 · 8 claims · 6 setups
A conservative genome-comparison method (MIA) identifies panorthologs — strict single-copy 1:1 orthologs containing only species divergences, no paralogy — to minimize errors from gene duplication in evolutionary sequence analysis.
-
Full-text index only
Inparanoid: a comprehensive database of eukaryotic orthologs.
PMID 15608241 · PMC540061 · Nucleic acids research · 2005 · 8 claims · 4 setups
The Inparanoid algorithm identifies true ortholog clusters by seeding on reciprocal best-matching pairs, gathering inparalogs (post-speciation duplicates) while excluding outparalogs (pre-speciation duplicates)
-
Has reproduction · 76
GeneSetCart: assembling, augmenting, combining, visualizing, and analyzing gene sets.
PMID 40208796 · PMC11984350 · GigaScience · 2025 · 8 claims · 8 setups
GeneSetCart is a web-based platform that lets users assemble, augment, combine, visualize, and analyze gene sets from multiple sources in one place
-
Full-text index only
The TIGR Gene Indices: clustering and assembling EST and known genes and integration with eukaryotic genomes.
PMID 15608288 · PMC540018 · Nucleic acids research · 2005 · 8 claims · 8 setups
The TIGR Gene Indices (TGI) are a collection of 77 species-specific databases that cluster and assemble EST and known gene sequences into tentative consensus (TC) sequences to identify and characterize expressed transcripts.
-
Full-text index only
EPGD: a comprehensive web resource for integrating and displaying eukaryotic paralog/paralogon information.
PMID 17984073 · PMC2238967 · Nucleic acids research · 2008 · 8 claims · 8 setups
EPGD is a gene-centered, internet-accessible database integrating paralog family and paralogon information for 26 eukaryotic genomes.
-
Full-text index only
InParanoid 6: eukaryotic ortholog clusters with inparalogs.
PMID 18055500 · PMC2238924 · Nucleic acids research · 2008 · 8 claims · 3 setups
InParanoid 6 is an updated eukaryotic ortholog database covering 35 species (34 eukaryotes plus E. coli as outgroup), providing pairwise ortholog clusters with inparalogs for all species pairs.
-
Full-text index only
The Princeton Protein Orthology Database (P-POD): a comparative genomics analysis tool for biologists.
PMID 17712414 · PMC1942082 · PloS one · 2007 · 8 claims · 5 setups
P-POD is the first comparative genomics database to combine results from multiple computational ortholog/homolog prediction methods with manually curated literature-derived experimental evidence of functional conservation.
-
Full-text index only
CLEAN: CLustering Enrichment ANalysis.
PMID 19640299 · PMC2734555 · BMC bioinformatics · 2009 · 8 claims · 4 setups
The gene-specific CLEAN score improves reproducibility of cluster analysis conclusions across independent datasets compared to the traditional cluster-wide score (cwCLEAN).
-
Has reproduction · 57
Diapause vs. reproductive programs: transcriptional phenotypes in a keystone copepod.
PMID 33782539 · PMC8007741 · Communications biology · 2021 · 8 claims · 7 setups
t-SNE clustering of all-gene expression data groups field-collected (diapause program) samples into one cluster while early and late culture (reproductive program) samples separate into two distinct phenotypes
-
Full-text index only
Developing a set of ancestry-sensitive DNA markers reflecting continental origins of humans.
PMID 19860882 · PMC2775748 · BMC genetics · 2009 · 8 claims · 8 setups
A set of 47 SNPs selected via the 4gen pairwise F_ST approach serves as an ASM panel distinguishing four continental groups (African, Eurasian, Asian/Oceanian, Native American)
-
Full-text index only
A protein interaction based model for schizophrenia study.
PMID 19091023 · PMC2638163 · BMC bioinformatics · 2008 · 8 claims · 4 setups
Products of 36 schizophrenia candidate genes cluster together into a single connected component within a PPI sub-network of 831 proteins
-
Full-text index only
Coiled-coil protein composition of 22 proteomes--differences and common themes in subcellular infrastructure and traffic control.
PMID 16288662 · PMC1322226 · BMC evolutionary biology · 2005 · 7 claims · 5 setups
Proteins with extended coiled-coil domains (>250 amino acids) are largely absent from bacterial genomes but present in archaea and eukaryotes.
-
Full-text index only
Application of functional genomics to primate endometrium: insights into biological processes.
PMID 17118168 · PMC1775064 · Reproductive biology and endocrinology : RB&E · 2006 · 8 claims · 7 setups
Gene expression profiles differ distinctly across proliferative, early-, mid-, and late-secretory phases of the human menstrual cycle, reflecting sequential estradiol and progesterone action.
-
Full-text index only
Meta-analysis of inter-species liver co-expression networks elucidates traits associated with common human diseases.
PMID 20019805 · PMC2787626 · PLoS computational biology · 2009 · 8 claims · 8 setups
A novel semi-parametric meta-analysis method (based on a gene-centric Glass's d effect size) outperforms existing parametric and non-parametric meta-analysis methods at identifying functionally coherent gene pairs across species.
-
Has reproduction · 89
DFAST and DAGA: web-based integrated genome annotation tools and resources.
PMID 27867804 · PMC5107635 · Bioscience of microbiota, food and health · 2016 · 8 claims · 7 setups
DFAST is a web-based bacterial genome annotation and DDBJ submission pipeline with integrated CheckM quality assessment and ANI taxonomic assessment.
-
Has reproduction · 71
Parsimonious Gene Correlation Network Analysis (PGCNA): a tool to define modular gene co-expression for refined molecular stratification in cancer.
PMID 30993001 · PMC6459838 · NPJ systems biology and applications · 2019 · 8 claims · 7 setups
Retaining only the top ~3 most correlated edges per gene (EPG3) combined with FastUnfold clustering (termed PGCNA) produces gene co-expression modules with significantly better separation and enrichment of known biology than using all edges or other clustering methods.
-
Full-text index only
Automatic discovery of cross-family sequence features associated with protein function.
PMID 16409628 · PMC1395344 · BMC bioinformatics · 2006 · 8 claims · 6 setups
A self-supervised data mining approach can find relationships between sequence features and functional annotations without preconceived functional categories.
-
Full-text index only
Variation in genetic admixture and population structure among Latinos: the Los Angeles Latino eye study (LALES).
PMID 19903357 · PMC3087512 · BMC genetics · 2009 · 7 claims · 6 setups
LALES Latinos show strong evidence of recent population admixture, primarily from Native American and European ancestries with smaller Asian and African contributions.
-
Full-text index only
OPTIC: orthologous and paralogous transcripts in clades.
PMID 17933761 · PMC2238935 · Nucleic acids research · 2008 · 6 claims · 7 setups
OPTIC is a database providing gene predictions and orthology assignments for three clades: amniotes (human, dog, mouse, opossum, platypus, chicken), 12 Drosophila species, and 4 Caenorhabditis nematodes.