Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
The TIGR Gene Indices: clustering and assembling EST and known genes and integration with eukaryotic genomes.
PMID 15608288 · PMC540018 · Nucleic acids research · 2005 · 8 claims · 8 setups
The TIGR Gene Indices (TGI) are a collection of 77 species-specific databases that cluster and assemble EST and known gene sequences into tentative consensus (TC) sequences to identify and characterize expressed transcripts.
-
Full-text index only
GeneSeer: a sage for gene names and genomic resources.
PMID 16176584 · PMC1266031 · BMC genomics · 2005 · 7 claims · 4 setups
GeneSeer aggregates gene name synonyms from GenBank, FlyBase, ExPASy, HUGO, ENSEMBL, UCSC and Gene Ontology into a name-translation database that maps any familiar name to a reference (SOFAR) identifier.
-
Full-text index only
Eighth major clade for hepatitis delta virus.
PMID 17073101 · PMC3294742 · Emerging infectious diseases · 2006 · 7 claims · 7 setups
Three HDV isolates (dFr644, dFr2072, dFr2736) form a monophyletic group distinct from HDV-1 through HDV-7, constituting a new eighth major clade (HDV-8) of the Deltavirus genus.
-
Full-text index only
Database resources of the National Center for Biotechnology Information.
PMID 17170002 · PMC1781113 · Nucleic acids research · 2007 · 8 claims · 8 setups
NCBI maintains an integrated suite of database resources (Entrez, PubMed, RefSeq, dbSNP, BLAST, etc.) for molecular biology data retrieval and analysis
-
Full-text index only
Identifying cis-regulatory sequences by word profile similarity.
PMID 19730735 · PMC2731932 · PloS one · 2009 · 8 claims · 8 setups
WPH-finder identifies putative co-regulated CRMs by scanning the genome for sequences with word profiles similar to a known CRM, without explicitly defining binding sites
-
Has reproduction · 49
oPOSSUM-3: advanced analysis of regulatory motif over-representation across genes or ChIP-Seq datasets.
PMID 22973536 · PMC3429929 · G3 (Bethesda, Md.) · 2012 · 8 claims · 6 setups
oPOSSUM-3 is a web-accessible system that identifies over-represented TFBS and TFBS families in DNA sequences of co-expressed genes or in sequences from high-throughput methods such as ChIP-Seq.
-
Full-text index only
Protein ranking by semi-supervised network propagation.
PMID 16723003 · PMC1810311 · BMC bioinformatics · 2006 · 8 claims · 5 setups
RankProp, a diffusion-based network propagation algorithm on a PSI-BLAST-derived protein similarity network, significantly outperforms local search methods (BLAST/PSI-BLAST) at detecting remote homologs.
-
Full-text index only
The Princeton Protein Orthology Database (P-POD): a comparative genomics analysis tool for biologists.
PMID 17712414 · PMC1942082 · PloS one · 2007 · 8 claims · 5 setups
P-POD is the first comparative genomics database to combine results from multiple computational ortholog/homolog prediction methods with manually curated literature-derived experimental evidence of functional conservation.
-
Full-text index only
KEGG for linking genomes to life and the environment.
PMID 18077471 · PMC2238879 · Nucleic acids research · 2008 · 8 claims · 4 setups
KEGG provides a reference knowledge base for linking genomes to life via PATHWAY mapping and to the environment via BRITE mapping.
-
Full-text index only
Ontological Discovery Environment: a system for integrating gene-phenotype associations.
PMID 19733230 · PMC2783409 · Genomics · 2009 · 8 claims · 8 setups
ODE is a web-based system for storing, sharing, retrieving and analyzing phenotype-centered genomic data sets across species and experimental systems
-
Full-text index only
Ensembl 2005.
PMID 15608235 · PMC540092 · Nucleic acids research · 2005 · 8 claims · 4 setups
Ensembl's automatic gene build system can flexibly and reliably annotate a wide variety of genomes with limited species-specific evidence.
-
Has reproduction · 76
GeneSetCart: assembling, augmenting, combining, visualizing, and analyzing gene sets.
PMID 40208796 · PMC11984350 · GigaScience · 2025 · 8 claims · 8 setups
GeneSetCart is a web-based platform that lets users assemble, augment, combine, visualize, and analyze gene sets from multiple sources in one place
-
Full-text index only
The biological function of some human transcription factor binding motifs varies with position relative to the transcription start site.
PMID 18367472 · PMC2377430 · Nucleic acids research · 2008 · 8 claims · 5 setups
1226 eight-letter DNA words show statistically significant positional preferences relative to the TSS across 7914 human promoter regions
-
Full-text index only
Comparative genomics of the syndecans defines an ancestral genomic context associated with matrilins in vertebrates.
PMID 16620374 · PMC1464127 · BMC genomics · 2006 · 8 claims · 6 setups
Syndecan-encoding sequences are present in Cnidaria and throughout the Bilateria, showing deep conservation of the family.
-
Has reproduction · 65
FusionQ: a novel approach for gene fusion detection and quantification from paired-end RNA-Seq.
PMID 23768108 · PMC3691734 · BMC bioinformatics · 2013 · 8 claims · 8 setups
FusionQ is a novel tool that detects gene fusions, constructs chimerical transcript structures, and estimates their abundances from paired-end RNA-Seq data.
-
Full-text index only
Peptide bioinformatics: peptide classification using peptide machines.
PMID 19065810 · PMC7122642 · Methods in molecular biology (Clifton, N.J.) · 2008 · 8 claims · 4 setups
The bio-basis function, which converts peptides into numerical vectors using nongapped pairwise homology alignment scores against indicator peptides, can statistically quantify peptide similarity for classification.
-
Full-text index only
Inparanoid: a comprehensive database of eukaryotic orthologs.
PMID 15608241 · PMC540061 · Nucleic acids research · 2005 · 8 claims · 4 setups
The Inparanoid algorithm identifies true ortholog clusters by seeding on reciprocal best-matching pairs, gathering inparalogs (post-speciation duplicates) while excluding outparalogs (pre-speciation duplicates)
-
Full-text index only
A genome-wide survey of segmental duplications that mediate common human genetic variation of chromosomal architecture.
PMID 15588494 · PMC3525102 · Human genomics · 2004 · 8 claims · 5 setups
PSD-mediated genomic architecture analogous to the 8p23/4p16 inversion regions is not unique to those loci but recurs genome-wide.
-
Full-text index only
SVC: structured visualization of evolutionary sequence conservation.
PMID 15991338 · PMC1160265 · Nucleic acids research · 2005 · 7 claims · 5 setups
SVC aligns protein-coding sequences of orthologous gene pairs and maps them back onto their encoding exons/introns to generate a scaffold of conserved gene structure.
-
Full-text index only
Empirical codon substitution matrix.
PMID 15927081 · PMC1173088 · BMC bioinformatics · 2005 · 8 claims · 5 setups
The authors present the first empirical codon substitution matrix built entirely from alignments of vertebrate coding DNA sequences.