Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
Polymorphix: a sequence polymorphism database.
PMID 15608242 · PMC540030 · Nucleic acids research · 2005 · 8 claims · 5 setups
Polymorphix is an ACNUC-structured database that organizes EMBL/GenBank sequences into within-species homologous sequence families using similarity and bibliographic criteria, with alignments, outgroups and phylogenetic trees provided.
-
Full-text index only
Evolutionary history of the UCP gene family: gene duplication and selection.
PMID 18980678 · PMC2584656 · BMC evolutionary biology · 2008 · 8 claims · 8 setups
The UCP gene family arose through two ancestral gene duplications early in vertebrate evolution, producing the UCP1, UCP2 and UCP3 lineages.
-
Full-text index only
The RHNumtS compilation: features and bioinformatics approaches to locate and quantify Human NumtS.
PMID 18522722 · PMC2447851 · BMC genomics · 2008 · 8 claims · 4 setups
A consensus Reference Human NumtS compilation (RHNumtS) was produced by comparing Blastn, MegaBlast and BLAT results with previously published compilations.
-
Full-text index only
Molecular evolution of Cide family proteins: novel domain formation in early vertebrates and the subsequent divergence.
PMID 18500987 · PMC2426694 · BMC evolutionary biology · 2008 · 8 claims · 5 setups
Sequences homologous to the CIDE-N domain/NCD show a wide phylogenetic distribution, from hydra and sea anemone to mammals, while true Cide proteins are restricted to vertebrates.
-
Full-text index only
The repertoire of G protein-coupled receptors in the sea squirt Ciona intestinalis.
PMID 18452600 · PMC2396169 · BMC evolutionary biology · 2008 · 8 claims · 5 setups
169 gene products in the Ciona genome were identified as putative GPCRs
-
Full-text index only
Inparanoid: a comprehensive database of eukaryotic orthologs.
PMID 15608241 · PMC540061 · Nucleic acids research · 2005 · 8 claims · 4 setups
The Inparanoid algorithm identifies true ortholog clusters by seeding on reciprocal best-matching pairs, gathering inparalogs (post-speciation duplicates) while excluding outparalogs (pre-speciation duplicates)
-
Full-text index only
Distribution and effects of nonsense polymorphisms in human genes.
PMID 18852891 · PMC2561068 · PloS one · 2008 · 8 claims · 8 setups
Nonsense SNPs occur at a lower density than nonsynonymous SNPs, indicating stronger purifying selection against premature stop codons than amino acid changes.
-
Has reproduction · 50
Ancient gene duplicates in Gossypium (cotton) exhibit near-complete expression divergence.
PMID 24558256 · PMC3971588 · Genome biology and evolution · 2014 · 8 claims · 8 setups
Nearly all (99.4%) ancient paralog pairs in Gossypium raimondii are differentially expressed in at least one of three tissues (petal, leaf, seed), indicating massive, near-complete expression-level divergence.
-
Has reproduction · 59
eDNAmap: A Metabarcoding Web Tool for Comparing Marine Biodiversity, With Special Reference to Teleost Fish.
PMID 41189540 · PMC12627913 · Molecular ecology resources · 2026 · 7 claims · 5 setups
eDNAmap is a web-based platform that maps sampling locations, generates heatmaps to evaluate batch effects, and performs nMDS and cluster analyses using similarity indices on uploaded eDNA composition data
-
Has reproduction · 94
Systematic assessment of pathway databases, based on a diverse collection of user-submitted experiments.
PMID 36088548 · PMC9487593 · Briefings in bioinformatics · 2022 · 8 claims · 6 setups
Well-established, hierarchically organized pathway annotation systems (e.g. GO, Reactome, KEGG) yield the best overall enrichment performance despite covering much of the human genome only in general terms.
-
Full-text index only
Molecular phylogeny of the kelch-repeat superfamily reveals an expansion of BTB/kelch proteins in animals.
PMID 13678422 · PMC222960 · BMC bioinformatics · 2003 · 8 claims · 8 setups
The human genome encodes at least 71 kelch-repeat proteins
-
Full-text index only
Phylogenetic profiling of the Arabidopsis thaliana proteome: what proteins distinguish plants from other organisms?
PMID 15287975 · PMC507878 · Genome biology · 2004 · 8 claims · 6 setups
3,848 Arabidopsis proteins were identified as likely plant-specific based on phylogenetic profiling and EST confirmation in multiple plant species
-
Full-text index only
Protein ranking by semi-supervised network propagation.
PMID 16723003 · PMC1810311 · BMC bioinformatics · 2006 · 8 claims · 5 setups
RankProp, a diffusion-based network propagation algorithm on a PSI-BLAST-derived protein similarity network, significantly outperforms local search methods (BLAST/PSI-BLAST) at detecting remote homologs.
-
Full-text index only
The global landscape of sequence diversity.
PMID 17996061 · PMC2258180 · Genome biology · 2007 · 7 claims · 5 setups
Eukaryotic sequence datasets show substantially greater genetic diversity (higher sequence/gene family discovery rates) than bacterial datasets, likely related to differences in modes of genetic inheritance.
-
Full-text index only
Origin and diversification of the basic helix-loop-helix gene family in metazoans: insights from comparative genomics.
PMID 17335570 · PMC1828162 · BMC evolutionary biology · 2007 · 8 claims · 4 setups
An initial diversification of bHLHs occurred in the pre-Cambrian, prior to metazoan cladogenesis
-
Full-text index only
VectorBase: a home for invertebrate vectors of human pathogens.
PMID 17145709 · PMC1751530 · Nucleic acids research · 2007 · 8 claims · 5 setups
VectorBase is a web-accessible data repository for information about invertebrate vectors of human pathogens
-
Full-text index only
Decoding of superimposed traces produced by direct sequencing of heterozygous indels.
PMID 18654614 · PMC2429969 · PLoS computational biology · 2008 · 7 claims · 3 setups
A dynamic programming method (implemented as web app Indelligent) can decode superimposed allelic sequences from a single mixed trace, using only the observed string of ambiguous peak calls, without a reference sequence or reverse trace.
-
Full-text index only
Columba: an integrated database of proteins, structures, and annotations.
PMID 15801979 · PMC1087474 · BMC bioinformatics · 2005 · 8 claims · 6 setups
COLUMBA physically integrates data from twelve protein structure-related databases (PDB, KEGG, Swiss-Prot, CATH, SCOP, Gene Ontology, ENZYME, etc.) into a single PostgreSQL data warehouse.
-
Full-text index only
NetworKIN: a resource for exploring cellular phosphorylation networks.
PMID 17981841 · PMC2238868 · Nucleic acids research · 2008 · 8 claims · 4 setups
NetworKIN integrates consensus substrate motifs with probabilistic network context modelling to predict cellular kinase-substrate relations.
-
Has reproduction · 87
De Novo Transcriptome Meta-Assembly of the Mixotrophic Freshwater Microalga Euglena gracilis.
PMID 34072576 · PMC8227486 · Genes · 2021 · 6 claims · 8 setups
A consensus transcriptome assembled by combining reads from five independent studies is the most complete E. gracilis transcriptome released to date, outperforming the two previously available transcriptomes (GEFR01 and GDJR01).