Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
A computational screen for type I polyketide synthases in metagenomics shotgun data.
PMID 18953415 · PMC2568958 · PloS one · 2008 · 8 claims · 6 setups
Combining HMM domain searches with maximum-likelihood phylogenetic trees can discriminate true PKS I sequences from evolutionarily related but functionally different enzymes (e.g., FAS I) in metagenomic data.
-
Full-text index only
Importance sampling for the infinite sites model.
PMID 18976228 · PMC2832804 · Statistical applications in genetics and molecular biology · 2008 · 7 claims · 2 setups
A new importance sampling proposal distribution for the ISM, derived from a new result on exact sampling from a single segregating site, generally shows greater efficiency than the GT and SD proposals.
-
Full-text index only
Evolutionary sequence analysis of complete eukaryote genomes.
PMID 15762985 · PMC1274250 · BMC bioinformatics · 2005 · 8 claims · 6 setups
A conservative genome-comparison method (MIA) identifies panorthologs — strict single-copy 1:1 orthologs containing only species divergences, no paralogy — to minimize errors from gene duplication in evolutionary sequence analysis.
-
Full-text index only
EGenBio: a data management system for evolutionary genomics and biodiversity.
PMID 17118150 · PMC1683573 · BMC bioinformatics · 2006 · 7 claims · 7 setups
EGenBio is a web-based system for integrated management, filtering, curation, and visualization of large-scale genomic sequences, alignments, and phylogenetic trees for evolutionary genomics and biodiversity research.
-
Full-text index only
A model-based approach to selection of tag SNPs.
PMID 16776821 · PMC1525207 · BMC bioinformatics · 2006 · 7 claims · 5 setups
The Li and Stephens hidden Markov model outperforms other tested models (simple Markov, two-state HMM, HMM-4D, greedy GR-1/GR-2) in description code-length, tag set information content, and prediction of tagged SNPs.
-
Full-text index only
Getting positive about selection.
PMID 12914654 · PMC193638 · Genome biology · 2003 · 8 claims · 4 setups
Purifying selection is the predominant form of molecular evolution, preserving fitness by eliminating deleterious mutations, while positive selection is rare but critical for adaptation.
-
Full-text index only
Filtering high-throughput protein-protein interaction data using a combination of genomic features.
PMID 15833142 · PMC1127019 · BMC bioinformatics · 2005 · 8 claims · 8 setups
A combination of three genomic features (interacting Pfam domains, GO annotations, sequence homology) using naive Bayesian networks predicts true protein-protein interactions with high sensitivity and good specificity.
-
Full-text index only
The other side of comparative genomics: genes with no orthologs between the cow and other mammalian species.
PMID 20003425 · PMC2808326 · BMC genomics · 2009 · 7 claims · 4 setups
3,801 bovine genes have no orthologs in human, mouse and dog, and 1,010 human genes have no orthologs in cow despite having orthologs in mouse and dog
-
Full-text index only
PhylomeDB: a database for genome-wide collections of gene phylogenies.
PMID 17962297 · PMC2238872 · Nucleic acids research · 2008 · 7 claims · 6 setups
PhylomeDB is a publicly accessible database storing complete, genome-wide collections of gene phylogenies (phylomes).
-
Full-text index only
Empirical codon substitution matrix.
PMID 15927081 · PMC1173088 · BMC bioinformatics · 2005 · 8 claims · 5 setups
The authors present the first empirical codon substitution matrix built entirely from alignments of vertebrate coding DNA sequences.
-
Full-text index only
Computational disease gene identification: a concert of methods prioritizes type 2 diabetes and obesity candidate genes.
PMID 16757574 · PMC1475747 · Nucleic acids research · 2006 · 6 claims · 8 setups
Applying seven independent computational disease-gene prioritization methods in concert to 9556 positional candidate genes identifies a prioritized set of likely T2D and obesity candidate genes
-
Has reproduction · 87
A target enrichment method for gathering phylogenetic information from hundreds of loci: An example from the Compositae.
PMID 25202605 · PMC4103609 · Applications in plant sciences · 2014 · 8 claims · 8 setups
A custom sequence capture probe set (9678 baits targeting 1061 orthologous genes) was designed to enrich COS loci across the Compositae.
-
Full-text index only
Genome-wide prioritization of disease genes and identification of disease-disease associations from an integrated human functional linkage network.
PMID 19728866 · PMC2768980 · Genome biology · 2009 · 6 claims · 6 setups
Integrating 16 genomic features (32 sub-features) via a naïve Bayes classifier produces a genome-scale FLN of 21,657 human genes and 22,388,609 weighted links that outperforms any individual data source for inferring functional linkages.
-
Full-text index only
Pathway-specific canalization and plasticity of gene expression during C. elegans dauer development.
PMID 41890963 · PMC13014971 · iScience · 2026 · 8 claims · 7 setups
Dauers induced by different environmental (starvation, pheromone, heat) or genetic (daf-2, daf-7, ilc-17.1, cep-1 OE) stimuli are transcriptionally distinct from each other and from continuously developing WT L2/L3 and adults
-
Full-text index only
Speeding disease gene discovery by sequence based candidate prioritization.
PMID 15766383 · PMC1274252 · BMC bioinformatics · 2005 · 7 claims · 8 setups
Disease genes (OMIM) differ significantly from non-disease genes in sequence-based features including gene/cDNA/protein size, exon number, homolog conservation, secretion signal, 3' UTR length, CpG islands, and distance to nearest gene.
-
Full-text index only
PeroxisomeDB: a database for the peroxisomal proteome, functional genomics and disease.
PMID 17135190 · PMC1747181 · Nucleic acids research · 2007 · 8 claims · 6 setups
PeroxisomeDB integrates the complete peroxisomal proteome of Homo sapiens and Saccharomyces cerevisiae into interrelated 'Genes', 'Functions', 'Metabolic pathways' and 'Diseases' sections with links to NCBI, ENSEMBL and UCSC
-
Has reproduction · 91
Genomic Description of 'Candidatus Abyssubacteria,' a Novel Subsurface Lineage Within the Candidate Phylum Hydrogenedentes.
PMID 30210471 · PMC6121073 · Frontiers in microbiology · 2018 · 8 claims · 7 setups
SURF_5 and SURF_17 are the first full genomes of a novel bacterial lineage, 'Candidatus Abyssubacteria,' within the candidate phylum Hydrogenedentes
-
Has reproduction
Genome-wide signatures of convergent evolution in echolocating mammals.
PMID 24005325 · PMC3836225 · Nature · 2013 · 8 claims · 8 setups
Genome-wide convergent sequence evolution between echolocating lineages is not rare but widespread and continuously distributed, with signatures consistent with convergence in nearly 200 loci out of 2,326 examined.
-
Has reproduction · 93
Elucidation of the molecular responses to waterlogging in Jatropha roots by transcriptome profiling.
PMID 25520726 · PMC4251292 · Frontiers in plant science · 2014 · 8 claims · 8 setups
24 h of waterlogging significantly alters mRNA abundance of 1968 genes in Jatropha roots (931 up, 1037 down).
-
Full-text index only
Single-cell sequencing reveals unexpected genetic diversity among Bodo spp. flagellates and their bacterial endosymbionts.
PMID 41848149 · PMC12999062 · Microbial genomics · 2026 · 8 claims · 8 setups
Seven single-cell genomes assembled from uncultured environmental Bodo cells represent three potentially novel Bodo species