Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
pTARGET: a web server for predicting protein subcellular localization.
PMID 16844995 · PMC1538910 · Nucleic acids research · 2006 · 7 claims · 3 setups
pTARGET web server predicts nine distinct subcellular localizations in eukaryotic non-plant proteins using an algorithm based on location-specific Pfam domain occurrence patterns and amino acid composition (AAC)
-
Full-text index only
Diversity of tRNA genes in eukaryotes.
PMID 17088292 · PMC1693877 · Nucleic acids research · 2006 · 8 claims · 6 setups
The number of tRNA genes having the same anticodon but different sequences elsewhere (isodecoder genes) varies significantly (10–246) across 11 eukaryotes despite isoacceptor numbers being similar (41–55)
-
Has reproduction · 38
Genomic capacities for Reactive Oxygen Species metabolism across marine phytoplankton.
PMID 37098087 · PMC10128935 · PloS one · 2023 · 8 claims · 3 setups
Genes encoding superoxide (O2•−) scavenging are ubiquitous across phytoplankton, but their fractional gene allocation decreases with increasing cell radius, consistent with a nearly fixed core gene set.
-
Full-text index only
SelenoDB 1.0 : a database of selenoprotein genes, proteins and SECIS elements.
PMID 18174224 · PMC2238826 · Nucleic acids research · 2008 · 6 claims · 5 setups
Standard genome annotation pipelines misannotate selenoprotein genes because they rely on UGA as a universal stop codon, failing to recognize its dual role as the selenocysteine-recoding codon.
-
Full-text index only
Automatic discovery of cross-family sequence features associated with protein function.
PMID 16409628 · PMC1395344 · BMC bioinformatics · 2006 · 8 claims · 6 setups
A self-supervised data mining approach can find relationships between sequence features and functional annotations without preconceived functional categories.
-
Full-text index only
Pegasys: software for executing and integrating analyses of biological sequences.
PMID 15096276 · PMC406494 · BMC bioinformatics · 2004 · 8 claims · 7 setups
Pegasys is a flexible, modular, customizable software system for executing and integrating heterogeneous biological sequence analysis tools
-
Full-text index only
PeroxisomeDB: a database for the peroxisomal proteome, functional genomics and disease.
PMID 17135190 · PMC1747181 · Nucleic acids research · 2007 · 8 claims · 6 setups
PeroxisomeDB integrates the complete peroxisomal proteome of Homo sapiens and Saccharomyces cerevisiae into interrelated 'Genes', 'Functions', 'Metabolic pathways' and 'Diseases' sections with links to NCBI, ENSEMBL and UCSC
-
Full-text index only
The global landscape of sequence diversity.
PMID 17996061 · PMC2258180 · Genome biology · 2007 · 7 claims · 5 setups
Eukaryotic sequence datasets show substantially greater genetic diversity (higher sequence/gene family discovery rates) than bacterial datasets, likely related to differences in modes of genetic inheritance.
-
Full-text index only
Pseudofam: the pseudogene families database.
PMID 18957444 · PMC2686518 · Nucleic acids research · 2009 · 8 claims · 7 setups
Pseudofam is an online database of pseudogene families built by mapping pseudogenes to Pfam protein families, providing query tools, statistics, and sequence alignments
-
Full-text index only
DBD--taxonomically broad transcription factor predictions: new content and functionality.
PMID 18073188 · PMC2238844 · Nucleic acids research · 2008 · 8 claims · 3 setups
DBD is a database of predicted sequence-specific DNA-binding transcription factors covering over 700 publicly available proteomes, up from 150 in the initial version.
-
Full-text index only
The human L-threonine 3-dehydrogenase gene is an expressed pseudogene.
PMID 12361482 · PMC131051 · BMC genetics · 2002 · 8 claims · 7 setups
The human TDH gene is located at chromosome 8p23-22, spans 10 kb, and has 8 exons that would be expected to encode a 369-residue ORF.
-
Full-text index only
Reconstructing the evolution of the mitochondrial ribosomal proteome.
PMID 17604309 · PMC1950548 · Nucleic acids research · 2007 · 8 claims · 6 setups
The ancestral mitoribosome was of alpha-proteobacterial descent and more than doubled its protein content in most eukaryotic lineages.
-
Full-text index only
Genome-wide in silico identification and analysis of cis natural antisense transcripts (cis-NATs) in ten species.
PMID 16849434 · PMC1524920 · Nucleic acids research · 2006 · 8 claims · 7 setups
A fast integrative in silico pipeline combining UniGene mRNA/EST mapping to GoldenPath genomes with CDS, poly(A) signal, poly(A) tail and splicing site evidence can reliably identify cis-NATs genome-wide across multiple species
-
Full-text index only
The TIGR Gene Indices: clustering and assembling EST and known genes and integration with eukaryotic genomes.
PMID 15608288 · PMC540018 · Nucleic acids research · 2005 · 8 claims · 8 setups
The TIGR Gene Indices (TGI) are a collection of 77 species-specific databases that cluster and assemble EST and known gene sequences into tentative consensus (TC) sequences to identify and characterize expressed transcripts.
-
Full-text index only
MODBASE, a database of annotated comparative protein structure models and associated resources.
PMID 18948282 · PMC2686492 · Nucleic acids research · 2009 · 8 claims · 8 setups
MODBASE contains 5,152,695 reliable comparative protein structure models for 1,593,209 unique protein sequences.
-
Full-text index only
Metagenomic analysis of respiratory tract DNA viral communities in cystic fibrosis and non-cystic fibrosis individuals.
PMID 19816605 · PMC2756586 · PloS one · 2009 · 8 claims · 8 setups
CF phage communities are highly similar to each other, whereas Non-CF individuals have more distinct, variable phage communities reflecting transient environmental sampling
-
Full-text index only
Comprehensive genome analysis of 203 genomes provides structural genomics with new insights into protein family space.
PMID 16481312 · PMC1373602 · Nucleic acids research · 2006 · 8 claims · 7 setups
The number of protein families continues to expand steadily as more genomes are sequenced, showing no sign of saturation.
-
Has reproduction · 85
What defines a photosynthetic microbial mat in western Antarctica?
PMID 40043057 · PMC11882083 · PloS one · 2025 · 6 claims · 8 setups
Taxonomic composition of Antarctic microbial mat communities is characterized by similar bacterial groups across regions
-
Has reproduction · 76
The genome and development-dependent transcriptomes of Pyronema confluens: a window into fungal evolution.
PMID 24068976 · PMC3778014 · PLoS genetics · 2013 · 8 claims · 8 setups
The 50 Mb P. confluens genome with 13,369 predicted protein-coding genes is more characteristic of higher filamentous ascomycetes than of the large, repeat-rich Tuber melanosporum genome, showing that the truffle's expanded genome is not typical of the Pezizales.
-
Full-text index only
Reconstruction of pathways associated with amino acid metabolism in human mitochondria.
PMID 18267298 · PMC5054205 · Genomics, proteomics & bioinformatics · 2007 · 8 claims · 5 setups
Out of 20 amino acids, the metabolic pathways of 17 utilize mitochondrial enzymes, and dysfunction of these enzymes causes over 40 known human mitochondrial diseases/disorders