Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
Protein coding potential of retroviruses and other transposable elements in vertebrate genomes.
PMID 15716312 · PMC549403 · Nucleic acids research · 2005 · 8 claims · 5 setups
About 1000 genes across four vertebrate gene sets analyzed contain at least one RETRA marker protein domain
-
Full-text index only
SysZNF: the C2H2 zinc finger gene database.
PMID 18974185 · PMC2686507 · Nucleic acids research · 2009 · 7 claims · 6 setups
SysZNF is a database that systematically catalogs C2H2-ZNF genes in human and mouse with physical location, gene models, expression probes, protein domains, homologs, and literature links
-
Full-text index only
Discovering cancer genes by integrating network and functional properties.
PMID 19765316 · PMC2758898 · BMC medical genomics · 2009 · 8 claims · 6 setups
Cancer genes have distinct PPI network topology (higher connectivity, higher clustering coefficient, shorter path length to known cancer genes) compared to non-cancer genes
-
Full-text index only
WebGestalt: an integrated system for exploring gene sets in various biological contexts.
PMID 15980575 · PMC1160236 · Nucleic acids research · 2005 · 8 claims · 6 setups
WebGestalt is an integrated web-based system composed of four modules: gene set management, information retrieval, organization/visualization, and statistics.
-
Full-text index only
Analysis of protein sequence and interaction data for candidate disease gene prediction.
PMID 17020920 · PMC1636487 · Nucleic acids research · 2006 · 8 claims · 7 setups
Combining CPS and CMP using known disease genes as input achieves sensitivity 0.52 and specificity 0.97, reducing candidate lists 13-fold
-
Full-text index only
What makes species unique? The contribution of proteins with obscure features.
PMID 16859532 · PMC1779552 · Genome biology · 2006 · 7 claims · 8 setups
POFs constitute 18-38% (average 26%) of a typical eukaryotic proteome
-
Full-text index only
Identification and functional analyses of 11,769 full-length human cDNAs focused on alternative splicing.
PMID 19880432 · PMC2780955 · DNA research : an international journal for rapid publication of reports on genes and genomes · 2009 · 8 claims · 5 setups
Identified 23,241 human genes transcribed into protein-coding mRNAs using full-length cDNA and 5'-EST sequence data
-
Full-text index only
Analysis of the glutathione S-transferase (GST) gene family.
PMID 15607001 · PMC3500200 · Human genomics · 2004 · 8 claims · 4 setups
The complete human GST gene family comprises 16 genes in six subfamilies: alpha (GSTA), mu (GSTM), omega (GSTO), pi (GSTP), theta (GSTT) and zeta (GSTZ).
-
Full-text index only
Genetic diversity among five T4-like bacteriophages.
PMID 16716236 · PMC1524935 · Virology journal · 2006 · 8 claims · 8 setups
A core set of 82 conserved genes (T4-like genes) is present in all five genomes analyzed, clustered in large collinear blocks.
-
Full-text index only
Dyneins across eukaryotes: a comparative genomic analysis.
PMID 17897317 · PMC2239267 · Traffic (Copenhagen, Denmark) · 2007 · 8 claims · 6 setups
Phylogenetic inference identified nine DHC families (two cytoplasmic, seven axonemal) and six IC families (one cytoplasmic)
-
Full-text index only
CanPredict: a computational tool for predicting cancer-associated missense mutations.
PMID 17537827 · PMC1933186 · Nucleic acids research · 2007 · 8 claims · 7 setups
CanPredict is a web application providing public access to a random forest classifier that combines SIFT, LogR.E-value, and GOSS scores to predict whether a missense mutation is cancer-associated
-
Full-text index only
Upgrades to StellaBase facilitate medical and genetic studies on the starlet sea anemone, Nematostella vectensis.
PMID 17982171 · PMC2238866 · Nucleic acids research · 2008 · 6 claims · 5 setups
StellaBase Disease houses homology data for 155,904 invertebrate isoforms of human disease genes across four model systems, including 14,874 predicted Nematostella genes
-
Full-text index only
Pseudofam: the pseudogene families database.
PMID 18957444 · PMC2686518 · Nucleic acids research · 2009 · 8 claims · 7 setups
Pseudofam is an online database of pseudogene families built by mapping pseudogenes to Pfam protein families, providing query tools, statistics, and sequence alignments
-
Full-text index only
Intrinsic structural disorder confers cellular viability on oncogenic fusion proteins.
PMID 19888473 · PMC2768585 · PLoS computational biology · 2009 · 8 claims · 5 setups
Translocation-related human proteins are significantly enriched in intrinsic structural disorder compared to all human proteins
-
Full-text index only
The genome of the simian and human malaria parasite Plasmodium knowlesi.
PMID 18843368 · PMC2656934 · Nature · 2008 · 8 claims · 7 setups
The P. knowlesi (H strain) nuclear genome was sequenced and assembled: 23.5 Mb across 14 chromosomes with 5,188 predicted protein-encoding genes.
-
Has reproduction · 78
Transcriptomic and physiological analysis of atractylodes chinensis in response to drought stress reveals the putative genes related to sesquiterpenoid biosynthesis.
PMID 38317086 · PMC10845750 · BMC plant biology · 2024 · 8 claims · 6 setups
Drought stress significantly increases MDA, proline, soluble sugar, and crude protein content and antioxidative enzyme (SOD, POD, CAT) activity in A. chinensis seedlings
-
Has reproduction · 92
Telomere-to-telomere reference genome for Panax ginseng highlights the evolution of saponin biosynthesis.
PMID 38883331 · PMC11179851 · Horticulture research · 2024 · 8 claims · 8 setups
A telomere-to-telomere reference genome of P. ginseng was assembled (3.45 Gb, 24 chromosomes, 77266 protein-coding genes)
-
Full-text index only
Molecular phylogeny of the kelch-repeat superfamily reveals an expansion of BTB/kelch proteins in animals.
PMID 13678422 · PMC222960 · BMC bioinformatics · 2003 · 8 claims · 8 setups
The human genome encodes at least 71 kelch-repeat proteins
-
Full-text index only
NovelFam3000--uncharacterized human protein domains conserved across model organisms.
PMID 16533400 · PMC1440326 · BMC genomics · 2006 · 8 claims · 7 setups
NovelFam3000 is an online data centre unifying bioinformatics resource links, news, comments, and user-submitted experimental data (including a Gene Characterization Index) for ~3000 uncharacterized Pfam-B/DUF domain families conserved across worm, fly, and human
-
Has reproduction · 94
Large-Scale Phylogenomics of the Lactobacillus casei Group Highlights Taxonomic Inconsistencies and Reveals Novel Clade-Associated Features.
PMID 28845461 · PMC5566788 · mSystems · 2017 · 8 claims · 8 setups
The L. casei group resolves into three distinct clades (A, B, C) supported by phylogeny, GC content, ANI, and TETRA, and many strains are misclassified relative to their nearest type strain.