Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
Optimality driven nearest centroid classification from genomic data.
PMID 17912341 · PMC1991588 · PloS one · 2007 · 7 claims · 5 setups
A theoretical result determines the subset of features of a given size that minimizes the misclassification rate for a nearest-centroid (LDA) classifier, based on equation (4).
-
Full-text index only
The DAVID Gene Functional Classification Tool: a novel biological module-centric algorithm to functionally analyze large gene lists.
PMID 17784955 · PMC2375021 · Genome biology · 2007 · 8 claims · 6 setups
Gene-gene functional similarity can be measured using kappa statistics applied to a binary gene-annotation-term matrix built from 14 annotation categories.
-
Full-text index only
Does distance matter? Variations in alternative 3' splicing regulation.
PMID 17704130 · PMC2018619 · Nucleic acids research · 2007 · 8 claims · 7 setups
Alternative 3' splice sites can be distinguished from constitutive splice sites by a combination of sequence/conservation properties that vary depending on the distance between the splice sites.
-
Full-text index only
Using comparative genomics to reorder the human genome sequence into a virtual sheep genome.
PMID 17663790 · PMC2323240 · Genome biology · 2007 · 8 claims · 6 setups
A sheep BAC library (CHORI-243) with ~13.5-fold genome coverage was constructed and end-sequenced.
-
Full-text index only
FatiGO +: a functional profiling tool for genomic data. Integration of functional annotation, regulatory motifs and interaction data with microarray experiments.
PMID 17478504 · PMC1933151 · Nucleic acids research · 2007 · 8 claims · 8 setups
FatiGO+ is a web-based tool for functional profiling of genome-scale experiments that integrates functional annotation, regulatory motifs and interaction data
-
Full-text index only
Patterns of somatic mutation in human cancer genomes.
PMID 17344846 · PMC2712719 · Nature · 2007 · 8 claims · 5 setups
Systematic resequencing of a large gene family (protein kinases) across diverse cancers reveals a larger repertoire of cancer genes than previously anticipated
-
Full-text index only
QuantiSNP: an Objective Bayes Hidden-Markov Model to detect and accurately map copy number variation using SNP genotyping data.
PMID 17341461 · PMC1874617 · Nucleic acids research · 2007 · 8 claims · 7 setups
QuantiSNP (OB-HMM) provides probabilistic quantification of copy number states and significantly improves accuracy of segmental aneuploidy identification and breakpoint mapping relative to existing tools (BeadStudio/Illumina)
-
Full-text index only
Synonymous substitution rates predict HIV disease progression as a result of underlying replication dynamics.
PMID 17305421 · PMC1797821 · PLoS computational biology · 2007 · 8 claims · 8 setups
The synonymous substitution rate (dS) of HIV env is strongly correlated with disease progression parameters (progression time, CD4+ decline rate, viral load increase rate), unlike the nonsynonymous rate (dN).
-
Full-text index only
Epidemiology of doublet/multiplet mutations in lung cancers: evidence that a subset arises by chronocoordinate events.
PMID 19005564 · PMC2579325 · PloS one · 2008 · 8 claims · 7 setups
Doublet mutations are significantly more frequent in EGFR (6.0%) and TP53 (2.3%) in human lung cancer than spontaneous doublets in mouse lacI (0.7%), about 8-fold and 3-fold higher respectively.
-
Full-text index only
Evolutionary modeling of rate shifts reveals specificity determinants in HIV-1 subtypes.
PMID 18989394 · PMC2566816 · PLoS computational biology · 2008 · 7 claims · 4 setups
A novel Bayesian method, RASER, can detect site-specific evolutionary rate shifts and the lineages in which they occurred without pre-specifying candidate lineages.
-
Full-text index only
High-throughput chromatin information enables accurate tissue-specific prediction of transcription factor binding sites.
PMID 18988630 · PMC2662491 · Nucleic acids research · 2009 · 8 claims · 8 setups
Incorporating H3K4me3 chromatin modification estimates greatly improves the accuracy of in silico prediction of in vivo TF binding for a wide range of TFs in human and mouse
-
Full-text index only
Proteomic profiling of Plasmodium sporozoite maturation identifies new proteins essential for parasite development and infectivity.
PMID 18974882 · PMC2570797 · PLoS pathogens · 2008 · 8 claims · 6 setups
Midgut (oocyst-derived) and salivary gland sporozoite proteomes are markedly different despite near-identical morphology, consistent with their differing hepatocyte infectivity
-
Full-text index only
Cataloging coding sequence variations in human genome databases.
PMID 18974781 · PMC2570488 · PloS one · 2008 · 8 claims · 7 setups
A significant proportion of CVs overlap between HGMD and dbSNP (4.36% of HGMD CVs registered in dbSNP; 8.11% of dbSNP CVs registered in HGMD), warranting caution when interpreting phenotypic relevance of concurrent CVs.
-
Full-text index only
MODBASE, a database of annotated comparative protein structure models and associated resources.
PMID 18948282 · PMC2686492 · Nucleic acids research · 2009 · 8 claims · 8 setups
MODBASE contains 5,152,695 reliable comparative protein structure models for 1,593,209 unique protein sequences.
-
Full-text index only
Genomics, molecular imaging, bioinformatics, and bio-nano-info integration are synergistic components of translational medicine and personalized healthcare research.
PMID 18831773 · PMC3226104 · BMC genomics · 2008 · 8 claims · 8 setups
Genomics, molecular imaging, bioinformatics, and bio-nano-info integration are synergistic components of translational medicine and personalized healthcare
-
Full-text index only
Adaptive discriminant function analysis and reranking of MS/MS database search results for improved peptide identification in shotgun proteomics.
PMID 18788775 · PMC3744223 · Journal of proteome research · 2008 · 7 claims · 4 setups
PeptideProphet's fixed LDA coefficients for combining search scores (Xcorr', ΔCn, SpRank) may not be optimal under all search/instrument conditions.
-
Full-text index only
Phylogenetic analysis of mRNA polyadenylation sites reveals a role of transposable elements in evolution of the 3'-end of genes.
PMID 18757892 · PMC2553571 · Nucleic acids research · 2008 · 8 claims · 6 setups
3'-most (L type) poly(A) sites are more conserved than upstream F/M type sites, while intronic (C/H type) sites are the least conserved
-
Full-text index only
Systems genetics of alcoholism.
PMID 23584748 · PMC3860445 · Alcohol research & health : the journal of the National Institute on Alcohol Abuse and Alcoholism · 2008 · 8 claims · 8 setups
Alcoholism is a multifactorial disease driven by interacting genetic, social, and environmental factors, with genetics accounting for 50-60% of risk.
-
Full-text index only
Prioritization of candidate cancer genes--an aid to oncogenomic studies.
PMID 18710882 · PMC2566894 · Nucleic acids research · 2008 · 8 claims · 8 setups
Computational classifiers using combinations of protein conservation, gene structure, protein domains, protein interactions, and regulatory data can distinguish known cancer genes (CD/CR) from unlabelled human genes
-
Full-text index only
A comparison of random sequence reads versus 16S rDNA sequences for estimating the biodiversity of a metagenomic library.
PMID 18682527 · PMC2532719 · Nucleic acids research · 2008 · 8 claims · 7 setups
Biodiversity observed by RSR analysis is consistent with that obtained by 16S rDNA analysis