Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
Genomic and epigenetic instability in colorectal cancer pathogenesis.
PMID 18773902 · PMC2866182 · Gastroenterology · 2008 · 8 claims · 7 setups
Genomic instability (CIN or MSI) is a key early molecular step in colorectal tumorigenesis that may initiate rather than merely accompany the adenoma-carcinoma sequence
-
Full-text index only
The Functional RNA Database 3.0: databases to support mining and annotation of functional RNAs.
PMID 18948287 · PMC2686472 · Nucleic acids research · 2009 · 8 claims · 5 setups
fRNAdb 3.0 is a completely rebuilt sequence database hosting a much larger collection of known/predicted non-coding RNA sequences with improved search functionality
-
Full-text index only
Cataloging coding sequence variations in human genome databases.
PMID 18974781 · PMC2570488 · PloS one · 2008 · 8 claims · 7 setups
A significant proportion of CVs overlap between HGMD and dbSNP (4.36% of HGMD CVs registered in dbSNP; 8.11% of dbSNP CVs registered in HGMD), warranting caution when interpreting phenotypic relevance of concurrent CVs.
-
Full-text index only
A SNP-centric database for the investigation of the human genome.
PMID 15046636 · PMC395999 · BMC bioinformatics · 2004 · 8 claims · 3 setups
SNPper is a web-based, integrated SNP database combining dbSNP, the Human Genome sequence (Goldenpath), LocusLink, GeneOntology, and SWISS-PROT data with querying, visualization, and export tools.
-
Full-text index only
Compressing DNA sequence databases with coil.
PMID 18489794 · PMC2426707 · BMC bioinformatics · 2008 · 8 claims · 1 setups
coil achieves higher compression ratio than state-of-the-art general-purpose compression tools on a large GenBank EST database file
-
Full-text index only
miRGen 2.0: a database of microRNA genomic information and regulation.
PMID 19850714 · PMC2808909 · Nucleic acids research · 2010 · 7 claims · 6 setups
miRGen 2.0 is a database providing comprehensive information about the genomic position of human and mouse microRNA coding transcripts and their regulation by transcription factors
-
Full-text index only
mtDB: Human Mitochondrial Genome Database, a resource for population genetics and medical sciences.
PMID 16381973 · PMC1347373 · Nucleic acids research · 2006 · 8 claims · 3 setups
mtDB is a comprehensive, actively maintained database of published human mitochondrial genome sequences, providing a common resource for population genetics and medical research
-
Full-text index only
Phenotypic categorization of genetic skin diseases reveals new relations between phenotypes, genes and pathways.
PMID 19744994 · PMC2773259 · Bioinformatics (Oxford, England) · 2009 · 8 claims · 5 setups
560 genetic skin diseases can be decomposed into 71 elementary phenotypic features (42 dermatologic, 29 systemic) that combine to represent each disease as a point in a multidimensional phenotype space
-
Full-text index only
Retroposition and evolution of the DNA-binding motifs of YY1, YY2 and REX1.
PMID 17478514 · PMC1904287 · Nucleic acids research · 2007 · 8 claims · 5 setups
62 YY1-related sequences were identified across genomes ranging from flying insects to humans, with high zinc finger domain conservation
-
Full-text index only
Kangaroo--a pattern-matching program for biological sequences.
PMID 12150718 · PMC119856 · BMC bioinformatics · 2002 · 7 claims · 2 setups
Kangaroo is a web-based regular expression pattern-matching program that searches DNA, protein, or coding-region sequences across ten organisms with no restriction on query length or complexity.
-
Full-text index only
EGenBio: a data management system for evolutionary genomics and biodiversity.
PMID 17118150 · PMC1683573 · BMC bioinformatics · 2006 · 7 claims · 7 setups
EGenBio is a web-based system for integrated management, filtering, curation, and visualization of large-scale genomic sequences, alignments, and phylogenetic trees for evolutionary genomics and biodiversity research.
-
Full-text index only
Satellog: a database for the identification and prioritization of satellite repeats in disease association studies.
PMID 15949044 · PMC1181805 · BMC bioinformatics · 2005 · 7 claims · 6 setups
Satellog is a database cataloging all pure 1-16 unit satellite repeats in the human genome with supplementary polymorphism, gene-location, and expression data for prioritizing repeats in disease-association studies.
-
Full-text index only
Molecular phylogeny of the antiangiogenic and neurotrophic serpin, pigment epithelium derived factor in vertebrates.
PMID 17020603 · PMC1609119 · BMC genomics · 2006 · 8 claims · 8 setups
A single PEDF gene is present in all examined vertebrate species but is absent from invertebrates (D. melanogaster, C. elegans, C. intestinalis)
-
Full-text index only
The Genographic Project public participation mitochondrial DNA database.
PMID 17604454 · PMC1904368 · PLoS genetics · 2007 · 7 claims · 4 setups
The Genographic Project created the largest standardized human mtDNA database to date, comprising 78,590 genotypes from the first 18 months of public participation.
-
Full-text index only
Genome mapping and expression analyses of human intronic noncoding RNAs reveal tissue-specific patterns and enrichment in genes related to regulation of transcription.
PMID 17386095 · PMC1868932 · Genome biology · 2007 · 8 claims · 4 setups
More than 55,000 totally intronic noncoding (TIN) RNAs are transcribed from the introns of 74% of unique RefSeq genes.
-
Full-text index only
NCBI Reference Sequence (RefSeq): a curated non-redundant sequence database of genomes, transcripts and proteins.
PMID 15608248 · PMC539979 · Nucleic acids research · 2005 · 7 claims · 5 setups
RefSeq provides a curated, non-redundant, explicitly linked collection of genomic, transcript and protein sequences spanning prokaryotes, eukaryotes and viruses.
-
Full-text index only
MODBASE: a database of annotated comparative protein structure models and associated resources.
PMID 16381869 · PMC1347422 · Nucleic acids research · 2006 · 8 claims · 7 setups
MODBASE is a database of automatically calculated comparative protein structure models covering all UniProt sequences matchable to a known structure
-
Full-text index only
Frameshift mutations in coding repeats of protein tyrosine phosphatase genes in colorectal tumors with microsatellite instability.
PMID 19000305 · PMC2586028 · BMC cancer · 2008 · 7 claims · 6 setups
16 PTP candidate genes containing coding mononucleotide repeats (cMNR) of at least 7 units were identified via bioinformatic analysis and screened in MSI-H cell lines, cancers, and adenomas
-
Full-text index only
The role of positive selection in determining the molecular cause of species differences in disease.
PMID 18837980 · PMC2576240 · BMC evolutionary biology · 2008 · 8 claims · 6 setups
Genes predicted to be under positive selection during human evolution are implicated in diseases (epithelial cancers, schizophrenia, autoimmune diseases, Alzheimer's disease) that differ in prevalence and symptomatology between humans and other mammals
-
Full-text index only
Genetic variation in an individual human exome.
PMID 18704161 · PMC2493042 · PLoS genetics · 2008 · 8 claims · 7 setups
The ~12,500 nonsilent coding variants in the HuRef exome can be reduced ~8-fold to a set of ~1,600 variants most likely to affect protein function.