Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
Function2Gene: a gene selection tool to increase the power of genetic association studies by utilizing public databases and expert knowledge.
PMID 18631403 · PMC2500032 · BMC bioinformatics · 2008 · 6 claims · 5 setups
Function2Gene is a set of Perl programs that queries public databases (NCBI, GeneCards, Harvester, with Uniprot/Ensembl also supported) using expert-selected keywords to rank genes by prior probability of disease association.
-
Full-text index only
BLASTO: a tool for searching orthologous groups.
PMID 17483516 · PMC1933156 · Nucleic acids research · 2007 · 7 claims · 2 setups
BLASTO treats each orthologous group as a unit and outputs a ranked list of orthologous groups instead of single sequences
-
Full-text index only
Integration of text- and data-mining using ontologies successfully selects disease gene candidates.
PMID 15767279 · PMC1065256 · Nucleic acids research · 2005 · 7 claims · 6 setups
Integrating eVOC anatomical ontology-based text-mining of PubMed abstracts with data-mining of gene expression annotation successfully selects and prioritizes candidate disease genes
-
Full-text index only
SNP selection for genes of iron metabolism in a study of genetic modifiers of hemochromatosis.
PMID 18366708 · PMC2289803 · BMC medical genetics · 2008 · 7 claims · 6 setups
Illumina validation/design scores above 0.6 are not strongly correlated with actual SNP genotyping performance (Gentrain score)
-
Has reproduction · 76
Bayesian prediction of microbial oxygen requirement.
PMID 26913185 · PMC4743139 · F1000Research · 2013 · 7 claims · 8 setups
A naive Bayesian classifier based on presence/absence of class-associated Pfam-A domains can distinguish three oxygen requirement classes (aerobe, anaerobe, facultative anaerobe) from genome sequence, unlike prior studies that only made pairwise distinctions.
-
Has reproduction · 53
PulmonDB: a curated lung disease gene expression database.
PMID 31949184 · PMC6965635 · Scientific reports · 2020 · 6 claims · 6 setups
PulmonDB is a curated, web-based gene expression database and R package integrating microarray and RNA-seq data for COPD and IPF with manually curated controlled-vocabulary annotation.
-
Full-text index only
A surrogate-based approach for post-genomic partner identification.
PMID 11602024 · PMC57814 · BMC biotechnology · 2001 · 8 claims · 5 setups
Peptide surrogates derived from random phage display libraries contain amino acid sequence information that identifies the natural biological partner of the panned target via database searching.
-
Full-text index only
The role of positive selection in determining the molecular cause of species differences in disease.
PMID 18837980 · PMC2576240 · BMC evolutionary biology · 2008 · 8 claims · 6 setups
Genes predicted to be under positive selection during human evolution are implicated in diseases (epithelial cancers, schizophrenia, autoimmune diseases, Alzheimer's disease) that differ in prevalence and symptomatology between humans and other mammals
-
Full-text index only
CGMIM: automated text-mining of Online Mendelian Inheritance in Man (OMIM) to identify genetically-associated cancers and candidate genes.
PMID 15796777 · PMC1274267 · BMC bioinformatics · 2005 · 8 claims · 2 setups
CGMIM is a Perl program that text-mines OMIM entries to identify cancer-gene associations and genetically-related cancer type pairs.
-
Full-text index only
CYCLONET--an integrated database on cell cycle regulation and carcinogenesis.
PMID 17202170 · PMC1899094 · Nucleic acids research · 2007 · 7 claims · 4 setups
Cyclonet is a web-based integrated database combining 'omics' and chemoinformatics data on mammalian cell cycle regulation in normal and pathological (cancer) states, built on a systems biology approach.
-
Has reproduction · 100
Integrative transcriptome sequencing identifies trans-splicing events with important roles in human embryonic stem cell pluripotency.
PMID 24131564 · PMC3875859 · Genome research · 2014 · 8 claims · 8 setups
TSscan, a computational pipeline integrating long- and short-read transcriptome sequencing from multiple hESC lines, can detect trans-splicing while minimizing false positives from experimental artifacts and genetic rearrangements.
-
Full-text index only
In silico and in vitro comparative analysis to select, validate and test SNPs for human identification.
PMID 18076761 · PMC2222643 · BMC genomics · 2007 · 8 claims · 7 setups
A panel of 24 SNPs was selected and validated for human identification using 1,040 unrelated samples from three populations (Italian, Benin Gulf, Mongolian)
-
Full-text index only
A survey of integral alpha-helical membrane proteins.
PMID 19760129 · PMC2780624 · Journal of structural and functional genomics · 2009 · 8 claims · 8 setups
An automated annotation pipeline defines the integral membrane genome and family associations for 21,379 proteins from 34 genomes, most belonging to 598 Pfam-derived membrane protein families.
-
Full-text index only
In silico and in vivo splicing analysis of MLH1 and MSH2 missense mutations shows exon- and tissue-specific effects.
PMID 16995940 · PMC1590028 · BMC genomics · 2006 · 8 claims · 6 setups
In silico ESE-prediction algorithms (ESEfinder, RescueESE, PESX) do not reliably predict actual in vivo splicing behavior of missense mutations
-
Full-text index only
Computational disease gene identification: a concert of methods prioritizes type 2 diabetes and obesity candidate genes.
PMID 16757574 · PMC1475747 · Nucleic acids research · 2006 · 6 claims · 8 setups
Applying seven independent computational disease-gene prioritization methods in concert to 9556 positional candidate genes identifies a prioritized set of likely T2D and obesity candidate genes
-
Has reproduction · 55
Genome-Wide Survey and Development of the First Microsatellite Markers Database (AnCorDB) in Anemone coronaria L.
PMID 35328546 · PMC8949970 · International journal of molecular sciences · 2022 · 8 claims · 8 setups
Generated the first draft genome assembly of A. coronaria by Illumina sequencing a haploid androgenetic plant
-
Full-text index only
The DNA sequence of the human X chromosome.
PMID 15772651 · PMC2665286 · Nature · 2005 · 8 claims · 8 setups
The euchromatic sequence of the human X chromosome was determined to 99.3% completeness (~155 Mb total)
-
Full-text index only
Indirect genomic effects on survival from gene expression data.
PMID 18358079 · PMC2397510 · Genome biology · 2008 · 7 claims · 6 setups
A novel methodology (dynamic path analysis combined with additive hazard survival regression) can detect and quantify indirect effects of gene expression on survival mediated through transcription factor target genes.
-
Full-text index only
Genomic analysis of the chromosome 15q11-q13 Prader-Willi syndrome region and characterization of transcripts for GOLGA8E and WHCD1L1 from the proximal breakpoint region.
PMID 18226259 · PMC2268926 · BMC genomics · 2008 · 8 claims · 7 setups
GOLGA8E and WHDC1L1 are characterized for the first time as protein-coding transcripts from the PWS proximal breakpoint region.
-
Has reproduction · 68
LaSSO, a strategy for genome-wide mapping of intronic lariats and branch points using RNA-seq.
PMID 24709818 · PMC4079972 · Genome research · 2014 · 8 claims · 8 setups
LaSSO (Lariat Sequence Site Origin) identifies intronic lariat reads and pinpoints branch points genome-wide from RNA-seq data by considering every intronic base as a potential branch point and including all possible exon-skipping lariats.