Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
Prediction of catalytic residues using Support Vector Machine with selected protein sequence and structural properties.
PMID 16790052 · PMC1534064 · BMC bioinformatics · 2006 · 8 claims · 7 setups
The Sequential Minimal Optimization (SMO) SVM algorithm was the best-performing classifier among 26 WEKA classifiers for predicting catalytic residues
-
Full-text index only
Environmental Burkholderia cepacia complex isolates in human infections.
PMID 17552100 · PMC2725883 · Emerging infectious diseases · 2007 · 6 claims · 4 setups
More than 20% of clinical Bcc isolates examined are indistinguishable by MLST from environmental isolates, linking the natural environment to emergence of clinical infections
-
Full-text index only
Random amino acid mutations and protein misfolding lead to Shannon limit in sequence-structure communication.
PMID 18769673 · PMC2518838 · PloS one · 2008 · 8 claims · 6 setups
The protein sequence-structure map behaves as a noisy digital communication channel whose capacity C exceeds the transmission rate R for native structures, satisfying Shannon's noisy channel theorem
-
Full-text index only
Novel CLCN1 mutations and clinical features of Korean patients with myotonia congenita.
PMID 19949657 · PMC2775849 · Journal of Korean medical science · 2009 · 7 claims · 8 setups
Sequencing of CLCN1 in 10 unrelated Korean MC patients identified nine different point mutations, six of which are novel (p.M128I, p.S189C, p.M373L, p.P480S, p.G523D, p.M609K).
-
Has reproduction · 77
Representing and querying disease networks using graph databases.
PMID 27462371 · PMC4960687 · BioData mining · 2016 · 7 claims · 8 setups
Graph databases are well suited for representing biological information that is highly connected, semi-structured, and unpredictable.
-
Full-text index only
Genetic diversity of clinical isolates of Bacillus cereus using multilocus sequence typing.
PMID 18990211 · PMC2585095 · BMC microbiology · 2008 · 8 claims · 7 setups
The 55 clinical B. cereus isolates were phylogenetically diverse, comprising 38 sequence types (STs) distributed across two of three previously described clades.
-
Full-text index only
DNA sequence variants in the LOXL1 gene are associated with pseudoexfoliation glaucoma in a U.S. clinic-based population with broad ethnic diversity.
PMID 18254956 · PMC2270804 · BMC medical genetics · 2008 · 8 claims · 5 setups
Three LOXL1 SNPs previously associated with pseudoexfoliation in Nordic populations are significantly associated with pseudoexfoliation syndrome and pseudoexfoliation glaucoma in a U.S. ethnically diverse population
-
Full-text index only
Leveraging human genomic information to identify nonhuman primate sequences for expression array development.
PMID 16288651 · PMC1314899 · BMC genomics · 2005 · 8 claims · 6 setups
Human genomic DNA sequence can be leveraged to obtain 3' end sequence of NHP orthologs, which can then be used to generate NHP oligonucleotide microarrays
-
Full-text index only
Cataloging coding sequence variations in human genome databases.
PMID 18974781 · PMC2570488 · PloS one · 2008 · 8 claims · 7 setups
A significant proportion of CVs overlap between HGMD and dbSNP (4.36% of HGMD CVs registered in dbSNP; 8.11% of dbSNP CVs registered in HGMD), warranting caution when interpreting phenotypic relevance of concurrent CVs.
-
Full-text index only
Identification of deleterious non-synonymous single nucleotide polymorphisms using sequence-derived information.
PMID 18588693 · PMC2446391 · BMC bioinformatics · 2008 · 8 claims · 5 setups
A decision tree built on 10 selected sequence-derived features classifies SAPs as Disease or Polymorphism with 82.6% accuracy and 0.607 MCC in cross-validation.
-
Full-text index only
Capturing genomic signatures of DNA sequence variation using a standard anonymous microarray platform.
PMID 17000641 · PMC1636412 · Nucleic acids research · 2006 · 8 claims · 6 setups
An anonymous SHyP oligonucleotide microarray can capture genomic signatures of DNA sequence variation from any organism, including a previously unsequenced species
-
Full-text index only
The diploid genome sequence of an Asian individual.
PMID 18987735 · PMC2716080 · Nature · 2008 · 8 claims · 8 setups
First diploid genome sequence of an Asian (Han Chinese) individual generated using massively parallel Illumina sequencing
-
Has reproduction · 92
Evaluation of core genome and whole genome multilocus sequence typing schemes for Campylobacter jejuni and Campylobacter coli outbreak detection in the USA.
PMID 37133905 · PMC10272873 · Microbial genomics · 2023 · 8 claims · 8 setups
cgMLST, wgMLST and hqSNP WGS-based analysis methods clustered C. jejuni and C. coli isolates in concordance with epidemiological data.
-
Full-text index only
Analysis of the glutathione S-transferase (GST) gene family.
PMID 15607001 · PMC3500200 · Human genomics · 2004 · 8 claims · 4 setups
The complete human GST gene family comprises 16 genes in six subfamilies: alpha (GSTA), mu (GSTM), omega (GSTO), pi (GSTP), theta (GSTT) and zeta (GSTZ).
-
Full-text index only
Characterisation of the genomic architecture of human chromosome 17q and evaluation of different methods for haplotype block definition.
PMID 15850495 · PMC1090572 · BMC genetics · 2005 · 8 claims · 6 setups
Haplotype block definitions based on LD measures (Definitions 1, 2, 3, 5) produce fewer, shorter blocks with limited sequence coverage compared to the haplotype diversity-based method (Definition 4)
-
Full-text index only
External contamination in single cell mtDNA analysis.
PMID 17668059 · PMC1930155 · PloS one · 2007 · 8 claims · 6 setups
External DNA contamination is a real and non-negligible problem in single-cell mtDNA sequence analysis
-
Full-text index only
The UCSC Genome Browser Database: 2008 update.
PMID 18086701 · PMC2238835 · Nucleic acids research · 2008 · 8 claims · 8 setups
The UCSC Genome Browser Database (GBD) provides integrated sequence and annotation data for a large collection of vertebrate and model organism genomes.
-
Full-text index only
Prioritization of candidate cancer genes--an aid to oncogenomic studies.
PMID 18710882 · PMC2566894 · Nucleic acids research · 2008 · 8 claims · 8 setups
Computational classifiers using combinations of protein conservation, gene structure, protein domains, protein interactions, and regulatory data can distinguish known cancer genes (CD/CR) from unlabelled human genes
-
Full-text index only
PhylomeDB: a database for genome-wide collections of gene phylogenies.
PMID 17962297 · PMC2238872 · Nucleic acids research · 2008 · 7 claims · 6 setups
PhylomeDB is a publicly accessible database storing complete, genome-wide collections of gene phylogenies (phylomes).
-
Has reproduction · 95
The archives are half-empty: an assessment of the availability of microbial community sequencing data.
PMID 32859925 · PMC7455719 · Communications biology · 2020 · 8 claims · 5 setups
More than half of surveyed amplicon sequencing studies were affected by lack of data deposition, improper file formatting, or inconsistent labeling that impede reuse.