Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Has reproduction · 77
Representing and querying disease networks using graph databases.
PMID 27462371 · PMC4960687 · BioData mining · 2016 · 7 claims · 8 setups
Graph databases are well suited for representing biological information that is highly connected, semi-structured, and unpredictable.
-
Has reproduction · 74
Wide-Open: Accelerating public data release by automating detection of overdue datasets.
PMID 28594819 · PMC5464523 · PLoS biology · 2017 · 7 claims · 5 setups
Wide-Open is a general text-mining approach that automatically detects overdue datasets by scanning PubMed articles for dataset accession identifiers and querying repositories to determine if the datasets remain private.
-
Full-text index only
Analysis of the glutathione S-transferase (GST) gene family.
PMID 15607001 · PMC3500200 · Human genomics · 2004 · 8 claims · 4 setups
The complete human GST gene family comprises 16 genes in six subfamilies: alpha (GSTA), mu (GSTM), omega (GSTO), pi (GSTP), theta (GSTT) and zeta (GSTZ).
-
Full-text index only
Towards precise classification of cancers based on robust gene functional expression profiles.
PMID 15774002 · PMC1274255 · BMC bioinformatics · 2005 · 6 claims · 7 setups
Functional expression profiles (FEPs) achieve comparable or better classification performance than conventional gene expression profiles (GEPs) across four public microarray datasets
-
Full-text index only
Construction of a nasopharyngeal carcinoma 2D/MS repository with Open Source XML database--Xindice.
PMID 16403238 · PMC1351203 · BMC bioinformatics · 2006 · 8 claims · 4 setups
No NPC proteome database existed prior to this work despite availability of other cancer proteome databases
-
Full-text index only
The Genographic Project public participation mitochondrial DNA database.
PMID 17604454 · PMC1904368 · PLoS genetics · 2007 · 7 claims · 4 setups
The Genographic Project created the largest standardized human mtDNA database to date, comprising 78,590 genotypes from the first 18 months of public participation.
-
Full-text index only
Cataloging coding sequence variations in human genome databases.
PMID 18974781 · PMC2570488 · PloS one · 2008 · 8 claims · 7 setups
A significant proportion of CVs overlap between HGMD and dbSNP (4.36% of HGMD CVs registered in dbSNP; 8.11% of dbSNP CVs registered in HGMD), warranting caution when interpreting phenotypic relevance of concurrent CVs.
-
Full-text index only
Integration with the human genome of peptide sequences obtained by high-throughput mass spectrometry.
PMID 15642101 · PMC549070 · Genome biology · 2005 · 8 claims · 4 setups
PeptideAtlas, a public database integrating MS/MS-derived peptide identifications with the human genome, was built as an expandable resource for proteomic data.
-
Has reproduction · 82
Landscape of allele-specific transcription factor binding in the human genome.
PMID 33980847 · PMC8115691 · Nature communications · 2021 · 8 claims · 6 setups
A novel statistical framework (ADASTRA) calls allele-specific TF binding from existing ChIP-Seq alignments by jointly correcting for background allelic dosage (BAD, from aneuploidy/CNVs) and reference mapping bias.
-
Full-text index only
Prediction of catalytic residues using Support Vector Machine with selected protein sequence and structural properties.
PMID 16790052 · PMC1534064 · BMC bioinformatics · 2006 · 8 claims · 7 setups
The Sequential Minimal Optimization (SMO) SVM algorithm was the best-performing classifier among 26 WEKA classifiers for predicting catalytic residues
-
Has reproduction · 62
Metatranscriptomics of the human oral microbiome during health and disease.
PMID 24692635 · PMC3977359 · mBio · 2014 · 8 claims · 8 setups
Disease-associated periodontal communities display conserved community-level metabolic gene expression profiles between patients, whereas the metabolic gene expression of individual species is highly variable between patients.
-
Full-text index only
Predicting the phenotypic effects of non-synonymous single nucleotide polymorphisms based on support vector machines.
PMID 18005451 · PMC2216041 · BMC bioinformatics · 2007 · 8 claims · 5 setups
Parepro, an SVM-based method integrating three attribute sets (RD, MI, IE) derived from evolutionary and residue-property information, predicts whether an nsSNP is deleterious or neutral.
-
Full-text index only
Computer-aided identification of polymorphism sets diagnostic for groups of bacterial and viral genetic variants.
PMID 17672919 · PMC1973086 · BMC bioinformatics · 2007 · 6 claims · 8 setups
The Not-N algorithm, incorporated into the Minimum SNPs program, identifies small marker sets diagnostic for user-defined subgroups of genetic variants with 0% false negatives
-
Full-text index only
Random amino acid mutations and protein misfolding lead to Shannon limit in sequence-structure communication.
PMID 18769673 · PMC2518838 · PloS one · 2008 · 8 claims · 6 setups
The protein sequence-structure map behaves as a noisy digital communication channel whose capacity C exceeds the transmission rate R for native structures, satisfying Shannon's noisy channel theorem
-
Full-text index only
A catalog of human cDNA expression clones and its application to structural genomics.
PMID 15345055 · PMC522878 · Genome biology · 2004 · 8 claims · 7 setups
A high-throughput screening approach can identify human cDNA clones from the hEx1 library that express soluble protein in E. coli
-
Full-text index only
Coverage and characteristics of the Affymetrix GeneChip Human Mapping 100K SNP set.
PMID 16680197 · PMC1456318 · PLoS genetics · 2006 · 7 claims · 7 setups
SNPs in the Affymetrix 100K set are undersampled from coding regions (both synonymous and nonsynonymous) and oversampled from regions outside genes, relative to HapMap SNPs
-
Full-text index only
Genetic diversity of clinical isolates of Bacillus cereus using multilocus sequence typing.
PMID 18990211 · PMC2585095 · BMC microbiology · 2008 · 8 claims · 7 setups
The 55 clinical B. cereus isolates were phylogenetically diverse, comprising 38 sequence types (STs) distributed across two of three previously described clades.
-
Full-text index only
Genetic variation in an individual human exome.
PMID 18704161 · PMC2493042 · PLoS genetics · 2008 · 8 claims · 7 setups
The ~12,500 nonsilent coding variants in the HuRef exome can be reduced ~8-fold to a set of ~1,600 variants most likely to affect protein function.
-
Full-text index only
Extraction of human kinase mutations from literature, databases and genotyping studies.
PMID 19758464 · PMC2745582 · BMC bioinformatics · 2009 · 7 claims · 6 setups
A literature mining pipeline combining MutationFinder, false-positive filtering, and SVM-based classification can extract and disambiguate single-point mutation mentions from abstracts and full text
-
Full-text index only
Combinatorial Mismatch Scan (CMS) for loci associated with dementia in the Amish.
PMID 16515697 · PMC1448207 · BMC medical genetics · 2006 · 8 claims · 7 setups
CMS compares IBS allele/genotype sharing between distantly related (beyond grandparental) affected and unaffected individuals from founder populations to detect disease loci while reducing confounding from population stratification and genetic heterogeneity.