Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
CGMIM: automated text-mining of Online Mendelian Inheritance in Man (OMIM) to identify genetically-associated cancers and candidate genes.
PMID 15796777 · PMC1274267 · BMC bioinformatics · 2005 · 8 claims · 2 setups
CGMIM is a Perl program that text-mines OMIM entries to identify cancer-gene associations and genetically-related cancer type pairs.
-
Full-text index only
PPC: an algorithm for accurate estimation of SNP allele frequencies in small equimolar pools of DNA using data from high density microarrays.
PMID 16199750 · PMC1240117 · Nucleic acids research · 2005 · 7 claims · 6 setups
The PPC algorithm, which applies a probe-pair-specific second-degree polynomial correction, increases the accuracy of allele frequency estimates from pooled DNA compared with previously described algorithms
-
Full-text index only
Optimal step length EM algorithm (OSLEM) for the estimation of haplotype frequency and its application in lipoprotein lipase genotyping.
PMID 12529185 · PMC149347 · BMC bioinformatics · 2003 · 5 claims · 4 setups
OSLEM (Optimal Step Length EM), which approximates an optimal step length via a fixed-point search (D_N = D_{N-1} + λ(D_preN - D_{N-1})), runs about twice as fast as standard EM while producing the same haplotype frequency estimates.
-
Has reproduction · 85
Digital sorting of complex tissues for cell type-specific gene expression profiles.
PMID 23497278 · PMC3626856 · BMC bioinformatics · 2013 · 8 claims · 8 setups
The Digital Sorting Algorithm (DSA) deconvolves mixed tissue expression into cell type-specific profiles using only marker genes, without requiring prior knowledge of cell type frequencies or in vitro pure-cell profiles.
-
Full-text index only
A third approach to gene prediction suggests thousands of additional human transcribed regions.
PMID 16543943 · PMC1391917 · PLoS computational biology · 2006 · 8 claims · 7 setups
A third basic concept for gene prediction exists, based on detecting strand-specific 'transcription footprints' (mutational and selectional biases) rather than gene structure or sequence similarity.
-
Full-text index only
Computer-aided identification of polymorphism sets diagnostic for groups of bacterial and viral genetic variants.
PMID 17672919 · PMC1973086 · BMC bioinformatics · 2007 · 6 claims · 8 setups
The Not-N algorithm, incorporated into the Minimum SNPs program, identifies small marker sets diagnostic for user-defined subgroups of genetic variants with 0% false negatives
-
Full-text index only
Assessing batch effects of genotype calling algorithm BRLMM for the Affymetrix GeneChip Human Mapping 500 K array set using 270 HapMap samples.
PMID 18793462 · PMC2537568 · BMC bioinformatics · 2008 · 8 claims · 4 setups
Batch size affects genotype calling results (call rate and concordance) and the resulting lists of significantly associated SNPs.
-
Has reproduction · 79
TSUNAMI: Translational Bioinformatics Tool Suite for Network Analysis and Mining.
PMID 33705981 · PMC9403021 · Genomics, proteomics & bioinformatics · 2021 · 8 claims · 6 setups
TSUNAMI is a freely accessible web-based tool suite that mines gene co-expression network (GCN) modules from public (GEO, TCGA) or user-uploaded numerical omics data and performs downstream gene set enrichment analysis.
-
Full-text index only
SNP haplotype tagging from DNA pools of two individuals.
PMID 12709267 · PMC156884 · BMC bioinformatics · 2003 · 8 claims · 3 setups
An algorithm can reconstruct haplotypes from pools of two individuals' DNA under very general conditions, without requiring Hardy-Weinberg equilibrium.
-
Full-text index only
Proteomics: characterizing the cogs in the machinery of life.
PMID 14630521 · PMC1241753 · Environmental health perspectives · 2003 · 8 claims · 5 setups
Protein expression patterns in blood serum, detected via SELDI-TOF mass spectrometry and analyzed with a genetic algorithm, can distinguish ovarian cancer patients from healthy individuals with very high sensitivity and specificity.
-
Full-text index only
The Princeton Protein Orthology Database (P-POD): a comparative genomics analysis tool for biologists.
PMID 17712414 · PMC1942082 · PloS one · 2007 · 8 claims · 5 setups
P-POD is the first comparative genomics database to combine results from multiple computational ortholog/homolog prediction methods with manually curated literature-derived experimental evidence of functional conservation.
-
Full-text index only
pTARGET: a web server for predicting protein subcellular localization.
PMID 16844995 · PMC1538910 · Nucleic acids research · 2006 · 7 claims · 3 setups
pTARGET web server predicts nine distinct subcellular localizations in eukaryotic non-plant proteins using an algorithm based on location-specific Pfam domain occurrence patterns and amino acid composition (AAC)
-
Full-text index only
SNiPer: improved SNP genotype calling for Affymetrix 10K GeneChip microarray data.
PMID 16262895 · PMC1280925 · BMC genomics · 2005 · 8 claims · 5 setups
Poorly performing SNPs (NoCall rate ≥25%) fail primarily due to inadequate training/localization of the MPAM statistical model call zone, not detection filter failure
-
Full-text index only
Genomic and proteomic approaches for studying human cancer: prospects for true patient-tailored therapy.
PMID 15601541 · PMC3525069 · Human genomics · 2004 · 8 claims · 6 setups
DNA microarray gene expression profiling generates robust molecular classifications for many tumour types (brain, breast, colon, gastric, kidney, leukaemia, lymphoma, lung, melanoma, ovary, prostate, etc.)
-
Full-text index only
Non-EST based prediction of exon skipping and intron retention events using Pfam information.
PMID 16204458 · PMC1243800 · Nucleic acids research · 2005 · 7 claims · 5 setups
A novel ab initio method predicts exon skipping and intron retention events using only Pfam domain annotation, via a Viterbi-like dynamic programming algorithm applied to the Pfam alignment.
-
Full-text index only
Characterisation of the genomic architecture of human chromosome 17q and evaluation of different methods for haplotype block definition.
PMID 15850495 · PMC1090572 · BMC genetics · 2005 · 8 claims · 6 setups
Haplotype block definitions based on LD measures (Definitions 1, 2, 3, 5) produce fewer, shorter blocks with limited sequence coverage compared to the haplotype diversity-based method (Definition 4)
-
Full-text index only
A high-throughput method for quantifying alleles and haplotypes of the malaria vaccine candidate Plasmodium falciparum merozoite surface protein-1 19 kDa.
PMID 16626494 · PMC1459863 · Malaria journal · 2006 · 8 claims · 5 setups
Pyrosequencing, after adjustment to a standard curve, provides accurate and precise estimates of allele frequencies in mixed MSP-1_19 infections
-
Full-text index only
Integrated algorithms for high-throughput examination of covalently labeled biomolecules by structural mass spectrometry.
PMID 19788317 · PMC2764328 · Analytical chemistry · 2009 · 7 claims · 4 setups
ProtMapMS automates data format conversion, mass spectrum interpretation, peptide detection/verification, modification site confirmation, and quantification of peptide oxidation extent from covalent labeling MS data
-
Full-text index only
Evidence for limited genetic compartmentalization of HIV-1 between lung and blood.
PMID 19759830 · PMC2736399 · PloS one · 2009 · 8 claims · 7 setups
Statistical evidence of genetic compartmentalization between lung and blood HIV-1 env sequences was found in 10 of 18 subjects.
-
Full-text index only
Local combinational variables: an approach used in DNA-binding helix-turn-helix motif prediction with sequence information.
PMID 19651875 · PMC2761287 · Nucleic acids research · 2009 · 8 claims · 7 setups
The LCV approach predicts HTH motifs with 93.29% accuracy, 93.93% sensitivity and 92.66% specificity using only primary sequence information