Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
Identification of deleterious non-synonymous single nucleotide polymorphisms using sequence-derived information.
PMID 18588693 · PMC2446391 · BMC bioinformatics · 2008 · 8 claims · 5 setups
A decision tree built on 10 selected sequence-derived features classifies SAPs as Disease or Polymorphism with 82.6% accuracy and 0.607 MCC in cross-validation.
-
Full-text index only
Decision forest analysis of 61 single nucleotide polymorphisms in a case-control study of esophageal cancer; a novel method.
PMID 16026601 · PMC1637030 · BMC bioinformatics · 2005 · 8 claims · 2 setups
DF-SNPs, a novel adaptation of the Decision Forest method, can classify esophageal cancer cases vs. controls based on SNP genotype data with high concordance, sensitivity, and specificity.
-
Full-text index only
Boosting accuracy of automated classification of fluorescence microscope images for location proteomics.
PMID 15207009 · PMC449699 · BMC bioinformatics · 2004 · 8 claims · 8 setups
New classifiers (SVMs, ensembles) and new wavelet-derived (Gabor, Daubechies) features improve recognition of protein subcellular location patterns over the previous neural network approach
-
Full-text index only
Towards precise classification of cancers based on robust gene functional expression profiles.
PMID 15774002 · PMC1274255 · BMC bioinformatics · 2005 · 6 claims · 7 setups
Functional expression profiles (FEPs) achieve comparable or better classification performance than conventional gene expression profiles (GEPs) across four public microarray datasets
-
Full-text index only
Constructing support vector machine ensembles for cancer classification based on proteomic profiling.
PMID 16689692 · PMC5173238 · Genomics, proteomics & bioinformatics · 2005 · 7 claims · 4 setups
CSVME, built by selecting a subset of base SVMs via SVM-RFE ranking and fusing them with a trained upper-layer SVM, achieves better classification performance than an ensemble of all base SVMs.
-
Full-text index only
Speeding disease gene discovery by sequence based candidate prioritization.
PMID 15766383 · PMC1274252 · BMC bioinformatics · 2005 · 7 claims · 8 setups
Disease genes (OMIM) differ significantly from non-disease genes in sequence-based features including gene/cDNA/protein size, exon number, homolog conservation, secretion signal, 3' UTR length, CpG islands, and distance to nearest gene.
-
Full-text index only
Machine-learning approaches for classifying haplogroup from Y chromosome STR data.
PMID 18551166 · PMC2396484 · PLoS computational biology · 2008 · 8 claims · 5 setups
Y-STR allelic variability is partitioned more by differences among haplogroups than by differences among populations, suggesting Y-STRs carry haplogroup information
-
Full-text index only
Ab initio identification of human microRNAs based on structure motifs.
PMID 18088431 · PMC2238772 · BMC bioinformatics · 2007 · 8 claims · 7 setups
MiRPred predicts miRNA precursors ab initio using only predicted secondary structure motifs, ignoring nucleotide sequence
-
Full-text index only
A comparison of classification methods for predicting Chronic Fatigue Syndrome based on genetic data.
PMID 19772600 · PMC2765429 · Journal of translational medicine · 2009 · 7 claims · 3 setups
The naive Bayes model with the wrapper-based feature selection approach performed best among all predictive models tested for distinguishing CFS from controls.
-
Has reproduction · 86
Molecular Classification Models for Triple Negative Breast Cancer Subtype Using Machine Learning.
PMID 34575658 · PMC8472680 · Journal of personalized medicine · 2021 · 7 claims · 4 setups
TNBC can be divided into four gene-expression-defined subtypes: BLIA, BLIS, MES, and LAR
-
Has reproduction · 53
Combining evidence of preferential gene-tissue relationships from multiple sources.
PMID 23950964 · PMC3741196 · PloS one · 2013 · 8 claims · 8 setups
A high-level integration approach combining three methods across four human microarray datasets, merged by consensus voting and a rule-based inner/total score, predicts preferentially expressed genes while reducing method- and study-specific bias.
-
Has reproduction · 83
Integrative transcriptomic and machine learning framework reveals candidate genes and potential mechanisms of aflatoxin B1 exposure in breast cancer.
PMID 41688730 · PMC12982753 · Scientific reports · 2026 · 7 claims · 8 setups
170 unique human AFB1 targets were identified by merging ChEMBL, SwissTargetPrediction, and PharmMapper predictions