Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
CanPredict: a computational tool for predicting cancer-associated missense mutations.
PMID 17537827 · PMC1933186 · Nucleic acids research · 2007 · 8 claims · 7 setups
CanPredict is a web application providing public access to a random forest classifier that combines SIFT, LogR.E-value, and GOSS scores to predict whether a missense mutation is cancer-associated
-
Full-text index only
SNAP: predict effect of non-synonymous polymorphisms on function.
PMID 17526529 · PMC1920242 · Nucleic acids research · 2007 · 7 claims · 8 setups
SNAP, a neural network-based method using sequence-derived information, predicts whether a non-synonymous SNP is neutral or non-neutral for protein function
-
Full-text index only
Cancer-specific high-throughput annotation of somatic mutations: computational prediction of driver missense mutations.
PMID 19654296 · PMC2763410 · Cancer research · 2009 · 7 claims · 7 setups
CHASM, a Random Forest-based computational method, was developed to identify and prioritize missense mutations likely to be functional drivers of tumor cell proliferation.
-
Full-text index only
Pol II promoter prediction using characteristic 4-mer motifs: a machine learning approach.
PMID 18834544 · PMC2575220 · BMC bioinformatics · 2008 · 8 claims · 8 setups
128 discriminating 4-mer motifs combined with an SVM (RBF kernel, LIBSVM) can distinguish promoter from non-promoter DNA sequences
-
Has reproduction · 59
Application of Machine Learning in Predicting Hepatic Metastasis or Primary Site in Gastroenteropancreatic Neuroendocrine Tumors.
PMID 37887568 · PMC10605255 · Current oncology (Toronto, Ont.) · 2023 · 8 claims · 7 setups
Multi-gene random forest models classify primary tumor vs. liver metastasis samples with 100% accuracy in training/test cohorts and >90% accuracy in an independent validation cohort
-
Full-text index only
Multilocus analysis of SNP and metabolic data within a given pathway.
PMID 16412218 · PMC1382210 · BMC genomics · 2006 · 8 claims · 7 setups
The combinatorial partitioning method (CPM) with optimal thresholds can identify SNPs associated with quantitative metabolite levels rather than only categorical traits.
-
Has reproduction · 81
Macrophages on the run: Exercise balances macrophage polarization for improved health.
PMID 39476967 · PMC11585839 · Molecular metabolism · 2024 · 8 claims · 7 setups
Immediate/acute exercise triggers an M1 (pro-inflammatory) macrophage polarization surge.
-
Full-text index only
ProMiR II: a web server for the probabilistic prediction of clustered, nonclustered, conserved and nonconserved microRNAs.
PMID 16845048 · PMC1538778 · Nucleic acids research · 2006 · 6 claims · 4 setups
ProMiR II improves on the original ProMiR by integrating free energy, G/C ratio, conservation score and entropy for more controllable miRNA prediction
-
Full-text index only
Splicing retention and enhancer divergence govern the evolutionary fate of ohnologues following whole-genome duplication in rainbow trout.
PMID 41832354 · PMC13106644 · Scientific reports · 2026 · 7 claims · 8 setups
Most rainbow trout ohnologues are retained through conservation (71.4%) rather than neo/sub-functionalization or specialization
-
Full-text index only
Genomic data sampling and its effect on classification performance assessment.
PMID 12553886 · PMC149349 · BMC bioinformatics · 2003 · 8 claims · 3 setups
Cross-validation, leave-one-out, and bootstrap are designed to reduce bias and variance in accuracy estimation from small samples.
-
Full-text index only
nsSNPAnalyzer: identifying disease-associated nonsynonymous single nucleotide polymorphisms.
PMID 15980516 · PMC1160133 · Nucleic acids research · 2005 · 6 claims · 4 setups
nsSNPAnalyzer is a web server that predicts whether a query nsSNP is disease-associated or functionally neutral using a Random Forest classifier combining structural and evolutionary information
-
Full-text index only
Predicting the phenotypic effects of non-synonymous single nucleotide polymorphisms based on support vector machines.
PMID 18005451 · PMC2216041 · BMC bioinformatics · 2007 · 8 claims · 5 setups
Parepro, an SVM-based method integrating three attribute sets (RD, MI, IE) derived from evolutionary and residue-property information, predicts whether an nsSNP is deleterious or neutral.
-
Full-text index only
MiPred: classification of real and pseudo microRNA precursors using random forest prediction model with combined features.
PMID 17553836 · PMC1933124 · Nucleic acids research · 2007 · 8 claims · 8 setups
A hybrid feature combining local contiguous triplet structure-sequence composition, MFE of the secondary structure, and P-value of a randomization test improves classification of real vs pseudo pre-miRNAs
-
Full-text index only
Logical Analysis of Data (LAD) model for the early diagnosis of acute ischemic stroke.
PMID 18616825 · PMC2492849 · BMC medical informatics and decision making · 2008 · 7 claims · 5 setups
An LAD classification model built from a support-set of 3 peptide peaks can distinguish stroke patients from controls with 75% accuracy on an independent validation set
-
Full-text index only
Boosting accuracy of automated classification of fluorescence microscope images for location proteomics.
PMID 15207009 · PMC449699 · BMC bioinformatics · 2004 · 8 claims · 8 setups
New classifiers (SVMs, ensembles) and new wavelet-derived (Gabor, Daubechies) features improve recognition of protein subcellular location patterns over the previous neural network approach
-
Full-text index only
Searching for interpretable rules for disease mutations: a simulated annealing bump hunting strategy.
PMID 16984653 · PMC1618409 · BMC bioinformatics · 2006 · 8 claims · 6 setups
The proposed feature set outperforms existing published feature sets for predicting effects of amino acid substitutions
-
Full-text index only
Identification of deleterious non-synonymous single nucleotide polymorphisms using sequence-derived information.
PMID 18588693 · PMC2446391 · BMC bioinformatics · 2008 · 8 claims · 5 setups
A decision tree built on 10 selected sequence-derived features classifies SAPs as Disease or Polymorphism with 82.6% accuracy and 0.607 MCC in cross-validation.
-
Full-text index only
Interaction profile-based protein classification of death domain.
PMID 15189571 · PMC459208 · BMC bioinformatics · 2004 · 7 claims · 6 setups
An SVM-based classifier using Residue Pair Interaction Profiles (RPIPs) can classify death domain superfamily members into subfamilies with 89% average cross-validation accuracy
-
Full-text index only
Prediction of catalytic residues using Support Vector Machine with selected protein sequence and structural properties.
PMID 16790052 · PMC1534064 · BMC bioinformatics · 2006 · 8 claims · 7 setups
The Sequential Minimal Optimization (SMO) SVM algorithm was the best-performing classifier among 26 WEKA classifiers for predicting catalytic residues
-
Has reproduction · 95
MetaMap: an atlas of metatranscriptomic reads in human disease-related RNA-seq data.
PMID 29901703 · PMC6025204 · GigaScience · 2018 · 8 claims · 7 setups
The MetaMap pipeline recapitulates known infection agents in bona fide dual RNA-seq validation studies (Salmonella, HPV, HSV, rhinovirus)