Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
CanPredict: a computational tool for predicting cancer-associated missense mutations.
PMID 17537827 · PMC1933186 · Nucleic acids research · 2007 · 8 claims · 7 setups
CanPredict is a web application providing public access to a random forest classifier that combines SIFT, LogR.E-value, and GOSS scores to predict whether a missense mutation is cancer-associated
-
Full-text index only
Genomic data sampling and its effect on classification performance assessment.
PMID 12553886 · PMC149349 · BMC bioinformatics · 2003 · 8 claims · 3 setups
Cross-validation, leave-one-out, and bootstrap are designed to reduce bias and variance in accuracy estimation from small samples.
-
Full-text index only
nsSNPAnalyzer: identifying disease-associated nonsynonymous single nucleotide polymorphisms.
PMID 15980516 · PMC1160133 · Nucleic acids research · 2005 · 6 claims · 4 setups
nsSNPAnalyzer is a web server that predicts whether a query nsSNP is disease-associated or functionally neutral using a Random Forest classifier combining structural and evolutionary information
-
Full-text index only
VIRGO: computational prediction of gene functions.
PMID 16845022 · PMC1538839 · Nucleic acids research · 2006 · 8 claims · 6 setups
VIRGO constructs a functional linkage network (FLN) from gene expression and molecular interaction data, labels genes with GO annotations, and propagates these labels to predict functions of unlabelled genes
-
Full-text index only
InSite: a computational method for identifying protein-protein interaction binding sites on a proteome-wide scale.
PMID 17868464 · PMC2375030 · Genome biology · 2007 · 8 claims · 8 setups
InSite predicts protein-pair-specific binding motifs ('Motif M on protein A binds to protein B') by integrating heterogeneous PPI and motif-motif interaction evidence within a Bayesian network trained by EM
-
Full-text index only
MiPred: classification of real and pseudo microRNA precursors using random forest prediction model with combined features.
PMID 17553836 · PMC1933124 · Nucleic acids research · 2007 · 8 claims · 8 setups
A hybrid feature combining local contiguous triplet structure-sequence composition, MFE of the secondary structure, and P-value of a randomization test improves classification of real vs pseudo pre-miRNAs
-
Full-text index only
SNAP: predict effect of non-synonymous polymorphisms on function.
PMID 17526529 · PMC1920242 · Nucleic acids research · 2007 · 7 claims · 8 setups
SNAP, a neural network-based method using sequence-derived information, predicts whether a non-synonymous SNP is neutral or non-neutral for protein function
-
Has reproduction · 74
SpaGene: A Deep Adversarial Framework for Spatial Gene Imputation.
PMID 42146899 · PMC13176606 · Computational and structural biotechnology journal · 2026 · 8 claims · 6 setups
SpaGene improves average PCC and SSIM and reduces RMSE compared to 6 baseline methods (SpaGE, gimVI, Tangram, VISTA, spRefine, stDiff) across 8 diverse ST-SC dataset pairs under gene-holdout evaluation.
-
Has reproduction · 90
A Decentralized Kidney Transplant Biopsy Classifier for Transplant Rejection Developed Using Genes of the Banff-Human Organ Transplant Panel.
PMID 35619722 · PMC9128066 · Frontiers in immunology · 2022 · 6 claims · 6 setups
A random forest model trained solely on B-HOT panel genes (B-HOT Model) accurately classifies kidney transplant biopsies as NR, ABMR, or TCMR.
-
Full-text index only
Interaction profile-based protein classification of death domain.
PMID 15189571 · PMC459208 · BMC bioinformatics · 2004 · 7 claims · 6 setups
An SVM-based classifier using Residue Pair Interaction Profiles (RPIPs) can classify death domain superfamily members into subfamilies with 89% average cross-validation accuracy
-
Full-text index only
Searching for interpretable rules for disease mutations: a simulated annealing bump hunting strategy.
PMID 16984653 · PMC1618409 · BMC bioinformatics · 2006 · 8 claims · 6 setups
The proposed feature set outperforms existing published feature sets for predicting effects of amino acid substitutions
-
Full-text index only
Predicting the phenotypic effects of non-synonymous single nucleotide polymorphisms based on support vector machines.
PMID 18005451 · PMC2216041 · BMC bioinformatics · 2007 · 8 claims · 5 setups
Parepro, an SVM-based method integrating three attribute sets (RD, MI, IE) derived from evolutionary and residue-property information, predicts whether an nsSNP is deleterious or neutral.
-
Full-text index only
Identification of deleterious non-synonymous single nucleotide polymorphisms using sequence-derived information.
PMID 18588693 · PMC2446391 · BMC bioinformatics · 2008 · 8 claims · 5 setups
A decision tree built on 10 selected sequence-derived features classifies SAPs as Disease or Polymorphism with 82.6% accuracy and 0.607 MCC in cross-validation.
-
Full-text index only
SNAP predicts effect of mutations on protein function.
PMID 18757876 · PMC2562009 · Bioinformatics (Oxford, England) · 2008 · 8 claims · 3 setups
SNAP is a publicly available web-server implementation predicting functional effects (neutral/non-neutral) of single amino acid substitutions.
-
Full-text index only
Prediction of catalytic residues using Support Vector Machine with selected protein sequence and structural properties.
PMID 16790052 · PMC1534064 · BMC bioinformatics · 2006 · 8 claims · 7 setups
The Sequential Minimal Optimization (SMO) SVM algorithm was the best-performing classifier among 26 WEKA classifiers for predicting catalytic residues
-
Full-text index only
Broad network-based predictability of Saccharomyces cerevisiae gene loss-of-function phenotypes.
PMID 18053250 · PMC2246260 · Genome biology · 2007 · 8 claims · 4 setups
Loss-of-function phenotypes in yeast are predictable from a gene's connections in a functional gene network via guilt-by-association.
-
Full-text index only
Boosting accuracy of automated classification of fluorescence microscope images for location proteomics.
PMID 15207009 · PMC449699 · BMC bioinformatics · 2004 · 8 claims · 8 setups
New classifiers (SVMs, ensembles) and new wavelet-derived (Gabor, Daubechies) features improve recognition of protein subcellular location patterns over the previous neural network approach
-
Full-text index only
Constructing support vector machine ensembles for cancer classification based on proteomic profiling.
PMID 16689692 · PMC5173238 · Genomics, proteomics & bioinformatics · 2005 · 7 claims · 4 setups
CSVME, built by selecting a subset of base SVMs via SVM-RFE ranking and fusing them with a trained upper-layer SVM, achieves better classification performance than an ensemble of all base SVMs.
-
Full-text index only
Prioritization of candidate cancer genes--an aid to oncogenomic studies.
PMID 18710882 · PMC2566894 · Nucleic acids research · 2008 · 8 claims · 8 setups
Computational classifiers using combinations of protein conservation, gene structure, protein domains, protein interactions, and regulatory data can distinguish known cancer genes (CD/CR) from unlabelled human genes
-
Full-text index only
Local combinational variables: an approach used in DNA-binding helix-turn-helix motif prediction with sequence information.
PMID 19651875 · PMC2761287 · Nucleic acids research · 2009 · 8 claims · 7 setups
The LCV approach predicts HTH motifs with 93.29% accuracy, 93.93% sensitivity and 92.66% specificity using only primary sequence information