Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
Predicting the phenotypic effects of non-synonymous single nucleotide polymorphisms based on support vector machines.
PMID 18005451 · PMC2216041 · BMC bioinformatics · 2007 · 8 claims · 5 setups
Parepro, an SVM-based method integrating three attribute sets (RD, MI, IE) derived from evolutionary and residue-property information, predicts whether an nsSNP is deleterious or neutral.
-
Has reproduction · 90
A Decentralized Kidney Transplant Biopsy Classifier for Transplant Rejection Developed Using Genes of the Banff-Human Organ Transplant Panel.
PMID 35619722 · PMC9128066 · Frontiers in immunology · 2022 · 6 claims · 6 setups
A random forest model trained solely on B-HOT panel genes (B-HOT Model) accurately classifies kidney transplant biopsies as NR, ABMR, or TCMR.
-
Full-text index only
MiPred: classification of real and pseudo microRNA precursors using random forest prediction model with combined features.
PMID 17553836 · PMC1933124 · Nucleic acids research · 2007 · 8 claims · 8 setups
A hybrid feature combining local contiguous triplet structure-sequence composition, MFE of the secondary structure, and P-value of a randomization test improves classification of real vs pseudo pre-miRNAs
-
Full-text index only
Identification of deleterious non-synonymous single nucleotide polymorphisms using sequence-derived information.
PMID 18588693 · PMC2446391 · BMC bioinformatics · 2008 · 8 claims · 5 setups
A decision tree built on 10 selected sequence-derived features classifies SAPs as Disease or Polymorphism with 82.6% accuracy and 0.607 MCC in cross-validation.
-
Full-text index only
Prediction of catalytic residues using Support Vector Machine with selected protein sequence and structural properties.
PMID 16790052 · PMC1534064 · BMC bioinformatics · 2006 · 8 claims · 7 setups
The Sequential Minimal Optimization (SMO) SVM algorithm was the best-performing classifier among 26 WEKA classifiers for predicting catalytic residues
-
Full-text index only
Broad network-based predictability of Saccharomyces cerevisiae gene loss-of-function phenotypes.
PMID 18053250 · PMC2246260 · Genome biology · 2007 · 8 claims · 4 setups
Loss-of-function phenotypes in yeast are predictable from a gene's connections in a functional gene network via guilt-by-association.
-
Full-text index only
Prioritization of candidate cancer genes--an aid to oncogenomic studies.
PMID 18710882 · PMC2566894 · Nucleic acids research · 2008 · 8 claims · 8 setups
Computational classifiers using combinations of protein conservation, gene structure, protein domains, protein interactions, and regulatory data can distinguish known cancer genes (CD/CR) from unlabelled human genes
-
Has reproduction · 76
Bayesian prediction of microbial oxygen requirement.
PMID 26913185 · PMC4743139 · F1000Research · 2013 · 7 claims · 8 setups
A naive Bayesian classifier based on presence/absence of class-associated Pfam-A domains can distinguish three oxygen requirement classes (aerobe, anaerobe, facultative anaerobe) from genome sequence, unlike prior studies that only made pairwise distinctions.
-
Full-text index only
SePaCS--a web-based application for classification of seroreactivity profiles.
PMID 17478503 · PMC1933220 · Nucleic acids research · 2007 · 8 claims · 4 setups
SePaCS is a freely available web-based tool that trains and applies multiple classification methods (4 Naive Bayes variants, SVM with RBF kernel, LDA, DLDA) to seroreactivity profiles and outputs results as a summary table plus a detailed PDF report
-
Full-text index only
An SVM-based system for predicting protein subnuclear localizations.
PMID 16336650 · PMC1325059 · BMC bioinformatics · 2005 · 7 claims · 3 setups
New kernels defined on k-peptide vectors mapped by BLOSUM62-based high-scored pair matrices (D1, D2, D3) improve SVM discrimination of protein subnuclear localization compared to conventional k-peptide encodings.
-
Has reproduction · 71
Gene Set Enrichment Analysis Reveals Individual Variability in Host Responses in Tuberculosis Patients.
PMID 34421903 · PMC8375662 · Frontiers in immunology · 2021 · 8 claims · 8 setups
TB patients show substantial individual variability in the intensity of hallmark IFN responses, as well as in complement system, metabolic, and other pathway responses.
-
Full-text index only
Analysis of protein sequence and interaction data for candidate disease gene prediction.
PMID 17020920 · PMC1636487 · Nucleic acids research · 2006 · 8 claims · 7 setups
Combining CPS and CMP using known disease genes as input achieves sensitivity 0.52 and specificity 0.97, reducing candidate lists 13-fold
-
Full-text index only
PlasmoDraft: a database of Plasmodium falciparum gene function predictions based on postgenomic data.
PMID 18925948 · PMC2605471 · BMC bioinformatics · 2008 · 8 claims · 4 setups
Gonna, a supervised k-nearest-neighbor Guilt-By-Association predictor, proposes GO annotations for a gene based on similarity of its transcriptome, proteome, or interactome profile to genes already annotated by GeneDB
-
Full-text index only
Potential biomarkers of human salivary function: a modified proteomic approach.
PMID 18804197 · PMC2633945 · Archives of oral biology · 2009 · 6 claims · 6 setups
Two SDS-PAGE bands, identified by MS-MS as statherin and a truncated (N-terminal 8-aa-missing) cystatin S, are the strongest and most consistent predictors of HAA/LAA group membership and clinical/microbiological outcomes
-
Full-text index only
Extraction of human kinase mutations from literature, databases and genotyping studies.
PMID 19758464 · PMC2745582 · BMC bioinformatics · 2009 · 7 claims · 6 setups
A literature mining pipeline combining MutationFinder, false-positive filtering, and SVM-based classification can extract and disambiguate single-point mutation mentions from abstracts and full text
-
Full-text index only
Prodepth: predict residue depth by support vector regression approach from protein sequences only.
PMID 19759917 · PMC2742725 · PloS one · 2009 · 8 claims · 8 setups
Residue depth can be reliably predicted solely from protein primary sequence using support vector regression on sequence-derived features.
-
Full-text index only
Long-range regulation is a major driving force in maintaining genome integrity.
PMID 19682388 · PMC2741452 · BMC evolutionary biology · 2009 · 7 claims · 5 setups
Long-range transcriptional regulation is a major driving force in maintaining genome integrity by constraining where chromosomal breakpoints can become fixed.
-
Has reproduction · 68
Constraints to gene flow increase the risk of genome erosion in the Ngorongoro Crater lion population.
PMID 40258987 · PMC12012037 · Communications biology · 2025 · 8 claims · 9 setups
200 years of quasi-isolation and the 1962 epizootic caused a two-fold increase in inbreeding and an excess of highly deleterious mutations in Crater lions relative to other Greater Serengeti populations
-
Full-text index only
Species-specific protein sequence and fold optimizations.
PMID 12487631 · PMC139977 · BMC bioinformatics · 2002 · 7 claims · 7 setups
Environmental niche is a significant factor explaining variability in amino acid composition across 100 complete genomes
-
Full-text index only
Filtering high-throughput protein-protein interaction data using a combination of genomic features.
PMID 15833142 · PMC1127019 · BMC bioinformatics · 2005 · 8 claims · 8 setups
A combination of three genomic features (interacting Pfam domains, GO annotations, sequence homology) using naive Bayesian networks predicts true protein-protein interactions with high sensitivity and good specificity.