Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
nsSNPAnalyzer: identifying disease-associated nonsynonymous single nucleotide polymorphisms.
PMID 15980516 · PMC1160133 · Nucleic acids research · 2005 · 6 claims · 4 setups
nsSNPAnalyzer is a web server that predicts whether a query nsSNP is disease-associated or functionally neutral using a Random Forest classifier combining structural and evolutionary information
-
Full-text index only
CanPredict: a computational tool for predicting cancer-associated missense mutations.
PMID 17537827 · PMC1933186 · Nucleic acids research · 2007 · 8 claims · 7 setups
CanPredict is a web application providing public access to a random forest classifier that combines SIFT, LogR.E-value, and GOSS scores to predict whether a missense mutation is cancer-associated
-
Full-text index only
MiPred: classification of real and pseudo microRNA precursors using random forest prediction model with combined features.
PMID 17553836 · PMC1933124 · Nucleic acids research · 2007 · 8 claims · 8 setups
A hybrid feature combining local contiguous triplet structure-sequence composition, MFE of the secondary structure, and P-value of a randomization test improves classification of real vs pseudo pre-miRNAs
-
Full-text index only
Prediction of candidate primary immunodeficiency disease genes using a support vector machine learning approach.
PMID 19801557 · PMC2780952 · DNA research : an international journal for rapid publication of reports on genes and genomes · 2009 · 6 claims · 3 setups
An SVM trained on 69 binary features of known PID genes can accurately classify PID vs non-PID genes and predict novel candidate PID genes
-
Has reproduction · 86
Molecular Classification Models for Triple Negative Breast Cancer Subtype Using Machine Learning.
PMID 34575658 · PMC8472680 · Journal of personalized medicine · 2021 · 6 claims · 4 setups
A training gene set of 719 unique upregulated DEGs (subtype-specific) can be used to build ML models that classify TNBC into BLIA, BLIS, MES, and LAR subtypes.
-
Full-text index only
Identification of deleterious non-synonymous single nucleotide polymorphisms using sequence-derived information.
PMID 18588693 · PMC2446391 · BMC bioinformatics · 2008 · 8 claims · 5 setups
A decision tree built on 10 selected sequence-derived features classifies SAPs as Disease or Polymorphism with 82.6% accuracy and 0.607 MCC in cross-validation.
-
Has reproduction · 74
SpaGene: A Deep Adversarial Framework for Spatial Gene Imputation.
PMID 42146899 · PMC13176606 · Computational and structural biotechnology journal · 2026 · 8 claims · 6 setups
SpaGene improves average PCC and SSIM and reduces RMSE compared to 6 baseline methods (SpaGE, gimVI, Tangram, VISTA, spRefine, stDiff) across 8 diverse ST-SC dataset pairs under gene-holdout evaluation.
-
Full-text index only
Boosting accuracy of automated classification of fluorescence microscope images for location proteomics.
PMID 15207009 · PMC449699 · BMC bioinformatics · 2004 · 8 claims · 8 setups
New classifiers (SVMs, ensembles) and new wavelet-derived (Gabor, Daubechies) features improve recognition of protein subcellular location patterns over the previous neural network approach
-
Full-text index only
Constructing support vector machine ensembles for cancer classification based on proteomic profiling.
PMID 16689692 · PMC5173238 · Genomics, proteomics & bioinformatics · 2005 · 7 claims · 4 setups
CSVME, built by selecting a subset of base SVMs via SVM-RFE ranking and fusing them with a trained upper-layer SVM, achieves better classification performance than an ensemble of all base SVMs.
-
Full-text index only
Speeding disease gene discovery by sequence based candidate prioritization.
PMID 15766383 · PMC1274252 · BMC bioinformatics · 2005 · 7 claims · 8 setups
Disease genes (OMIM) differ significantly from non-disease genes in sequence-based features including gene/cDNA/protein size, exon number, homolog conservation, secretion signal, 3' UTR length, CpG islands, and distance to nearest gene.
-
Full-text index only
Ab initio identification of human microRNAs based on structure motifs.
PMID 18088431 · PMC2238772 · BMC bioinformatics · 2007 · 8 claims · 7 setups
MiRPred predicts miRNA precursors ab initio using only predicted secondary structure motifs, ignoring nucleotide sequence
-
Has reproduction · 83
Hierarchical classification-based pan-cancer methylation analysis to classify primary cancer.
PMID 38066424 · PMC10709847 · BMC bioinformatics · 2023 · 8 claims · 5 setups
CHCT, a hierarchical classification tool, splits classification of 30 cancer types into ten smaller subproblems using a two-tier architecture to classify primary cancer by methylation profile
-
Full-text index only
Statistical learning of peptide retention behavior in chromatographic separations: a new kernel-based approach for computational proteomics.
PMID 18053132 · PMC2254445 · BMC bioinformatics · 2007 · 6 claims · 5 setups
The paired oligo-border kernel (POBK) combined with SVMs predicts peptide adsorption/elution in SAX-SPE and retention time in IP-RP-HPLC more accurately than existing methods.
-
Full-text index only
Cancer-specific high-throughput annotation of somatic mutations: computational prediction of driver missense mutations.
PMID 19654296 · PMC2763410 · Cancer research · 2009 · 7 claims · 7 setups
CHASM, a Random Forest-based computational method, was developed to identify and prioritize missense mutations likely to be functional drivers of tumor cell proliferation.
-
Has reproduction · 70
Predicting enhancers in mammalian genomes using supervised hidden Markov models.
PMID 30917778 · PMC6437899 · BMC bioinformatics · 2019 · 8 claims · 8 setups
eHMM predicts enhancers with high precision and recall comparable to state-of-the-art methods and consistently outperforms them in accuracy and resolution
-
Full-text index only
Machine-learning approaches for classifying haplogroup from Y chromosome STR data.
PMID 18551166 · PMC2396484 · PLoS computational biology · 2008 · 8 claims · 5 setups
Y-STR allelic variability is partitioned more by differences among haplogroups than by differences among populations, suggesting Y-STRs carry haplogroup information
-
Has reproduction · 78
Enhancing chemotherapy response prediction via matched colorectal tumor-organoid gene expression analysis and network-based biomarker selection.
PMID 39754813 · PMC11754497 · Translational oncology · 2025 · 6 claims · 8 setups
A consensus WGCNA approach combining matched tumor-organoid and independent organoid drug-response expression data identifies gene modules and hub genes predictive of 5-FU chemotherapy response
-
Full-text index only
Using ESTs to improve the accuracy of de novo gene prediction.
PMID 16817966 · PMC1534067 · BMC bioinformatics · 2006 · 8 claims · 8 setups
TWINSCAN_EST combines EST alignments with TWINSCAN via a trainable 'ESTseq' representation and improves exact gene structure prediction accuracy on the whole C. elegans genome