Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
Pol II promoter prediction using characteristic 4-mer motifs: a machine learning approach.
PMID 18834544 · PMC2575220 · BMC bioinformatics · 2008 · 8 claims · 8 setups
128 discriminating 4-mer motifs combined with an SVM (RBF kernel, LIBSVM) can distinguish promoter from non-promoter DNA sequences
-
Full-text index only
MiPred: classification of real and pseudo microRNA precursors using random forest prediction model with combined features.
PMID 17553836 · PMC1933124 · Nucleic acids research · 2007 · 8 claims · 8 setups
A hybrid feature combining local contiguous triplet structure-sequence composition, MFE of the secondary structure, and P-value of a randomization test improves classification of real vs pseudo pre-miRNAs
-
Full-text index only
Identification of diagnostic markers for tuberculosis by proteomic fingerprinting of serum.
PMID 16980117 · PMC7159276 · Lancet (London, England) · 2006 · 8 claims · 5 setups
An SVM classifier trained on serum proteomic profiles discriminated patients with active tuberculosis from controls with clinically overlapping conditions
-
Full-text index only
An SVM-based system for predicting protein subnuclear localizations.
PMID 16336650 · PMC1325059 · BMC bioinformatics · 2005 · 7 claims · 3 setups
New kernels defined on k-peptide vectors mapped by BLOSUM62-based high-scored pair matrices (D1, D2, D3) improve SVM discrimination of protein subnuclear localization compared to conventional k-peptide encodings.
-
Full-text index only
Predicting the phenotypic effects of non-synonymous single nucleotide polymorphisms based on support vector machines.
PMID 18005451 · PMC2216041 · BMC bioinformatics · 2007 · 8 claims · 5 setups
Parepro, an SVM-based method integrating three attribute sets (RD, MI, IE) derived from evolutionary and residue-property information, predicts whether an nsSNP is deleterious or neutral.
-
Has reproduction · 32
Identifying COVID-19-Specific Transcriptomic Biomarkers with Machine Learning Methods.
PMID 34307679 · PMC8272456 · BioMed research international · 2021 · 7 claims · 2 setups
A pipeline combining Boruta and mRMR feature selection with incremental feature selection (IFS) was used to identify COVID-19-specific transcriptomic biomarkers from blood gene expression data.
-
Has reproduction · 89
miRge 2.0 for comprehensive analysis of microRNA sequencing data.
PMID 30153801 · PMC6112139 · BMC bioinformatics · 2018 · 8 claims · 6 setups
miRge 2.0 introduces a novel SVM-based miRNA detection method using both hairpin structure and isomiR composition, yielding higher specificity for miRNA identification
-
Full-text index only
Statistical learning of peptide retention behavior in chromatographic separations: a new kernel-based approach for computational proteomics.
PMID 18053132 · PMC2254445 · BMC bioinformatics · 2007 · 6 claims · 5 setups
The paired oligo-border kernel (POBK) combined with SVMs predicts peptide adsorption/elution in SAX-SPE and retention time in IP-RP-HPLC more accurately than existing methods.
-
Full-text index only
Exhaustive prediction of disease susceptibility to coding base changes in the human genome.
PMID 18793467 · PMC2537574 · BMC bioinformatics · 2008 · 8 claims · 7 setups
Inter-species conservation is the strongest single predictor of disease-associated coding mutations among the factors tested.
-
Has reproduction · 86
Molecular Classification Models for Triple Negative Breast Cancer Subtype Using Machine Learning.
PMID 34575658 · PMC8472680 · Journal of personalized medicine · 2021 · 6 claims · 4 setups
A training gene set of 719 unique upregulated DEGs (subtype-specific) can be used to build ML models that classify TNBC into BLIA, BLIS, MES, and LAR subtypes.
-
Full-text index only
A comprehensive sensitivity analysis of microarray breast cancer classification under feature variability.
PMID 19941644 · PMC2789744 · BMC bioinformatics · 2009 · 7 claims · 4 setups
Feature variability strongly influences breast cancer signature composition even when array platform and patient stratification are identical.
-
Full-text index only
Local combinational variables: an approach used in DNA-binding helix-turn-helix motif prediction with sequence information.
PMID 19651875 · PMC2761287 · Nucleic acids research · 2009 · 8 claims · 7 setups
The LCV approach predicts HTH motifs with 93.29% accuracy, 93.93% sensitivity and 92.66% specificity using only primary sequence information
-
Has reproduction · 49
Integrative transcriptomics and single-cell transcriptomics analyses reveal potential biomarkers and mechanisms of action in papillary thyroid carcinoma.
PMID 40520228 · PMC12162626 · Frontiers in genetics · 2025 · 8 claims · 8 setups
ENTPD1, SERPINA1, and TACSTD2 are potential transcriptomic biomarkers for PTC
-
Full-text index only
Boosting accuracy of automated classification of fluorescence microscope images for location proteomics.
PMID 15207009 · PMC449699 · BMC bioinformatics · 2004 · 8 claims · 8 setups
New classifiers (SVMs, ensembles) and new wavelet-derived (Gabor, Daubechies) features improve recognition of protein subcellular location patterns over the previous neural network approach
-
Full-text index only
Discovering cancer genes by integrating network and functional properties.
PMID 19765316 · PMC2758898 · BMC medical genomics · 2009 · 8 claims · 6 setups
Cancer genes have distinct PPI network topology (higher connectivity, higher clustering coefficient, shorter path length to known cancer genes) compared to non-cancer genes
-
Full-text index only
Emerging translational bioinformatics: knowledge-guided biomarker identification for cancer diagnostics.
PMID 19964620 · PMC5003034 · Annual International Conference of the IEEE Engineering in Medicine and Biology Society. IEEE Engineering in Medicine and Biology Society. Annual International Conference · 2009 · 8 claims · 3 setups
omniBiomarker, a web-based application, uses prior biological knowledge to identify the most biologically relevant gene ranking metric for a given clinical problem
-
Full-text index only
Supervised learning-based tagSNP selection for genome-wide disease classifications.
PMID 18366619 · PMC2386071 · BMC genomics · 2008 · 7 claims · 2 setups
SRFA (Supervised Recursive Feature Addition) is a novel feature selection method combining supervised learning and statistical redundancy measures for SNP selection
-
Has reproduction · 50
Wx: a neural network-based feature selection algorithm for transcriptomic data.
PMID 31324856 · PMC6642261 · Scientific reports · 2019 · 8 claims · 8 setups
The Wx algorithm ranks genes by a discriminative index (DI) score representing classification power for distinguishing given groups, enabling intuitive selection of optimal biomarker genes.