Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Has reproduction · 80
Colorectal Cancer Prediction Based on Weighted Gene Co-Expression Network Analysis and Variational Auto-Encoder.
PMID 32825264 · PMC7563725 · Biomolecules · 2020 · 6 claims · 7 setups
Combining WGCNA-derived hub genes with a VAE-derived 10-dimensional representation as features for an SVM classifier achieves high accuracy (0.9692) and AUC (0.9981) for colorectal cancer prediction.
-
Full-text index only
Exhaustive prediction of disease susceptibility to coding base changes in the human genome.
PMID 18793467 · PMC2537574 · BMC bioinformatics · 2008 · 8 claims · 7 setups
Inter-species conservation is the strongest single predictor of disease-associated coding mutations among the factors tested.
-
Full-text index only
Pol II promoter prediction using characteristic 4-mer motifs: a machine learning approach.
PMID 18834544 · PMC2575220 · BMC bioinformatics · 2008 · 8 claims · 8 setups
128 discriminating 4-mer motifs combined with an SVM (RBF kernel, LIBSVM) can distinguish promoter from non-promoter DNA sequences
-
Full-text index only
Accurate splice site prediction using support vector machines.
PMID 18269701 · PMC2230508 · BMC bioinformatics · 2007 · 8 claims · 5 setups
Weighted degree (WD) kernel SVMs outperform Markov Chains, GeneSplicer and SpliceMachine for genome-wide splice site recognition
-
Full-text index only
Prediction of catalytic residues using Support Vector Machine with selected protein sequence and structural properties.
PMID 16790052 · PMC1534064 · BMC bioinformatics · 2006 · 8 claims · 7 setups
The Sequential Minimal Optimization (SMO) SVM algorithm was the best-performing classifier among 26 WEKA classifiers for predicting catalytic residues
-
Has reproduction · 59
Application of Machine Learning in Predicting Hepatic Metastasis or Primary Site in Gastroenteropancreatic Neuroendocrine Tumors.
PMID 37887568 · PMC10605255 · Current oncology (Toronto, Ont.) · 2023 · 8 claims · 7 setups
Multi-gene random forest models classify primary tumor vs. liver metastasis samples with 100% accuracy in training/test cohorts and >90% accuracy in an independent validation cohort
-
Full-text index only
All systems GO for understanding mouse gene function.
PMID 15610553 · PMC549721 · Journal of biology · 2004 · 7 claims · 4 setups
Quantitative, multivariate cross-tissue expression measurements are powerfully predictive of gene function
-
Full-text index only
Predicting the phenotypic effects of non-synonymous single nucleotide polymorphisms based on support vector machines.
PMID 18005451 · PMC2216041 · BMC bioinformatics · 2007 · 8 claims · 5 setups
Parepro, an SVM-based method integrating three attribute sets (RD, MI, IE) derived from evolutionary and residue-property information, predicts whether an nsSNP is deleterious or neutral.
-
Full-text index only
Mining novel biomarkers for prognosis of gastric cancer with serum proteomics.
PMID 19740432 · PMC2753349 · Journal of experimental & clinical cancer research : CR · 2009 · 7 claims · 4 setups
A 5-peak prognosis pattern (4474, 4542, 6443/6643, 4988, 6685 Da) predicts poor vs good prognosis in GC with higher sensitivity/specificity than CEA and TNM stage
-
Has reproduction · 50
STAT1 and IL-7 as potential diagnostic biomarkers for distinguishing high-grade from low-grade serous ovarian cancer: a multi-cohort analysis.
PMID 42058211 · PMC13120972 · Frontiers in immunology · 2026 · 7 claims · 6 setups
STAT1 and IL-7 are differentially expressed immune-related genes that can distinguish HGSOC from LGSOC and may serve as ancillary diagnostic biomarkers.
-
Has reproduction · 49
Integration of Transcriptomics With Interpretable Artificial Intelligence for Identifying Molecular Signatures of Physiological Stress in Sleep Deprivation.
PMID 42216239 · PMC13240488 · Journal of cellular and molecular medicine · 2026 · 8 claims · 8 setups
S100A3 is a robust candidate biomarker showing consistent discriminatory performance across the acute sleep deprivation training cohort, an independent sleep deprivation cohort, and a chronic insomnia cohort.
-
Full-text index only
Emerging translational bioinformatics: knowledge-guided biomarker identification for cancer diagnostics.
PMID 19964620 · PMC5003034 · Annual International Conference of the IEEE Engineering in Medicine and Biology Society. IEEE Engineering in Medicine and Biology Society. Annual International Conference · 2009 · 8 claims · 3 setups
omniBiomarker, a web-based application, uses prior biological knowledge to identify the most biologically relevant gene ranking metric for a given clinical problem
-
Full-text index only
Supervised learning-based tagSNP selection for genome-wide disease classifications.
PMID 18366619 · PMC2386071 · BMC genomics · 2008 · 7 claims · 2 setups
SRFA (Supervised Recursive Feature Addition) is a novel feature selection method combining supervised learning and statistical redundancy measures for SNP selection
-
Full-text index only
A modified T-test feature selection method and its application on the HapMap genotype data.
PMID 18267305 · PMC5054219 · Genomics, proteomics & bioinformatics · 2007 · 7 claims · 4 setups
A modified t-test ranking measure, extended to handle nominal SNP genotype data via vector transformation, can effectively rank SNPs by their discriminative capability for population classification.
-
Full-text index only
Local combinational variables: an approach used in DNA-binding helix-turn-helix motif prediction with sequence information.
PMID 19651875 · PMC2761287 · Nucleic acids research · 2009 · 8 claims · 7 setups
The LCV approach predicts HTH motifs with 93.29% accuracy, 93.93% sensitivity and 92.66% specificity using only primary sequence information
-
Has reproduction · 62
Predicting Bone Metastasis Using Gene Expression-Based Machine Learning Models.
PMID 34858485 · PMC8631472 · Frontiers in genetics · 2021 · 7 claims · 5 setups
A DNN model using the top 34 betweenness-centrality-ranked hub genes predicts bone metastasis with AUC of 92.11% on the GEO validation data.
-
Has reproduction · 66
Integrative bioinformatics and artificial intelligence analyses of transcriptomics data identified genes associated with major depressive disorders including NRG1.
PMID 37583471 · PMC10423927 · Neurobiology of stress · 2023 · 7 claims · 5 setups
Differentially expressed genes in MDD patients are enriched in immune response, inflammatory response, neurodegeneration, and cerebellar atrophy pathways.
-
Has reproduction · 96
Scalable Prediction of Acute Myeloid Leukemia Using High-Dimensional Machine Learning and Blood Transcriptomics.
PMID 31918046 · PMC6992905 · iScience · 2020 · 8 claims · 8 setups
Data-driven, high-dimensional ML approaches that learn multivariate signatures directly from genome-wide transcriptomic data (no prior gene selection) yield accurate and robust AML classifiers.
-
Full-text index only
A comprehensive sensitivity analysis of microarray breast cancer classification under feature variability.
PMID 19941644 · PMC2789744 · BMC bioinformatics · 2009 · 7 claims · 4 setups
Feature variability strongly influences breast cancer signature composition even when array platform and patient stratification are identical.
-
Full-text index only
Sequence and structure signatures of cancer mutation hotspots in protein kinases.
PMID 19834613 · PMC2759519 · PloS one · 2009 · 8 claims · 6 setups
Developed CKMD (Composite Kinase Mutation Database), an integrated bioinformatics resource mapping genetic variation in protein kinase genes to sequence, structural, and functional data