Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
Classification of real and pseudo microRNA precursors using local structure-sequence features and support vector machine.
PMID 16381612 · PMC1360673 · BMC bioinformatics · 2005 · 7 claims · 7 setups
A 32-dimensional triplet structure-sequence feature vector combined with SVM (triplet-SVM) can distinguish real human pre-miRNAs from pseudo pre-miRNA hairpins with ~90% accuracy.
-
Full-text index only
Discovery and identification of potential biomarkers of papillary thyroid carcinoma.
PMID 19785722 · PMC2761863 · Molecular cancer · 2009 · 8 claims · 7 setups
A 3-peak (m/z 9190, 6631, 8697 Da) SVM classification model discriminates PTC from non-cancer controls with high sensitivity and specificity
-
Has reproduction · 86
Molecular Classification Models for Triple Negative Breast Cancer Subtype Using Machine Learning.
PMID 34575658 · PMC8472680 · Journal of personalized medicine · 2021 · 6 claims · 4 setups
A training gene set of 719 unique upregulated DEGs (subtype-specific) can be used to build ML models that classify TNBC into BLIA, BLIS, MES, and LAR subtypes.
-
Has reproduction · 83
Hierarchical classification-based pan-cancer methylation analysis to classify primary cancer.
PMID 38066424 · PMC10709847 · BMC bioinformatics · 2023 · 8 claims · 5 setups
CHCT, a hierarchical classification tool, splits classification of 30 cancer types into ten smaller subproblems using a two-tier architecture to classify primary cancer by methylation profile
-
Has reproduction · 60
TRAPID 2.0: a web application for taxonomic and functional analysis of de novo transcriptomes.
PMID 34197621 · PMC8464036 · Nucleic acids research · 2021 · 8 claims · 8 setups
TRAPID 2.0 is a web application performing global characterization of de novo transcriptomes via structural, functional, and taxonomic annotation in an initial processing phase, followed by an exploratory phase of downstream analyses.
-
Has reproduction · 99
Evaluation of taxonomic classification and profiling methods for long-read shotgun metagenomic sequencing datasets.
PMID 36513983 · PMC9749362 · BMC bioinformatics · 2022 · 8 claims · 7 setups
Long-read classifiers generally performed best among the 11 methods tested
-
Full-text index only
Boosting accuracy of automated classification of fluorescence microscope images for location proteomics.
PMID 15207009 · PMC449699 · BMC bioinformatics · 2004 · 8 claims · 8 setups
New classifiers (SVMs, ensembles) and new wavelet-derived (Gabor, Daubechies) features improve recognition of protein subcellular location patterns over the previous neural network approach
-
Full-text index only
Towards precise classification of cancers based on robust gene functional expression profiles.
PMID 15774002 · PMC1274255 · BMC bioinformatics · 2005 · 6 claims · 7 setups
Functional expression profiles (FEPs) achieve comparable or better classification performance than conventional gene expression profiles (GEPs) across four public microarray datasets
-
Full-text index only
Predicting the phenotypic effects of non-synonymous single nucleotide polymorphisms based on support vector machines.
PMID 18005451 · PMC2216041 · BMC bioinformatics · 2007 · 8 claims · 5 setups
Parepro, an SVM-based method integrating three attribute sets (RD, MI, IE) derived from evolutionary and residue-property information, predicts whether an nsSNP is deleterious or neutral.
-
Full-text index only
MiPred: classification of real and pseudo microRNA precursors using random forest prediction model with combined features.
PMID 17553836 · PMC1933124 · Nucleic acids research · 2007 · 8 claims · 8 setups
A hybrid feature combining local contiguous triplet structure-sequence composition, MFE of the secondary structure, and P-value of a randomization test improves classification of real vs pseudo pre-miRNAs
-
Full-text index only
Exhaustive prediction of disease susceptibility to coding base changes in the human genome.
PMID 18793467 · PMC2537574 · BMC bioinformatics · 2008 · 8 claims · 7 setups
Inter-species conservation is the strongest single predictor of disease-associated coding mutations among the factors tested.
-
Full-text index only
Logical Analysis of Data (LAD) model for the early diagnosis of acute ischemic stroke.
PMID 18616825 · PMC2492849 · BMC medical informatics and decision making · 2008 · 7 claims · 5 setups
An LAD classification model built from a support-set of 3 peptide peaks can distinguish stroke patients from controls with 75% accuracy on an independent validation set
-
Full-text index only
Genomic data sampling and its effect on classification performance assessment.
PMID 12553886 · PMC149349 · BMC bioinformatics · 2003 · 8 claims · 3 setups
Cross-validation, leave-one-out, and bootstrap are designed to reduce bias and variance in accuracy estimation from small samples.
-
Has reproduction · 22
Machine learning-based DNA microarray analysis for disease detection using the MICRO-AI framework.
PMID 41925147 · PMC13051084 · Science progress · 2026 · 8 claims · 1 setups
MICRO-AI's attention-weighted feature fusion reduces dimensionality by over 99% (from ~20,000 to ~127 genes) without loss of biological significance
-
Full-text index only
A scale space approach for unsupervised feature selection in mass spectra classification for ovarian cancer detection.
PMID 19828085 · PMC2762074 · BMC bioinformatics · 2009 · 7 claims · 1 setups
A scale-space based unsupervised feature extraction method combined with SVM classification achieves high accuracy in ovarian cancer detection from serum mass spectra.
-
Has reproduction · 87
A robust data scaling algorithm to improve classification accuracies in biomedical data.
PMID 27612635 · PMC5016890 · BMC bioinformatics · 2016 · 8 claims · 2 setups
Models trained on data scaled by the GL algorithm outperform models trained on data scaled by the Min-max or Z-score algorithms across 16 binary classification tasks, measured by AUROC and percentage of correct classification
-
Has reproduction · 76
Topologically inferring pathway activity toward precise cancer classification via integrating genomic and metabolomic data: prostate cancer as a case.
PMID 26286638 · PMC4541321 · Scientific reports · 2015 · 6 claims · 4 setups
DRW-GM integrates gene expression and metabolomic profiles via directed random walk on a global gene–metabolite pathway graph to weight genes by topological importance and infer reproducible pathway activities
-
Has reproduction · 78
T Cells With Activated STAT4 Drive the High-Risk Rejection State to Renal Allograft Failure After Kidney Transplantation.
PMID 35844542 · PMC9283858 · Frontiers in immunology · 2022 · 8 claims · 8 setups
Unsupervised UMAP/Leiden clustering of 2,611 microarray datasets reveals 6 rejection states that diverge from traditional Banff clinical classification
-
Has reproduction · 65
Interpretable and integrative analysis of single-cell multiomics with scMKL.
PMID 40770488 · PMC12328712 · Communications biology · 2025 · 8 claims · 7 setups
scMKL combines multiple kernel learning with random Fourier features and group Lasso to jointly model transcriptomic and epigenomic single-cell data interpretably
-
Has reproduction · 42
CanCellCap: robust cancer cell capture across tissue types on single-cell RNA-seq data by multi-domain learning.
PMID 40739511 · PMC12312500 · BMC biology · 2025 · 8 claims · 8 setups
CanCellCap, a multi-domain learning framework integrating domain adversarial learning and Mixture of Experts, identifies cancer cells across all tissues, cancers, and sequencing platforms by extracting tissue-common and tissue-specific gene expression patterns.