Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
A comprehensive sensitivity analysis of microarray breast cancer classification under feature variability.
PMID 19941644 · PMC2789744 · BMC bioinformatics · 2009 · 7 claims · 4 setups
Feature variability strongly influences breast cancer signature composition even when array platform and patient stratification are identical.
-
Full-text index only
Statistical challenges in preprocessing in microarray experiments in cancer.
PMID 18829474 · PMC3529914 · Clinical cancer research : an official journal of the American Association for Cancer Research · 2008 · 8 claims · 7 setups
Choice of pre-processing method materially changes which features are found significantly associated with survival in the Beer et al. lung cancer microarray dataset
-
Has reproduction · 69
Automatic discovery of 100-miRNA signature for cancer classification using ensemble feature selection.
PMID 31533612 · PMC6751684 · BMC bioinformatics · 2019 · 8 claims · 6 setups
An ensemble feature selection strategy using consensus of feature relevance across 8 classifier types identifies a 100-miRNA signature from a 1046-feature TCGA dataset
-
Has reproduction · 72
Prediction of prognostic signatures in triple-negative breast cancer based on the differential expression analysis via NanoString nCounter immune panel.
PMID 33138797 · PMC7607642 · BMC cancer · 2020 · 8 claims · 8 setups
edgeR-based DEG selection is more appropriate for feature selection than Elastic Net when sample sizes are small.
-
Full-text index only
Swarm intelligence based wavelet coefficient feature selection for mass spectral classification: an application to proteomics data.
PMID 19733729 · PMC2748225 · Analytica chimica acta · 2009 · 8 claims · 4 setups
ACA-based wavelet coefficient feature selection can achieve up to 100% classification accuracy on training, validating, and independent testing sets using only 5 selected features.
-
Full-text index only
On consensus biomarker selection.
PMID 17570864 · PMC1892093 · BMC bioinformatics · 2007 · 7 claims · 2 setups
Four popular feature ranking criteria (t-statistic, mutual information, peak probability contrasts, random forest variable importance) produce different rankings of the same features on the same dataset.
-
Full-text index only
A scale space approach for unsupervised feature selection in mass spectra classification for ovarian cancer detection.
PMID 19828085 · PMC2762074 · BMC bioinformatics · 2009 · 7 claims · 1 setups
A scale-space based unsupervised feature extraction method combined with SVM classification achieves high accuracy in ovarian cancer detection from serum mass spectra.
-
Has reproduction
Artificial Intelligence Meets Whole Slide Images: Deep Learning Model Shapes an Immune-Hot Tumor and Guides Precision Therapy in Bladder Cancer.
PMID 36245985 · PMC9553530 · Journal of oncology · 2022 · 6 claims · 8 setups
A deep learning WSI cluster (three-class mini batch K-means on Inception V3 features) is associated with overall survival (P<0.001) and is an independent prognostic predictor (P=0.031) in BLCA.
-
Full-text index only
Discovering cancer genes by integrating network and functional properties.
PMID 19765316 · PMC2758898 · BMC medical genomics · 2009 · 8 claims · 6 setups
Cancer genes have distinct PPI network topology (higher connectivity, higher clustering coefficient, shorter path length to known cancer genes) compared to non-cancer genes
-
Has reproduction · 59
Refining breast cancer biomarker discovery and drug targeting through an advanced data-driven approach.
PMID 38253993 · PMC10810249 · BMC bioinformatics · 2024 · 8 claims · 8 setups
The BGWO_SA_Ens algorithm (hybrid BGWO + simulated annealing with an ensemble classifier objective function) selects predictive breast cancer biomarker genes with high classification performance
-
Has reproduction · 100
Smart spatial omics (S2-omics) optimizes region of interest selection to capture molecular heterogeneity in diverse tissues.
PMID 41298871 · PMC12662399 · Nature cell biology · 2025 · 7 claims · 6 setups
S2-omics is an end-to-end workflow that automatically selects ROIs from H&E histology images to maximize molecular information content for spatial omics profiling.
-
Has reproduction · 81
Enabling Single-Cell Drug Response Annotations from Bulk RNA-Seq Using SCAD.
PMID 36762572 · PMC10104628 · Advanced science (Weinheim, Baden-Wurttemberg, Germany) · 2023 · 7 claims · 7 setups
SCAD, a transfer learning framework integrating adversarial discriminative domain adaptation (ADDA), can infer single-cell drug sensitivities by transferring knowledge from bulk RNA-seq pharmacogenomic data (GDSC) to scRNA-seq target domains
-
Has reproduction · 62
scATD: a high-throughput and interpretable framework for single-cell cancer drug resistance prediction and biomarker identification.
PMID 40501071 · PMC12159290 · Briefings in bioinformatics · 2025 · 8 claims · 6 setups
scATD enables high-throughput single-cell drug sensitivity prediction for new patients without model parameter retraining via bidirectional Bi-AdaIN style transfer
-
Has reproduction · 65
Interpretable and integrative analysis of single-cell multiomics with scMKL.
PMID 40770488 · PMC12328712 · Communications biology · 2025 · 7 claims · 7 setups
scMKL outperforms SVM, EasyMKL, MLP, and XGBoost in classification accuracy (AUROC) across multiple single-cell cancer datasets
-
Full-text index only
Constructing support vector machine ensembles for cancer classification based on proteomic profiling.
PMID 16689692 · PMC5173238 · Genomics, proteomics & bioinformatics · 2005 · 7 claims · 4 setups
CSVME, built by selecting a subset of base SVMs via SVM-RFE ranking and fusing them with a trained upper-layer SVM, achieves better classification performance than an ensemble of all base SVMs.
-
Full-text index only
Proteomics as a tool for biomarker discovery.
PMID 18057524 · PMC3851415 · Disease markers · 2007 · 8 claims · 7 setups
A useful clinical biomarker must be easily attainable, have adequate sensitivity, have adequate specificity, and lead to patient benefit through intervention
-
Full-text index only
Prioritization of candidate cancer genes--an aid to oncogenomic studies.
PMID 18710882 · PMC2566894 · Nucleic acids research · 2008 · 8 claims · 8 setups
Computational classifiers using combinations of protein conservation, gene structure, protein domains, protein interactions, and regulatory data can distinguish known cancer genes (CD/CR) from unlabelled human genes
-
Has reproduction · 75
stDyer-image improves clustering analysis of spatially resolved transcriptomics and proteomics with morphological images.
PMID 41692960 · PMC12960910 · Bioinformatics (Oxford, England) · 2026 · 7 claims · 5 setups
stDyer-image is an end-to-end deep learning framework that directly associates image features with predicted cluster labels to improve clustering of SRT and SRP data with images.
-
Full-text index only
Predicting positive p53 cancer rescue regions using Most Informative Positive (MIP) active learning.
PMID 19756158 · PMC2742196 · PLoS computational biology · 2009 · 8 claims · 4 setups
MIP active learning is a novel active learning method that preferentially seeks informative Positive (functionally active) examples rather than only maximizing classifier accuracy.
-
Has reproduction · 83
Gene-expression patterns in peripheral blood classify familial breast cancer susceptibility.
PMID 26538066 · PMC4634735 · BMC medical genomics · 2015 · 8 claims · 5 setups
A multigene peripheral-blood gene-expression biomarker accurately classifies which women from high-risk families develop familial breast cancer.