Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Has reproduction · 81
Enabling Single-Cell Drug Response Annotations from Bulk RNA-Seq Using SCAD.
PMID 36762572 · PMC10104628 · Advanced science (Weinheim, Baden-Wurttemberg, Germany) · 2023 · 7 claims · 7 setups
SCAD, a transfer learning framework integrating adversarial discriminative domain adaptation (ADDA), can infer single-cell drug sensitivities by transferring knowledge from bulk RNA-seq pharmacogenomic data (GDSC) to scRNA-seq target domains
-
Has reproduction · 44
An OMICs-based meta-analysis to support infection state stratification.
PMID 33560295 · PMC8388022 · Bioinformatics (Oxford, England) · 2021 · 7 claims · 6 setups
Multi-class machine learning models built from cross-platform microarray meta-analysis can distinguish bacterial, viral and no-infection states with high accuracy (best model: 93% bacterial, 89% viral correct).
-
Full-text index only
A comprehensive sensitivity analysis of microarray breast cancer classification under feature variability.
PMID 19941644 · PMC2789744 · BMC bioinformatics · 2009 · 7 claims · 4 setups
Feature variability strongly influences breast cancer signature composition even when array platform and patient stratification are identical.
-
Has reproduction · 86
Molecular Classification Models for Triple Negative Breast Cancer Subtype Using Machine Learning.
PMID 34575658 · PMC8472680 · Journal of personalized medicine · 2021 · 6 claims · 4 setups
A training gene set of 719 unique upregulated DEGs (subtype-specific) can be used to build ML models that classify TNBC into BLIA, BLIS, MES, and LAR subtypes.
-
Full-text index only
Statistical challenges in preprocessing in microarray experiments in cancer.
PMID 18829474 · PMC3529914 · Clinical cancer research : an official journal of the American Association for Cancer Research · 2008 · 8 claims · 7 setups
Choice of pre-processing method materially changes which features are found significantly associated with survival in the Beer et al. lung cancer microarray dataset
-
Has reproduction · 72
Prediction of prognostic signatures in triple-negative breast cancer based on the differential expression analysis via NanoString nCounter immune panel.
PMID 33138797 · PMC7607642 · BMC cancer · 2020 · 8 claims · 7 setups
edgeR identifies 9 DEGs associated with pCR and 13 DEGs associated with relapse from 579 immune genes in a small TNBC sample set (n=55)
-
Has reproduction · 83
Multimodal data integration for biologically-relevant artificial intelligence to guide adjuvant chemotherapy in stage II colorectal cancer.
PMID 40472802 · PMC12171563 · EBioMedicine · 2025 · 6 claims · 7 setups
AI-derived radiological clustering identifies stage II CRC patients with significantly different survival benefit from adjuvant chemotherapy
-
Has reproduction · 89
Graph Random Forest: A Graph Embedded Algorithm for Identifying Highly Connected Important Features.
PMID 37509188 · PMC10377046 · Biomolecules · 2023 · 8 claims · 3 setups
Graph Random Forest (GRF) embeds graph/network information directly into the decision-tree building process by splitting on features in the k-hop neighborhood of a data-driven head-splitting node.
-
Full-text index only
PA-GOSUB: a searchable database of model organism protein sequences with their predicted Gene Ontology molecular function and subcellular localization.
PMID 15608166 · PMC540074 · Nucleic acids research · 2005 · 7 claims · 4 setups
PA-GOSUB significantly extends the coverage of GO molecular function and subcellular localization annotations for 10 model organism proteomes compared with existing databases (GOA, Swiss-Prot).
-
Has reproduction · 44
Dynamic Gene Attention Focus (DyGAF): Enhancing Biomarker Identification Through Dual-Model Attention Networks.
PMID 40160891 · PMC11951896 · Bioinformatics and biology insights · 2025 · 6 claims · 5 setups
DyGAF, a dual-model attention neural network (independent Model A + dependent Model B), identifies and ranks genes by significance for COVID-19 biomarker discovery more effectively than differential expression analysis (DEA) and random forest (RF) feature selection
-
Has reproduction · 79
Interpretable prediction models for widespread m6A RNA modification across cell lines and tissues.
PMID 37995291 · PMC10697738 · Bioinformatics (Oxford, England) · 2023 · 7 claims · 6 setups
CLSM6A, a CNN-based model set, predicts single-nucleotide-resolution m6A RNA modification sites across eight cell lines and three tissues in H. sapiens
-
Has reproduction · 83
Gene-expression patterns in peripheral blood classify familial breast cancer susceptibility.
PMID 26538066 · PMC4634735 · BMC medical genomics · 2015 · 8 claims · 5 setups
A multigene peripheral-blood gene-expression biomarker accurately classifies which women from high-risk families develop familial breast cancer.
-
Has reproduction · 40
On the holobiont 'predictome' of immunocompetence in pigs.
PMID 37127575 · PMC10150480 · Genetics, selection, evolution : GSE · 2023 · 8 claims · 8 setups
Holobiont (combined genotype + microbiome) models performed better than partial models (genotype-only or microbiome-only) overall
-
Has reproduction · 85
Predicting the pathogenicity of missense variants using features derived from AlphaFold2.
PMID 37084271 · PMC10203375 · Bioinformatics (Oxford, England) · 2023 · 6 claims · 8 setups
AlphaFold2-derived structural features (solvent accessibility, amino acid network features, physicochemical environment, pLDDT) can be used to train a random forest classifier (AlphScore) that distinguishes proxy-benign from proxy-pathogenic missense variants.
-
Full-text index only
Interaction profile-based protein classification of death domain.
PMID 15189571 · PMC459208 · BMC bioinformatics · 2004 · 7 claims · 6 setups
An SVM-based classifier using Residue Pair Interaction Profiles (RPIPs) can classify death domain superfamily members into subfamilies with 89% average cross-validation accuracy
-
Has reproduction · 88
Human methylome variation across Infinium 450K data on the Gene Expression Omnibus.
PMID 33937763 · PMC8061458 · NAR genomics and bioinformatics · 2021 · 8 claims · 6 setups
Approximately two-thirds of compiled HM450K samples are from blood, one-quarter from brain, and roughly one-third from cancer patients.
-
Has reproduction · 80
Comprehensive analysis of transcriptomics and radiomics revealed the potential of TEDC2 as a diagnostic marker for lung adenocarcinoma.
PMID 39553728 · PMC11569783 · PeerJ · 2024 · 8 claims · 8 setups
WGCNA identified 214 key genes in the blue module most correlated with LUAD
-
Has reproduction · 83
Integrative transcriptomic and machine learning framework reveals candidate genes and potential mechanisms of aflatoxin B1 exposure in breast cancer.
PMID 41688730 · PMC12982753 · Scientific reports · 2026 · 7 claims · 8 setups
Twenty-two genes lie at the intersection of AFB1-predicted targets and breast cancer-associated co-expression modules/DEGs
-
Has reproduction · 67
Research and experimental verification on the mechanisms of cellular senescence in triple-negative breast cancer.
PMID 38435998 · PMC10909353 · PeerJ · 2024 · 8 claims · 8 setups
TNBC can be classified into three molecular subtypes (clusters 1, 2, 3) based on cellular senescence-related pathways, with distinct prognoses (cluster 1 best, then 2, then 3).