Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Has reproduction · 84
An integrated in silico-in vitro approach for identifying therapeutic targets against osteoarthritis.
PMID 36352408 · PMC9648005 · BMC biology · 2022 · 7 claims · 5 setups
A signal transduction/gene regulatory network model of the articular chondrocyte was built combining knowledge-based curation and data-driven (machine learning) network inference
-
Has reproduction
Using random walks to identify cancer-associated modules in expression data.
PMID 24128261 · PMC4015830 · BioData mining · 2013 · 8 claims · 8 setups
Walktrap-GM, a random-walk community detection algorithm adapted with stopping criteria (maximum modularity, maximum size, maximum module score), identifies modules significantly enriched with cancer genes in expression-weighted interaction networks.
-
Has reproduction · 90
A Decentralized Kidney Transplant Biopsy Classifier for Transplant Rejection Developed Using Genes of the Banff-Human Organ Transplant Panel.
PMID 35619722 · PMC9128066 · Frontiers in immunology · 2022 · 6 claims · 6 setups
A random forest model trained solely on B-HOT panel genes (B-HOT Model) accurately classifies kidney transplant biopsies as NR, ABMR, or TCMR.
-
Has reproduction · 44
An OMICs-based meta-analysis to support infection state stratification.
PMID 33560295 · PMC8388022 · Bioinformatics (Oxford, England) · 2021 · 7 claims · 6 setups
Multi-class machine learning models built from cross-platform microarray meta-analysis can distinguish bacterial, viral and no-infection states with high accuracy (best model: 93% bacterial, 89% viral correct).
-
Has reproduction
Predicting favorable landing pads for targeted integrations in Chinese hamster ovary cell lines by learning stability characteristics from random transgene integrations.
PMID 33304461 · PMC7710658 · Computational and structural biotechnology journal · 2020 · 7 claims · 6 setups
Expression stability in CHO cell lines is controlled at three levels: choice of integration site, integrity/concatemerization pattern of the transgene, and stress-related cellular processes.
-
Full-text index only
Computational verification of protein-protein interactions by orthologous co-expression.
PMID 15740634 · PMC555590 · BMC bioinformatics · 2005 · 7 claims · 8 setups
Co-expression of orthologous protein pairs across multiple species can verify/predict S. cerevisiae PPIs with better performance than S. cerevisiae co-expression alone.
-
Full-text index only
A global definition of expression context is conserved between orthologs, but does not correlate with sequence conservation.
PMID 16423292 · PMC1382217 · BMC genomics · 2006 · 7 claims · 6 setups
Expression context is largely conserved between orthologs across four eukaryote species.
-
Full-text index only
A simple and robust method for connecting small-molecule drugs using gene-expression signatures.
PMID 18518950 · PMC2464610 · BMC bioinformatics · 2008 · 8 claims · 4 setups
A new method for building reference gene-expression profiles and scoring/testing connections improves on the original Connectivity Map by enabling statistical significance testing of connections.
-
Has reproduction · 96
Scalable Prediction of Acute Myeloid Leukemia Using High-Dimensional Machine Learning and Blood Transcriptomics.
PMID 31918046 · PMC6992905 · iScience · 2020 · 8 claims · 8 setups
Data-driven, high-dimensional ML approaches that learn multivariate signatures directly from genome-wide transcriptomic data (no prior gene selection) yield accurate and robust AML classifiers.
-
Has reproduction · 69
Automatic discovery of 100-miRNA signature for cancer classification using ensemble feature selection.
PMID 31533612 · PMC6751684 · BMC bioinformatics · 2019 · 7 claims · 8 setups
An ensemble feature selection method based on classifier consensus identifies a robust 100-miRNA signature from TCGA data.
-
Has reproduction · 59
Application of Machine Learning in Predicting Hepatic Metastasis or Primary Site in Gastroenteropancreatic Neuroendocrine Tumors.
PMID 37887568 · PMC10605255 · Current oncology (Toronto, Ont.) · 2023 · 8 claims · 7 setups
Multi-gene random forest models classify primary tumor vs. liver metastasis samples with 100% accuracy in training/test cohorts and >90% accuracy in an independent validation cohort
-
Has reproduction · 78
Enhancing chemotherapy response prediction via matched colorectal tumor-organoid gene expression analysis and network-based biomarker selection.
PMID 39754813 · PMC11754497 · Translational oncology · 2025 · 6 claims · 8 setups
A consensus WGCNA approach combining matched tumor-organoid and independent organoid drug-response expression data identifies gene modules and hub genes predictive of 5-FU chemotherapy response
-
Has reproduction · 68
Machine learning algorithm predicts fibrosis-related blood diagnosis markers of intervertebral disc degeneration.
PMID 37915003 · PMC10619283 · BMC medical genomics · 2023 · 7 claims · 7 setups
CEP120 and SPDL1 are fibrosis-related diagnostic genes for IDD, identified via a random forest model from 29 differentially expressed fibrosis-related genes
-
Has reproduction · 85
A mechanistic model captures the emergence and implications of non-genetic heterogeneity and reversible drug resistance in ER+ breast cancer cells.
PMID 34316714 · PMC8271219 · NAR cancer · 2021 · 7 claims · 8 setups
EMT and tamoxifen-resistance (TamR) regulatory axes can drive one another, enabling non-genetic heterogeneity via six co-existing phenotypes (ES, ER, HS, HR, MS, MR)
-
Has reproduction · 81
SEMdag: Fast learning of Directed Acyclic Graphs via node or layer ordering.
PMID 39775401 · PMC11709272 · PloS one · 2025 · 8 claims · 5 setups
SEMdag() is a two-step order-based algorithm for fast learning of high-dimensional linear SEMs, using knowledge-based (KB) or data-driven bottom-up (BU) node/layer ordering followed by penalized (L1) DAG estimation
-
Has reproduction · 87
CoINcIDE: A framework for discovery of patient subtypes across multiple datasets.
PMID 26961683 · PMC4784276 · Genome medicine · 2016 · 8 claims · 6 setups
CoINcIDE is a methodological framework that discovers replicable patient subtypes (meta-clusters) across multiple datasets by finding consensus across dataset-specific clusterings, requiring no between-dataset transformations.
-
Has reproduction · 72
Prediction of prognostic signatures in triple-negative breast cancer based on the differential expression analysis via NanoString nCounter immune panel.
PMID 33138797 · PMC7607642 · BMC cancer · 2020 · 8 claims · 7 setups
edgeR identifies 9 DEGs associated with pCR and 13 DEGs associated with relapse from 579 immune genes in a small TNBC sample set (n=55)
-
Has reproduction · 50
Exploiting convergent phenotypes to derive a pan-cancer cisplatin response gene expression signature.
PMID 37076665 · PMC10115855 · NPJ precision oncology · 2023 · 8 claims · 8 setups
A convergent-phenotype-based seed gene/co-expression method can extract consensus gene expression signatures predictive of response to chemotherapeutic drugs in the GDSC database
-
Has reproduction · 83
A temporal classifier predicts histopathology state and parses acute-chronic phasing in inflammatory bowel disease patients.
PMID 36694043 · PMC9873918 · Communications biology · 2023 · 8 claims · 7 setups
The DSS phenotype-by-time interaction defines parsimonious temporal (dynamic) expression and splicing signatures of acute and chronic colitis distinct from time-specific differential expression.
-
Has reproduction · 44
Dynamic Gene Attention Focus (DyGAF): Enhancing Biomarker Identification Through Dual-Model Attention Networks.
PMID 40160891 · PMC11951896 · Bioinformatics and biology insights · 2025 · 6 claims · 5 setups
DyGAF, a dual-model attention neural network (independent Model A + dependent Model B), identifies and ranks genes by significance for COVID-19 biomarker discovery more effectively than differential expression analysis (DEA) and random forest (RF) feature selection