Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
CLEAN: CLustering Enrichment ANalysis.
PMID 19640299 · PMC2734555 · BMC bioinformatics · 2009 · 8 claims · 4 setups
The gene-specific CLEAN score improves reproducibility of cluster analysis conclusions across independent datasets compared to the traditional cluster-wide score (cwCLEAN).
-
Has reproduction · 91
A reference profile-free deconvolution method to infer cancer cell-intrinsic subtypes and tumor-type-specific stromal profiles.
PMID 32111252 · PMC7049190 · Genome medicine · 2020 · 8 claims · 8 setups
DeClust is a reference profile-free deconvolution method that simultaneously deconvolves bulk tumor expression into cancer, immune, and stromal compartments and clusters samples into cancer cell-intrinsic molecular subtypes, outputting subtype-specific reference profiles for the cohort rather than for individuals.
-
Has reproduction · 51
SGCP: a spectral self-learning method for clustering genes in co-expression networks.
PMID 38956463 · PMC11221046 · BMC bioinformatics · 2024 · 7 claims · 4 setups
SGCP, a spectral self-learning method, yields gene co-expression modules with higher GO enrichment than WGCNA, CoExpNets, and CEMiTool across 12 real gene expression datasets.
-
Has reproduction · 87
CoINcIDE: A framework for discovery of patient subtypes across multiple datasets.
PMID 26961683 · PMC4784276 · Genome medicine · 2016 · 8 claims · 6 setups
CoINcIDE is a methodological framework that discovers replicable patient subtypes (meta-clusters) across multiple datasets by finding consensus across dataset-specific clusterings, requiring no between-dataset transformations.
-
Has reproduction · 71
Comprehensive analysis of a novel RNA modifications-related model in the prognostic characterization, immune landscape and drug therapy of bladder cancer.
PMID 37124622 · PMC10131083 · Frontiers in genetics · 2023 · 7 claims · 8 setups
Two distinct RNA modification patterns exist among BCa samples with radically different clinical outcomes and biological characteristics.
-
Has reproduction · 71
Parsimonious Gene Correlation Network Analysis (PGCNA): a tool to define modular gene co-expression for refined molecular stratification in cancer.
PMID 30993001 · PMC6459838 · NPJ systems biology and applications · 2019 · 8 claims · 7 setups
Retaining only the top ~3 most correlated edges per gene (EPG3) combined with FastUnfold clustering (termed PGCNA) produces gene co-expression modules with significantly better separation and enrichment of known biology than using all edges or other clustering methods.
-
Has reproduction · 50
BiRNA-BERT allows efficient RNA language modeling with adaptive tokenization.
PMID 41266599 · PMC12635123 · Communications biology · 2025 · 8 claims · 8 setups
BiRNA-BERT uses adaptive dual-tokenization that dynamically selects nucleotide-level (NUC) or byte-pair encoding (BPE) tokens based on input sequence length
-
Has reproduction · 85
ScLRTC: imputation for single-cell RNA-seq data via low-rank tensor completion.
PMID 34844559 · PMC8628418 · BMC genomics · 2021 · 8 claims · 8 setups
scLRTC imputes dropout entries closest to the original expression values on simulated datasets, outperforming other state-of-the-art methods by SSE and PCC.
-
Full-text index only
Iterative class discovery and feature selection using Minimal Spanning Trees.
PMID 15355552 · PMC520744 · BMC bioinformatics · 2004 · 7 claims · 5 setups
Iterating between MST-based clustering and t-statistic feature selection removes noise genes step-wise while sharpening the sample clustering
-
Full-text index only
Multi-organ expression profiling uncovers a gene module in coronary artery disease involving transendothelial migration of leukocytes and LIM domain binding 2: the Stockholm Atherosclerosis Gene Expression (STAGE) study.
PMID 19997623 · PMC2780352 · PLoS genetics · 2009 · 8 claims · 6 setups
Functionally associated gene modules, not individual genes, underlie CAD development and can be identified via multi-organ expression clustering
-
Has reproduction · 74
Evaluation of classification and forecasting methods on time series gene expression data.
PMID 33156855 · PMC7647064 · PloS one · 2020 · 6 claims · 3 setups
Deep learning based methods generally outperform traditional approaches for time series gene expression classification
-
Has reproduction · 80
PanglaoDB: a web server for exploration of mouse and human single-cell RNA sequencing data.
PMID 30951143 · PMC6450036 · Database : the journal of biological databases and curation · 2019 · 7 claims · 7 setups
PanglaoDB is a web server providing pre-processed and pre-computed analyses of >1054 single-cell experiments (>4 million cells) from mouse and human across many tissues and platforms.
-
Has reproduction · 70
Bulk and single-cell characterisation of the immune heterogeneity of atherosclerosis identifies novel targets for immunotherapy.
PMID 36855107 · PMC9974063 · BMC biology · 2023 · 8 claims · 8 setups
Integration of scRNA-seq datasets from human atherosclerosis samples identifies 28 distinct immune cell subpopulations with heterogeneity in tissue preference, genetics, function, immune dynamics, transcriptional regulators, metabolism, and cell communication.
-
Has reproduction · 63
Creation of a Single Cell RNASeq Meta-Atlas to Define Human Liver Immune Homeostasis.
PMID 34335581 · PMC8322955 · Frontiers in immunology · 2021 · 7 claims · 7 setups
Independent human liver immune scRNA-seq datasets can be combined into an integrated meta-atlas in which all datasets co-cluster, despite differing cell-type proportions between studies.
-
Has reproduction · 67
Generative and integrative modeling for transcriptomics with formalin fixed paraffin embedded material.
PMID 41029822 · PMC12486589 · Journal of translational medicine · 2025 · 8 claims · 5 setups
fRNA-seq transcript counts are best fit by the negative binomial distribution, with little evidence supporting zero-inflated extensions
-
Has reproduction · 24
MiGPC: a comprehensive catalog of enzybiotics from environmental metagenomes.
PMID 41888223 · PMC13172421 · Scientific reports · 2026 · 8 claims · 8 setups
MiGPC is the first genome-resolved metagenomic gene and protein catalog specifically targeted to enzybiotics
-
Full-text index only
The use of edge-betweenness clustering to investigate biological function in protein interaction networks.
PMID 15740614 · PMC555937 · BMC bioinformatics · 2005 · 8 claims · 7 setups
Edge-Betweenness clustering separates protein interaction graphs into subgraphs whose GO term distributions show significant correlations, revealing biologically meaningful functional modules.
-
Full-text index only
Unravelling the hidden heterogeneities of diffuse large B-cell lymphoma based on coupled two-way clustering.
PMID 17888167 · PMC2082044 · BMC genomics · 2007 · 8 claims · 6 setups
A proposed coupled two-way clustering (CTWC/SPC) method combined with a GO-based functional concept consistency score can identify compact, robust gene subsets that define clinically meaningful DLBCL subtypes
-
Full-text index only
Bayesian survival analysis in genetic association studies.
PMID 18617538 · PMC2530885 · Bioinformatics (Oxford, England) · 2008 · 7 claims · 5 setups
A novel Bayesian method (BETA-Surv) extends prior case-control haplotype-clustering work to censored survival outcomes by clustering haplotypes via gene tree/perfect phylogeny topology and relative mutation age.
-
Has reproduction · 89
Spatial information matters: are traditional imputation methods effective for spatial transcriptomics data?
PMID 41627342 · PMC12862982 · Briefings in bioinformatics · 2026 · 7 claims · 3 setups
No single existing SOTA imputation method consistently performs well across newer SRT platforms/datasets