Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
A comprehensive sensitivity analysis of microarray breast cancer classification under feature variability.
PMID 19941644 · PMC2789744 · BMC bioinformatics · 2009 · 7 claims · 4 setups
Feature variability strongly influences breast cancer signature composition even when array platform and patient stratification are identical.
-
Has reproduction · 69
Automatic discovery of 100-miRNA signature for cancer classification using ensemble feature selection.
PMID 31533612 · PMC6751684 · BMC bioinformatics · 2019 · 7 claims · 8 setups
An ensemble feature selection method based on classifier consensus identifies a robust 100-miRNA signature from TCGA data.
-
Has reproduction · 92
Prognostic biomarker discovery in pancreatic cancer through hybrid ensemble feature selection and multi-omics data.
PMID 41957754 · PMC13188360 · BioData mining · 2026 · 7 claims · 3 setups
The hEFS framework integrates data subsampling with multiple prognostic models (embedded and wrapper-based), aggregates feature rankings via a voting-theory-inspired approach, and selects the optimal feature subset via Pareto front optimization, eliminating user-defined thresholds.
-
Full-text index only
Optimality driven nearest centroid classification from genomic data.
PMID 17912341 · PMC1991588 · PloS one · 2007 · 7 claims · 5 setups
A theoretical result determines the subset of features of a given size that minimizes the misclassification rate for a nearest-centroid (LDA) classifier, based on equation (4).
-
Has reproduction · 83
Analyzing biomarker discovery: Estimating the reproducibility of biomarker sets.
PMID 35901020 · PMC9333302 · PloS one · 2022 · 7 claims · 3 setups
A Reproducibility Score, RS(D,BD), defined as the average Jaccard overlap between biomarker sets found by the same discovery process on comparable datasets from the same distribution, quantifies biomarker reproducibility on a 0-1 scale
-
Has reproduction · 100
Betacoronavirus-specific alternate splicing.
PMID 35074468 · PMC8782732 · Genomics · 2022 · 8 claims · 8 setups
Genes showing differential alternative splicing in SARS-CoV-2 have a similar functional profile to those in SARS-CoV and MERS, affecting a diverse set of genes and biological functions related to virus biology.
-
Has reproduction · 67
binny: an automated binning algorithm to recover high-quality genomes from complex metagenomic datasets.
PMID 36239393 · PMC9677464 · Briefings in bioinformatics · 2022 · 8 claims · 8 setups
binny outperforms or is highly competitive with commonly used and state-of-the-art binning methods (MetaBAT2, MaxBin2, CONCOCT, VAMB, SemiBin, MetaDecoder)
-
Full-text index only
Classification of real and pseudo microRNA precursors using local structure-sequence features and support vector machine.
PMID 16381612 · PMC1360673 · BMC bioinformatics · 2005 · 7 claims · 7 setups
A 32-dimensional triplet structure-sequence feature vector combined with SVM (triplet-SVM) can distinguish real human pre-miRNAs from pseudo pre-miRNA hairpins with ~90% accuracy.
-
Has reproduction · 97
CellFishing.jl: an ultrafast and scalable cell search method for single-cell RNA sequencing.
PMID 30744683 · PMC6371477 · Genome biology · 2019 · 8 claims · 8 setups
CellFishing.jl searches prebuilt databases for cells with similar expression patterns with high accuracy and throughput using locality-sensitive hashing.
-
Has reproduction · 79
Interpretable prediction models for widespread m6A RNA modification across cell lines and tissues.
PMID 37995291 · PMC10697738 · Bioinformatics (Oxford, England) · 2023 · 7 claims · 6 setups
CLSM6A, a CNN-based model set, predicts single-nucleotide-resolution m6A RNA modification sites across eight cell lines and three tissues in H. sapiens
-
Has reproduction · 44
An OMICs-based meta-analysis to support infection state stratification.
PMID 33560295 · PMC8388022 · Bioinformatics (Oxford, England) · 2021 · 7 claims · 6 setups
Multi-class machine learning models built from cross-platform microarray meta-analysis can distinguish bacterial, viral and no-infection states with high accuracy (best model: 93% bacterial, 89% viral correct).
-
Has reproduction · 50
The molecular landscape of sepsis severity in infants: enhanced coagulation, innate immunity, and T cell repression.
PMID 38817614 · PMC11137207 · Frontiers in immunology · 2024 · 7 claims · 7 setups
Most published adult/other-cohort sepsis gene signatures have limited utility for infant sepsis; only 2 of 7 achieved >80% accuracy in infants
-
Has reproduction · 78
Machine learning and free energy clustering reveal PAH protein binding linked to AD risk.
PMID 41953002 · PMC13053772 · iScience · 2026 · 7 claims · 8 setups
An integrated framework of bioinformatics, machine learning, and ΔG clustering can prioritize PAHs for AD-associated neurotoxicity.
-
Has reproduction · 72
Prediction of prognostic signatures in triple-negative breast cancer based on the differential expression analysis via NanoString nCounter immune panel.
PMID 33138797 · PMC7607642 · BMC cancer · 2020 · 8 claims · 7 setups
edgeR identifies 9 DEGs associated with pCR and 13 DEGs associated with relapse from 579 immune genes in a small TNBC sample set (n=55)
-
Has reproduction · 80
Specific signature biomarkers highlight the potential mechanisms of circulating neutrophils in aneurysmal subarachnoid hemorrhage.
PMID 36438795 · PMC9685413 · Frontiers in pharmacology · 2022 · 7 claims · 8 setups
Six genes (CST7, HSP90AB1, PADI4, PLBD1, RAB32, SLAMF6) are signature diagnostic biomarkers for aSAH identified by LASSO and SVM-RFE.
-
Has reproduction · 66
Integrative bioinformatics and artificial intelligence analyses of transcriptomics data identified genes associated with major depressive disorders including NRG1.
PMID 37583471 · PMC10423927 · Neurobiology of stress · 2023 · 7 claims · 5 setups
Differentially expressed genes in MDD patients are enriched in immune response, inflammatory response, neurodegeneration, and cerebellar atrophy pathways.
-
Has reproduction · 62
scATD: a high-throughput and interpretable framework for single-cell cancer drug resistance prediction and biomarker identification.
PMID 40501071 · PMC12159290 · Briefings in bioinformatics · 2025 · 8 claims · 6 setups
scATD enables high-throughput single-cell drug sensitivity prediction for new patients without model parameter retraining via bidirectional Bi-AdaIN style transfer
-
Full-text index only
Sequence variation in G-protein-coupled receptors: analysis of single nucleotide polymorphisms.
PMID 15784611 · PMC1069129 · Nucleic acids research · 2005 · 7 claims · 8 setups
Position-specific phylogenetic features describing evolutionary conservation at a site (e.g. SIFT score, normalized site entropy, residue frequency change) are the best individual discriminators of disease-causing versus neutral GPCR mutations.
-
Has reproduction · 55
Identification and verification of diagnostic biomarkers in recurrent pregnancy loss via machine learning algorithm and WGCNA.
PMID 37691920 · PMC10485775 · Frontiers in immunology · 2023 · 8 claims · 8 setups
352 DEGs (198 up-regulated, 154 down-regulated) were identified between RPL and control endometrial samples
-
Has reproduction · 85
Predicting the pathogenicity of missense variants using features derived from AlphaFold2.
PMID 37084271 · PMC10203375 · Bioinformatics (Oxford, England) · 2023 · 6 claims · 8 setups
AlphaFold2-derived structural features (solvent accessibility, amino acid network features, physicochemical environment, pLDDT) can be used to train a random forest classifier (AlphScore) that distinguishes proxy-benign from proxy-pathogenic missense variants.