Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
Supervised learning-based tagSNP selection for genome-wide disease classifications.
PMID 18366619 · PMC2386071 · BMC genomics · 2008 · 7 claims · 2 setups
SRFA (Supervised Recursive Feature Addition) is a novel feature selection method combining supervised learning and statistical redundancy measures for SNP selection
-
Full-text index only
Information-theoretic identification of predictive SNPs and supervised visualization of genome-wide association studies.
PMID 16899448 · PMC1557808 · Nucleic acids research · 2006 · 7 claims · 4 setups
3D VizStruct (DFT-based radial mapping + KLD as z-axis) can identify SNPs/polymorphic markers that are predictive of underlying biological class distinctions across diverse datasets
-
Full-text index only
A comprehensive sensitivity analysis of microarray breast cancer classification under feature variability.
PMID 19941644 · PMC2789744 · BMC bioinformatics · 2009 · 7 claims · 4 setups
Feature variability strongly influences breast cancer signature composition even when array platform and patient stratification are identical.
-
Has reproduction · 87
CoINcIDE: A framework for discovery of patient subtypes across multiple datasets.
PMID 26961683 · PMC4784276 · Genome medicine · 2016 · 8 claims · 6 setups
CoINcIDE is a methodological framework that discovers replicable patient subtypes (meta-clusters) across multiple datasets by finding consensus across dataset-specific clusterings, requiring no between-dataset transformations.
-
Has reproduction · 95
Mouse-Geneformer: A deep learning model for mouse single-cell transcriptome and its cross-species utility.
PMID 40106407 · PMC11964219 · PLoS genetics · 2025 · 7 claims · 6 setups
Mouse-Geneformer, a Transformer Encoder model pre-trained via masked-token self-supervised learning on mouse-Genecorpus-20M, was successfully constructed following the original human Geneformer architecture.
-
Full-text index only
Identification of diagnostic markers for tuberculosis by proteomic fingerprinting of serum.
PMID 16980117 · PMC7159276 · Lancet (London, England) · 2006 · 8 claims · 5 setups
An SVM classifier trained on serum proteomic profiles discriminated patients with active tuberculosis from controls with clinically overlapping conditions
-
Full-text index only
Integrative genomics analysis of chromosome 5p gain in cervical cancer reveals target over-expressed genes, including Drosha.
PMID 18559093 · PMC2440550 · Molecular cancer · 2008 · 7 claims · 6 setups
Gain of chromosome 5p is the most frequent genomic alteration in invasive cervical cancer
-
Full-text index only
MALDI profiling of human lung cancer subtypes.
PMID 19890392 · PMC2767501 · PloS one · 2009 · 8 claims · 8 setups
PIMAC/MALDI-TOF peptide profiles combined with classification models can distinguish normal lung from tumor and differentiate NSCLC histological subtypes
-
Full-text index only
Serum protein profile in systemic-onset juvenile idiopathic arthritis differentiates response versus nonresponse to therapy.
PMID 15987476 · PMC1175022 · Arthritis research & therapy · 2005 · 8 claims · 8 setups
SELDI-TOF MS can differentiate serum protein profiles of active versus well-controlled SJIA
-
Full-text index only
Extraction of human kinase mutations from literature, databases and genotyping studies.
PMID 19758464 · PMC2745582 · BMC bioinformatics · 2009 · 7 claims · 6 setups
A literature mining pipeline combining MutationFinder, false-positive filtering, and SVM-based classification can extract and disambiguate single-point mutation mentions from abstracts and full text
-
Full-text index only
On consensus biomarker selection.
PMID 17570864 · PMC1892093 · BMC bioinformatics · 2007 · 7 claims · 2 setups
Four popular feature ranking criteria (t-statistic, mutual information, peak probability contrasts, random forest variable importance) produce different rankings of the same features on the same dataset.