Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
Prediction of catalytic residues using Support Vector Machine with selected protein sequence and structural properties.
PMID 16790052 · PMC1534064 · BMC bioinformatics · 2006 · 8 claims · 7 setups
The Sequential Minimal Optimization (SMO) SVM algorithm was the best-performing classifier among 26 WEKA classifiers for predicting catalytic residues
-
Full-text index only
Genomic data sampling and its effect on classification performance assessment.
PMID 12553886 · PMC149349 · BMC bioinformatics · 2003 · 8 claims · 3 setups
Cross-validation, leave-one-out, and bootstrap are designed to reduce bias and variance in accuracy estimation from small samples.
-
Full-text index only
Supervised learning-based tagSNP selection for genome-wide disease classifications.
PMID 18366619 · PMC2386071 · BMC genomics · 2008 · 7 claims · 2 setups
SRFA (Supervised Recursive Feature Addition) is a novel feature selection method combining supervised learning and statistical redundancy measures for SNP selection
-
Full-text index only
Automatic discovery of cross-family sequence features associated with protein function.
PMID 16409628 · PMC1395344 · BMC bioinformatics · 2006 · 8 claims · 6 setups
A self-supervised data mining approach can find relationships between sequence features and functional annotations without preconceived functional categories.
-
Full-text index only
Information-theoretic identification of predictive SNPs and supervised visualization of genome-wide association studies.
PMID 16899448 · PMC1557808 · Nucleic acids research · 2006 · 7 claims · 4 setups
3D VizStruct (DFT-based radial mapping + KLD as z-axis) can identify SNPs/polymorphic markers that are predictive of underlying biological class distinctions across diverse datasets
-
Full-text index only
Prediction of candidate primary immunodeficiency disease genes using a support vector machine learning approach.
PMID 19801557 · PMC2780952 · DNA research : an international journal for rapid publication of reports on genes and genomes · 2009 · 6 claims · 3 setups
An SVM trained on 69 binary features of known PID genes can accurately classify PID vs non-PID genes and predict novel candidate PID genes
-
Has reproduction · 95
Mouse-Geneformer: A deep learning model for mouse single-cell transcriptome and its cross-species utility.
PMID 40106407 · PMC11964219 · PLoS genetics · 2025 · 7 claims · 6 setups
Mouse-Geneformer, a Transformer Encoder model pre-trained via masked-token self-supervised learning on mouse-Genecorpus-20M, was successfully constructed following the original human Geneformer architecture.
-
Full-text index only
SePaCS--a web-based application for classification of seroreactivity profiles.
PMID 17478503 · PMC1933220 · Nucleic acids research · 2007 · 8 claims · 4 setups
SePaCS is a freely available web-based tool that trains and applies multiple classification methods (4 Naive Bayes variants, SVM with RBF kernel, LDA, DLDA) to seroreactivity profiles and outputs results as a summary table plus a detailed PDF report
-
Has reproduction · 87
CoINcIDE: A framework for discovery of patient subtypes across multiple datasets.
PMID 26961683 · PMC4784276 · Genome medicine · 2016 · 8 claims · 6 setups
CoINcIDE is a methodological framework that discovers replicable patient subtypes (meta-clusters) across multiple datasets by finding consensus across dataset-specific clusterings, requiring no between-dataset transformations.
-
Full-text index only
Identification of diagnostic markers for tuberculosis by proteomic fingerprinting of serum.
PMID 16980117 · PMC7159276 · Lancet (London, England) · 2006 · 8 claims · 5 setups
An SVM classifier trained on serum proteomic profiles discriminated patients with active tuberculosis from controls with clinically overlapping conditions
-
Full-text index only
Zebrafish whole-adult-organism chemogenomics for large-scale predictive and discovery chemical biology.
PMID 18618001 · PMC2442223 · PLoS genetics · 2008 · 8 claims · 6 setups
Zebrafish whole-adult-organism chemogenomics generates robust prediction models that discriminate P(H)AHs from ECs across independent experiments
-
Full-text index only
Expression genomics in breast cancer research: microarrays at the crossroads of biology and medicine.
PMID 17397520 · PMC1868923 · Breast cancer research : BCR · 2007 · 8 claims · 8 setups
Genome-wide expression microarray studies reveal transcriptional networks/signatures that explain breast cancer biological and clinical heterogeneity
-
Has reproduction · 73
Detecting aberrant DNA methylation in Illumina DNA methylation arrays: a toolbox and recommendations for its use.
PMID 37218167 · PMC10208159 · Epigenetics · 2023 · 8 claims · 7 setups
Probe-specific upper and lower thresholds for flagging aberrant DNA methylation can be derived from a reference database of >2,000 normal and tumour-adjacent normal samples spanning 25 tissue types.
-
Full-text index only
A DNA microarray survey of gene expression in normal human tissues.
PMID 15774023 · PMC1088941 · Genome biology · 2005 · 6 claims · 6 setups
Unsupervised hierarchical clustering of gene expression groups normal tissue samples largely according to anatomic location, cellular composition, or physiologic function.
-
Full-text index only
Analysis of a set of missense, frameshift, and in-frame deletion variants of BRCA1.
PMID 18992264 · PMC2682550 · Mutation research · 2009 · 8 claims · 8 setups
A combined functional assay, bioinformatics prediction, and structural modeling approach can classify BRCA1 variants of uncertain significance
-
Full-text index only
Bayesian survival analysis in genetic association studies.
PMID 18617538 · PMC2530885 · Bioinformatics (Oxford, England) · 2008 · 7 claims · 5 setups
A novel Bayesian method (BETA-Surv) extends prior case-control haplotype-clustering work to censored survival outcomes by clustering haplotypes via gene tree/perfect phylogeny topology and relative mutation age.
-
Full-text index only
Integrative genomics analysis of chromosome 5p gain in cervical cancer reveals target over-expressed genes, including Drosha.
PMID 18559093 · PMC2440550 · Molecular cancer · 2008 · 7 claims · 6 setups
Gain of chromosome 5p is the most frequent genomic alteration in invasive cervical cancer
-
Full-text index only
Machine-learning approaches for classifying haplogroup from Y chromosome STR data.
PMID 18551166 · PMC2396484 · PLoS computational biology · 2008 · 8 claims · 5 setups
Y-STR allelic variability is partitioned more by differences among haplogroups than by differences among populations, suggesting Y-STRs carry haplogroup information
-
Full-text index only
Glioblastoma subclasses can be defined by activity among signal transduction pathways and associated genomic alterations.
PMID 19915670 · PMC2771920 · PloS one · 2009 · 8 claims · 6 setups
Proteomic analysis of glioma samples reveals three signaling subclasses of GBM associated with predominant EGFR activation, PDGFR activation, or loss of NF1
-
Full-text index only
Genomic transcriptional profiling identifies a candidate blood biomarker signature for the diagnosis of septicemic melioidosis.
PMID 19903332 · PMC3091321 · Genome biology · 2009 · 6 claims · 5 setups
A candidate 37-transcript diagnostic signature distinguishes septicemic melioidosis from sepsis caused by other organisms with 100% accuracy in the training set and 78%/80% accuracy in two independent validation sets