Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
Supervised learning-based tagSNP selection for genome-wide disease classifications.
PMID 18366619 · PMC2386071 · BMC genomics · 2008 · 7 claims · 2 setups
SRFA (Supervised Recursive Feature Addition) is a novel feature selection method combining supervised learning and statistical redundancy measures for SNP selection
-
Full-text index only
Disease-aging network reveals significant roles of aging genes in connecting genetic diseases.
PMID 19779549 · PMC2739292 · PLoS computational biology · 2009 · 8 claims · 8 setups
Human disease genes are much closer to aging genes in the PPI network than expected by chance
-
Full-text index only
The process chain for peptidomic biomarker discovery.
PMID 16410650 · PMC3850862 · Disease markers · 2006 · 8 claims · 3 setups
Peptidomics (comprehensive analysis of peptides and small proteins <20 kDa) fills a methodological gap left by standard proteomics, which mainly addresses proteins in the ~10-200 kDa range.
-
Full-text index only
SysPIMP: the web-based systematical platform for identifying human disease-related mutated sequences from mass spectrometry.
PMID 19036792 · PMC2686442 · Nucleic acids research · 2009 · 8 claims · 7 setups
SysPIMP is a web-based platform integrating disease mutation databases with X!Tandem and BLAST to identify disease-related mutated proteins from MS results
-
Has reproduction · 81
SEMdag: Fast learning of Directed Acyclic Graphs via node or layer ordering.
PMID 39775401 · PMC11709272 · PloS one · 2025 · 8 claims · 5 setups
SEMdag() is a two-step order-based algorithm for fast learning of high-dimensional linear SEMs, using knowledge-based (KB) or data-driven bottom-up (BU) node/layer ordering followed by penalized (L1) DAG estimation
-
Has reproduction · 22
Machine learning-based DNA microarray analysis for disease detection using the MICRO-AI framework.
PMID 41925147 · PMC13051084 · Science progress · 2026 · 8 claims · 1 setups
MICRO-AI's attention-weighted feature fusion reduces dimensionality by over 99% (from ~20,000 to ~127 genes) without loss of biological significance
-
Full-text index only
Utah's Family High Risk Program: bridging the gap between genomics and public health.
PMID 15888235 · PMC1327718 · Preventing chronic disease · 2005 · 8 claims · 6 setups
Collection of family history through the Family High Risk Program (FHRP) is a cost-effective method for identifying and intervening with high-risk populations for chronic disease
-
Full-text index only
Using structural bioinformatics to investigate the impact of non synonymous SNPs and disease mutations: scope and limitations.
PMID 19758473 · PMC2745591 · BMC bioinformatics · 2009 · 8 claims · 8 setups
None of 39 tested structural properties can be used as a sole classification criterion to separate neutral SNPs from disease mutations.
-
Has reproduction · 81
Assessing personalized molecular portraits underlying endothelial-to-mesenchymal transition within pulmonary arterial hypertension.
PMID 39462326 · PMC11513636 · Molecular medicine (Cambridge, Mass.) · 2024 · 8 claims · 8 setups
scRNA-seq of PAH and control lung tissue identifies nine distinct cell populations with high heterogeneity in composition, function, distribution, and communication
-
Full-text index only
Identification of deleterious non-synonymous single nucleotide polymorphisms using sequence-derived information.
PMID 18588693 · PMC2446391 · BMC bioinformatics · 2008 · 8 claims · 5 setups
A decision tree built on 10 selected sequence-derived features classifies SAPs as Disease or Polymorphism with 82.6% accuracy and 0.607 MCC in cross-validation.
-
Full-text index only
Predicting phenotype and emerging strains among Chlamydia trachomatis infections.
PMID 19788805 · PMC2819883 · Emerging infectious diseases · 2009 · 8 claims · 7 setups
A 7-locus MLST scheme selected from conserved housekeeping genes shared across 4 Chlamydiaceae species (7 genomes) can genotype diverse C. trachomatis reference and clinical isolates.
-
Full-text index only
Exhaustive prediction of disease susceptibility to coding base changes in the human genome.
PMID 18793467 · PMC2537574 · BMC bioinformatics · 2008 · 8 claims · 7 setups
Inter-species conservation is the strongest single predictor of disease-associated coding mutations among the factors tested.
-
Full-text index only
nsSNPAnalyzer: identifying disease-associated nonsynonymous single nucleotide polymorphisms.
PMID 15980516 · PMC1160133 · Nucleic acids research · 2005 · 6 claims · 4 setups
nsSNPAnalyzer is a web server that predicts whether a query nsSNP is disease-associated or functionally neutral using a Random Forest classifier combining structural and evolutionary information
-
Full-text index only
Searching for interpretable rules for disease mutations: a simulated annealing bump hunting strategy.
PMID 16984653 · PMC1618409 · BMC bioinformatics · 2006 · 8 claims · 6 setups
The proposed feature set outperforms existing published feature sets for predicting effects of amino acid substitutions
-
Full-text index only
Biomarkers that discriminate multiple myeloma patients with or without skeletal involvement detected using SELDI-TOF mass spectrometry and statistical and machine learning tools.
PMID 17124346 · PMC3862287 · Disease markers · 2006 · 8 claims · 5 setups
SELDI-TOF MS serum profiling can discriminate MM patients with vs without skeletal (bone lesion) involvement using peak biomarkers
-
Full-text index only
Diagnostic proteomics: serum proteomic patterns for the detection of early stage cancers.
PMID 15258335 · PMC3851082 · Disease markers · 2003 · 8 claims · 8 setups
Proteomic pattern analysis of serum mass spectra, without identifying the underlying proteins, can distinguish cancer patients from healthy controls with high sensitivity and specificity.
-
Full-text index only
Speeding disease gene discovery by sequence based candidate prioritization.
PMID 15766383 · PMC1274252 · BMC bioinformatics · 2005 · 7 claims · 8 setups
Disease genes (OMIM) differ significantly from non-disease genes in sequence-based features including gene/cDNA/protein size, exon number, homolog conservation, secretion signal, 3' UTR length, CpG islands, and distance to nearest gene.
-
Full-text index only
Predicting deleterious nsSNPs: an analysis of sequence and structural attributes.
PMID 16630345 · PMC1489951 · BMC bioinformatics · 2006 · 8 claims · 7 setups
Sequence conservation (PSIC score difference) at the nsSNP position is the single most useful attribute for predicting deleterious vs neutral status.
-
Has reproduction · 62
Metatranscriptomics of the human oral microbiome during health and disease.
PMID 24692635 · PMC3977359 · mBio · 2014 · 8 claims · 8 setups
Disease-associated periodontal communities display conserved community-level metabolic gene expression profiles between patients, whereas the metabolic gene expression of individual species is highly variable between patients.
-
Full-text index only
A scale space approach for unsupervised feature selection in mass spectra classification for ovarian cancer detection.
PMID 19828085 · PMC2762074 · BMC bioinformatics · 2009 · 7 claims · 1 setups
A scale-space based unsupervised feature extraction method combined with SVM classification achieves high accuracy in ovarian cancer detection from serum mass spectra.