Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
In silico analysis of missense substitutions using sequence-alignment based methods.
PMID 18951440 · PMC3431198 · Human mutation · 2008 · 8 claims · 7 setups
Carefully validated PMSA-based computational algorithms can achieve predictive values of ~75-95% for classifying missense substitutions as pathogenic or neutral.
-
Full-text index only
Distribution and effects of nonsense polymorphisms in human genes.
PMID 18852891 · PMC2561068 · PloS one · 2008 · 8 claims · 8 setups
Nonsense SNPs occur at a lower density than nonsynonymous SNPs, indicating stronger purifying selection against premature stop codons than amino acid changes.
-
Full-text index only
Pol II promoter prediction using characteristic 4-mer motifs: a machine learning approach.
PMID 18834544 · PMC2575220 · BMC bioinformatics · 2008 · 8 claims · 8 setups
128 discriminating 4-mer motifs combined with an SVM (RBF kernel, LIBSVM) can distinguish promoter from non-promoter DNA sequences
-
Full-text index only
Prioritization of candidate cancer genes--an aid to oncogenomic studies.
PMID 18710882 · PMC2566894 · Nucleic acids research · 2008 · 8 claims · 8 setups
Computational classifiers using combinations of protein conservation, gene structure, protein domains, protein interactions, and regulatory data can distinguish known cancer genes (CD/CR) from unlabelled human genes
-
Full-text index only
Estimation of relevant variables on high-dimensional biological patterns using iterated weighted kernel functions.
PMID 18509521 · PMC2396875 · PloS one · 2008 · 7 claims · 6 setups
wKIERA combines a weighted-kernel discriminant (kernel perceptron) with an iterative stochastic probability estimation-of-distribution algorithm to estimate a relevance distribution over variables
-
Full-text index only
MALDI profiling of human lung cancer subtypes.
PMID 19890392 · PMC2767501 · PloS one · 2009 · 8 claims · 8 setups
PIMAC/MALDI-TOF peptide profiles combined with classification models can distinguish normal lung from tumor and differentiate NSCLC histological subtypes
-
Full-text index only
Swarm intelligence based wavelet coefficient feature selection for mass spectral classification: an application to proteomics data.
PMID 19733729 · PMC2748225 · Analytica chimica acta · 2009 · 8 claims · 4 setups
ACA-based wavelet coefficient feature selection can achieve up to 100% classification accuracy on training, validating, and independent testing sets using only 5 selected features.
-
Full-text index only
Prediction of candidate primary immunodeficiency disease genes using a support vector machine learning approach.
PMID 19801557 · PMC2780952 · DNA research : an international journal for rapid publication of reports on genes and genomes · 2009 · 6 claims · 3 setups
An SVM trained on 69 binary features of known PID genes can accurately classify PID vs non-PID genes and predict novel candidate PID genes
-
Full-text index only
Supervised learning-based tagSNP selection for genome-wide disease classifications.
PMID 18366619 · PMC2386071 · BMC genomics · 2008 · 7 claims · 2 setups
SRFA (Supervised Recursive Feature Addition) is a novel feature selection method combining supervised learning and statistical redundancy measures for SNP selection
-
Has reproduction · 69
Automatic discovery of 100-miRNA signature for cancer classification using ensemble feature selection.
PMID 31533612 · PMC6751684 · BMC bioinformatics · 2019 · 7 claims · 8 setups
An ensemble feature selection method based on classifier consensus identifies a robust 100-miRNA signature from TCGA data.
-
Has reproduction · 71
IRSN-23 gene diagnosis enhances breast cancer subtype classification and predicts response to neoadjuvant chemotherapy: new validation analyses.
PMID 40128415 · PMC11993443 · Breast cancer (Tokyo, Japan) · 2025 · 8 claims · 8 setups
IRSN-23 Gp-R patients have significantly higher pCR rates than Gp-NR patients without anti-HER2 therapy, across the OUH cohort and multiple independent public datasets
-
Has reproduction · 85
What defines a photosynthetic microbial mat in western Antarctica?
PMID 40043057 · PMC11882083 · PloS one · 2025 · 6 claims · 8 setups
Taxonomic composition of Antarctic microbial mat communities is characterized by similar bacterial groups across regions
-
Full-text index only
Serum diagnosis of diffuse large B-cell lymphomas and further identification of response to therapy using SELDI-TOF-MS and tree analysis patterning.
PMID 18163913 · PMC2242801 · BMC cancer · 2007 · 8 claims · 8 setups
SELDI-TOF-MS serum proteomic patterns analyzed by decision tree classification (Biomarker Pattern Software) can discriminate DLBCL patients from healthy controls with high sensitivity and specificity.
-
Full-text index only
The protein-phosphatome of the human malaria parasite Plasmodium falciparum.
PMID 18793411 · PMC2559854 · BMC genomics · 2008 · 8 claims · 8 setups
P. falciparum possesses 27 putative protein phosphatase sequences across the four major PP families (PPP, PPM, PTP, NIF), plus 7 additional sequences predicted to dephosphorylate non-protein substrates, totaling 34.
-
Full-text index only
A comprehensive sensitivity analysis of microarray breast cancer classification under feature variability.
PMID 19941644 · PMC2789744 · BMC bioinformatics · 2009 · 7 claims · 4 setups
Feature variability strongly influences breast cancer signature composition even when array platform and patient stratification are identical.
-
Full-text index only
Bcipep: a database of B-cell epitopes.
PMID 15921533 · PMC1173103 · BMC genomics · 2005 · 8 claims · 2 setups
Bcipep is a comprehensive database of experimentally determined linear B-cell epitopes compiled from literature and other public databases
-
Has reproduction · 50
Wx: a neural network-based feature selection algorithm for transcriptomic data.
PMID 31324856 · PMC6642261 · Scientific reports · 2019 · 8 claims · 8 setups
The Wx algorithm ranks genes by a discriminative index (DI) score representing classification power for distinguishing given groups, enabling intuitive selection of optimal biomarker genes.
-
Has reproduction · 68
Cell-type annotation with accurate unseen cell-type identification using multiple references.
PMID 37379341 · PMC10335708 · PLoS computational biology · 2023 · 8 claims · 4 setups
mtANN integrates multiple reference datasets and eight gene selection methods via ensemble learning (multiple deep classification models + majority voting) to improve cell-type annotation accuracy
-
Has reproduction · 59
Application of Machine Learning in Predicting Hepatic Metastasis or Primary Site in Gastroenteropancreatic Neuroendocrine Tumors.
PMID 37887568 · PMC10605255 · Current oncology (Toronto, Ont.) · 2023 · 8 claims · 7 setups
Multi-gene random forest models classify primary tumor vs. liver metastasis samples with 100% accuracy in training/test cohorts and >90% accuracy in an independent validation cohort
-
Has reproduction · 71
Qualitative Transcriptional Signature for the Pathological Diagnosis of Pancreatic Cancer.
PMID 33173782 · PMC7538791 · Frontiers in molecular biosciences · 2020 · 8 claims · 7 setups
A 12-gene-pair (17-gene) REO-based signature discriminates PC and cancer-adjacent normal tissue from non-tumor (healthy/pancreatitis) pancreatic tissue.