Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
Protein ranking by semi-supervised network propagation.
PMID 16723003 · PMC1810311 · BMC bioinformatics · 2006 · 8 claims · 5 setups
RankProp, a diffusion-based network propagation algorithm on a PSI-BLAST-derived protein similarity network, significantly outperforms local search methods (BLAST/PSI-BLAST) at detecting remote homologs.
-
Full-text index only
Integrating complex genomic datasets and tumour cell sensitivity profiles to address a 'simple' question: which patients should get this drug?
PMID 20003409 · PMC2799438 · BMC medicine · 2009 · 8 claims · 5 setups
A panel of 48 genomically characterized breast cancer cell lines can model patient tumour heterogeneity to identify biomarkers predicting response to PG-11047
-
Full-text index only
SNP selection for genes of iron metabolism in a study of genetic modifiers of hemochromatosis.
PMID 18366708 · PMC2289803 · BMC medical genetics · 2008 · 7 claims · 6 setups
Illumina validation/design scores above 0.6 are not strongly correlated with actual SNP genotyping performance (Gentrain score)
-
Full-text index only
Targeted next-generation sequencing of a cancer transcriptome enhances detection of sequence variants and novel fusion transcripts.
PMID 19835606 · PMC2784330 · Genome biology · 2009 · 7 claims · 2 setups
Hybrid selection of cDNA dramatically increases the specificity of sequencing reads mapping to targeted cancer-related transcripts.
-
Has reproduction · 76
Bayesian prediction of microbial oxygen requirement.
PMID 26913185 · PMC4743139 · F1000Research · 2013 · 7 claims · 8 setups
A naive Bayesian classifier based on presence/absence of class-associated Pfam-A domains can distinguish three oxygen requirement classes (aerobe, anaerobe, facultative anaerobe) from genome sequence, unlike prior studies that only made pairwise distinctions.
-
Has reproduction · 50
Cost-effectively dissecting the genetic architecture of complex wool traits in rabbits by low-coverage sequencing.
PMID 36401180 · PMC9673297 · Genetics, selection, evolution : GSE · 2022 · 8 claims · 8 setups
BaseVar + STITCH at 1.0X sequencing depth with a sample size >300 achieves the highest genotyping accuracy among tested imputation strategies (genotype concordance >98.8%, genotype accuracy >0.97).
-
Has reproduction · 92
Prognostic biomarker discovery in pancreatic cancer through hybrid ensemble feature selection and multi-omics data.
PMID 41957754 · PMC13188360 · BioData mining · 2026 · 7 claims · 3 setups
The hEFS framework integrates data subsampling with multiple prognostic models (embedded and wrapper-based), aggregates feature rankings via a voting-theory-inspired approach, and selects the optimal feature subset via Pareto front optimization, eliminating user-defined thresholds.
-
Has reproduction
Genomic prediction based on selective linkage disequilibrium pruning of low-coverage whole-genome sequence variants in a pure Duroc population.
PMID 37853325 · PMC10583454 · Genetics, selection, evolution : GSE · 2023 · 8 claims · 6 setups
Selective linkage disequilibrium pruning (SLDP) refines whole-genome SNP sets using GWAS prior information to improve genomic prediction accuracy.
-
Full-text index only
The HIV positive selection mutation database.
PMID 17108357 · PMC1669717 · Nucleic acids research · 2007 · 8 claims · 5 setups
The database provides codon-level Ka/Ks selection pressure maps for HIV protease and the first 381 codons of RT, built from a novel ~50,000-sample clinical dataset.
-
Full-text index only
Prediction of catalytic residues using Support Vector Machine with selected protein sequence and structural properties.
PMID 16790052 · PMC1534064 · BMC bioinformatics · 2006 · 8 claims · 7 setups
The Sequential Minimal Optimization (SMO) SVM algorithm was the best-performing classifier among 26 WEKA classifiers for predicting catalytic residues
-
Has reproduction · 85
Optimisation of the core subset for the APY approximation of genomic relationships.
PMID 36418945 · PMC9682752 · Genetics, selection, evolution : GSE · 2022 · 7 claims · 3 setups
APY approximates the full genomic relationship matrix by splitting genotyped animals into a core subset (fully dependent, direct inverse) and a non-core subset (conditionally independent given core), reducing inversion cost.
-
Has reproduction · 86
Molecular Classification Models for Triple Negative Breast Cancer Subtype Using Machine Learning.
PMID 34575658 · PMC8472680 · Journal of personalized medicine · 2021 · 6 claims · 4 setups
A training gene set of 719 unique upregulated DEGs (subtype-specific) can be used to build ML models that classify TNBC into BLIA, BLIS, MES, and LAR subtypes.
-
Has reproduction · 100
Smart spatial omics (S2-omics) optimizes region of interest selection to capture molecular heterogeneity in diverse tissues.
PMID 41298871 · PMC12662399 · Nature cell biology · 2025 · 7 claims · 6 setups
S2-omics is an end-to-end workflow that automatically selects ROIs from H&E histology images to maximize molecular information content for spatial omics profiling.
-
Has reproduction · 55
Natural clines and human management impact the genetic structure of Algerian honey bee populations.
PMID 38114899 · PMC10729559 · Genetics, selection, evolution : GSE · 2023 · 7 claims · 8 setups
No significant admixture from European subspecies was detected in Algerian honey bees, suggesting large-scale queen imports have not occurred in Algeria.
-
Has reproduction · 90
A Decentralized Kidney Transplant Biopsy Classifier for Transplant Rejection Developed Using Genes of the Banff-Human Organ Transplant Panel.
PMID 35619722 · PMC9128066 · Frontiers in immunology · 2022 · 6 claims · 6 setups
A random forest model trained solely on B-HOT panel genes (B-HOT Model) accurately classifies kidney transplant biopsies as NR, ABMR, or TCMR.
-
Full-text index only
Identification of deleterious non-synonymous single nucleotide polymorphisms using sequence-derived information.
PMID 18588693 · PMC2446391 · BMC bioinformatics · 2008 · 8 claims · 5 setups
A decision tree built on 10 selected sequence-derived features classifies SAPs as Disease or Polymorphism with 82.6% accuracy and 0.607 MCC in cross-validation.
-
Full-text index only
A comparison of classification methods for predicting Chronic Fatigue Syndrome based on genetic data.
PMID 19772600 · PMC2765429 · Journal of translational medicine · 2009 · 7 claims · 3 setups
The naive Bayes model with the wrapper-based feature selection approach performed best among all predictive models tested for distinguishing CFS from controls.
-
Has reproduction · 83
Gene-expression patterns in peripheral blood classify familial breast cancer susceptibility.
PMID 26538066 · PMC4634735 · BMC medical genomics · 2015 · 8 claims · 5 setups
A multigene peripheral-blood gene-expression biomarker accurately classifies which women from high-risk families develop familial breast cancer.
-
Full-text index only
Boosting accuracy of automated classification of fluorescence microscope images for location proteomics.
PMID 15207009 · PMC449699 · BMC bioinformatics · 2004 · 8 claims · 8 setups
New classifiers (SVMs, ensembles) and new wavelet-derived (Gabor, Daubechies) features improve recognition of protein subcellular location patterns over the previous neural network approach
-
Full-text index only
Prioritization of candidate cancer genes--an aid to oncogenomic studies.
PMID 18710882 · PMC2566894 · Nucleic acids research · 2008 · 8 claims · 8 setups
Computational classifiers using combinations of protein conservation, gene structure, protein domains, protein interactions, and regulatory data can distinguish known cancer genes (CD/CR) from unlabelled human genes