Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
Automatic discovery of cross-family sequence features associated with protein function.
PMID 16409628 · PMC1395344 · BMC bioinformatics · 2006 · 8 claims · 6 setups
A self-supervised data mining approach can find relationships between sequence features and functional annotations without preconceived functional categories.
-
Full-text index only
Prediction of candidate primary immunodeficiency disease genes using a support vector machine learning approach.
PMID 19801557 · PMC2780952 · DNA research : an international journal for rapid publication of reports on genes and genomes · 2009 · 6 claims · 3 setups
An SVM trained on 69 binary features of known PID genes can accurately classify PID vs non-PID genes and predict novel candidate PID genes
-
Full-text index only
A cellular epigenetic classification system for glioblastoma.
PMID 41499453 · PMC13128495 · Neuro-oncology · 2026 · 8 claims · 8 setups
ITHresolveGBM, a hierarchical two-step NMF method, deconvolutes bulk GBM DNA methylation profiles into three non-malignant (immune, glial, neuronal) and three malignant components
-
Full-text index only
Non-EST-based prediction of novel alternatively spliced cassette exons with cell signaling function in Caenorhabditis elegans and human.
PMID 17452356 · PMC1904267 · Nucleic acids research · 2007 · 8 claims · 7 setups
PASE (Prediction of Alternative Signaling Exons) is a computational algorithm combining Markov splice-site models, a Bayesian classifier, species conservation, and Scansite motif scoring to identify novel alternative cassette exons involved in cell signaling.
-
Full-text index only
Machine-learning approaches for classifying haplogroup from Y chromosome STR data.
PMID 18551166 · PMC2396484 · PLoS computational biology · 2008 · 8 claims · 5 setups
Y-STR allelic variability is partitioned more by differences among haplogroups than by differences among populations, suggesting Y-STRs carry haplogroup information
-
Full-text index only
Prediction of myeloid malignant cells in Fanconi anemia using machine learning.
PMID 41557613 · PMC12818649 · PloS one · 2026 · 6 claims · 7 setups
A DNN classifier trained on AML scRNA-seq data accurately predicts AML-like transcriptional profiles at single-cell resolution
-
Full-text index only
Ab initio identification of human microRNAs based on structure motifs.
PMID 18088431 · PMC2238772 · BMC bioinformatics · 2007 · 8 claims · 7 setups
MiRPred predicts miRNA precursors ab initio using only predicted secondary structure motifs, ignoring nucleotide sequence
-
Full-text index only
MiPred: classification of real and pseudo microRNA precursors using random forest prediction model with combined features.
PMID 17553836 · PMC1933124 · Nucleic acids research · 2007 · 8 claims · 8 setups
A hybrid feature combining local contiguous triplet structure-sequence composition, MFE of the secondary structure, and P-value of a randomization test improves classification of real vs pseudo pre-miRNAs
-
Full-text index only
CanPredict: a computational tool for predicting cancer-associated missense mutations.
PMID 17537827 · PMC1933186 · Nucleic acids research · 2007 · 8 claims · 7 setups
CanPredict is a web application providing public access to a random forest classifier that combines SIFT, LogR.E-value, and GOSS scores to predict whether a missense mutation is cancer-associated
-
Full-text index only
DeCAF defines clinical fibroblast subtypes and multidimensional tumor-stroma crosstalk shaping prognosis and immunotherapy response.
PMID 41707654 · PMC12923980 · Cell reports. Medicine · 2026 · 8 claims · 8 setups
DeCAF is a single-sample kTSP classifier using 9 TSP gene pairs that predicts proCAF vs restCAF subtypes from bulk expression data
-
Full-text index only
Predicting positive p53 cancer rescue regions using Most Informative Positive (MIP) active learning.
PMID 19756158 · PMC2742196 · PLoS computational biology · 2009 · 8 claims · 4 setups
MIP active learning is a novel active learning method that preferentially seeks informative Positive (functionally active) examples rather than only maximizing classifier accuracy.
-
Full-text index only
FLYNC: a machine-learning-driven framework for discovering long noncoding RNAs in Drosophila melanogaster.
PMID 41551930 · PMC12805895 · NAR genomics and bioinformatics · 2026 · 7 claims · 8 setups
FLYNC, an explainable boosting machine (EBM) model, accurately predicts the probability that a newly identified RNA transcript in D. melanogaster is a lncRNA
-
Full-text index only
PA-GOSUB: a searchable database of model organism protein sequences with their predicted Gene Ontology molecular function and subcellular localization.
PMID 15608166 · PMC540074 · Nucleic acids research · 2005 · 7 claims · 4 setups
PA-GOSUB significantly extends the coverage of GO molecular function and subcellular localization annotations for 10 model organism proteomes compared with existing databases (GOA, Swiss-Prot).
-
Has reproduction · 89
miRge 2.0 for comprehensive analysis of microRNA sequencing data.
PMID 30153801 · PMC6112139 · BMC bioinformatics · 2018 · 8 claims · 6 setups
An SVM-based novel miRNA detection model achieves an average MCC of 0.939 across 32 human cell datasets and outperforms miRDeep2 and miRAnalyzer on phylogenetic conservation of predicted miRNAs
-
Full-text index only
A non-parametric meta-analysis approach for combining independent microarray datasets: application using two microarray datasets pertaining to chronic allograft nephropathy.
PMID 18302764 · PMC2276496 · BMC genomics · 2008 · 8 claims · 6 setups
A novel non-parametric meta-analysis approach for combining independent microarray datasets is presented, requiring no distributional assumptions and being logically intuitive.
-
Has reproduction · 83
Hierarchical classification-based pan-cancer methylation analysis to classify primary cancer.
PMID 38066424 · PMC10709847 · BMC bioinformatics · 2023 · 8 claims · 8 setups
CHCT, a two-tier hierarchical classification tool built from methylation data, accurately classifies primary cancer type across 30 cancer types.
-
Full-text index only
Discovering cancer genes by integrating network and functional properties.
PMID 19765316 · PMC2758898 · BMC medical genomics · 2009 · 8 claims · 6 setups
Cancer genes have distinct PPI network topology (higher connectivity, higher clustering coefficient, shorter path length to known cancer genes) compared to non-cancer genes
-
Full-text index only
The evolution of gene regulation in mammalian cerebellum development.
PMID 41610256 · PMC7618896 · Science (New York, N.Y.) · 2026 · 8 claims · 8 setups
Combined single-nucleus RNA-seq and ATAC-seq atlases of cerebellum development were generated/integrated across six mammals (human, bonobo, macaque, marmoset, mouse, opossum)
-
Full-text index only
Integrating complex genomic datasets and tumour cell sensitivity profiles to address a 'simple' question: which patients should get this drug?
PMID 20003409 · PMC2799438 · BMC medicine · 2009 · 8 claims · 5 setups
A panel of 48 genomically characterized breast cancer cell lines can model patient tumour heterogeneity to identify biomarkers predicting response to PG-11047
-
Full-text index only
Detection of alternative splicing: deep sequencing or deep learning?
PMID 41520225 · PMC12790623 · Briefings in bioinformatics · 2026 · 8 claims · 8 setups
Sequence-based deep learning tools (AlphaGenome, SpliceAI, DeepSplice) show potential for initial hypothesis development and as additional filters in standard RNA-seq pipelines, especially when sequencing depth is limited.