Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
Prioritization of candidate cancer genes--an aid to oncogenomic studies.
PMID 18710882 · PMC2566894 · Nucleic acids research · 2008 · 8 claims · 8 setups
Computational classifiers using combinations of protein conservation, gene structure, protein domains, protein interactions, and regulatory data can distinguish known cancer genes (CD/CR) from unlabelled human genes
-
Full-text index only
MiPred: classification of real and pseudo microRNA precursors using random forest prediction model with combined features.
PMID 17553836 · PMC1933124 · Nucleic acids research · 2007 · 8 claims · 8 setups
A hybrid feature combining local contiguous triplet structure-sequence composition, MFE of the secondary structure, and P-value of a randomization test improves classification of real vs pseudo pre-miRNAs
-
Full-text index only
CanPredict: a computational tool for predicting cancer-associated missense mutations.
PMID 17537827 · PMC1933186 · Nucleic acids research · 2007 · 8 claims · 7 setups
CanPredict is a web application providing public access to a random forest classifier that combines SIFT, LogR.E-value, and GOSS scores to predict whether a missense mutation is cancer-associated
-
Has reproduction · 65
A urine extracellular vesicle lncRNA classifier for high-grade prostate cancer and increased risk of progression: A multi-center study.
PMID 37852185 · PMC10591064 · Cell reports. Medicine · 2023 · 8 claims · 8 setups
A 3-lncRNA urine extracellular vesicle classifier (Clnc: AC015987.1, CTD-2589M5.4, RP11-363E6.3) detects high-grade PCa with higher accuracy than PCA3, mpMRI, PCPT-RC 2.0, and ERSPC-RC
-
Full-text index only
Machine-learning approaches for classifying haplogroup from Y chromosome STR data.
PMID 18551166 · PMC2396484 · PLoS computational biology · 2008 · 8 claims · 5 setups
Y-STR allelic variability is partitioned more by differences among haplogroups than by differences among populations, suggesting Y-STRs carry haplogroup information
-
Full-text index only
Prediction of catalytic residues using Support Vector Machine with selected protein sequence and structural properties.
PMID 16790052 · PMC1534064 · BMC bioinformatics · 2006 · 8 claims · 7 setups
The Sequential Minimal Optimization (SMO) SVM algorithm was the best-performing classifier among 26 WEKA classifiers for predicting catalytic residues
-
Full-text index only
DNA methylation biomarkers-based pan-cancer classifier: predictive modeling for cancer classification.
PMID 42152108 · PMC13185202 · Genome medicine · 2026 · 8 claims · 5 setups
Relatively simple ML models (logistic regression) outperform complex algorithms such as deep neural networks for methylation-based cancer classification
-
Full-text index only
DeCAF defines clinical fibroblast subtypes and multidimensional tumor-stroma crosstalk shaping prognosis and immunotherapy response.
PMID 41707654 · PMC12923980 · Cell reports. Medicine · 2026 · 8 claims · 8 setups
DeCAF is a single-sample kTSP classifier using 9 TSP gene pairs that predicts proCAF vs restCAF subtypes from bulk expression data
-
Full-text index only
Prediction of candidate primary immunodeficiency disease genes using a support vector machine learning approach.
PMID 19801557 · PMC2780952 · DNA research : an international journal for rapid publication of reports on genes and genomes · 2009 · 6 claims · 3 setups
An SVM trained on 69 binary features of known PID genes can accurately classify PID vs non-PID genes and predict novel candidate PID genes
-
Full-text index only
Constructing support vector machine ensembles for cancer classification based on proteomic profiling.
PMID 16689692 · PMC5173238 · Genomics, proteomics & bioinformatics · 2005 · 7 claims · 4 setups
CSVME, built by selecting a subset of base SVMs via SVM-RFE ranking and fusing them with a trained upper-layer SVM, achieves better classification performance than an ensemble of all base SVMs.
-
Full-text index only
FLYNC: a machine-learning-driven framework for discovering long noncoding RNAs in Drosophila melanogaster.
PMID 41551930 · PMC12805895 · NAR genomics and bioinformatics · 2026 · 7 claims · 8 setups
FLYNC, an explainable boosting machine (EBM) model, accurately predicts the probability that a newly identified RNA transcript in D. melanogaster is a lncRNA
-
Has reproduction
Fast, accurate, and racially unbiased pan-cancer tumor-only variant calling with tabular machine learning.
PMID 36611079 · PMC9825621 · NPJ precision oncology · 2023 · 8 claims · 8 setups
Tree-based (XGBoost, LightGBM) and deep-learning (TabNet) tabular ML classifiers achieve state-of-the-art somatic vs germline classification in tumor-only WES samples, outperforming PureCN.
-
Full-text index only
Ab initio identification of human microRNAs based on structure motifs.
PMID 18088431 · PMC2238772 · BMC bioinformatics · 2007 · 8 claims · 7 setups
MiRPred predicts miRNA precursors ab initio using only predicted secondary structure motifs, ignoring nucleotide sequence
-
Full-text index only
Speeding disease gene discovery by sequence based candidate prioritization.
PMID 15766383 · PMC1274252 · BMC bioinformatics · 2005 · 7 claims · 8 setups
Disease genes (OMIM) differ significantly from non-disease genes in sequence-based features including gene/cDNA/protein size, exon number, homolog conservation, secretion signal, 3' UTR length, CpG islands, and distance to nearest gene.
-
Full-text index only
Metagenomic study of the oral microbiota by Illumina high-throughput sequencing.
PMID 19796657 · PMC3568755 · Journal of microbiological methods · 2009 · 8 claims · 6 setups
The 16S rRNA V5 hypervariable region, amplified as a short ~82-base segment, provides reliable taxonomic identification of oral bacteria against public databases like HOMD.
-
Full-text index only
Systematic annotation of orphan RNAs reveals blood-accessible molecular barcodes of cancer identity and cancer-emergent oncogenic drivers.
PMID 41579861 · PMC12923976 · Cell reports. Medicine · 2026 · 8 claims · 8 setups
oncRNA binary presence-absence patterns constitute digital molecular barcodes that capture cancer type and subtype identity
-
Has reproduction · 89
miRge 2.0 for comprehensive analysis of microRNA sequencing data.
PMID 30153801 · PMC6112139 · BMC bioinformatics · 2018 · 8 claims · 6 setups
An SVM-based novel miRNA detection model achieves an average MCC of 0.939 across 32 human cell datasets and outperforms miRDeep2 and miRAnalyzer on phylogenetic conservation of predicted miRNAs
-
Has reproduction · 76
Topologically inferring pathway activity toward precise cancer classification via integrating genomic and metabolomic data: prostate cancer as a case.
PMID 26286638 · PMC4541321 · Scientific reports · 2015 · 6 claims · 6 setups
DRW-GM evaluates gene topological importance by directed random walk on a global gene–metabolite pathway graph integrating gene expression and metabolomic profiles to infer pathway activities
-
Has reproduction · 83
Hierarchical classification-based pan-cancer methylation analysis to classify primary cancer.
PMID 38066424 · PMC10709847 · BMC bioinformatics · 2023 · 8 claims · 8 setups
CHCT, a two-tier hierarchical classification tool built from methylation data, accurately classifies primary cancer type across 30 cancer types.
-
Full-text index only
Widespread dysregulation of MiRNAs by MYCN amplification and chromosomal imbalances in neuroblastoma: association of miRNA expression with survival.
PMID 19924232 · PMC2773120 · PloS one · 2009 · 8 claims · 5 setups
37 miRNAs are significantly differentially expressed between MYCN-amplified (MNA) and non-MNA neuroblastoma tumors, suggesting direct or indirect regulation by MYCN