Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
MiPred: classification of real and pseudo microRNA precursors using random forest prediction model with combined features.
PMID 17553836 · PMC1933124 · Nucleic acids research · 2007 · 8 claims · 8 setups
A hybrid feature combining local contiguous triplet structure-sequence composition, MFE of the secondary structure, and P-value of a randomization test improves classification of real vs pseudo pre-miRNAs
-
Full-text index only
nsSNPAnalyzer: identifying disease-associated nonsynonymous single nucleotide polymorphisms.
PMID 15980516 · PMC1160133 · Nucleic acids research · 2005 · 6 claims · 4 setups
nsSNPAnalyzer is a web server that predicts whether a query nsSNP is disease-associated or functionally neutral using a Random Forest classifier combining structural and evolutionary information
-
Full-text index only
Searching for interpretable rules for disease mutations: a simulated annealing bump hunting strategy.
PMID 16984653 · PMC1618409 · BMC bioinformatics · 2006 · 8 claims · 6 setups
The proposed feature set outperforms existing published feature sets for predicting effects of amino acid substitutions
-
Has reproduction · 90
A Decentralized Kidney Transplant Biopsy Classifier for Transplant Rejection Developed Using Genes of the Banff-Human Organ Transplant Panel.
PMID 35619722 · PMC9128066 · Frontiers in immunology · 2022 · 6 claims · 6 setups
A random forest model trained solely on B-HOT panel genes (B-HOT Model) accurately classifies kidney transplant biopsies as NR, ABMR, or TCMR.
-
Full-text index only
CanPredict: a computational tool for predicting cancer-associated missense mutations.
PMID 17537827 · PMC1933186 · Nucleic acids research · 2007 · 8 claims · 7 setups
CanPredict is a web application providing public access to a random forest classifier that combines SIFT, LogR.E-value, and GOSS scores to predict whether a missense mutation is cancer-associated
-
Full-text index only
Cancer-specific high-throughput annotation of somatic mutations: computational prediction of driver missense mutations.
PMID 19654296 · PMC2763410 · Cancer research · 2009 · 7 claims · 7 setups
CHASM, a Random Forest-based computational method, was developed to identify and prioritize missense mutations likely to be functional drivers of tumor cell proliferation.
-
Has reproduction · 48
Improved epigenetic age prediction models by combining sex chromosome and autosomal markers.
PMID 40665390 · PMC12261677 · Epigenetics & chromatin · 2025 · 7 claims · 5 setups
Combining sex chromosomal DNAm markers with autosomal age-informative markers can produce a high-accuracy age prediction model competitive with autosomal-only models
-
Full-text index only
Boosting accuracy of automated classification of fluorescence microscope images for location proteomics.
PMID 15207009 · PMC449699 · BMC bioinformatics · 2004 · 8 claims · 8 setups
New classifiers (SVMs, ensembles) and new wavelet-derived (Gabor, Daubechies) features improve recognition of protein subcellular location patterns over the previous neural network approach
-
Full-text index only
Ab initio identification of human microRNAs based on structure motifs.
PMID 18088431 · PMC2238772 · BMC bioinformatics · 2007 · 8 claims · 7 setups
MiRPred predicts miRNA precursors ab initio using only predicted secondary structure motifs, ignoring nucleotide sequence
-
Full-text index only
Prioritization of candidate cancer genes--an aid to oncogenomic studies.
PMID 18710882 · PMC2566894 · Nucleic acids research · 2008 · 8 claims · 8 setups
Computational classifiers using combinations of protein conservation, gene structure, protein domains, protein interactions, and regulatory data can distinguish known cancer genes (CD/CR) from unlabelled human genes
-
Has reproduction · 71
Gene Set Enrichment Analysis Reveals Individual Variability in Host Responses in Tuberculosis Patients.
PMID 34421903 · PMC8375662 · Frontiers in immunology · 2021 · 8 claims · 8 setups
TB patients show substantial individual variability in the intensity of hallmark IFN responses, as well as in complement system, metabolic, and other pathway responses.
-
Has reproduction · 83
Hierarchical classification-based pan-cancer methylation analysis to classify primary cancer.
PMID 38066424 · PMC10709847 · BMC bioinformatics · 2023 · 8 claims · 5 setups
CHCT, a hierarchical classification tool, splits classification of 30 cancer types into ten smaller subproblems using a two-tier architecture to classify primary cancer by methylation profile
-
Full-text index only
Speeding disease gene discovery by sequence based candidate prioritization.
PMID 15766383 · PMC1274252 · BMC bioinformatics · 2005 · 7 claims · 8 setups
Disease genes (OMIM) differ significantly from non-disease genes in sequence-based features including gene/cDNA/protein size, exon number, homolog conservation, secretion signal, 3' UTR length, CpG islands, and distance to nearest gene.
-
Full-text index only
Analysis of protein sequence and interaction data for candidate disease gene prediction.
PMID 17020920 · PMC1636487 · Nucleic acids research · 2006 · 8 claims · 7 setups
Combining CPS and CMP using known disease genes as input achieves sensitivity 0.52 and specificity 0.97, reducing candidate lists 13-fold
-
Full-text index only
Genomic variation in myeloma: design, content, and initial application of the Bank On A Cure SNP Panel to detect associations with progression-free survival.
PMID 18778477 · PMC2553089 · BMC medicine · 2008 · 7 claims · 7 setups
A custom BOAC SNP panel of 3404 SNPs in 983 genes was developed using the Affymetrix GeneChip Targeted Genotyping Platform, focused on non-synonymous coding SNPs and regulatory-region SNPs in candidate genes.
-
Full-text index only
Extraction of human kinase mutations from literature, databases and genotyping studies.
PMID 19758464 · PMC2745582 · BMC bioinformatics · 2009 · 7 claims · 6 setups
A literature mining pipeline combining MutationFinder, false-positive filtering, and SVM-based classification can extract and disambiguate single-point mutation mentions from abstracts and full text
-
Full-text index only
Genome-wide prioritization of disease genes and identification of disease-disease associations from an integrated human functional linkage network.
PMID 19728866 · PMC2768980 · Genome biology · 2009 · 6 claims · 6 setups
Integrating 16 genomic features (32 sub-features) via a naïve Bayes classifier produces a genome-scale FLN of 21,657 human genes and 22,388,609 weighted links that outperforms any individual data source for inferring functional linkages.
-
Full-text index only
MatchMiner: a tool for batch navigation among gene and gene product identifiers.
PMID 12702208 · PMC154578 · Genome biology · 2003 · 8 claims · 3 setups
MatchMiner's LookUp function automates batch translation of an input list of gene identifiers into a matching list of a different identifier type.
-
Full-text index only
Biocomputing enters its adolescence.
PMID 15960815 · PMC1175967 · Genome biology · 2005 · 8 claims · 8 setups
A 'match augmentation' algorithm efficiently matches structural motifs by prioritizing functionally significant residues, enabling function prediction between evolutionarily unrelated proteins
-
Has reproduction · 78
Enhancing chemotherapy response prediction via matched colorectal tumor-organoid gene expression analysis and network-based biomarker selection.
PMID 39754813 · PMC11754497 · Translational oncology · 2025 · 6 claims · 8 setups
A consensus WGCNA approach combining matched tumor-organoid and independent organoid drug-response expression data identifies gene modules and hub genes predictive of 5-FU chemotherapy response