Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
In Silico screening for functional candidates amongst hypothetical proteins.
PMID 19754976 · PMC2758874 · BMC bioinformatics · 2009 · 7 claims · 6 setups
An in silico selection strategy combining subcellular targeting-signal prediction with protein domain identification can enrich for true functional proteins among hypothetical proteins
-
Has reproduction · 76
Bayesian prediction of microbial oxygen requirement.
PMID 26913185 · PMC4743139 · F1000Research · 2013 · 7 claims · 8 setups
A naive Bayesian classifier based on presence/absence of class-associated Pfam-A domains can distinguish three oxygen requirement classes (aerobe, anaerobe, facultative anaerobe) from genome sequence, unlike prior studies that only made pairwise distinctions.
-
Full-text index only
EpiToolKit--a web server for computational immunomics.
PMID 18440979 · PMC2447732 · Nucleic acids research · 2008 · 7 claims · 3 setups
EpiToolKit is a web server integrating five MHC class I and two MHC class II epitope prediction methods in a unified, user-friendly interface.
-
Has reproduction · 44
An OMICs-based meta-analysis to support infection state stratification.
PMID 33560295 · PMC8388022 · Bioinformatics (Oxford, England) · 2021 · 7 claims · 6 setups
Multi-class machine learning models built from cross-platform microarray meta-analysis can distinguish bacterial, viral and no-infection states with high accuracy (best model: 93% bacterial, 89% viral correct).
-
Has reproduction · 90
Streaming Long-Read Sequence Alignments for HLA Predictions Using HLAminer.
PMID 40145684 · PMC11948951 · Current protocols · 2025 · 6 claims · 5 setups
Streaming minimap2 alignment output directly into HLAminer via Unix pipe (-a stream mode) enables HLA class I and II allele prediction without storing bulky SAM alignment files on disk
-
Full-text index only
An SVM-based system for predicting protein subnuclear localizations.
PMID 16336650 · PMC1325059 · BMC bioinformatics · 2005 · 7 claims · 3 setups
New kernels defined on k-peptide vectors mapped by BLOSUM62-based high-scored pair matrices (D1, D2, D3) improve SVM discrimination of protein subnuclear localization compared to conventional k-peptide encodings.
-
Has reproduction
Artificial Intelligence Meets Whole Slide Images: Deep Learning Model Shapes an Immune-Hot Tumor and Guides Precision Therapy in Bladder Cancer.
PMID 36245985 · PMC9553530 · Journal of oncology · 2022 · 6 claims · 8 setups
A deep learning WSI cluster (three-class mini batch K-means on Inception V3 features) is associated with overall survival (P<0.001) and is an independent prognostic predictor (P=0.031) in BLCA.
-
Full-text index only
miRGator: an integrated system for functional annotation of microRNAs.
PMID 17942429 · PMC2238850 · Nucleic acids research · 2008 · 8 claims · 8 setups
miRGator integrates target prediction, functional enrichment analysis (GO/pathway/disease), and expression data (miRNA/mRNA/protein) into one system for functional annotation of miRNAs
-
Full-text index only
GeneMark: web software for gene finding in prokaryotes, eukaryotes and viruses.
PMID 15980510 · PMC1160247 · Nucleic acids research · 2005 · 8 claims · 2 setups
The GeneMark website provides web interfaces to the GeneMark family of ab initio gene-finding programs for prokaryotic, eukaryotic and viral genomic sequences
-
Has reproduction · 56
Comparative Metagenomic Analysis of Biosynthetic Diversity across Sponge Microbiomes Highlights Metabolic Novelty, Conservation, and Diversification.
PMID 35862823 · PMC9426513 · mSystems · 2022 · 8 claims · 5 setups
The vast majority of recovered gene cluster families (GCFs) in sponge microbiomes show no similarity to any characterized BGC, revealing extreme biosynthetic novelty
-
Full-text index only
GeneAlign: a coding exon prediction tool based on phylogenetical comparisons.
PMID 16845010 · PMC1538901 · Nucleic acids research · 2006 · 8 claims · 5 setups
GeneAlign predicts coding exons by using signal detection (GeneSplicer/WMM) combined with CORAL, a heuristic linear-time alignment tool, to align candidate signal-flanked regions against annotated exons of a homologous organism's genes
-
Full-text index only
Zebrafish whole-adult-organism chemogenomics for large-scale predictive and discovery chemical biology.
PMID 18618001 · PMC2442223 · PLoS genetics · 2008 · 8 claims · 6 setups
Zebrafish whole-adult-organism chemogenomics generates robust prediction models that discriminate P(H)AHs from ECs across independent experiments
-
Full-text index only
Automatic discovery of cross-family sequence features associated with protein function.
PMID 16409628 · PMC1395344 · BMC bioinformatics · 2006 · 8 claims · 6 setups
A self-supervised data mining approach can find relationships between sequence features and functional annotations without preconceived functional categories.
-
Full-text index only
PA-GOSUB: a searchable database of model organism protein sequences with their predicted Gene Ontology molecular function and subcellular localization.
PMID 15608166 · PMC540074 · Nucleic acids research · 2005 · 7 claims · 4 setups
PA-GOSUB significantly extends the coverage of GO molecular function and subcellular localization annotations for 10 model organism proteomes compared with existing databases (GOA, Swiss-Prot).
-
Full-text index only
Genomic transcriptional profiling identifies a candidate blood biomarker signature for the diagnosis of septicemic melioidosis.
PMID 19903332 · PMC3091321 · Genome biology · 2009 · 6 claims · 5 setups
A candidate 37-transcript diagnostic signature distinguishes septicemic melioidosis from sepsis caused by other organisms with 100% accuracy in the training set and 78%/80% accuracy in two independent validation sets
-
Has reproduction · 85
Predicting the pathogenicity of missense variants using features derived from AlphaFold2.
PMID 37084271 · PMC10203375 · Bioinformatics (Oxford, England) · 2023 · 6 claims · 8 setups
AlphaFold2-derived structural features (solvent accessibility, amino acid network features, physicochemical environment, pLDDT) can be used to train a random forest classifier (AlphScore) that distinguishes proxy-benign from proxy-pathogenic missense variants.
-
Full-text index only
Optimality driven nearest centroid classification from genomic data.
PMID 17912341 · PMC1991588 · PloS one · 2007 · 7 claims · 5 setups
A theoretical result determines the subset of features of a given size that minimizes the misclassification rate for a nearest-centroid (LDA) classifier, based on equation (4).
-
Full-text index only
Expression genomics in breast cancer research: microarrays at the crossroads of biology and medicine.
PMID 17397520 · PMC1868923 · Breast cancer research : BCR · 2007 · 8 claims · 8 setups
Genome-wide expression microarray studies reveal transcriptional networks/signatures that explain breast cancer biological and clinical heterogeneity
-
Full-text index only
A non-parametric meta-analysis approach for combining independent microarray datasets: application using two microarray datasets pertaining to chronic allograft nephropathy.
PMID 18302764 · PMC2276496 · BMC genomics · 2008 · 8 claims · 6 setups
A novel non-parametric meta-analysis approach for combining independent microarray datasets is presented, requiring no distributional assumptions and being logically intuitive.
-
Full-text index only
Prodepth: predict residue depth by support vector regression approach from protein sequences only.
PMID 19759917 · PMC2742725 · PloS one · 2009 · 8 claims · 8 setups
Residue depth can be reliably predicted solely from protein primary sequence using support vector regression on sequence-derived features.