Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Has reproduction · 85
Predicting the pathogenicity of missense variants using features derived from AlphaFold2.
PMID 37084271 · PMC10203375 · Bioinformatics (Oxford, England) · 2023 · 6 claims · 8 setups
AlphaFold2-derived structural features (solvent accessibility, amino acid network features, physicochemical environment, pLDDT) can be used to train a random forest classifier (AlphScore) that distinguishes proxy-benign from proxy-pathogenic missense variants.
-
Full-text index only
MiPred: classification of real and pseudo microRNA precursors using random forest prediction model with combined features.
PMID 17553836 · PMC1933124 · Nucleic acids research · 2007 · 8 claims · 8 setups
A hybrid feature combining local contiguous triplet structure-sequence composition, MFE of the secondary structure, and P-value of a randomization test improves classification of real vs pseudo pre-miRNAs
-
Has reproduction · 86
Molecular Classification Models for Triple Negative Breast Cancer Subtype Using Machine Learning.
PMID 34575658 · PMC8472680 · Journal of personalized medicine · 2021 · 6 claims · 4 setups
A training gene set of 719 unique upregulated DEGs (subtype-specific) can be used to build ML models that classify TNBC into BLIA, BLIS, MES, and LAR subtypes.
-
Full-text index only
Reconstruction of human protein interolog network using evolutionary conserved network.
PMID 17493278 · PMC1885812 · BMC bioinformatics · 2007 · 8 claims · 7 setups
A relative conservation score derived from maximal quasi-cliques in protein interaction networks, combined with other interaction features, can score and rank predicted human interologs for confidence.
-
Has reproduction · 83
Integrative transcriptomic and machine learning framework reveals candidate genes and potential mechanisms of aflatoxin B1 exposure in breast cancer.
PMID 41688730 · PMC12982753 · Scientific reports · 2026 · 7 claims · 8 setups
Twenty-two genes lie at the intersection of AFB1-predicted targets and breast cancer-associated co-expression modules/DEGs
-
Full-text index only
PA-GOSUB: a searchable database of model organism protein sequences with their predicted Gene Ontology molecular function and subcellular localization.
PMID 15608166 · PMC540074 · Nucleic acids research · 2005 · 7 claims · 4 setups
PA-GOSUB significantly extends the coverage of GO molecular function and subcellular localization annotations for 10 model organism proteomes compared with existing databases (GOA, Swiss-Prot).
-
Full-text index only
Interaction profile-based protein classification of death domain.
PMID 15189571 · PMC459208 · BMC bioinformatics · 2004 · 7 claims · 6 setups
An SVM-based classifier using Residue Pair Interaction Profiles (RPIPs) can classify death domain superfamily members into subfamilies with 89% average cross-validation accuracy
-
Full-text index only
MACSIMS: multiple alignment of complete sequences information management system.
PMID 16792820 · PMC1539025 · BMC bioinformatics · 2006 · 8 claims · 5 setups
MACSIMS is a multiple alignment-based information management system combining knowledge-based database mining with ab initio sequence predictions
-
Has reproduction · 78
Machine learning and free energy clustering reveal PAH protein binding linked to AD risk.
PMID 41953002 · PMC13053772 · iScience · 2026 · 7 claims · 8 setups
An integrated framework of bioinformatics, machine learning, and ΔG clustering can prioritize PAHs for AD-associated neurotoxicity.
-
Full-text index only
miRBase: tools for microRNA genomics.
PMID 17991681 · PMC2238936 · Nucleic acids research · 2008 · 8 claims · 6 setups
miRBase release 10.0 contains 5071 miRNA hairpin loci from 58 species, expressing 5922 distinct mature miRNA sequences, a growth of over 2000 sequences in 2 years
-
Full-text index only
BioHealthBase: informatics support in the elucidation of influenza virus host pathogen interactions and virulence.
PMID 17965094 · PMC2238987 · Nucleic acids research · 2008 · 7 claims · 5 setups
BioHealthBase BRC is a public integrated bioinformatics database and analysis resource for influenza virus, Francisella tularensis, Mycobacterium tuberculosis, Microsporidia species and ricin toxin.
-
Full-text index only
Predicting positive p53 cancer rescue regions using Most Informative Positive (MIP) active learning.
PMID 19756158 · PMC2742196 · PLoS computational biology · 2009 · 8 claims · 4 setups
MIP active learning is a novel active learning method that preferentially seeks informative Positive (functionally active) examples rather than only maximizing classifier accuracy.
-
Has reproduction · 70
Predicting enhancers in mammalian genomes using supervised hidden Markov models.
PMID 30917778 · PMC6437899 · BMC bioinformatics · 2019 · 8 claims · 8 setups
eHMM predicts enhancers with high precision and recall comparable to state-of-the-art methods and consistently outperforms them in accuracy and resolution
-
Has reproduction · 63
Comparative Genomics of Borderline Oxacillin-Resistant Staphylococcus aureus Detected during a Pseudo-outbreak of Methicillin-Resistant S. aureus in a Neonatal Intensive Care Unit.
PMID 35038924 · PMC8764539 · mBio · 2022 · 7 claims · 8 setups
Of 42 isolates flagged as MRSA by screening agar, only 9 were PBP2a- and mecA-positive true MRSA, while the remaining 33 were mecA-negative and largely met criteria for BORSA
-
Has reproduction · 81
Enabling Single-Cell Drug Response Annotations from Bulk RNA-Seq Using SCAD.
PMID 36762572 · PMC10104628 · Advanced science (Weinheim, Baden-Wurttemberg, Germany) · 2023 · 7 claims · 7 setups
SCAD, a transfer learning framework integrating adversarial discriminative domain adaptation (ADDA), can infer single-cell drug sensitivities by transferring knowledge from bulk RNA-seq pharmacogenomic data (GDSC) to scRNA-seq target domains
-
Has reproduction · 90
A Decentralized Kidney Transplant Biopsy Classifier for Transplant Rejection Developed Using Genes of the Banff-Human Organ Transplant Panel.
PMID 35619722 · PMC9128066 · Frontiers in immunology · 2022 · 6 claims · 6 setups
A random forest model trained solely on B-HOT panel genes (B-HOT Model) accurately classifies kidney transplant biopsies as NR, ABMR, or TCMR.
-
Full-text index only
Anopheles gambiae genome reannotation through synthesis of ab initio and comparative gene prediction algorithms.
PMID 16569258 · PMC1557760 · Genome biology · 2006 · 8 claims · 7 setups
An exon-gene-union (EGU) algorithm followed by an open-reading-frame-selection algorithm can synthesize ab initio (GENSCAN, GeneMark, SNAP) and comparative (Ensembl/Genewise) predictions into a single, more complete CDS set
-
Full-text index only
SePaCS--a web-based application for classification of seroreactivity profiles.
PMID 17478503 · PMC1933220 · Nucleic acids research · 2007 · 8 claims · 4 setups
SePaCS is a freely available web-based tool that trains and applies multiple classification methods (4 Naive Bayes variants, SVM with RBF kernel, LDA, DLDA) to seroreactivity profiles and outputs results as a summary table plus a detailed PDF report
-
Full-text index only
High-throughput crystallography for structural genomics.
PMID 19765976 · PMC2764548 · Current opinion in structural biology · 2009 · 8 claims · 8 setups
SG programs use genomic sequence data to select structurally novel protein targets, avoiding proteins with known structural homologues
-
Full-text index only
Discovering cancer genes by integrating network and functional properties.
PMID 19765316 · PMC2758898 · BMC medical genomics · 2009 · 8 claims · 6 setups
Cancer genes have distinct PPI network topology (higher connectivity, higher clustering coefficient, shorter path length to known cancer genes) compared to non-cancer genes