Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Has reproduction · 86
Molecular Classification Models for Triple Negative Breast Cancer Subtype Using Machine Learning.
PMID 34575658 · PMC8472680 · Journal of personalized medicine · 2021 · 6 claims · 4 setups
A training gene set of 719 unique upregulated DEGs (subtype-specific) can be used to build ML models that classify TNBC into BLIA, BLIS, MES, and LAR subtypes.
-
Has reproduction · 70
Predicting enhancers in mammalian genomes using supervised hidden Markov models.
PMID 30917778 · PMC6437899 · BMC bioinformatics · 2019 · 8 claims · 8 setups
eHMM predicts enhancers with high precision and recall comparable to state-of-the-art methods and consistently outperforms them in accuracy and resolution
-
Full-text index only
Identification of deleterious non-synonymous single nucleotide polymorphisms using sequence-derived information.
PMID 18588693 · PMC2446391 · BMC bioinformatics · 2008 · 8 claims · 5 setups
A decision tree built on 10 selected sequence-derived features classifies SAPs as Disease or Polymorphism with 82.6% accuracy and 0.607 MCC in cross-validation.
-
Full-text index only
Boosting accuracy of automated classification of fluorescence microscope images for location proteomics.
PMID 15207009 · PMC449699 · BMC bioinformatics · 2004 · 8 claims · 8 setups
New classifiers (SVMs, ensembles) and new wavelet-derived (Gabor, Daubechies) features improve recognition of protein subcellular location patterns over the previous neural network approach
-
Full-text index only
Constructing support vector machine ensembles for cancer classification based on proteomic profiling.
PMID 16689692 · PMC5173238 · Genomics, proteomics & bioinformatics · 2005 · 7 claims · 4 setups
CSVME, built by selecting a subset of base SVMs via SVM-RFE ranking and fusing them with a trained upper-layer SVM, achieves better classification performance than an ensemble of all base SVMs.
-
Full-text index only
Ab initio identification of human microRNAs based on structure motifs.
PMID 18088431 · PMC2238772 · BMC bioinformatics · 2007 · 8 claims · 7 setups
MiRPred predicts miRNA precursors ab initio using only predicted secondary structure motifs, ignoring nucleotide sequence
-
Has reproduction · 100
Gene signature discovery and systematic validation across diverse clinical cohorts for TB prognosis and response to treatment.
PMID 37471455 · PMC10393163 · PLoS computational biology · 2023 · 8 claims · 8 setups
A network-based meta-analysis across studies identifies a common 45-gene signature specific to active TB disease that accounts for cohort/population heterogeneity
-
Has reproduction · 68
Machine learning algorithm predicts fibrosis-related blood diagnosis markers of intervertebral disc degeneration.
PMID 37915003 · PMC10619283 · BMC medical genomics · 2023 · 7 claims · 7 setups
CEP120 and SPDL1 are fibrosis-related diagnostic genes for IDD, identified via a random forest model from 29 differentially expressed fibrosis-related genes
-
Has reproduction · 74
SpaGene: A Deep Adversarial Framework for Spatial Gene Imputation.
PMID 42146899 · PMC13176606 · Computational and structural biotechnology journal · 2026 · 8 claims · 6 setups
SpaGene improves average PCC and SSIM and reduces RMSE compared to 6 baseline methods (SpaGE, gimVI, Tangram, VISTA, spRefine, stDiff) across 8 diverse ST-SC dataset pairs under gene-holdout evaluation.
-
Full-text index only
Widespread dysregulation of MiRNAs by MYCN amplification and chromosomal imbalances in neuroblastoma: association of miRNA expression with survival.
PMID 19924232 · PMC2773120 · PloS one · 2009 · 8 claims · 5 setups
37 miRNAs are significantly differentially expressed between MYCN-amplified (MNA) and non-MNA neuroblastoma tumors, suggesting direct or indirect regulation by MYCN
-
Full-text index only
Cancer-specific high-throughput annotation of somatic mutations: computational prediction of driver missense mutations.
PMID 19654296 · PMC2763410 · Cancer research · 2009 · 7 claims · 7 setups
CHASM, a Random Forest-based computational method, was developed to identify and prioritize missense mutations likely to be functional drivers of tumor cell proliferation.
-
Full-text index only
Predicting failure rate of PCR in large genomes.
PMID 18492719 · PMC2441781 · Nucleic acids research · 2008 · 7 claims · 8 setups
The number of predicted primer-binding sites in genomic DNA is the most important factor determining PCR failure.
-
Has reproduction · 72
Prediction of prognostic signatures in triple-negative breast cancer based on the differential expression analysis via NanoString nCounter immune panel.
PMID 33138797 · PMC7607642 · BMC cancer · 2020 · 8 claims · 7 setups
edgeR identifies 9 DEGs associated with pCR and 13 DEGs associated with relapse from 579 immune genes in a small TNBC sample set (n=55)
-
Has reproduction · 63
Comparative Genomics of Borderline Oxacillin-Resistant Staphylococcus aureus Detected during a Pseudo-outbreak of Methicillin-Resistant S. aureus in a Neonatal Intensive Care Unit.
PMID 35038924 · PMC8764539 · mBio · 2022 · 7 claims · 8 setups
Of 42 isolates flagged as MRSA by screening agar, only 9 were PBP2a- and mecA-positive true MRSA, while the remaining 33 were mecA-negative and largely met criteria for BORSA
-
Full-text index only
Computational disease gene identification: a concert of methods prioritizes type 2 diabetes and obesity candidate genes.
PMID 16757574 · PMC1475747 · Nucleic acids research · 2006 · 6 claims · 8 setups
Applying seven independent computational disease-gene prioritization methods in concert to 9556 positional candidate genes identifies a prioritized set of likely T2D and obesity candidate genes
-
Full-text index only
Machine-learning approaches for classifying haplogroup from Y chromosome STR data.
PMID 18551166 · PMC2396484 · PLoS computational biology · 2008 · 8 claims · 5 setups
Y-STR allelic variability is partitioned more by differences among haplogroups than by differences among populations, suggesting Y-STRs carry haplogroup information
-
Has reproduction · 88
Comprehensive benchmarking of large language models for RNA secondary structure prediction.
PMID 40205851 · PMC11982019 · Briefings in bioinformatics · 2025 · 7 claims · 4 setups
Existing RNA-LLMs had not previously been evaluated for secondary structure prediction in a unified, fair experimental setup with the same datasets and prediction model.