Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Has reproduction · 92
Prognostic biomarker discovery in pancreatic cancer through hybrid ensemble feature selection and multi-omics data.
PMID 41957754 · PMC13188360 · BioData mining · 2026 · 7 claims · 3 setups
The hEFS framework integrates data subsampling with multiple prognostic models (embedded and wrapper-based), aggregates feature rankings via a voting-theory-inspired approach, and selects the optimal feature subset via Pareto front optimization, eliminating user-defined thresholds.
-
Full-text index only
Evolutionary algorithms for the selection of single nucleotide polymorphisms.
PMID 12875658 · PMC183839 · BMC bioinformatics · 2003 · 8 claims · 3 setups
Evolutionary algorithms are well suited to multiobjective optimization problems with large, intractable search spaces such as SNP selection, unlike exact methods (exhaustive enumeration) or single-objective search techniques (tabu search, simulated annealing).
-
Full-text index only
htSNPer1.0: software for haplotype block partition and htSNPs selection.
PMID 15740612 · PMC1274247 · BMC bioinformatics · 2005 · 6 claims · 1 setups
The GBB algorithm finds the globally optimal minimal htSNP set with far less computing time than exhaustive/enumeration search.
-
Has reproduction · 68
Cell-type annotation with accurate unseen cell-type identification using multiple references.
PMID 37379341 · PMC10335708 · PLoS computational biology · 2023 · 8 claims · 4 setups
mtANN integrates multiple reference datasets and eight gene selection methods via ensemble learning (multiple deep classification models + majority voting) to improve cell-type annotation accuracy
-
Full-text index only
Prediction of catalytic residues using Support Vector Machine with selected protein sequence and structural properties.
PMID 16790052 · PMC1534064 · BMC bioinformatics · 2006 · 8 claims · 7 setups
The Sequential Minimal Optimization (SMO) SVM algorithm was the best-performing classifier among 26 WEKA classifiers for predicting catalytic residues
-
Has reproduction · 76
Bayesian prediction of microbial oxygen requirement.
PMID 26913185 · PMC4743139 · F1000Research · 2013 · 7 claims · 8 setups
A naive Bayesian classifier based on presence/absence of class-associated Pfam-A domains can distinguish three oxygen requirement classes (aerobe, anaerobe, facultative anaerobe) from genome sequence, unlike prior studies that only made pairwise distinctions.
-
Has reproduction · 87
R2DT is a framework for predicting and visualising RNA secondary structure using templates.
PMID 34108470 · PMC8190129 · Nature communications · 2021 · 8 claims · 6 setups
R2DT is a template-based computational framework/pipeline that predicts and visualises RNA 2D structure in standardised, community-accepted layouts
-
Full-text index only
Function2Gene: a gene selection tool to increase the power of genetic association studies by utilizing public databases and expert knowledge.
PMID 18631403 · PMC2500032 · BMC bioinformatics · 2008 · 6 claims · 5 setups
Function2Gene is a set of Perl programs that queries public databases (NCBI, GeneCards, Harvester, with Uniprot/Ensembl also supported) using expert-selected keywords to rank genes by prior probability of disease association.
-
Has reproduction · 100
Smart spatial omics (S2-omics) optimizes region of interest selection to capture molecular heterogeneity in diverse tissues.
PMID 41298871 · PMC12662399 · Nature cell biology · 2025 · 7 claims · 6 setups
S2-omics is an end-to-end workflow that automatically selects ROIs from H&E histology images to maximize molecular information content for spatial omics profiling.
-
Has reproduction · 75
Sequencing of human genomes with nanopore technology.
PMID 31015479 · PMC6478738 · Nature communications · 2019 · 8 claims · 7 setups
A novel single-sample, reference panel-free, read-based phasing algorithm built on the STITCH model improves nanopore SNV calling from modest baseline levels.
-
Has reproduction · 67
Generative and integrative modeling for transcriptomics with formalin fixed paraffin embedded material.
PMID 41029822 · PMC12486589 · Journal of translational medicine · 2025 · 8 claims · 5 setups
fRNA-seq transcript counts are best fit by the negative binomial distribution, with little evidence supporting zero-inflated extensions
-
Full-text index only
Computational tradeoffs in multiplex PCR assay design for SNP genotyping.
PMID 16042802 · PMC1190169 · BMC genomics · 2005 · 7 claims · 6 setups
Achieving high-multiplexing/high-coverage multiplex PCR designs is subject to a computational phase transition as the SNP-pair compatibility probability crosses a critical threshold
-
Full-text index only
GoMiner: a resource for biological interpretation of genomic and proteomic data.
PMID 12702209 · PMC154579 · Genome biology · 2003 · 8 claims · 4 setups
GoMiner organizes 'interesting' gene lists (e.g., differentially expressed genes) into the Gene Ontology hierarchy for biological interpretation, displaying results as both a tree and a directed acyclic graph (DAG).
-
Full-text index only
Identification of deleterious non-synonymous single nucleotide polymorphisms using sequence-derived information.
PMID 18588693 · PMC2446391 · BMC bioinformatics · 2008 · 8 claims · 5 setups
A decision tree built on 10 selected sequence-derived features classifies SAPs as Disease or Polymorphism with 82.6% accuracy and 0.607 MCC in cross-validation.
-
Full-text index only
Analysis of protein sequence and interaction data for candidate disease gene prediction.
PMID 17020920 · PMC1636487 · Nucleic acids research · 2006 · 8 claims · 7 setups
Combining CPS and CMP using known disease genes as input achieves sensitivity 0.52 and specificity 0.97, reducing candidate lists 13-fold