Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
Automatic discovery of cross-family sequence features associated with protein function.
PMID 16409628 · PMC1395344 · BMC bioinformatics · 2006 · 8 claims · 6 setups
A self-supervised data mining approach can find relationships between sequence features and functional annotations without preconceived functional categories.
-
Full-text index only
Prediction of catalytic residues using Support Vector Machine with selected protein sequence and structural properties.
PMID 16790052 · PMC1534064 · BMC bioinformatics · 2006 · 8 claims · 7 setups
The Sequential Minimal Optimization (SMO) SVM algorithm was the best-performing classifier among 26 WEKA classifiers for predicting catalytic residues
-
Full-text index only
Genomic data sampling and its effect on classification performance assessment.
PMID 12553886 · PMC149349 · BMC bioinformatics · 2003 · 8 claims · 3 setups
Cross-validation, leave-one-out, and bootstrap are designed to reduce bias and variance in accuracy estimation from small samples.
-
Full-text index only
Does distance matter? Variations in alternative 3' splicing regulation.
PMID 17704130 · PMC2018619 · Nucleic acids research · 2007 · 8 claims · 7 setups
Alternative 3' splice sites can be distinguished from constitutive splice sites by a combination of sequence/conservation properties that vary depending on the distance between the splice sites.
-
Full-text index only
MALDI profiling of human lung cancer subtypes.
PMID 19890392 · PMC2767501 · PloS one · 2009 · 8 claims · 8 setups
PIMAC/MALDI-TOF peptide profiles combined with classification models can distinguish normal lung from tumor and differentiate NSCLC histological subtypes
-
Has reproduction · 88
Comprehensive benchmarking of large language models for RNA secondary structure prediction.
PMID 40205851 · PMC11982019 · Briefings in bioinformatics · 2025 · 7 claims · 4 setups
Existing RNA-LLMs had not previously been evaluated for secondary structure prediction in a unified, fair experimental setup with the same datasets and prediction model.
-
Full-text index only
Serum protein profile in systemic-onset juvenile idiopathic arthritis differentiates response versus nonresponse to therapy.
PMID 15987476 · PMC1175022 · Arthritis research & therapy · 2005 · 8 claims · 8 setups
SELDI-TOF MS can differentiate serum protein profiles of active versus well-controlled SJIA