Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
Peptide bioinformatics: peptide classification using peptide machines.
PMID 19065810 · PMC7122642 · Methods in molecular biology (Clifton, N.J.) · 2008 · 8 claims · 4 setups
The bio-basis function, which converts peptides into numerical vectors using nongapped pairwise homology alignment scores against indicator peptides, can statistically quantify peptide similarity for classification.
-
Full-text index only
MODBASE, a database of annotated comparative protein structure models and associated resources.
PMID 18948282 · PMC2686492 · Nucleic acids research · 2009 · 8 claims · 8 setups
MODBASE contains 5,152,695 reliable comparative protein structure models for 1,593,209 unique protein sequences.
-
Full-text index only
Evolutionary trace annotation of protein function in the structural proteome.
PMID 20036248 · PMC2831211 · Journal of molecular biology · 2010 · 8 claims · 7 setups
ET-ranked residue clusters can be used to build 3D templates that predict GO function in enzymes and non-enzymes alike, without prior knowledge of functional mechanism.
-
Full-text index only
Swarm intelligence based wavelet coefficient feature selection for mass spectral classification: an application to proteomics data.
PMID 19733729 · PMC2748225 · Analytica chimica acta · 2009 · 8 claims · 4 setups
ACA-based wavelet coefficient feature selection can achieve up to 100% classification accuracy on training, validating, and independent testing sets using only 5 selected features.
-
Has reproduction · 67
Adaptive learning embedding features to improve the predictive performance of SARS-CoV-2 phosphorylation sites.
PMID 37847658 · PMC10628388 · Bioinformatics (Oxford, England) · 2023 · 7 claims · 3 setups
PSPred-ALE, a deep learning predictor using a self-adaptive learning embedding algorithm, automatically extracts contextual sequence features and identifies SARS-CoV-2 phosphorylation sites without feature engineering.
-
Has reproduction · 83
Analyzing biomarker discovery: Estimating the reproducibility of biomarker sets.
PMID 35901020 · PMC9333302 · PloS one · 2022 · 7 claims · 3 setups
A Reproducibility Score, RS(D,BD), defined as the average Jaccard overlap between biomarker sets found by the same discovery process on comparable datasets from the same distribution, quantifies biomarker reproducibility on a 0-1 scale
-
Full-text index only
Assessment of serum proteomics to detect large colon adenomas.
PMID 18708413 · PMC2561171 · Cancer epidemiology, biomarkers & prevention : a publication of the American Association for Cancer Research, cosponsored by the American Society of Preventive Oncology · 2008 · 7 claims · 1 setups
In the primary, pre-specified blinded validation analysis, the serum proteomics (ProteomeQuest) model showed no significant discrimination between large adenoma and normal subjects (accuracy 51%).
-
Has reproduction · 89
MirDIP 5.2: tissue context annotation and novel microRNA curation.
PMID 36453996 · PMC9825511 · Nucleic acids research · 2023 · 7 claims · 6 setups
mirDIP 5.2 removed eight outdated resources, added miRNATIP, and ran five prediction algorithms against miRBase and mirGeneDB miRNAs to expand and improve interaction coverage
-
Has reproduction
Different approaches to Imaging Mass Cytometry data analysis.
PMID 37092034 · PMC10115470 · Bioinformatics advances · 2023 · 8 claims · 5 setups
Imaging Mass Cytometry (IMC) is a high-multiplexing imaging platform capable of simultaneously detecting and visualizing up to 40 different protein targets in tissue sections.
-
Full-text index only
Report of the 9th HLPP Workshop October 2007, Seoul, Korea.
PMID 18683817 · PMC4601560 · Proteomics · 2008 · 8 claims · 8 setups
An integrated separating-identifying platform identified 6788 proteins (≥2 peptides, 95% confidence) in Chinese human liver samples, including 3721 new to liver and 977 hypothetical proteins
-
Has reproduction · 81
Comparing the utility of in vivo transposon mutagenesis approaches in yeast species to infer gene essentiality.
PMID 32681306 · PMC7599172 · Current genetics · 2020 · 7 claims · 7 setups
A Random Forest machine-learning approach can predict gene essentiality from in vivo transposon insertion data across multiple yeast species and transposon systems
-
Has reproduction
Artificial Intelligence Approach in Machine Learning-Based Modeling and Networking of the Coronavirus Pathogenesis Pathway.
PMID 40699865 · PMC12191508 · Current issues in molecular biology · 2025 · 8 claims · 8 setups
The coronavirus pathogenesis pathway is activated in SARS-CoV-2-infected iPSC-derived cardiac cells and in SARS-CoV/SARS-CoV-2-infected LUAD cells
-
Has reproduction · 85
PowerBacGWAS: a computational pipeline to perform power calculations for bacterial genome-wide association studies.
PMID 35338232 · PMC8956664 · Communications biology · 2022 · 8 claims · 8 setups
Two computational approaches (sub-sampling and phenotype-simulation) can be implemented to perform power calculations for bacterial GWAS using existing genome collections, packaged as the PowerBacGWAS pipeline
-
Full-text index only
Gene Prospector: an evidence gateway for evaluating potential susceptibility genes and interacting risk factors for human diseases.
PMID 19063745 · PMC2613935 · BMC bioinformatics · 2008 · 8 claims · 5 setups
Gene Prospector is a Web-based application that selects and prioritizes potential disease-related genes using a curated, updated literature database of genetic association studies