Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
PA-GOSUB: a searchable database of model organism protein sequences with their predicted Gene Ontology molecular function and subcellular localization.
PMID 15608166 · PMC540074 · Nucleic acids research · 2005 · 7 claims · 4 setups
PA-GOSUB significantly extends the coverage of GO molecular function and subcellular localization annotations for 10 model organism proteomes compared with existing databases (GOA, Swiss-Prot).
-
Full-text index only
Predicting deleterious nsSNPs: an analysis of sequence and structural attributes.
PMID 16630345 · PMC1489951 · BMC bioinformatics · 2006 · 8 claims · 7 setups
Sequence conservation (PSIC score difference) at the nsSNP position is the single most useful attribute for predicting deleterious vs neutral status.
-
Full-text index only
Sequence and structure signatures of cancer mutation hotspots in protein kinases.
PMID 19834613 · PMC2759519 · PloS one · 2009 · 8 claims · 6 setups
Developed CKMD (Composite Kinase Mutation Database), an integrated bioinformatics resource mapping genetic variation in protein kinase genes to sequence, structural, and functional data
-
Full-text index only
Prediction of catalytic residues using Support Vector Machine with selected protein sequence and structural properties.
PMID 16790052 · PMC1534064 · BMC bioinformatics · 2006 · 8 claims · 7 setups
The Sequential Minimal Optimization (SMO) SVM algorithm was the best-performing classifier among 26 WEKA classifiers for predicting catalytic residues
-
Has reproduction · 87
Machine learning reveals microbial interactions driving plastic degradation across plastisphere environments.
PMID 41657981 · PMC12876002 · Frontiers in microbiology · 2025 · 6 claims · 7 setups
Wastewater plastispheres harbor the most diverse and compositionally even microbial communities among the three habitats.
-
Has reproduction · 55
Identification and verification of diagnostic biomarkers in recurrent pregnancy loss via machine learning algorithm and WGCNA.
PMID 37691920 · PMC10485775 · Frontiers in immunology · 2023 · 8 claims · 8 setups
352 DEGs (198 up-regulated, 154 down-regulated) were identified between RPL and control endometrial samples
-
Full-text index only
Cancer-specific high-throughput annotation of somatic mutations: computational prediction of driver missense mutations.
PMID 19654296 · PMC2763410 · Cancer research · 2009 · 7 claims · 7 setups
CHASM, a Random Forest-based computational method, was developed to identify and prioritize missense mutations likely to be functional drivers of tumor cell proliferation.
-
Has reproduction · 81
Identification of Proteins Deregulated by Platinum-Based Chemotherapy as Novel Biomarkers and Therapeutic Targets in Non-Small Cell Lung Cancer.
PMID 33777753 · PMC7991912 · Frontiers in oncology · 2021 · 7 claims · 8 setups
Cisplatin exposure induces significant deregulation of protein expression networks in NSCLC cells
-
Full-text index only
Identification of diagnostic markers for tuberculosis by proteomic fingerprinting of serum.
PMID 16980117 · PMC7159276 · Lancet (London, England) · 2006 · 8 claims · 5 setups
An SVM classifier trained on serum proteomic profiles discriminated patients with active tuberculosis from controls with clinically overlapping conditions
-
Full-text index only
The impact of peptide abundance and dynamic range on stable-isotope-based quantitative proteomic analyses.
PMID 18798661 · PMC2746028 · Journal of proteome research · 2008 · 8 claims · 7 setups
Over half of confidently identified peptides in complex mixtures have S/N ratios below 10 on both FT-ICR and Orbitrap instruments
-
Full-text index only
Extraction of human kinase mutations from literature, databases and genotyping studies.
PMID 19758464 · PMC2745582 · BMC bioinformatics · 2009 · 7 claims · 6 setups
A literature mining pipeline combining MutationFinder, false-positive filtering, and SVM-based classification can extract and disambiguate single-point mutation mentions from abstracts and full text
-
Full-text index only
Integrated proteomic and transcriptomic profiling of mouse lung development and Nmyc target genes.
PMID 17486137 · PMC2673710 · Molecular systems biology · 2007 · 8 claims · 7 setups
Global MudPIT-based proteomic profiling across six mouse lung developmental time points (E13.5–P56) identifies thousands of proteins and captures developmental/cell-biological expression patterns.