Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
Searching for interpretable rules for disease mutations: a simulated annealing bump hunting strategy.
PMID 16984653 · PMC1618409 · BMC bioinformatics · 2006 · 8 claims · 6 setups
The proposed feature set outperforms existing published feature sets for predicting effects of amino acid substitutions
-
Has reproduction · 88
CIRCprimerXL: Convenient and High-Throughput PCR Primer Design for Circular RNA Quantification.
PMID 36304334 · PMC9580850 · Frontiers in bioinformatics · 2022 · 8 claims · 5 setups
CIRCprimerXL is a high-throughput, user-friendly circRNA RT-qPCR primer design pipeline available as both a web tool and a standalone Nextflow/Docker pipeline
-
Full-text index only
Prediction of catalytic residues using Support Vector Machine with selected protein sequence and structural properties.
PMID 16790052 · PMC1534064 · BMC bioinformatics · 2006 · 8 claims · 7 setups
The Sequential Minimal Optimization (SMO) SVM algorithm was the best-performing classifier among 26 WEKA classifiers for predicting catalytic residues
-
Full-text index only
Boosting accuracy of automated classification of fluorescence microscope images for location proteomics.
PMID 15207009 · PMC449699 · BMC bioinformatics · 2004 · 8 claims · 8 setups
New classifiers (SVMs, ensembles) and new wavelet-derived (Gabor, Daubechies) features improve recognition of protein subcellular location patterns over the previous neural network approach
-
Full-text index only
Constructing support vector machine ensembles for cancer classification based on proteomic profiling.
PMID 16689692 · PMC5173238 · Genomics, proteomics & bioinformatics · 2005 · 7 claims · 4 setups
CSVME, built by selecting a subset of base SVMs via SVM-RFE ranking and fusing them with a trained upper-layer SVM, achieves better classification performance than an ensemble of all base SVMs.
-
Full-text index only
Integrated analysis of genetic and proteomic data identifies biomarkers associated with adverse events following smallpox vaccination.
PMID 18923431 · PMC2692715 · Genes and immunity · 2009 · 7 claims · 6 setups
A two-stage strategy (Random Forest filtering followed by decision tree modeling) can integrate categorical genetic and continuous proteomic data to identify biomarkers of AE risk
-
Has reproduction · 69
Automatic discovery of 100-miRNA signature for cancer classification using ensemble feature selection.
PMID 31533612 · PMC6751684 · BMC bioinformatics · 2019 · 7 claims · 8 setups
An ensemble feature selection method based on classifier consensus identifies a robust 100-miRNA signature from TCGA data.
-
Has reproduction · 71
Gene Set Enrichment Analysis Reveals Individual Variability in Host Responses in Tuberculosis Patients.
PMID 34421903 · PMC8375662 · Frontiers in immunology · 2021 · 8 claims · 8 setups
TB patients show substantial individual variability in the intensity of hallmark IFN responses, as well as in complement system, metabolic, and other pathway responses.
-
Full-text index only
SePaCS--a web-based application for classification of seroreactivity profiles.
PMID 17478503 · PMC1933220 · Nucleic acids research · 2007 · 8 claims · 4 setups
SePaCS is a freely available web-based tool that trains and applies multiple classification methods (4 Naive Bayes variants, SVM with RBF kernel, LDA, DLDA) to seroreactivity profiles and outputs results as a summary table plus a detailed PDF report
-
Full-text index only
Effect of the assignment of ancestral CpG state on the estimation of nucleotide substitution rates in mammals.
PMID 18826599 · PMC2576242 · BMC evolutionary biology · 2008 · 7 claims · 4 setups
CpG/non-CpG assignment based on presence/absence of a CpG dinucleotide seriously biases substitution rate estimates, overestimating CpG changes and underestimating non-CpG changes.
-
Full-text index only
Network-assisted protein identification and data interpretation in shotgun proteomics.
PMID 19690572 · PMC2736651 · Molecular systems biology · 2009 · 7 claims · 7 setups
Confidently identified proteins in a sample form tightly connected sub-networks in the protein interaction network, with significantly higher clustering coefficients than random or topology-matched random sub-networks.
-
Full-text index only
Filtering high-throughput protein-protein interaction data using a combination of genomic features.
PMID 15833142 · PMC1127019 · BMC bioinformatics · 2005 · 8 claims · 8 setups
A combination of three genomic features (interacting Pfam domains, GO annotations, sequence homology) using naive Bayesian networks predicts true protein-protein interactions with high sensitivity and good specificity.
-
Full-text index only
Identification of serum biomarkers for colon cancer by proteomic analysis.
PMID 16755300 · PMC2361335 · British journal of cancer · 2006 · 8 claims · 8 setups
Complement C3a des-arg, α1-antitrypsin and transferrin were identified as serum proteins with diagnostic potential for CRC.
-
Full-text index only
High incidence of human bocavirus infection in children in Spain.
PMID 17904416 · PMC7108365 · Journal of clinical virology : the official publication of the Pan American Society for Clinical Virology · 2007 · 8 claims · 6 setups
HBoV is a significant contributor to acute lower respiratory tract infection in young children in Spain
-
Full-text index only
Predicting failure rate of PCR in large genomes.
PMID 18492719 · PMC2441781 · Nucleic acids research · 2008 · 7 claims · 8 setups
The number of predicted primer-binding sites in genomic DNA is the most important factor determining PCR failure.
-
Full-text index only
Aerobic nonylphenol degradation and nitro-nonylphenol formation by microbial cultures from sediments.
PMID 20043151 · PMC2825322 · Applied microbiology and biotechnology · 2010 · 8 claims · 8 setups
Aerobic biodegradation of branched NP in polluted river sediment occurs within 8 days after a 2-day lag phase at 30°C
-
Full-text index only
Discovering cancer genes by integrating network and functional properties.
PMID 19765316 · PMC2758898 · BMC medical genomics · 2009 · 8 claims · 6 setups
Cancer genes have distinct PPI network topology (higher connectivity, higher clustering coefficient, shorter path length to known cancer genes) compared to non-cancer genes
-
Full-text index only
A comparison of classification methods for predicting Chronic Fatigue Syndrome based on genetic data.
PMID 19772600 · PMC2765429 · Journal of translational medicine · 2009 · 7 claims · 3 setups
The naive Bayes model with the wrapper-based feature selection approach performed best among all predictive models tested for distinguishing CFS from controls.
-
Full-text index only
Decision forest analysis of 61 single nucleotide polymorphisms in a case-control study of esophageal cancer; a novel method.
PMID 16026601 · PMC1637030 · BMC bioinformatics · 2005 · 8 claims · 2 setups
DF-SNPs, a novel adaptation of the Decision Forest method, can classify esophageal cancer cases vs. controls based on SNP genotype data with high concordance, sensitivity, and specificity.