Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
Automatic discovery of cross-family sequence features associated with protein function.
PMID 16409628 · PMC1395344 · BMC bioinformatics · 2006 · 8 claims · 6 setups
A self-supervised data mining approach can find relationships between sequence features and functional annotations without preconceived functional categories.
-
Has reproduction · 95
Pathway-targeting gene matrix for Drosophila gene set enrichment analysis.
PMID 34710184 · PMC8553153 · PloS one · 2021 · 8 claims · 4 setups
Gene matrix files for GSEA are largely unavailable for Drosophila, limiting pathway-level enrichment analysis in this model organism
-
Full-text index only
An SVM-based system for predicting protein subnuclear localizations.
PMID 16336650 · PMC1325059 · BMC bioinformatics · 2005 · 7 claims · 3 setups
New kernels defined on k-peptide vectors mapped by BLOSUM62-based high-scored pair matrices (D1, D2, D3) improve SVM discrimination of protein subnuclear localization compared to conventional k-peptide encodings.
-
Full-text index only
AutoCSA, an algorithm for high throughput DNA sequence variant detection in cancer genomes.
PMID 17485433 · PMC5947781 · Bioinformatics (Oxford, England) · 2007 · 7 claims · 2 setups
AutoCSA is an automated algorithm, extended from the CSA protocol, that detects DNA sequence variants in cancer genomes with minimal manual intervention
-
Full-text index only
Peptide bioinformatics: peptide classification using peptide machines.
PMID 19065810 · PMC7122642 · Methods in molecular biology (Clifton, N.J.) · 2008 · 8 claims · 4 setups
The bio-basis function, which converts peptides into numerical vectors using nongapped pairwise homology alignment scores against indicator peptides, can statistically quantify peptide similarity for classification.
-
Full-text index only
Bayesian model accounting for within-class biological variability in Serial Analysis of Gene Expression (SAGE).
PMID 15339345 · PMC517707 · BMC bioinformatics · 2004 · 7 claims · 5 setups
A Bayesian mixture model is proposed to account for within-class biological variability in SAGE/Digital-Northern/MPSS tag counting data.
-
Full-text index only
Genomic transcriptional profiling identifies a candidate blood biomarker signature for the diagnosis of septicemic melioidosis.
PMID 19903332 · PMC3091321 · Genome biology · 2009 · 6 claims · 5 setups
A candidate 37-transcript diagnostic signature distinguishes septicemic melioidosis from sepsis caused by other organisms with 100% accuracy in the training set and 78%/80% accuracy in two independent validation sets
-
Has reproduction · 67
Cyrface: An interface from Cytoscape to R that provides a user interface to R packages.
PMID 24715956 · PMC3962008 · F1000Research · 2013 · 8 claims · 6 setups
Cyrface is a Cytoscape app/Java library providing a general interface from Cytoscape (Java) to any R function or package.
-
Full-text index only
PA-GOSUB: a searchable database of model organism protein sequences with their predicted Gene Ontology molecular function and subcellular localization.
PMID 15608166 · PMC540074 · Nucleic acids research · 2005 · 7 claims · 4 setups
PA-GOSUB significantly extends the coverage of GO molecular function and subcellular localization annotations for 10 model organism proteomes compared with existing databases (GOA, Swiss-Prot).
-
Has reproduction · 60
Integrating herbarium specimen observations into global phenology data systems.
PMID 30937223 · PMC6426164 · Applications in plant sciences · 2019 · 7 claims · 5 setups
A new PPO release adds terms and properties to relate observations of parts of plants to whole plants, enabling integration of herbarium phenology data with field observation data.
-
Full-text index only
Optimality driven nearest centroid classification from genomic data.
PMID 17912341 · PMC1991588 · PloS one · 2007 · 7 claims · 5 setups
A theoretical result determines the subset of features of a given size that minimizes the misclassification rate for a nearest-centroid (LDA) classifier, based on equation (4).
-
Full-text index only
ARED 3.0: the large and diverse AU-rich transcriptome.
PMID 16381826 · PMC1347415 · Nucleic acids research · 2006 · 7 claims · 6 setups
ARED 3.0 computationally mapped more than 4000 ARE-mRNAs to the human genome, representing 5-8% of human genes.
-
Full-text index only
Integrated multi-level quality control for proteomic profiling studies using mass spectrometry.
PMID 19055809 · PMC2657802 · BMC bioinformatics · 2008 · 7 claims · 5 setups
QC processes for identifying and removing low-quality spectra are often overlooked in proteomic profiling studies
-
Full-text index only
Genome-wide analysis of human disease alleles reveals that their locations are correlated in paralogous proteins.
PMID 18989397 · PMC2565504 · PLoS computational biology · 2008 · 7 claims · 5 setups
The locations of sequence variants are correlated between paralogous human proteins more than expected by chance.
-
Has reproduction · 49
Insights into the differentiation and adaptation within Circaeasteraceae from Circaeaster agrestis genome sequencing and resequencing.
PMID 36895650 · PMC9988679 · iScience · 2023 · 8 claims · 8 setups
C. agrestis and K. uniflora are sister species with contrasting reproductive modes, providing a natural system to test effects of sexual vs asexual reproduction on genome evolution
-
Full-text index only
Identification of the REST regulon reveals extensive transposable element-mediated binding site duplication.
PMID 16899447 · PMC1557810 · Nucleic acids research · 2006 · 8 claims · 8 setups
The RE1 PSSM identifies functional RE1 binding sites with greater sensitivity and selectivity than the previously used RE1 consensus sequence
-
Full-text index only
Metabolomic profiling in LRRK2-related Parkinson's disease.
PMID 19847307 · PMC2761616 · PloS one · 2009 · 7 claims · 3 setups
Metabolomic profiles of both idiopathic PD and LRRK2 PD are clearly separated from controls
-
Full-text index only
The UCSC genome browser database: update 2007.
PMID 17142222 · PMC1669757 · Nucleic acids research · 2007 · 8 claims · 8 setups
The UCSC Genome Browser Database provides sequence and annotation data for 13 vertebrate and 19 invertebrate species as of September 2006.
-
Full-text index only
Bias of selection on human copy-number variants.
PMID 16482228 · PMC1366494 · PLoS genetics · 2006 · 8 claims · 8 setups
Human CNVs are significantly overrepresented near telomeres and centromeres and enriched in simple tandem repeats relative to the genome as a whole
-
Full-text index only
Prioritization of candidate cancer genes--an aid to oncogenomic studies.
PMID 18710882 · PMC2566894 · Nucleic acids research · 2008 · 8 claims · 8 setups
Computational classifiers using combinations of protein conservation, gene structure, protein domains, protein interactions, and regulatory data can distinguish known cancer genes (CD/CR) from unlabelled human genes