Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
A modified T-test feature selection method and its application on the HapMap genotype data.
PMID 18267305 · PMC5054219 · Genomics, proteomics & bioinformatics · 2007 · 7 claims · 4 setups
A modified t-test ranking measure, extended to handle nominal SNP genotype data via vector transformation, can effectively rank SNPs by their discriminative capability for population classification.
-
Full-text index only
MiPred: classification of real and pseudo microRNA precursors using random forest prediction model with combined features.
PMID 17553836 · PMC1933124 · Nucleic acids research · 2007 · 8 claims · 8 setups
A hybrid feature combining local contiguous triplet structure-sequence composition, MFE of the secondary structure, and P-value of a randomization test improves classification of real vs pseudo pre-miRNAs
-
Full-text index only
Speeding disease gene discovery by sequence based candidate prioritization.
PMID 15766383 · PMC1274252 · BMC bioinformatics · 2005 · 7 claims · 8 setups
Disease genes (OMIM) differ significantly from non-disease genes in sequence-based features including gene/cDNA/protein size, exon number, homolog conservation, secretion signal, 3' UTR length, CpG islands, and distance to nearest gene.
-
Full-text index only
Next-generation high-density self-assembling functional protein arrays.
PMID 18469824 · PMC3070491 · Nature methods · 2008 · 8 claims · 7 setups
A next-generation NAPPA method produces high-density protein microarrays displaying over 1500 unique proteins with >90% expression success
-
Has reproduction · 50
Workflow sharing with automated metadata validation and test execution to improve the reusability of published workflows.
PMID 36810800 · PMC9944229 · GigaScience · 2022 · 8 claims · 5 setups
Yevis is a system that builds a workflow registry which automatically validates and tests workflows prior to publication, ensuring they are 'reusable with confidence'.
-
Full-text index only
Fast-evolving noncoding sequences in the human genome.
PMID 17578567 · PMC2394770 · Genome biology · 2007 · 8 claims · 6 setups
1,356 conserved noncoding sequences show human-specific accelerated substitution rates (ANC sequences) relative to chimpanzee
-
Full-text index only
Supervised learning-based tagSNP selection for genome-wide disease classifications.
PMID 18366619 · PMC2386071 · BMC genomics · 2008 · 7 claims · 2 setups
SRFA (Supervised Recursive Feature Addition) is a novel feature selection method combining supervised learning and statistical redundancy measures for SNP selection
-
Full-text index only
Classification of real and pseudo microRNA precursors using local structure-sequence features and support vector machine.
PMID 16381612 · PMC1360673 · BMC bioinformatics · 2005 · 7 claims · 7 setups
A 32-dimensional triplet structure-sequence feature vector combined with SVM (triplet-SVM) can distinguish real human pre-miRNAs from pseudo pre-miRNA hairpins with ~90% accuracy.
-
Full-text index only
Application of machine learning in SNP discovery.
PMID 16398931 · PMC1955739 · BMC bioinformatics · 2006 · 8 claims · 6 setups
PolyBayes produces high false-positive SNP predictions even with stringent parameters
-
Full-text index only
Optimality driven nearest centroid classification from genomic data.
PMID 17912341 · PMC1991588 · PloS one · 2007 · 7 claims · 5 setups
A theoretical result determines the subset of features of a given size that minimizes the misclassification rate for a nearest-centroid (LDA) classifier, based on equation (4).
-
Full-text index only
Detecting purely epistatic multi-locus interactions by an omnibus permutation test on ensembles of two-locus analyses.
PMID 19761607 · PMC2759961 · BMC bioinformatics · 2009 · 8 claims · 5 setups
2LOmb performs an omnibus permutation test on ensembles of two-locus analyses via a four-step algorithm (two-locus analysis, permutation test, global p-value determination, progressive ensemble search)
-
Has reproduction · 86
Molecular Classification Models for Triple Negative Breast Cancer Subtype Using Machine Learning.
PMID 34575658 · PMC8472680 · Journal of personalized medicine · 2021 · 6 claims · 4 setups
A training gene set of 719 unique upregulated DEGs (subtype-specific) can be used to build ML models that classify TNBC into BLIA, BLIS, MES, and LAR subtypes.
-
Has reproduction · 48
Comparative analysis of molecular signatures reveals a hybrid approach in breast cancer: Combining the Nottingham Prognostic Index with gene expressions into a hybrid signature.
PMID 35143511 · PMC8830616 · PloS one · 2022 · 8 claims · 6 setups
A hybrid signature combining the Nottingham Prognostic Index with SIS-selected gene expressions can be built in a data-driven fashion (NPI treated as a gene expression during feature selection).
-
Has reproduction · 89
Graph Random Forest: A Graph Embedded Algorithm for Identifying Highly Connected Important Features.
PMID 37509188 · PMC10377046 · Biomolecules · 2023 · 8 claims · 3 setups
Graph Random Forest (GRF) embeds graph/network information directly into the decision-tree building process by splitting on features in the k-hop neighborhood of a data-driven head-splitting node.
-
Has reproduction · 83
Analyzing biomarker discovery: Estimating the reproducibility of biomarker sets.
PMID 35901020 · PMC9333302 · PloS one · 2022 · 7 claims · 3 setups
A Reproducibility Score, RS(D,BD), defined as the average Jaccard overlap between biomarker sets found by the same discovery process on comparable datasets from the same distribution, quantifies biomarker reproducibility on a 0-1 scale
-
Has reproduction · 83
Gene-expression patterns in peripheral blood classify familial breast cancer susceptibility.
PMID 26538066 · PMC4634735 · BMC medical genomics · 2015 · 8 claims · 5 setups
A multigene peripheral-blood gene-expression biomarker accurately classifies which women from high-risk families develop familial breast cancer.
-
Full-text index only
MACSIMS: multiple alignment of complete sequences information management system.
PMID 16792820 · PMC1539025 · BMC bioinformatics · 2006 · 8 claims · 5 setups
MACSIMS is a multiple alignment-based information management system combining knowledge-based database mining with ab initio sequence predictions
-
Full-text index only
Identification of deleterious non-synonymous single nucleotide polymorphisms using sequence-derived information.
PMID 18588693 · PMC2446391 · BMC bioinformatics · 2008 · 8 claims · 5 setups
A decision tree built on 10 selected sequence-derived features classifies SAPs as Disease or Polymorphism with 82.6% accuracy and 0.607 MCC in cross-validation.
-
Has reproduction · 72
Prediction of prognostic signatures in triple-negative breast cancer based on the differential expression analysis via NanoString nCounter immune panel.
PMID 33138797 · PMC7607642 · BMC cancer · 2020 · 8 claims · 7 setups
edgeR identifies 9 DEGs associated with pCR and 13 DEGs associated with relapse from 579 immune genes in a small TNBC sample set (n=55)
-
Full-text index only
Discovering cancer genes by integrating network and functional properties.
PMID 19765316 · PMC2758898 · BMC medical genomics · 2009 · 8 claims · 6 setups
Cancer genes have distinct PPI network topology (higher connectivity, higher clustering coefficient, shorter path length to known cancer genes) compared to non-cancer genes