Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Has reproduction · 63
Target identification for repurposed drugs active against SARS-CoV-2 via high-throughput inverse docking.
PMID 34825285 · PMC8616721 · Journal of computer-aided molecular design · 2022 · 8 claims · 6 setups
Combining Vinardo, Ledock, and Korp-PL scoring functions (via averaged Z-scores) improves correct target identification over any single scoring function.
-
Full-text index only
Information extraction from full text scientific articles: where are the keywords?
PMID 12775220 · PMC166134 · BMC bioinformatics · 2003 · 8 claims · 5 setups
The keyword content of the five article sections (A, I, M, R, D) is heterogeneous, i.e., different sections carry different kinds of information.
-
Full-text index only
Shotgun haplotyping: a novel method for surveying allelic sequence variation.
PMID 16221968 · PMC1253838 · Nucleic acids research · 2005 · 8 claims · 7 setups
A novel shotgun haplotyping method generates haplotypic sequences from long PCR products by shotgun sequencing both alleles concurrently and using read-pair information to separate alleles during assembly
-
Full-text index only
A clustering property of highly-degenerate transcription factor binding sites in the mammalian genome.
PMID 16670430 · PMC1456330 · Nucleic acids research · 2006 · 8 claims · 7 setups
Highly-degenerate RE1 sites are significantly enriched in promoters of validated and putative REST target genes compared to control promoters
-
Has reproduction · 90
A Decentralized Kidney Transplant Biopsy Classifier for Transplant Rejection Developed Using Genes of the Banff-Human Organ Transplant Panel.
PMID 35619722 · PMC9128066 · Frontiers in immunology · 2022 · 6 claims · 6 setups
A random forest model trained solely on B-HOT panel genes (B-HOT Model) accurately classifies kidney transplant biopsies as NR, ABMR, or TCMR.
-
Full-text index only
Inparanoid: a comprehensive database of eukaryotic orthologs.
PMID 15608241 · PMC540061 · Nucleic acids research · 2005 · 8 claims · 4 setups
The Inparanoid algorithm identifies true ortholog clusters by seeding on reciprocal best-matching pairs, gathering inparalogs (post-speciation duplicates) while excluding outparalogs (pre-speciation duplicates)
-
Has reproduction · 43
StatsDB: platform-agnostic storage and understanding of next generation sequencing run metrics.
PMID 24627795 · PMC3938176 · F1000Research · 2013 · 8 claims · 6 setups
StatsDB is an open-source software package for storage and analysis of next generation sequencing run metrics, backed by an SQL (MySQL) database with Perl and Java APIs.
-
Full-text index only
Predicting the phenotypic effects of non-synonymous single nucleotide polymorphisms based on support vector machines.
PMID 18005451 · PMC2216041 · BMC bioinformatics · 2007 · 8 claims · 5 setups
Parepro, an SVM-based method integrating three attribute sets (RD, MI, IE) derived from evolutionary and residue-property information, predicts whether an nsSNP is deleterious or neutral.
-
Full-text index only
MiPred: classification of real and pseudo microRNA precursors using random forest prediction model with combined features.
PMID 17553836 · PMC1933124 · Nucleic acids research · 2007 · 8 claims · 8 setups
A hybrid feature combining local contiguous triplet structure-sequence composition, MFE of the secondary structure, and P-value of a randomization test improves classification of real vs pseudo pre-miRNAs
-
Full-text index only
Logical Analysis of Data (LAD) model for the early diagnosis of acute ischemic stroke.
PMID 18616825 · PMC2492849 · BMC medical informatics and decision making · 2008 · 7 claims · 5 setups
An LAD classification model built from a support-set of 3 peptide peaks can distinguish stroke patients from controls with 75% accuracy on an independent validation set
-
Full-text index only
Identification of deleterious non-synonymous single nucleotide polymorphisms using sequence-derived information.
PMID 18588693 · PMC2446391 · BMC bioinformatics · 2008 · 8 claims · 5 setups
A decision tree built on 10 selected sequence-derived features classifies SAPs as Disease or Polymorphism with 82.6% accuracy and 0.607 MCC in cross-validation.
-
Full-text index only
SNP500Cancer: a public resource for sequence validation, assay development, and frequency analysis for genetic variation in candidate genes.
PMID 16381944 · PMC1347513 · Nucleic acids research · 2006 · 7 claims · 4 setups
SNP500Cancer provides sequence and genotype assay information for candidate cancer-related SNPs to support molecular epidemiology and complex disease mapping studies
-
Has reproduction · 50
Estimating and Correcting for Off-Target Cellular Contamination in Brain Cell Type Specific RNA-Seq Data.
PMID 33746712 · PMC7966716 · Frontiers in molecular neuroscience · 2021 · 6 claims · 7 setups
A computational method using high-quality scRNA-seq reference data can estimate per-sample, per-cell-type off-target contamination coefficients in sctRNA-seq datasets.
-
Has reproduction · 73
Genetic polyploid phasing from low-depth progeny samples.
PMID 35692633 · PMC9184567 · iScience · 2022 · 8 claims · 7 setups
WH-PPG phases polyploid parental samples by scoring informative variant pairs with a Bayesian log-likelihood model of progeny allele depths, clustering alleles by co-occurrence likelihood, and assigning clusters to haplotypes via interval scheduling
-
Full-text index only
SNiPer: improved SNP genotype calling for Affymetrix 10K GeneChip microarray data.
PMID 16262895 · PMC1280925 · BMC genomics · 2005 · 8 claims · 5 setups
Poorly performing SNPs (NoCall rate ≥25%) fail primarily due to inadequate training/localization of the MPAM statistical model call zone, not detection filter failure
-
Full-text index only
Prediction of catalytic residues using Support Vector Machine with selected protein sequence and structural properties.
PMID 16790052 · PMC1534064 · BMC bioinformatics · 2006 · 8 claims · 7 setups
The Sequential Minimal Optimization (SMO) SVM algorithm was the best-performing classifier among 26 WEKA classifiers for predicting catalytic residues
-
Full-text index only
Broad network-based predictability of Saccharomyces cerevisiae gene loss-of-function phenotypes.
PMID 18053250 · PMC2246260 · Genome biology · 2007 · 8 claims · 4 setups
Loss-of-function phenotypes in yeast are predictable from a gene's connections in a functional gene network via guilt-by-association.
-
Has reproduction · 95
MetaMap: an atlas of metatranscriptomic reads in human disease-related RNA-seq data.
PMID 29901703 · PMC6025204 · GigaScience · 2018 · 6 claims · 7 setups
A two-step 'omni' RNA-seq pipeline (MetaMap) combining STAR human alignment with CLARK-S metagenomic classification can quantify archaeal, bacterial, and viral reads from the non-human read fraction of human RNA-seq data
-
Has reproduction · 86
Single-cell transcriptome maps of myeloid blood cell lineages in Drosophila.
PMID 32900993 · PMC7479620 · Nature communications · 2020 · 8 claims · 8 setups
Single-cell RNA-seq of developing Drosophila lymph glands resolves heterogeneity of hemocytes and identifies major and sub cell types.
-
Full-text index only
siRNA screen of the human signaling proteome identifies the PtdIns(3,4,5)P3-mTOR signaling pathway as a primary regulator of transferrin uptake.
PMID 17640392 · PMC2323231 · Genome biology · 2007 · 8 claims · 8 setups
The PtdIns(3,4,5)P3-mTOR signaling pathway is a primary positive regulator of transferrin uptake.