Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
CGMIM: automated text-mining of Online Mendelian Inheritance in Man (OMIM) to identify genetically-associated cancers and candidate genes.
PMID 15796777 · PMC1274267 · BMC bioinformatics · 2005 · 8 claims · 2 setups
CGMIM is a Perl program that text-mines OMIM entries to identify cancer-gene associations and genetically-related cancer type pairs.
-
Full-text index only
CanPredict: a computational tool for predicting cancer-associated missense mutations.
PMID 17537827 · PMC1933186 · Nucleic acids research · 2007 · 8 claims · 7 setups
CanPredict is a web application providing public access to a random forest classifier that combines SIFT, LogR.E-value, and GOSS scores to predict whether a missense mutation is cancer-associated
-
Has reproduction · 78
Single duplex DNA sequencing with CODEC detects mutations with high sensitivity.
PMID 37106072 · PMC10181940 · Nature genetics · 2023 · 8 claims · 8 setups
CODEC concatenates both strands of an original DNA duplex into a single NGS read pair via an adapter quadruplex and strand-displacing extension, enabling single-duplex resolution
-
Has reproduction · 87
CoINcIDE: A framework for discovery of patient subtypes across multiple datasets.
PMID 26961683 · PMC4784276 · Genome medicine · 2016 · 8 claims · 4 setups
CoINcIDE is a novel framework for discovering patient subtypes across multiple datasets that requires no between-dataset transformations (e.g., batch correction)
-
Has reproduction · 87
Genetic demultiplexing of pooled single-cell RNA-sequencing samples in cancer facilitates effective experimental design.
PMID 34553212 · PMC8458035 · GigaScience · 2021 · 8 claims · 6 setups
Genetic variation-based demultiplexing tools can be effectively deployed on pooled scRNA-seq experimental designs in cancer tissue (HGSOC and lung adenocarcinoma) despite somatic variation.
-
Full-text index only
BayesRare: Bayesian mixture model for population-level rare cell type detection in multi-subject single-cell RNA sequencing data.
PMID 41632592 · PMC12867491 · Briefings in bioinformatics · 2026 · 8 claims · 4 setups
BayesRare is a hierarchical Bayesian mixture model framework for population-level rare cell type detection in multi-subject scRNA-seq data
-
Full-text index only
CARAT: a novel method for allelic detection of DNA copy number changes using high density oligonucleotide arrays.
PMID 16504045 · PMC1402331 · BMC bioinformatics · 2006 · 8 claims · 5 setups
CARAT is a novel algorithm that uses SNP probe intensity and genotype-based allelic dosage response in a regression framework to estimate allele-specific copy number genome-wide.
-
Full-text index only
BreakDancer: an algorithm for high-resolution mapping of genomic structural variation.
PMID 19668202 · PMC3661775 · Nature methods · 2009 · 8 claims · 8 setups
BreakDancer (BreakDancerMax + BreakDancerMini) is a software package that predicts a wide variety of structural variants including deletions, insertions, inversions, and intra/inter-chromosomal translocations from paired-end short-insert sequencing reads.
-
Full-text index only
AutoCSA, an algorithm for high throughput DNA sequence variant detection in cancer genomes.
PMID 17485433 · PMC5947781 · Bioinformatics (Oxford, England) · 2007 · 7 claims · 2 setups
AutoCSA is an automated algorithm, extended from the CSA protocol, that detects DNA sequence variants in cancer genomes with minimal manual intervention
-
Full-text index only
Function2Gene: a gene selection tool to increase the power of genetic association studies by utilizing public databases and expert knowledge.
PMID 18631403 · PMC2500032 · BMC bioinformatics · 2008 · 6 claims · 5 setups
Function2Gene is a set of Perl programs that queries public databases (NCBI, GeneCards, Harvester, with Uniprot/Ensembl also supported) using expert-selected keywords to rank genes by prior probability of disease association.
-
Full-text index only
PreTSA: computationally efficient modeling of temporal and spatial gene expression patterns.
PMID 41673899 · PMC12998178 · Genome biology · 2026 · 7 claims · 8 setups
PreTSA dramatically reduces computational time and memory versus GAM (Monocle, TSCAN) and PseudotimeDE for identifying temporally variable genes (TVGs) while producing highly similar results