Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
SPRINT: a new parallel framework for R.
PMID 19114001 · PMC2628907 · BMC bioinformatics · 2008 · 8 claims · 1 setups
SPRINT is a prototype R framework that wraps parallelised functions, requiring minimal modification to existing sequential R scripts and no parallel programming expertise from the user
-
Has reproduction · 94
Deep learning from phylogenies to uncover the epidemiological dynamics of outbreaks.
PMID 35794110 · PMC9258765 · Nature communications · 2022 · 8 claims · 5 setups
Deep learning (FFNN-SS and CNN-CBLV) enables accurate and fast likelihood-free estimation of epidemiological parameters and model selection from phylogenies
-
Has reproduction · 42
CanCellCap: robust cancer cell capture across tissue types on single-cell RNA-seq data by multi-domain learning.
PMID 40739511 · PMC12312500 · BMC biology · 2025 · 8 claims · 8 setups
CanCellCap, a multi-domain learning framework integrating domain adversarial learning and Mixture of Experts, identifies cancer cells across all tissues, cancers, and sequencing platforms by extracting tissue-common and tissue-specific gene expression patterns.
-
Full-text index only
Detecting purely epistatic multi-locus interactions by an omnibus permutation test on ensembles of two-locus analyses.
PMID 19761607 · PMC2759961 · BMC bioinformatics · 2009 · 8 claims · 5 setups
2LOmb performs an omnibus permutation test on ensembles of two-locus analyses via a four-step algorithm (two-locus analysis, permutation test, global p-value determination, progressive ensemble search)
-
Full-text index only
Screening large-scale association study data: exploiting interactions using random forests.
PMID 15588316 · PMC545646 · BMC genetics · 2004 · 7 claims · 3 setups
Random forest importance measure significantly outperforms the Fisher Exact test as a screening tool when risk SNPs interact.
-
Has reproduction · 85
ScLRTC: imputation for single-cell RNA-seq data via low-rank tensor completion.
PMID 34844559 · PMC8628418 · BMC genomics · 2021 · 8 claims · 8 setups
scLRTC imputes dropout entries closest to the original expression values on simulated datasets, outperforming other state-of-the-art methods by SSE and PCC.
-
Has reproduction · 65
FusionQ: a novel approach for gene fusion detection and quantification from paired-end RNA-Seq.
PMID 23768108 · PMC3691734 · BMC bioinformatics · 2013 · 8 claims · 8 setups
FusionQ is a novel tool that detects gene fusions, constructs chimerical transcript structures, and estimates their abundances from paired-end RNA-Seq data.
-
Has reproduction · 61
TEMP: a computational method for analyzing transposable element polymorphism in populations.
PMID 24753423 · PMC4066757 · Nucleic acids research · 2014 · 8 claims · 8 setups
TEMP combines pair-end (discordant) read and split (soft-clipped) read information to identify both presence and absence of TE insertions in genomic DNA from heterogeneous/pooled samples.
-
Full-text index only
Inferring human colonization history using a copying model.
PMID 18497854 · PMC2367454 · PLoS genetics · 2008 · 8 claims · 6 setups
A copying-model approach using SNP haplotype sharing can infer both the order of population founding and the donor populations contributing ancestry to each new population.
-
Has reproduction · 68
Bayesian transcriptome assembly.
PMID 25367074 · PMC4397945 · Genome biology · 2014 · 8 claims · 8 setups
Bayesembler, a probabilistic transcriptome assembler built on a Bayesian model of the RNA sequencing process with Gibbs sampling over expressed candidates, abundances and read assignments, is introduced.
-
Full-text index only
PedGenie: an analysis approach for genetic association testing in extended pedigrees and genealogies of arbitrary size.
PMID 16620382 · PMC1459209 · BMC bioinformatics · 2006 · 7 claims · 3 setups
PedGenie is a valid, flexible statistical tool for genetic association analysis in pedigrees of arbitrary size and structure using Monte Carlo significance testing
-
Full-text index only
Testing groups of genomic locations for enrichment in disease loci using linkage scan data: a method for hypothesis testing.
PMID 16848972 · PMC3525155 · Human genomics · 2006 · 8 claims · 2 setups
A method testing enrichment of a group of genomic locations for disease loci by comparing the average NPL score of the group to a null distribution from randomly drawn groups of equal size
-
Full-text index only
Optimized mixed Markov models for motif identification.
PMID 16749929 · PMC1534070 · BMC bioinformatics · 2006 · 8 claims · 4 setups
OMiMa can incorporate more than NNSplice's pairwise dependencies
-
Full-text index only
On the analysis of glycomics mass spectrometry data via the regularized area under the ROC curve.
PMID 18076765 · PMC2211327 · BMC bioinformatics · 2007 · 8 claims · 4 setups
The TGDR-AUC algorithm regularizes the empirical AUC by replacing the non-differentiable 0-1 loss with a smooth sigmoid surrogate function and applies constrained threshold gradient descent regularization
-
Full-text index only
Optimality driven nearest centroid classification from genomic data.
PMID 17912341 · PMC1991588 · PloS one · 2007 · 7 claims · 5 setups
A theoretical result determines the subset of features of a given size that minimizes the misclassification rate for a nearest-centroid (LDA) classifier, based on equation (4).
-
Full-text index only
Operon information improves gene expression estimation for cDNA microarrays.
PMID 16630355 · PMC1513396 · BMC genomics · 2006 · 7 claims · 3 setups
A hierarchical Bayesian model that borrows expression information from other genes within the same operon improves estimation of relative transcript levels.
-
Full-text index only
Predicting survival outcomes using subsets of significant genes in prognostic marker studies with microarrays.
PMID 16549007 · PMC1544357 · BMC bioinformatics · 2006 · 7 claims · 2 setups
A methodology combining Cox proportional hazards models with a compound covariate, cross-validated log partial likelihood (ACVL) for predictive accuracy, and permutation-based significance testing can identify an optimal subset of significant genes for survival prediction
-
Full-text index only
Expression profiling of drug response--from genes to pathways.
PMID 17117610 · PMC3181826 · Dialogues in clinical neuroscience · 2006 · 8 claims · 8 setups
Understanding individual response to a drug (efficacy/tolerability) is the major bottleneck in current drug development and clinical trials.
-
Full-text index only
Quantitative analysis of single nucleotide polymorphisms within copy number variation.
PMID 19093001 · PMC2600609 · PloS one · 2008 · 8 claims · 2 setups
Copy number variation is a major factor in HWE violation for SNPs with small minor allele frequency, large sample size, and 0-1% genotyping error rate
-
Has reproduction · 87
Forseti: a mechanistic and predictive model of the splicing status of scRNA-seq reads.
PMID 38940130 · PMC11256924 · Bioinformatics (Oxford, England) · 2024 · 7 claims · 5 setups
Forseti is the first probabilistic model for resolving the splicing status of exonic scRNA-seq reads by scoring putative fragments linking read alignments to proximate priming sites