Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Has reproduction · 65
FusionQ: a novel approach for gene fusion detection and quantification from paired-end RNA-Seq.
PMID 23768108 · PMC3691734 · BMC bioinformatics · 2013 · 8 claims · 8 setups
FusionQ is a novel tool that detects gene fusions, constructs chimerical transcript structures, and estimates their abundances from paired-end RNA-Seq data.
-
Full-text index only
Meta-analysis of inter-species liver co-expression networks elucidates traits associated with common human diseases.
PMID 20019805 · PMC2787626 · PLoS computational biology · 2009 · 8 claims · 8 setups
A novel semi-parametric meta-analysis method (based on a gene-centric Glass's d effect size) outperforms existing parametric and non-parametric meta-analysis methods at identifying functionally coherent gene pairs across species.
-
Full-text index only
Evolutionary algorithms for the selection of single nucleotide polymorphisms.
PMID 12875658 · PMC183839 · BMC bioinformatics · 2003 · 8 claims · 3 setups
Evolutionary algorithms are well suited to multiobjective optimization problems with large, intractable search spaces such as SNP selection, unlike exact methods (exhaustive enumeration) or single-objective search techniques (tabu search, simulated annealing).
-
Full-text index only
Analysis of concordance of different haplotype block partitioning algorithms.
PMID 16356172 · PMC1343594 · BMC bioinformatics · 2005 · 7 claims · 7 setups
Each block partitioning algorithm infers blocks differing in number, size, and coverage under different SNP density and allele frequency conditions.
-
Has reproduction · 61
TEMP: a computational method for analyzing transposable element polymorphism in populations.
PMID 24753423 · PMC4066757 · Nucleic acids research · 2014 · 8 claims · 8 setups
TEMP combines pair-end (discordant) read and split (soft-clipped) read information to identify both presence and absence of TE insertions in genomic DNA from heterogeneous/pooled samples.
-
Full-text index only
Antigenic diversity, transmission mechanisms, and the evolution of pathogens.
PMID 19847288 · PMC2759524 · PLoS computational biology · 2009 · 8 claims · 3 setups
Three distinct infection types (A, B, C) emerge as maxima in the pathogen fitness landscape, each with characteristic within-host dynamics, contact network structure, and transmission mode
-
Full-text index only
A response to Yu et al. "A forward-backward fragment assembling algorithm for the identification of genomic amplification and deletion breakpoints using high-density single nucleotide polymorphism (SNP) array", BMC Bioinformatics 2007, 8: 145.
PMID 17939873 · PMC2222656 · BMC bioinformatics · 2007 · 8 claims · 4 setups
Yu et al.'s original comparison ran RJaCGH's MCMC sampler for a severely insufficient number of iterations (50 burn-in, 500 total)
-
Full-text index only
Size matters: just how big is BIG?: Quantifying realistic sample size requirements for human genome epidemiology.
PMID 18676414 · PMC2639365 · International journal of epidemiology · 2009 · 7 claims · 2 setups
Conventional power calculations for case-control studies disregard analytic complexity (e.g. clinical assessment errors, unmeasured aetiological determinants) and can seriously underestimate true sample size requirements
-
Full-text index only
Computational tradeoffs in multiplex PCR assay design for SNP genotyping.
PMID 16042802 · PMC1190169 · BMC genomics · 2005 · 7 claims · 6 setups
Achieving high-multiplexing/high-coverage multiplex PCR designs is subject to a computational phase transition as the SNP-pair compatibility probability crosses a critical threshold
-
Full-text index only
Application of two machine learning algorithms to genetic association studies in the presence of covariates.
PMID 19014573 · PMC2620353 · BMC genetics · 2008 · 8 claims · 3 setups
The relative performance of RF and MARS for detecting genotype-trait associations depends on both the strategy used to handle covariates and the true underlying model of association (e.g., confounding vs. mediation vs. interaction).
-
Has reproduction · 75
Revealing the critical state and identifying individualized dynamic network biomarker for type 2 diabetes through advanced analysis methods on individual basis.
PMID 39890881 · PMC11785715 · Scientific reports · 2025 · 8 claims · 5 setups
sJSD, NIG, and TNFE methods can detect critical states/tipping points before disease deterioration using only a single sample
-
Has reproduction · 76
Tracing human genetic histories and natural selection with precise local ancestry inference.
PMID 40379651 · PMC12084304 · Nature communications · 2025 · 7 claims · 7 setups
Orchestra, a two-stage LAI method combining a recombination-distance base layer with a deep learning (convolutional + attention) smoothing module, outperforms RFmix, FLARE and Gnomix in precision and recall across simulated admixture generations.
-
Full-text index only
Benchmarking tools for the alignment of functional noncoding DNA.
PMID 14736341 · PMC344529 · BMC bioinformatics · 2004 · 8 claims · 4 setups
Global alignment tools (Avid, ClustalW, Lagan, Needle, DiAlign-G) typically have higher sensitivity over entire noncoding sequences and within constrained blocks than local tools
-
Full-text index only
On the analysis of glycomics mass spectrometry data via the regularized area under the ROC curve.
PMID 18076765 · PMC2211327 · BMC bioinformatics · 2007 · 8 claims · 4 setups
The TGDR-AUC algorithm regularizes the empirical AUC by replacing the non-differentiable 0-1 loss with a smooth sigmoid surrogate function and applies constrained threshold gradient descent regularization
-
Full-text index only
Direct maximum parsimony phylogeny reconstruction from genotype data.
PMID 18053244 · PMC2222657 · BMC bioinformatics · 2007 · 6 claims · 4 setups
The paper presents the first practical method for computing maximum parsimony phylogenies directly from genotype data, using integer linear programming.
-
Full-text index only
Optimality driven nearest centroid classification from genomic data.
PMID 17912341 · PMC1991588 · PloS one · 2007 · 7 claims · 5 setups
A theoretical result determines the subset of features of a given size that minimizes the misclassification rate for a nearest-centroid (LDA) classifier, based on equation (4).
-
Full-text index only
Simultaneous analysis of all SNPs in genome-wide and re-sequencing association studies.
PMID 18654633 · PMC2464715 · PLoS genetics · 2008 · 8 claims · 5 setups
A Bayesian-inspired penalised maximum likelihood stochastic search method can simultaneously analyse all SNPs (up to 500K) from a GWA study in a few hours on a desktop workstation
-
Has reproduction · 68
Bayesian transcriptome assembly.
PMID 25367074 · PMC4397945 · Genome biology · 2014 · 8 claims · 8 setups
Bayesembler, a probabilistic transcriptome assembler built on a Bayesian model of the RNA sequencing process with Gibbs sampling over expressed candidates, abundances and read assignments, is introduced.
-
Full-text index only
BFAST: an alignment tool for large scale genome resequencing.
PMID 19907642 · PMC2770639 · PloS one · 2009 · 7 claims · 4 setups
BFAST is a new algorithm and freely available software tool for aligning large-scale short-read sequencing data to large reference genomes with user-customizable speed and accuracy
-
Full-text index only
Operon information improves gene expression estimation for cDNA microarrays.
PMID 16630355 · PMC1513396 · BMC genomics · 2006 · 7 claims · 3 setups
A hierarchical Bayesian model that borrows expression information from other genes within the same operon improves estimation of relative transcript levels.