Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
Simultaneous analysis of all SNPs in genome-wide and re-sequencing association studies.
PMID 18654633 · PMC2464715 · PLoS genetics · 2008 · 8 claims · 5 setups
A Bayesian-inspired penalised maximum likelihood stochastic search method can simultaneously analyse all SNPs (up to 500K) from a GWA study in a few hours on a desktop workstation
-
Full-text index only
Testing groups of genomic locations for enrichment in disease loci using linkage scan data: a method for hypothesis testing.
PMID 16848972 · PMC3525155 · Human genomics · 2006 · 8 claims · 2 setups
A method testing enrichment of a group of genomic locations for disease loci by comparing the average NPL score of the group to a null distribution from randomly drawn groups of equal size
-
Has reproduction · 89
TrEMOLO: accurate transposable element allele frequency estimation using long-read sequencing data combining assembly and mapping-based approaches.
PMID 37013657 · PMC10069131 · Genome biology · 2023 · 6 claims · 6 setups
TrEMOLO combines an assembly-based INSIDER module and a mapping-based OUTSIDER module to detect TE insertions/deletions from long-read sequencing data and estimate their allele frequency
-
Full-text index only
Duplication count distributions in DNA sequences.
PMID 19256873 · PMC3121164 · Physical review. E, Statistical, nonlinear, and soft matter physics · 2008 · 8 claims · 8 setups
Duplication count distributions N(c) for complex 40-mers show power-law-like decay for c roughly 3 to 50 (or higher) across human, C. elegans, A. thaliana, and D. melanogaster genomes.
-
Full-text index only
A statistical approach designed for finding mathematically defined repeats in shotgun data and determining the length distribution of clone-inserts.
PMID 15626332 · PMC5172250 · Genomics, proteomics & bioinformatics · 2003 · 8 claims · 6 setups
Repeats of different copy number have distinct probabilities of appearance in shotgun data, which can be modeled statistically to define recognition thresholds (MDRs) at different shotgun coverages.
-
Has reproduction · 43
TransFlow: a Snakemake workflow for transmission analysis of Mycobacterium tuberculosis whole-genome sequencing data.
PMID 36469333 · PMC9825751 · Bioinformatics (Oxford, England) · 2023 · 8 claims · 8 setups
TransFlow is a Snakemake- and Conda-based workflow that combines state-of-the-art tools into a single, fast, scalable pipeline for MTBC WGS-based transmission analysis.
-
Full-text index only
Assessing the gene space in draft genomes.
PMID 19042974 · PMC2615622 · Nucleic acids research · 2009 · 6 claims · 7 setups
The proportion of mapped CEGs in a draft genome assembly is a useful metric for describing gene space completeness, complementing N50 and x-fold coverage.
-
Has reproduction · 45
Identifying and classifying trait linked polymorphisms in non-reference species by walking coloured de bruijn graphs.
PMID 23536903 · PMC3607606 · PloS one · 2013 · 8 claims · 9 setups
Bubbleparse detects sequence variants directly from NGS reads without a reference genome, using the coloured de Bruijn graph implementation of Cortex plus a new depth-first bubble-finding module.
-
Full-text index only
PBAT: a comprehensive software package for genome-wide association analysis of complex family-based studies.
PMID 15814068 · PMC3525120 · Human genomics · 2005 · 8 claims · 1 setups
PBAT provides comprehensive tools for family-based association analysis, including nuclear families with missing parental genotypes, extended pedigrees, SNP and haplotype analysis, quantitative/qualitative/multivariate/longitudinal traits and time-to-onset phenotypes
-
Full-text index only
Searching for SNPs with cloud computing.
PMID 19930550 · PMC3091327 · Genome biology · 2009 · 8 claims · 4 setups
Crossbow combines the Bowtie short-read aligner and SOAPsnp SNP caller into a seamless, automatic Hadoop/MapReduce pipeline for whole-genome resequencing analysis
-
Full-text index only
BFAST: an alignment tool for large scale genome resequencing.
PMID 19907642 · PMC2770639 · PloS one · 2009 · 7 claims · 4 setups
BFAST is a new algorithm and freely available software tool for aligning large-scale short-read sequencing data to large reference genomes with user-customizable speed and accuracy
-
Full-text index only
The signal in the genomes.
PMID 16683016 · PMC1447653 · PLoS computational biology · 2006 · 7 claims · 3 setups
A high breakpoint reuse rate in the output of rearrangement algorithms indicates loss of historical signal, not good evidence for genomic fragile regions
-
Has reproduction · 61
TEMP: a computational method for analyzing transposable element polymorphism in populations.
PMID 24753423 · PMC4066757 · Nucleic acids research · 2014 · 8 claims · 8 setups
TEMP combines pair-end (discordant) read and split (soft-clipped) read information to identify both presence and absence of TE insertions in genomic DNA from heterogeneous/pooled samples.
-
Full-text index only
Iterative pruning PCA improves resolution of highly structured populations.
PMID 19930644 · PMC2790469 · BMC bioinformatics · 2009 · 7 claims · 7 setups
ipPCA is a novel algorithm that assigns individuals to subpopulations and infers the total number of subpopulations (K) present in genotypic data
-
Full-text index only
BreakDancer: an algorithm for high-resolution mapping of genomic structural variation.
PMID 19668202 · PMC3661775 · Nature methods · 2009 · 8 claims · 8 setups
BreakDancer (BreakDancerMax + BreakDancerMini) is a software package that predicts a wide variety of structural variants including deletions, insertions, inversions, and intra/inter-chromosomal translocations from paired-end short-insert sequencing reads.
-
Full-text index only
A note on generalized Genome Scan Meta-Analysis statistics.
PMID 15717930 · PMC551600 · BMC bioinformatics · 2005 · 7 claims · 3 setups
An Edgeworth series approximation to the null distribution of the weighted GSMA statistic provides a more accurate representation than the normal approximation, especially in the tails
-
Full-text index only
Evolutionary algorithms for the selection of single nucleotide polymorphisms.
PMID 12875658 · PMC183839 · BMC bioinformatics · 2003 · 8 claims · 3 setups
Evolutionary algorithms are well suited to multiobjective optimization problems with large, intractable search spaces such as SNP selection, unlike exact methods (exhaustive enumeration) or single-objective search techniques (tabu search, simulated annealing).
-
Has reproduction · 67
binny: an automated binning algorithm to recover high-quality genomes from complex metagenomic datasets.
PMID 36239393 · PMC9677464 · Briefings in bioinformatics · 2022 · 8 claims · 8 setups
binny outperforms or is highly competitive with commonly used and state-of-the-art binning methods (MetaBAT2, MaxBin2, CONCOCT, VAMB, SemiBin, MetaDecoder)
-
Full-text index only
Imputation of missing genotypes: an empirical evaluation of IMPUTE.
PMID 19077279 · PMC2636842 · BMC genetics · 2008 · 8 claims · 7 setups
IMPUTE achieves 97% median genotype imputation accuracy in Caucasian (NNC) subjects when <10% of SNPs are untyped
-
Has reproduction · 42
KAGE: fast alignment-free graph-based genotyping of SNPs and short indels.
PMID 36195962 · PMC9531401 · Genome biology · 2022 · 7 claims · 7 setups
KAGE combines population-based kmer count modeling with single-variant prior adjustment into an alignment-free genotyper that matches the accuracy of the best existing alignment-free genotypers while being an order of magnitude faster.