Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
Bayesian survival analysis in genetic association studies.
PMID 18617538 · PMC2530885 · Bioinformatics (Oxford, England) · 2008 · 7 claims · 5 setups
A novel Bayesian method (BETA-Surv) extends prior case-control haplotype-clustering work to censored survival outcomes by clustering haplotypes via gene tree/perfect phylogeny topology and relative mutation age.
-
Has reproduction · 76
Tracing human genetic histories and natural selection with precise local ancestry inference.
PMID 40379651 · PMC12084304 · Nature communications · 2025 · 7 claims · 7 setups
Orchestra, a two-stage LAI method combining a recombination-distance base layer with a deep learning (convolutional + attention) smoothing module, outperforms RFmix, FLARE and Gnomix in precision and recall across simulated admixture generations.
-
Full-text index only
The origins of lactase persistence in Europe.
PMID 19714206 · PMC2722739 · PLoS computational biology · 2009 · 8 claims · 5 setups
The −13,910*T allele first underwent selection among dairying farmers around 7,500 years ago in a region between the central Balkans and central Europe, possibly linked to the Linearbandkeramik culture.
-
Full-text index only
Selecting additional tag SNPs for tolerating missing data in genotyping.
PMID 16259642 · PMC1316880 · BMC bioinformatics · 2005 · 7 claims · 6 setups
There exists a subset of SNPs (robust tag SNPs) that can distinguish all distinct haplotypes even when up to m SNPs are missing
-
Full-text index only
Cubic exact solutions for the estimation of pairwise haplotype frequencies: implications for linkage disequilibrium analyses and a web tool 'CubeX'.
PMID 17980034 · PMC2180187 · BMC bioinformatics · 2007 · 6 claims · 4 setups
CubeX, a Python program/web tool, computes the exact algebraic (Cardan/Nickalls) solution(s) of Hill's cubic equation to estimate pairwise haplotype frequencies, D', r2 and chi-square for each solution
-
Full-text index only
A statistical approach designed for finding mathematically defined repeats in shotgun data and determining the length distribution of clone-inserts.
PMID 15626332 · PMC5172250 · Genomics, proteomics & bioinformatics · 2003 · 8 claims · 6 setups
Repeats of different copy number have distinct probabilities of appearance in shotgun data, which can be modeled statistically to define recognition thresholds (MDRs) at different shotgun coverages.
-
Full-text index only
Analyses and comparison of accuracy of different genotype imputation methods.
PMID 18958166 · PMC2569208 · PloS one · 2008 · 8 claims · 3 setups
Stronger LD produces higher imputation accuracy rates for all five methods
-
Full-text index only
QuantiSNP: an Objective Bayes Hidden-Markov Model to detect and accurately map copy number variation using SNP genotyping data.
PMID 17341461 · PMC1874617 · Nucleic acids research · 2007 · 8 claims · 7 setups
QuantiSNP (OB-HMM) provides probabilistic quantification of copy number states and significantly improves accuracy of segmental aneuploidy identification and breakpoint mapping relative to existing tools (BeadStudio/Illumina)
-
Has reproduction · 67
binny: an automated binning algorithm to recover high-quality genomes from complex metagenomic datasets.
PMID 36239393 · PMC9677464 · Briefings in bioinformatics · 2022 · 8 claims · 8 setups
binny outperforms or is highly competitive with commonly used and state-of-the-art binning methods (MetaBAT2, MaxBin2, CONCOCT, VAMB, SemiBin, MetaDecoder)
-
Full-text index only
Absence of the TAP2 human recombination hotspot in chimpanzees.
PMID 15208713 · PMC423135 · PLoS biology · 2004 · 6 claims · 7 setups
The human TAP2 recombination hotspot is absent from the homologous region in western chimpanzees.
-
Full-text index only
An analysis of the feasibility of short read sequencing.
PMID 16275781 · PMC1278949 · Nucleic acids research · 2005 · 8 claims · 8 setups
Re-sequencing and de novo sequencing of the majority of a bacterial genome is possible with read lengths of 20-30 nt.
-
Full-text index only
The use of edge-betweenness clustering to investigate biological function in protein interaction networks.
PMID 15740614 · PMC555937 · BMC bioinformatics · 2005 · 8 claims · 7 setups
Edge-Betweenness clustering separates protein interaction graphs into subgraphs whose GO term distributions show significant correlations, revealing biologically meaningful functional modules.
-
Full-text index only
BOAT: Basic Oligonucleotide Alignment Tool.
PMID 19958483 · PMC2788372 · BMC genomics · 2009 · 7 claims · 3 setups
BOAT can accurately and efficiently map sequencing reads to a reference genome while handling several substitutions and indels simultaneously
-
Full-text index only
Population history and natural selection shape patterns of genetic variation in 132 genes.
PMID 15361935 · PMC515367 · PLoS biology · 2004 · 7 claims · 5 setups
Developed a rigorous computational approach that corrects for multiple hypothesis testing and models population demographic history to test for natural selection
-
Full-text index only
SNPAnalyzer: a web-based integrated workbench for single-nucleotide polymorphism analysis.
PMID 15980517 · PMC1160189 · Nucleic acids research · 2005 · 8 claims · 4 setups
SNPAnalyzer is an integrated web-based workbench that performs four statistical SNP analyses (Hardy-Weinberg equilibrium, haplotype estimation, linkage disequilibrium, and QTL analysis) in one common computational environment.
-
Full-text index only
G-quadruplexes: the beginning and end of UTRs.
PMID 18832370 · PMC2577360 · Nucleic acids research · 2008 · 8 claims · 5 setups
UTRs show significant strand asymmetry with C-PQS more common than G-PQS, consistent with general depletion of G-quadruplex-forming RNA
-
Full-text index only
What can genome-wide association studies tell us about the genetics of common disease?
PMID 18454206 · PMC2323402 · PLoS genetics · 2008 · 8 claims · 4 setups
Apparent patterns of common, low-effect disease-associated alleles largely reflect statistical power of studies rather than the true underlying distribution of disease variants
-
Full-text index only
Imputation-based analysis of association studies: candidate regions and quantitative traits.
PMID 17676998 · PMC1934390 · PLoS genetics · 2007 · 8 claims · 2 setups
Imputation-based Bayesian regression increases power to detect association compared with standard single-SNP tests, even when the causal variant is directly typed
-
Full-text index only
Multiplexed discovery of sequence polymorphisms using base-specific cleavage and MALDI-TOF MS.
PMID 15731331 · PMC549577 · Nucleic acids research · 2005 · 8 claims · 7 setups
Multiplexed base-specific cleavage/MALDI-TOF MS (Multiplexed Comparative Sequence Analysis) enables simultaneous discovery of sequence polymorphisms across multiple target regions
-
Full-text index only
High resolution array-CGH analysis of single cells.
PMID 17178751 · PMC1807964 · Nucleic acids research · 2007 · 7 claims · 7 setups
Single copy number changes as small as 8.3 Mb can be detected reliably in single cells using GenomePlex WGA combined with high-resolution tiling-path array-CGH.