Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
Calculating expected DNA remnants from ancient founding events in human population genetics.
PMID 18928554 · PMC2588638 · BMC genetics · 2008 · 8 claims · 3 setups
Genetic parameters (native/migrant population size, mutation rate, generations since admixture) strongly determine the final frequency of migrant alleles detectable today.
-
Full-text index only
Is replication the gold standard for validating genome-wide association findings?
PMID 19112512 · PMC2605260 · PloS one · 2008 · 8 claims · 4 setups
The probability of replicating a specific GWA-identified variant decreases as the number of independent GWA/replication studies increases, when individual study power is less than 100%.
-
Full-text index only
Accuracy of predicting the genetic risk of disease using a genome-wide approach.
PMID 18852893 · PMC2561058 · PloS one · 2008 · 8 claims · 4 setups
Deterministic equations can predict the accuracy (r_gĝ) of genome-wide genetic risk/value prediction for continuous, dichotomous, and case-control study designs.
-
Full-text index only
Bayesian survival analysis in genetic association studies.
PMID 18617538 · PMC2530885 · Bioinformatics (Oxford, England) · 2008 · 7 claims · 5 setups
A novel Bayesian method (BETA-Surv) extends prior case-control haplotype-clustering work to censored survival outcomes by clustering haplotypes via gene tree/perfect phylogeny topology and relative mutation age.
-
Full-text index only
Screening large-scale association study data: exploiting interactions using random forests.
PMID 15588316 · PMC545646 · BMC genetics · 2004 · 7 claims · 3 setups
Random forest importance measure significantly outperforms the Fisher Exact test as a screening tool when risk SNPs interact.
-
Has reproduction · 98
Uncertainty in the mating strategy of honeybees causes bias and unreliability in the estimates of genetic parameters.
PMID 38632535 · PMC11022492 · Genetics, selection, evolution : GSE · 2024 · 7 claims · 3 setups
The most precise estimates of genetic parameters and genetic trends are obtained when breeding queens are mated with drones of a single DPQ that is correctly assigned in the pedigree (SS mating).
-
Full-text index only
PedGenie: an analysis approach for genetic association testing in extended pedigrees and genealogies of arbitrary size.
PMID 16620382 · PMC1459209 · BMC bioinformatics · 2006 · 7 claims · 3 setups
PedGenie is a valid, flexible statistical tool for genetic association analysis in pedigrees of arbitrary size and structure using Monte Carlo significance testing
-
Full-text index only
Application of two machine learning algorithms to genetic association studies in the presence of covariates.
PMID 19014573 · PMC2620353 · BMC genetics · 2008 · 8 claims · 3 setups
The relative performance of RF and MARS for detecting genotype-trait associations depends on both the strategy used to handle covariates and the true underlying model of association (e.g., confounding vs. mediation vs. interaction).
-
Full-text index only
Stability analysis of mixtures of mutagenetic trees.
PMID 18366778 · PMC2335279 · BMC bioinformatics · 2008 · 7 claims · 5 setups
Mutagenetic trees mixture models capture multiple alternative pathways of ordered accumulation of genetic events (e.g., HIV resistance mutations, cancer chromosomal aberrations).
-
Full-text index only
A method for detecting epistasis in genome-wide studies using case-control multi-locus association analysis.
PMID 18667089 · PMC2533022 · BMC genomics · 2008 · 7 claims · 2 setups
HFCC is a method/software for genome-wide epistasis detection using case-control multi-locus association analysis, combining a fast computing algorithm with flexibility to test a variety of epistatic models.
-
Has reproduction · 83
SIRE 2.0: a novel method for estimating polygenic host effects underlying infectious disease transmission, and analytical expressions for prediction accuracies.
PMID 40169992 · PMC11963337 · Genetics, selection, evolution : GSE · 2025 · 8 claims · 2 setups
SIRE 2.0 is a novel Bayesian methodology and software tool for estimating polygenic contributions (variance components and additive genetic effects) to host susceptibility, infectivity and recoverability from temporal epidemic data using pedigree/genomic relationship matrices.
-
Full-text index only
Imputation of missing genotypes: an empirical evaluation of IMPUTE.
PMID 19077279 · PMC2636842 · BMC genetics · 2008 · 8 claims · 7 setups
IMPUTE achieves 97% median genotype imputation accuracy in Caucasian (NNC) subjects when <10% of SNPs are untyped
-
Full-text index only
Simultaneous analysis of all SNPs in genome-wide and re-sequencing association studies.
PMID 18654633 · PMC2464715 · PLoS genetics · 2008 · 8 claims · 5 setups
A Bayesian-inspired penalised maximum likelihood stochastic search method can simultaneously analyse all SNPs (up to 500K) from a GWA study in a few hours on a desktop workstation
-
Full-text index only
The origins of lactase persistence in Europe.
PMID 19714206 · PMC2722739 · PLoS computational biology · 2009 · 8 claims · 5 setups
The −13,910*T allele first underwent selection among dairying farmers around 7,500 years ago in a region between the central Balkans and central Europe, possibly linked to the Linearbandkeramik culture.
-
Full-text index only
Size matters: just how big is BIG?: Quantifying realistic sample size requirements for human genome epidemiology.
PMID 18676414 · PMC2639365 · International journal of epidemiology · 2009 · 7 claims · 2 setups
Conventional power calculations for case-control studies disregard analytic complexity (e.g. clinical assessment errors, unmeasured aetiological determinants) and can seriously underestimate true sample size requirements
-
Has reproduction · 93
A comparative study on recombination activity in cattle.
PMID 41942849 · PMC13067647 · Genetics, selection, evolution : GSE · 2026 · 8 claims · 8 setups
Genotype data with high systematic missingness across breeds and arrays can be streamlined and analysed with three complementary recombination-estimation approaches (HMM-based LINKPHASE3, deterministic hsphase, likelihood-based hsrecombi)
-
Has reproduction · 61
TEMP: a computational method for analyzing transposable element polymorphism in populations.
PMID 24753423 · PMC4066757 · Nucleic acids research · 2014 · 8 claims · 8 setups
TEMP combines pair-end (discordant) read and split (soft-clipped) read information to identify both presence and absence of TE insertions in genomic DNA from heterogeneous/pooled samples.
-
Full-text index only
Calibrating the performance of SNP arrays for whole-genome association studies.
PMID 18584036 · PMC2432039 · PLoS genetics · 2008 · 8 claims · 7 setups
Previous SNP array genetic coverage estimates are inflated due to SNP overfitting and sample overfitting, since they were evaluated on the same HapMap SNPs/individuals used to design the arrays.
-
Full-text index only
Inferring human colonization history using a copying model.
PMID 18497854 · PMC2367454 · PLoS genetics · 2008 · 8 claims · 6 setups
A copying-model approach using SNP haplotype sharing can infer both the order of population founding and the donor populations contributing ancestry to each new population.
-
Full-text index only
Genome-wide prediction of functional gene-gene interactions inferred from patterns of genetic differentiation in mice and men.
PMID 18270580 · PMC2217631 · PloS one · 2008 · 8 claims · 6 setups
Pairs of unlinked SNPs showing excess genetic differentiation (LD in mouse RILs, Fst in human populations) beyond what simulations/coalescent models predict by chance represent candidate functionally interacting (epistatic) gene pairs.