Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
Predicting survival outcomes using subsets of significant genes in prognostic marker studies with microarrays.
PMID 16549007 · PMC1544357 · BMC bioinformatics · 2006 · 7 claims · 2 setups
A methodology combining Cox proportional hazards models with a compound covariate, cross-validated log partial likelihood (ACVL) for predictive accuracy, and permutation-based significance testing can identify an optimal subset of significant genes for survival prediction
-
Full-text index only
Size matters: just how big is BIG?: Quantifying realistic sample size requirements for human genome epidemiology.
PMID 18676414 · PMC2639365 · International journal of epidemiology · 2009 · 7 claims · 2 setups
Conventional power calculations for case-control studies disregard analytic complexity (e.g. clinical assessment errors, unmeasured aetiological determinants) and can seriously underestimate true sample size requirements
-
Has reproduction · 98
Uncertainty in the mating strategy of honeybees causes bias and unreliability in the estimates of genetic parameters.
PMID 38632535 · PMC11022492 · Genetics, selection, evolution : GSE · 2024 · 7 claims · 3 setups
The most precise estimates of genetic parameters and genetic trends are obtained when breeding queens are mated with drones of a single DPQ that is correctly assigned in the pedigree (SS mating).
-
Full-text index only
Application of two machine learning algorithms to genetic association studies in the presence of covariates.
PMID 19014573 · PMC2620353 · BMC genetics · 2008 · 8 claims · 3 setups
The relative performance of RF and MARS for detecting genotype-trait associations depends on both the strategy used to handle covariates and the true underlying model of association (e.g., confounding vs. mediation vs. interaction).
-
Full-text index only
Quantitative analysis of single nucleotide polymorphisms within copy number variation.
PMID 19093001 · PMC2600609 · PloS one · 2008 · 8 claims · 2 setups
Copy number variation is a major factor in HWE violation for SNPs with small minor allele frequency, large sample size, and 0-1% genotyping error rate
-
Has reproduction · 53
spliceJAC: transition genes and state-specific gene regulation from single-cell transcriptome data.
PMID 36321549 · PMC9627675 · Molecular systems biology · 2022 · 8 claims · 6 setups
spliceJAC quantifies multivariate mRNA splicing from unspliced/spliced count matrices to construct cell state-specific gene-gene (Jacobian) interaction matrices.
-
Has reproduction · 40
DeepGSEA: explainable deep gene set enrichment analysis for single-cell transcriptomic data.
PMID 38950178 · PMC11236288 · Bioinformatics (Oxford, England) · 2024 · 8 claims · 2 setups
DeepGSEA is an explainable deep gene set enrichment analysis method built on interpretable, prototype-based neural networks.
-
Has reproduction · 10
RADAR: differential analysis of MeRIP-seq data with a random effect model.
PMID 31870409 · PMC6927177 · Genome biology · 2019 · 8 claims · 6 setups
RADAR is a novel analytical tool for differential methylation analysis of MeRIP-seq data combining gene-level INPUT normalization with a Poisson random effect model.
-
Full-text index only
A note on generalized Genome Scan Meta-Analysis statistics.
PMID 15717930 · PMC551600 · BMC bioinformatics · 2005 · 7 claims · 3 setups
An Edgeworth series approximation to the null distribution of the weighted GSMA statistic provides a more accurate representation than the normal approximation, especially in the tails
-
Full-text index only
Inferring human colonization history using a copying model.
PMID 18497854 · PMC2367454 · PLoS genetics · 2008 · 8 claims · 6 setups
A copying-model approach using SNP haplotype sharing can infer both the order of population founding and the donor populations contributing ancestry to each new population.
-
Has reproduction · 85
PowerBacGWAS: a computational pipeline to perform power calculations for bacterial genome-wide association studies.
PMID 35338232 · PMC8956664 · Communications biology · 2022 · 8 claims · 8 setups
Two computational approaches (sub-sampling and phenotype-simulation) can be implemented to perform power calculations for bacterial GWAS using existing genome collections, packaged as the PowerBacGWAS pipeline
-
Full-text index only
A statistical model to identify differentially expressed proteins in 2D PAGE gels.
PMID 19763172 · PMC2734266 · PLoS computational biology · 2009 · 7 claims · 5 setups
A mixture likelihood model incorporating both detected and non-detected proteins has higher statistical power to detect differential expression than standard approaches like the Student's t-test.
-
Has reproduction · 62
Equivalent change enrichment analysis: assessing equivalent and inverse change in biological pathways between diverse experiments.
PMID 32093613 · PMC7041296 · BMC genomics · 2020 · 7 claims · 3 setups
The Equivalent Change Index (ECI), a gene-level statistic ranging from -1 to 1, quantifies whether a gene was changed to the same (1) or completely opposite (-1) degree across two experiments relative to their controls.
-
Full-text index only
Digital evolution.
PMID 14551915 · PMC212697 · PLoS biology · 2003 · 7 claims · 5 setups
Digital organisms (Avidians) evolved from simple self-replicators, through an unexpected transitional form, to complex performers of many logic functions, with the full genealogy traceable and no missing links.
-
Full-text index only
Population history and natural selection shape patterns of genetic variation in 132 genes.
PMID 15361935 · PMC515367 · PLoS biology · 2004 · 7 claims · 5 setups
Developed a rigorous computational approach that corrects for multiple hypothesis testing and models population demographic history to test for natural selection
-
Full-text index only
Design and analysis issues in genome-wide somatic mutation studies of cancer.
PMID 18692126 · PMC2820387 · Genomics · 2009 · 6 claims · 4 setups
Two-stage (discovery + validation) sequencing designs efficiently allocate resources and can produce highly informative candidate driver gene lists even with relatively small sample sizes.
-
Full-text index only
Stability analysis of mixtures of mutagenetic trees.
PMID 18366778 · PMC2335279 · BMC bioinformatics · 2008 · 7 claims · 5 setups
Mutagenetic trees mixture models capture multiple alternative pathways of ordered accumulation of genetic events (e.g., HIV resistance mutations, cancer chromosomal aberrations).
-
Has reproduction · 83
SIRE 2.0: a novel method for estimating polygenic host effects underlying infectious disease transmission, and analytical expressions for prediction accuracies.
PMID 40169992 · PMC11963337 · Genetics, selection, evolution : GSE · 2025 · 8 claims · 2 setups
SIRE 2.0 is a novel Bayesian methodology and software tool for estimating polygenic contributions (variance components and additive genetic effects) to host susceptibility, infectivity and recoverability from temporal epidemic data using pedigree/genomic relationship matrices.
-
Has reproduction
Genomic prediction based on selective linkage disequilibrium pruning of low-coverage whole-genome sequence variants in a pure Duroc population.
PMID 37853325 · PMC10583454 · Genetics, selection, evolution : GSE · 2023 · 8 claims · 6 setups
Selective linkage disequilibrium pruning (SLDP) refines whole-genome SNP sets using GWAS prior information to improve genomic prediction accuracy.
-
Full-text index only
Genome-wide scans for loci under selection in humans.
PMID 16004726 · PMC3525256 · Human genomics · 2005 · 8 claims · 4 setups
Natural selection and population demographic history both distort patterns of genetic variation relative to the standard neutral model, so single-locus tests cannot unambiguously distinguish selection from demography.