Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
Size matters: just how big is BIG?: Quantifying realistic sample size requirements for human genome epidemiology.
PMID 18676414 · PMC2639365 · International journal of epidemiology · 2009 · 7 claims · 2 setups
Conventional power calculations for case-control studies disregard analytic complexity (e.g. clinical assessment errors, unmeasured aetiological determinants) and can seriously underestimate true sample size requirements
-
Full-text index only
Quantitative analysis of single nucleotide polymorphisms within copy number variation.
PMID 19093001 · PMC2600609 · PloS one · 2008 · 8 claims · 2 setups
Copy number variation is a major factor in HWE violation for SNPs with small minor allele frequency, large sample size, and 0-1% genotyping error rate
-
Full-text index only
Sequence polymorphisms cause many false cis eQTLs.
PMID 17637838 · PMC1906859 · PloS one · 2007 · 8 claims · 7 setups
Many reported local/cis eQTLs are false positives caused by probe-region sequence polymorphisms affecting hybridization rather than true cis-regulatory expression differences.
-
Full-text index only
Screening large-scale association study data: exploiting interactions using random forests.
PMID 15588316 · PMC545646 · BMC genetics · 2004 · 7 claims · 3 setups
Random forest importance measure significantly outperforms the Fisher Exact test as a screening tool when risk SNPs interact.
-
Full-text index only
Incorporation of genetic model parameters for cost-effective designs of genetic association studies using DNA pooling.
PMID 17634103 · PMC1947971 · BMC genomics · 2007 · 8 claims · 4 setups
A closed-form approximation to the F-test non-centrality parameter (NCP) incorporating genetic model parameters (disease allele frequency, marker allele frequency, prevalence, genotype relative risk, sample size, genetic model, number of pools/replicates, machine variability) can be used to compute power for DNA pooling association studies
-
Full-text index only
The evolutionary dynamics of a rapidly mutating virus within and between hosts: the case of hepatitis C virus.
PMID 19911046 · PMC2768904 · PLoS computational biology · 2009 · 8 claims · 3 setups
The replication rate of the strain that initiates an infection has a strong effect on the fitness of the infection at the between-host level, even though the virus evolves rapidly within the host.
-
Has reproduction · 85
PowerBacGWAS: a computational pipeline to perform power calculations for bacterial genome-wide association studies.
PMID 35338232 · PMC8956664 · Communications biology · 2022 · 8 claims · 8 setups
Two computational approaches (sub-sampling and phenotype-simulation) can be implemented to perform power calculations for bacterial GWAS using existing genome collections, packaged as the PowerBacGWAS pipeline
-
Full-text index only
A non-parametric meta-analysis approach for combining independent microarray datasets: application using two microarray datasets pertaining to chronic allograft nephropathy.
PMID 18302764 · PMC2276496 · BMC genomics · 2008 · 8 claims · 6 setups
A novel non-parametric meta-analysis approach for combining independent microarray datasets is presented, requiring no distributional assumptions and being logically intuitive.
-
Full-text index only
Population history and natural selection shape patterns of genetic variation in 132 genes.
PMID 15361935 · PMC515367 · PLoS biology · 2004 · 7 claims · 5 setups
Developed a rigorous computational approach that corrects for multiple hypothesis testing and models population demographic history to test for natural selection
-
Full-text index only
A simple and efficient algorithm for genome-wide homozygosity analysis in disease.
PMID 19756043 · PMC2758715 · Molecular systems biology · 2009 · 8 claims · 4 setups
A genome-wide AH analysis (GAHA) algorithm can identify disease-associated loci by comparing frequencies of homozygous segments between cases and controls using a z-statistic proportion test
-
Has reproduction · 100
Computational modeling demonstrates that glioblastoma cells can survive spatial environmental challenges through exploratory adaptation.
PMID 31836713 · PMC6911112 · Nature communications · 2019 · 8 claims · 6 setups
Exploratory adaptation (stochastic gene-regulatory network perturbation) explains how GBM cells adapt phenotypically across spatially distinct tumor regions
-
Has reproduction · 100
Intratumoral heterogeneity in microsatellite instability status at single-cell resolution.
PMID 41767255 · PMC12936829 · iScience · 2026 · 8 claims · 7 setups
A novel computational (Snakemake) pipeline quantifies intratumoral heterogeneity in MSI status at single-cell resolution
-
Full-text index only
Disease-aging network reveals significant roles of aging genes in connecting genetic diseases.
PMID 19779549 · PMC2739292 · PLoS computational biology · 2009 · 8 claims · 8 setups
Human disease genes are much closer to aging genes in the PPI network than expected by chance
-
Full-text index only
Discovery and characterization of gene-by-environment and epistatic genetic effects in a vertebrate model.
PMID 41672067 · PMC13174225 · Cell genomics · 2026 · 7 claims · 7 setups
A segregation analysis in an F2 medaka cross identified 16 QTLs linked to embryonic heart rate variation across temperatures
-
Has reproduction · 87
CoINcIDE: A framework for discovery of patient subtypes across multiple datasets.
PMID 26961683 · PMC4784276 · Genome medicine · 2016 · 8 claims · 4 setups
CoINcIDE is a novel framework for discovering patient subtypes across multiple datasets that requires no between-dataset transformations (e.g., batch correction)
-
Has reproduction · 88
AuPairWise: A Method to Estimate RNA-Seq Replicability through Co-expression.
PMID 27082953 · PMC4833304 · PLoS computational biology · 2016 · 7 claims · 6 setups
Sample-sample correlation of transcript abundances is a misleading measure of replicability for assessing differential expression, because it is dominated by gene-specific dynamic ranges rather than condition-dependent variation.
-
Full-text index only
A genome-wide approach to identify genetic loci with a signature of natural selection in the Irish population.
PMID 16904005 · PMC1779589 · Genome biology · 2006 · 8 claims · 7 setups
Eight SNPs with extreme European-branch locus-specific branch length (LSBL) were selected from a genome-wide FST dataset as candidates for selection in Europe.
-
Full-text index only
Insights into the coupling of duplication events and macroevolution from an age profile of animal transmembrane gene families.
PMID 16895434 · PMC1534073 · PLoS computational biology · 2006 · 8 claims · 7 setups
The density of transmembrane gene duplicates positively correlates with the estimated maximum number of cell types of common ancestors
-
Full-text index only
Multilocus analysis of SNP and metabolic data within a given pathway.
PMID 16412218 · PMC1382210 · BMC genomics · 2006 · 8 claims · 7 setups
The combinatorial partitioning method (CPM) with optimal thresholds can identify SNPs associated with quantitative metabolite levels rather than only categorical traits.
-
Full-text index only
What can genome-wide association studies tell us about the genetics of common disease?
PMID 18454206 · PMC2323402 · PLoS genetics · 2008 · 8 claims · 4 setups
Apparent patterns of common, low-effect disease-associated alleles largely reflect statistical power of studies rather than the true underlying distribution of disease variants