Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
Periodicity of SNP distribution around transcription start sites.
PMID 16579865 · PMC1448210 · BMC genomics · 2006 · 8 claims · 6 setups
SNP density around TSS shows a 146-nucleotide periodicity
-
Full-text index only
Power analysis for genome-wide association studies.
PMID 17725844 · PMC2042984 · BMC genetics · 2007 · 8 claims · 6 setups
Developed a method to compute genome-wide association study power using tag SNPs and representative population genotype data (HapMap), equivalent to the cumulative r2-adjusted power of Jorgenson and Witte.
-
Full-text index only
Incorporation of genetic model parameters for cost-effective designs of genetic association studies using DNA pooling.
PMID 17634103 · PMC1947971 · BMC genomics · 2007 · 8 claims · 4 setups
A closed-form approximation to the F-test non-centrality parameter (NCP) incorporating genetic model parameters (disease allele frequency, marker allele frequency, prevalence, genotype relative risk, sample size, genetic model, number of pools/replicates, machine variability) can be used to compute power for DNA pooling association studies
-
Full-text index only
CorGen--measuring and generating long-range correlations for DNA sequence analysis.
PMID 16845099 · PMC1538783 · Nucleic acids research · 2006 · 8 claims · 3 setups
CorGen is a web server that measures long-range correlations in DNA sequences and generates random sequences with the same (or user-specified) correlation and composition parameters
-
Has reproduction · 95
Energy, power, and infrastructure demands from electrifying airport ground support equipment at United States airports.
PMID 41912553 · PMC13199391 · Nature communications · 2026 · 8 claims · 4 setups
Electrifying major GSE at U.S. airports can significantly increase grid power and energy demand, with peak power at large airports reaching up to 20 MW and annual consumption approaching 51,000 MWh.
-
Full-text index only
Is replication the gold standard for validating genome-wide association findings?
PMID 19112512 · PMC2605260 · PloS one · 2008 · 8 claims · 4 setups
The probability of replicating a specific GWA-identified variant decreases as the number of independent GWA/replication studies increases, when individual study power is less than 100%.
-
Full-text index only
Size matters: just how big is BIG?: Quantifying realistic sample size requirements for human genome epidemiology.
PMID 18676414 · PMC2639365 · International journal of epidemiology · 2009 · 7 claims · 2 setups
Conventional power calculations for case-control studies disregard analytic complexity (e.g. clinical assessment errors, unmeasured aetiological determinants) and can seriously underestimate true sample size requirements
-
Has reproduction · 88
pwrEWAS: a user-friendly tool for comprehensive power estimation for epigenome wide association studies (EWAS).
PMID 31035919 · PMC6489300 · BMC bioinformatics · 2019 · 8 claims · 8 setups
pwrEWAS is a user-friendly tool for comprehensive power estimation for two-group EWAS comparisons using Illumina Human Methylation BeadChip data.
-
Has reproduction · 86
Multi-INTACT: integrative analysis of the genome, transcriptome, and proteome identifies causal mechanisms of complex traits.
PMID 39901160 · PMC11789355 · Genome biology · 2025 · 8 claims · 2 setups
Multi-INTACT achieves higher power than existing single-gene-product methods while maintaining calibrated false discovery rates in simulations.
-
Full-text index only
A method for detecting epistasis in genome-wide studies using case-control multi-locus association analysis.
PMID 18667089 · PMC2533022 · BMC genomics · 2008 · 7 claims · 2 setups
HFCC is a method/software for genome-wide epistasis detection using case-control multi-locus association analysis, combining a fast computing algorithm with flexibility to test a variety of epistatic models.
-
Has reproduction · 85
PowerBacGWAS: a computational pipeline to perform power calculations for bacterial genome-wide association studies.
PMID 35338232 · PMC8956664 · Communications biology · 2022 · 8 claims · 8 setups
Two computational approaches (sub-sampling and phenotype-simulation) can be implemented to perform power calculations for bacterial GWAS using existing genome collections, packaged as the PowerBacGWAS pipeline
-
Full-text index only
Genome-wide copy number profiling on high-density bacterial artificial chromosomes, single-nucleotide polymorphisms, and oligonucleotide microarrays: a platform comparison based on statistical power analysis.
PMID 17363414 · PMC2779891 · DNA research : an international journal for rapid publication of reports on genes and genomes · 2007 · 8 claims · 6 setups
High-density oligonucleotide/SNP platforms are superior to the BAC platform for genome-wide detection of copy-number variations smaller than 1 Mb
-
Full-text index only
Calibrating the performance of SNP arrays for whole-genome association studies.
PMID 18584036 · PMC2432039 · PLoS genetics · 2008 · 8 claims · 7 setups
Previous SNP array genetic coverage estimates are inflated due to SNP overfitting and sample overfitting, since they were evaluated on the same HapMap SNPs/individuals used to design the arrays.
-
Full-text index only
What can genome-wide association studies tell us about the genetics of common disease?
PMID 18454206 · PMC2323402 · PLoS genetics · 2008 · 8 claims · 4 setups
Apparent patterns of common, low-effect disease-associated alleles largely reflect statistical power of studies rather than the true underlying distribution of disease variants
-
Full-text index only
A statistical model to identify differentially expressed proteins in 2D PAGE gels.
PMID 19763172 · PMC2734266 · PLoS computational biology · 2009 · 7 claims · 5 setups
A mixture likelihood model incorporating both detected and non-detected proteins has higher statistical power to detect differential expression than standard approaches like the Student's t-test.
-
Full-text index only
A simple and efficient algorithm for genome-wide homozygosity analysis in disease.
PMID 19756043 · PMC2758715 · Molecular systems biology · 2009 · 8 claims · 4 setups
A genome-wide AH analysis (GAHA) algorithm can identify disease-associated loci by comparing frequencies of homozygous segments between cases and controls using a z-statistic proportion test
-
Full-text index only
Neuropixels Opto: combining high-resolution electrophysiology and optogenetics.
PMID 42225964 · PMC13259958 · Nature methods · 2026 · 8 claims · 7 setups
Neuropixels Opto probes integrate 960 recording sites and two sets of 14 light emitters (blue and red) on a 70-μm-wide, 1-cm-long shank via monolithic CMOS+photonics integration
-
Full-text index only
DBD--taxonomically broad transcription factor predictions: new content and functionality.
PMID 18073188 · PMC2238844 · Nucleic acids research · 2008 · 8 claims · 3 setups
DBD is a database of predicted sequence-specific DNA-binding transcription factors covering over 700 publicly available proteomes, up from 150 in the initial version.
-
Has reproduction · 73
GREIN: An Interactive Web Platform for Re-analyzing GEO RNA-seq Data.
PMID 31110304 · PMC6527554 · Scientific reports · 2019 · 8 claims · 7 setups
GREIN is a web application providing user-friendly interfaces to manipulate, visualize, and analyze GEO RNA-seq data.
-
Has reproduction · 83
Metabolite-Centric Reporter Pathway and Tripartite Network Analysis of Arabidopsis Under Cold Stress.
PMID 30258841 · PMC6143811 · Frontiers in bioengineering and biotechnology · 2018 · 8 claims · 8 setups
Metabolite-centric reporter pathway analysis (RPAm) computes reporter metabolites and reporter pathways from transcriptome P-values by aggregating Z-scores of neighboring genes in a genome-scale metabolic network