Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
Iterative pruning PCA improves resolution of highly structured populations.
PMID 19930644 · PMC2790469 · BMC bioinformatics · 2009 · 7 claims · 7 setups
ipPCA is a novel algorithm that assigns individuals to subpopulations and infers the total number of subpopulations (K) present in genotypic data
-
Full-text index only
Automated recognition of retroviral sequences in genomic data--RetroTector.
PMID 17636050 · PMC1976444 · Nucleic acids research · 2007 · 8 claims · 8 setups
RetroTector uses 'fragment threading' (detection of chains of conserved retroviral motifs satisfying distance constraints) combined with LTR detection and protein reconstruction to identify ERVs in genomic sequences
-
Full-text index only
Is replication the gold standard for validating genome-wide association findings?
PMID 19112512 · PMC2605260 · PloS one · 2008 · 8 claims · 4 setups
The probability of replicating a specific GWA-identified variant decreases as the number of independent GWA/replication studies increases, when individual study power is less than 100%.
-
Full-text index only
Calibrating the performance of SNP arrays for whole-genome association studies.
PMID 18584036 · PMC2432039 · PLoS genetics · 2008 · 8 claims · 7 setups
Previous SNP array genetic coverage estimates are inflated due to SNP overfitting and sample overfitting, since they were evaluated on the same HapMap SNPs/individuals used to design the arrays.
-
Has reproduction · 94
Deep learning from phylogenies to uncover the epidemiological dynamics of outbreaks.
PMID 35794110 · PMC9258765 · Nature communications · 2022 · 8 claims · 5 setups
Deep learning (FFNN-SS and CNN-CBLV) enables accurate and fast likelihood-free estimation of epidemiological parameters and model selection from phylogenies
-
Full-text index only
Association testing of novel type 2 diabetes risk alleles in the JAZF1, CDC123/CAMK1D, TSPAN8, THADA, ADAMTS9, and NOTCH2 loci with insulin release, insulin sensitivity, and obesity in a population-based sample of 4,516 glucose-tolerant middle-aged Danes.
PMID 18567820 · PMC2518507 · Diabetes · 2008 · 8 claims · 5 setups
CDC123/CAMK1D rs12779790 risk allele (homozygous) is associated with decreased insulinogenic index, corrected insulin response (CIR), and AUC-insulin/AUC-glucose ratio, indicating impaired insulin release
-
Has reproduction · 68
Constraints to gene flow increase the risk of genome erosion in the Ngorongoro Crater lion population.
PMID 40258987 · PMC12012037 · Communications biology · 2025 · 8 claims · 9 setups
200 years of quasi-isolation and the 1962 epizootic caused a two-fold increase in inbreeding and an excess of highly deleterious mutations in Crater lions relative to other Greater Serengeti populations
-
Full-text index only
Inferring human colonization history using a copying model.
PMID 18497854 · PMC2367454 · PLoS genetics · 2008 · 8 claims · 6 setups
A copying-model approach using SNP haplotype sharing can infer both the order of population founding and the donor populations contributing ancestry to each new population.
-
Full-text index only
Optimized mixed Markov models for motif identification.
PMID 16749929 · PMC1534070 · BMC bioinformatics · 2006 · 8 claims · 4 setups
OMiMa can incorporate more than NNSplice's pairwise dependencies
-
Full-text index only
Modeling the amplification dynamics of human Alu retrotransposons.
PMID 16201008 · PMC1239904 · PLoS computational biology · 2005 · 8 claims · 4 setups
Combining sequence diversity (π) and insertion polymorphism level (IPL) statistics can statistically exclude implausible Alu amplification scenarios and narrow the range of plausible ones for individual subfamilies.
-
Full-text index only
Stability analysis of mixtures of mutagenetic trees.
PMID 18366778 · PMC2335279 · BMC bioinformatics · 2008 · 7 claims · 5 setups
Mutagenetic trees mixture models capture multiple alternative pathways of ordered accumulation of genetic events (e.g., HIV resistance mutations, cancer chromosomal aberrations).
-
Full-text index only
Evolutionary distance estimation and fidelity of pair wise sequence alignment.
PMID 15840174 · PMC1087827 · BMC bioinformatics · 2005 · 8 claims · 8 setups
Evolutionary distance estimation is relatively unaffected by alignment error as long as 50% or more of homologous sites remain identical between sequences
-
Full-text index only
Predicting survival outcomes using subsets of significant genes in prognostic marker studies with microarrays.
PMID 16549007 · PMC1544357 · BMC bioinformatics · 2006 · 7 claims · 2 setups
A methodology combining Cox proportional hazards models with a compound covariate, cross-validated log partial likelihood (ACVL) for predictive accuracy, and permutation-based significance testing can identify an optimal subset of significant genes for survival prediction
-
Full-text index only
Importance sampling for the infinite sites model.
PMID 18976228 · PMC2832804 · Statistical applications in genetics and molecular biology · 2008 · 7 claims · 2 setups
A new importance sampling proposal distribution for the ISM, derived from a new result on exact sampling from a single segregating site, generally shows greater efficiency than the GT and SD proposals.
-
Has reproduction · 87
Forseti: a mechanistic and predictive model of the splicing status of scRNA-seq reads.
PMID 38940130 · PMC11256924 · Bioinformatics (Oxford, England) · 2024 · 7 claims · 5 setups
Forseti is the first probabilistic model for resolving the splicing status of exonic scRNA-seq reads by scoring putative fragments linking read alignments to proximate priming sites
-
Full-text index only
A statistical approach for array CGH data analysis.
PMID 15705208 · PMC549559 · BMC bioinformatics · 2005 · 8 claims · 4 setups
Existing model-selection criteria (AIC, BIC, and prior ad hoc penalties) are not well adapted to estimating the number of segments in array CGH data
-
Full-text index only
Antigenic diversity, transmission mechanisms, and the evolution of pathogens.
PMID 19847288 · PMC2759524 · PLoS computational biology · 2009 · 8 claims · 3 setups
Three distinct infection types (A, B, C) emerge as maxima in the pathogen fitness landscape, each with characteristic within-host dynamics, contact network structure, and transmission mode
-
Full-text index only
Quantitative analysis of single nucleotide polymorphisms within copy number variation.
PMID 19093001 · PMC2600609 · PloS one · 2008 · 8 claims · 2 setups
Copy number variation is a major factor in HWE violation for SNPs with small minor allele frequency, large sample size, and 0-1% genotyping error rate
-
Has reproduction · 83
SIRE 2.0: a novel method for estimating polygenic host effects underlying infectious disease transmission, and analytical expressions for prediction accuracies.
PMID 40169992 · PMC11963337 · Genetics, selection, evolution : GSE · 2025 · 8 claims · 2 setups
SIRE 2.0 is a novel Bayesian methodology and software tool for estimating polygenic contributions (variance components and additive genetic effects) to host susceptibility, infectivity and recoverability from temporal epidemic data using pedigree/genomic relationship matrices.
-
Full-text index only
A statistical approach designed for finding mathematically defined repeats in shotgun data and determining the length distribution of clone-inserts.
PMID 15626332 · PMC5172250 · Genomics, proteomics & bioinformatics · 2003 · 8 claims · 6 setups
Repeats of different copy number have distinct probabilities of appearance in shotgun data, which can be modeled statistically to define recognition thresholds (MDRs) at different shotgun coverages.