Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
Importance sampling for the infinite sites model.
PMID 18976228 · PMC2832804 · Statistical applications in genetics and molecular biology · 2008 · 7 claims · 2 setups
A new importance sampling proposal distribution for the ISM, derived from a new result on exact sampling from a single segregating site, generally shows greater efficiency than the GT and SD proposals.
-
Full-text index only
The fragile breakage versus random breakage models of chromosome evolution.
PMID 16501665 · PMC1378107 · PLoS computational biology · 2006 · 8 claims · 6 setups
Sankoff and Trinh's synteny block identification algorithm (ST-Synteny) is flawed, producing erroneous block identifications even in small toy examples.
-
Full-text index only
Visualization of shared genomic regions and meiotic recombination in high-density SNP data.
PMID 19696932 · PMC2725774 · PloS one · 2009 · 8 claims · 7 setups
SNPduo is a command-line (SNPduo++) and web-accessible tool that analyzes and visualizes relatedness between two individuals using identity by state (IBS) from SNP genotypes.
-
Has reproduction · 84
A scalable, open-source implementation of a large-scale mechanistic model for single cell proliferation and death signaling.
PMID 35729113 · PMC9213456 · Nature communications · 2022 · 7 claims · 5 setups
Developed a python-based, scalable model creation and simulation pipeline that converts structured text files into an SBML-standard model and is high-performance/cloud-computing ready.
-
Full-text index only
Incorporation of genetic model parameters for cost-effective designs of genetic association studies using DNA pooling.
PMID 17634103 · PMC1947971 · BMC genomics · 2007 · 8 claims · 4 setups
A closed-form approximation to the F-test non-centrality parameter (NCP) incorporating genetic model parameters (disease allele frequency, marker allele frequency, prevalence, genotype relative risk, sample size, genetic model, number of pools/replicates, machine variability) can be used to compute power for DNA pooling association studies
-
Full-text index only
Stability analysis of mixtures of mutagenetic trees.
PMID 18366778 · PMC2335279 · BMC bioinformatics · 2008 · 7 claims · 5 setups
Mutagenetic trees mixture models capture multiple alternative pathways of ordered accumulation of genetic events (e.g., HIV resistance mutations, cancer chromosomal aberrations).
-
Has reproduction · 89
Graph Random Forest: A Graph Embedded Algorithm for Identifying Highly Connected Important Features.
PMID 37509188 · PMC10377046 · Biomolecules · 2023 · 8 claims · 3 setups
Graph Random Forest (GRF) embeds graph/network information directly into the decision-tree building process by splitting on features in the k-hop neighborhood of a data-driven head-splitting node.
-
Has reproduction · 74
Software JimenaE allows efficient dynamic simulations of Boolean networks, centrality and system state analysis.
PMID 36725967 · PMC9892028 · Scientific reports · 2023 · 8 claims · 4 setups
JimenaE simulates Boolean networks dynamically and systematically calculates all system states rather than heuristically as SQUAD does.
-
Has reproduction · 89
TrEMOLO: accurate transposable element allele frequency estimation using long-read sequencing data combining assembly and mapping-based approaches.
PMID 37013657 · PMC10069131 · Genome biology · 2023 · 6 claims · 6 setups
TrEMOLO combines an assembly-based INSIDER module and a mapping-based OUTSIDER module to detect TE insertions/deletions from long-read sequencing data and estimate their allele frequency
-
Full-text index only
DiagHunter and GenoPix2D: programs for genomic comparisons, large-scale homology discovery and visualization.
PMID 14519203 · PMC328457 · Genome biology · 2003 · 7 claims · 5 setups
DiagHunter identifies large-scale synteny blocks within or between genomes efficiently despite background noise and genomic discontinuities, without performing sequence alignment
-
Full-text index only
Iterative class discovery and feature selection using Minimal Spanning Trees.
PMID 15355552 · PMC520744 · BMC bioinformatics · 2004 · 7 claims · 5 setups
Iterating between MST-based clustering and t-statistic feature selection removes noise genes step-wise while sharpening the sample clustering
-
Full-text index only
Multiplexed discovery of sequence polymorphisms using base-specific cleavage and MALDI-TOF MS.
PMID 15731331 · PMC549577 · Nucleic acids research · 2005 · 8 claims · 7 setups
Multiplexed base-specific cleavage/MALDI-TOF MS (Multiplexed Comparative Sequence Analysis) enables simultaneous discovery of sequence polymorphisms across multiple target regions
-
Full-text index only
Cubic exact solutions for the estimation of pairwise haplotype frequencies: implications for linkage disequilibrium analyses and a web tool 'CubeX'.
PMID 17980034 · PMC2180187 · BMC bioinformatics · 2007 · 6 claims · 4 setups
CubeX, a Python program/web tool, computes the exact algebraic (Cardan/Nickalls) solution(s) of Hill's cubic equation to estimate pairwise haplotype frequencies, D', r2 and chi-square for each solution
-
Full-text index only
Is replication the gold standard for validating genome-wide association findings?
PMID 19112512 · PMC2605260 · PloS one · 2008 · 8 claims · 4 setups
The probability of replicating a specific GWA-identified variant decreases as the number of independent GWA/replication studies increases, when individual study power is less than 100%.
-
Full-text index only
The origins of lactase persistence in Europe.
PMID 19714206 · PMC2722739 · PLoS computational biology · 2009 · 8 claims · 5 setups
The −13,910*T allele first underwent selection among dairying farmers around 7,500 years ago in a region between the central Balkans and central Europe, possibly linked to the Linearbandkeramik culture.
-
Has reproduction · 53
spliceJAC: transition genes and state-specific gene regulation from single-cell transcriptome data.
PMID 36321549 · PMC9627675 · Molecular systems biology · 2022 · 8 claims · 6 setups
spliceJAC quantifies multivariate mRNA splicing from unspliced/spliced count matrices to construct cell state-specific gene-gene (Jacobian) interaction matrices.
-
Has reproduction · 50
MoDLE: high-performance stochastic modeling of DNA loop extrusion interactions.
PMID 36451166 · PMC9710047 · Genome biology · 2022 · 7 claims · 6 setups
MoDLE is a high-performance stochastic model that simulates DNA-DNA contacts from loop extrusion genome-wide in minutes using less than 1 GB of RAM
-
Has reproduction · 89
Improved eukaryotic detection compatible with large-scale automated analysis of metagenomes.
PMID 37032329 · PMC10084625 · Microbiome · 2023 · 8 claims · 7 setups
MAPQ ≥30 filtering improves precision but substantially reduces recall, especially for unrepresented/divergent eukaryotic taxa
-
Full-text index only
Benchmarking tools for the alignment of functional noncoding DNA.
PMID 14736341 · PMC344529 · BMC bioinformatics · 2004 · 8 claims · 4 setups
Global alignment tools (Avid, ClustalW, Lagan, Needle, DiAlign-G) typically have higher sensitivity over entire noncoding sequences and within constrained blocks than local tools
-
Full-text index only
Design and analysis issues in genome-wide somatic mutation studies of cancer.
PMID 18692126 · PMC2820387 · Genomics · 2009 · 6 claims · 4 setups
Two-stage (discovery + validation) sequencing designs efficiently allocate resources and can produce highly informative candidate driver gene lists even with relatively small sample sizes.