Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
Design and analysis issues in genome-wide somatic mutation studies of cancer.
PMID 18692126 · PMC2820387 · Genomics · 2009 · 6 claims · 4 setups
Two-stage (discovery + validation) sequencing designs efficiently allocate resources and can produce highly informative candidate driver gene lists even with relatively small sample sizes.
-
Full-text index only
Microsatellites and SNPs linkage analysis in a Sardinian genetic isolate confirms several essential hypertension loci previously identified in different populations.
PMID 19715579 · PMC2741446 · BMC medical genetics · 2009 · 8 claims · 6 setups
Three loci (2q24, 11q23.1-25, 13q14.11-21.33) were identified by both the microsatellite and SNP genome-wide scans
-
Has reproduction · 50
MoDLE: high-performance stochastic modeling of DNA loop extrusion interactions.
PMID 36451166 · PMC9710047 · Genome biology · 2022 · 7 claims · 6 setups
MoDLE is a high-performance stochastic model that simulates DNA-DNA contacts from loop extrusion genome-wide in minutes using less than 1 GB of RAM
-
Full-text index only
Simultaneous analysis of all SNPs in genome-wide and re-sequencing association studies.
PMID 18654633 · PMC2464715 · PLoS genetics · 2008 · 8 claims · 5 setups
A Bayesian-inspired penalised maximum likelihood stochastic search method can simultaneously analyse all SNPs (up to 500K) from a GWA study in a few hours on a desktop workstation
-
Has reproduction · 89
TrEMOLO: accurate transposable element allele frequency estimation using long-read sequencing data combining assembly and mapping-based approaches.
PMID 37013657 · PMC10069131 · Genome biology · 2023 · 6 claims · 6 setups
TrEMOLO combines an assembly-based INSIDER module and a mapping-based OUTSIDER module to detect TE insertions/deletions from long-read sequencing data and estimate their allele frequency
-
Has reproduction · 85
PowerBacGWAS: a computational pipeline to perform power calculations for bacterial genome-wide association studies.
PMID 35338232 · PMC8956664 · Communications biology · 2022 · 8 claims · 8 setups
Two computational approaches (sub-sampling and phenotype-simulation) can be implemented to perform power calculations for bacterial GWAS using existing genome collections, packaged as the PowerBacGWAS pipeline
-
Has reproduction · 96
A bioinformatic pipeline for simulating viral integration data.
PMID 35496474 · PMC9046613 · Data in brief · 2022 · 7 claims · 3 setups
A snakemake-based pipeline was developed to simulate integration of a viral or vector genome into a host genome, including sub-genomic fragment integration, structural variation, and host-site deletions.
-
Has reproduction · 84
Improving recombinant protein production by yeast through genome-scale modeling using proteome constraints.
PMID 35624178 · PMC9142503 · Nature communications · 2022 · 7 claims · 5 setups
pcSecYeast, a proteome-constrained genome-scale model integrating metabolism, translation, and detailed secretory pathway processing (translocation, PTMs, folding, misfolding, degradation), was constructed for S. cerevisiae
-
Full-text index only
A note on generalized Genome Scan Meta-Analysis statistics.
PMID 15717930 · PMC551600 · BMC bioinformatics · 2005 · 7 claims · 3 setups
An Edgeworth series approximation to the null distribution of the weighted GSMA statistic provides a more accurate representation than the normal approximation, especially in the tails
-
Full-text index only
The signal in the genomes.
PMID 16683016 · PMC1447653 · PLoS computational biology · 2006 · 7 claims · 3 setups
A high breakpoint reuse rate in the output of rearrangement algorithms indicates loss of historical signal, not good evidence for genomic fragile regions
-
Has reproduction · 98
Massively parallel genomic perturbations with multi-target CRISPR interrogates Cas9 activity and DNA repair at endogenous sites.
PMID 36064968 · PMC9481459 · Nature cell biology · 2022 · 8 claims · 6 setups
Multi-target gRNAs (mgRNAs) can direct Cas9 to over a hundred well-mapped endogenous genomic sites simultaneously, enabling massively parallel, high-throughput interrogation of Cas9 activity via short-read sequencing
-
Full-text index only
A simple and efficient algorithm for genome-wide homozygosity analysis in disease.
PMID 19756043 · PMC2758715 · Molecular systems biology · 2009 · 8 claims · 4 setups
A genome-wide AH analysis (GAHA) algorithm can identify disease-associated loci by comparing frequencies of homozygous segments between cases and controls using a z-statistic proportion test
-
Full-text index only
Hit selection with false discovery rate control in genome-scale RNAi screens.
PMID 18628291 · PMC2504311 · Nucleic acids research · 2008 · 8 claims · 3 setups
A Bayesian FDR-controlling methodology for hit selection in genome-scale RNAi HTS is proposed, using a direct posterior probability approach analogous to Newton et al.
-
Full-text index only
A genome-wide approach to identify genetic loci with a signature of natural selection in the Irish population.
PMID 16904005 · PMC1779589 · Genome biology · 2006 · 8 claims · 7 setups
Eight SNPs with extreme European-branch locus-specific branch length (LSBL) were selected from a genome-wide FST dataset as candidates for selection in Europe.
-
Full-text index only
A third approach to gene prediction suggests thousands of additional human transcribed regions.
PMID 16543943 · PMC1391917 · PLoS computational biology · 2006 · 8 claims · 7 setups
A third basic concept for gene prediction exists, based on detecting strand-specific 'transcription footprints' (mutational and selectional biases) rather than gene structure or sequence similarity.
-
Full-text index only
Sequence occurrence and structural uniqueness of a G-quadruplex in the human c-kit promoter.
PMID 17720713 · PMC2034477 · Nucleic acids research · 2007 · 8 claims · 4 setups
The native 22-nt c-kit87 sequence occurs only once in the entire human genome.
-
Full-text index only
Simple models of genomic variation in human SNP density.
PMID 17553150 · PMC1919371 · BMC genomics · 2007 · 6 claims · 4 setups
Hierarchical Poisson model B, which allows both the mutation-rate proxy (Beta-distributed Λ) and the ARG-size proxy (Gamma-distributed T) to vary, fits the observed SNP density distribution significantly better than models with only one or neither varying.
-
Full-text index only
Duplication count distributions in DNA sequences.
PMID 19256873 · PMC3121164 · Physical review. E, Statistical, nonlinear, and soft matter physics · 2008 · 8 claims · 8 setups
Duplication count distributions N(c) for complex 40-mers show power-law-like decay for c roughly 3 to 50 (or higher) across human, C. elegans, A. thaliana, and D. melanogaster genomes.
-
Full-text index only
G-quadruplexes: the beginning and end of UTRs.
PMID 18832370 · PMC2577360 · Nucleic acids research · 2008 · 8 claims · 5 setups
UTRs show significant strand asymmetry with C-PQS more common than G-PQS, consistent with general depletion of G-quadruplex-forming RNA
-
Full-text index only
The whole alignment and nothing but the alignment: the problem of spurious alignment flanks.
PMID 18796526 · PMC2566872 · Nucleic acids research · 2008 · 8 claims · 4 setups
Some common scoring schemes tend to overextend alignments, generating spurious alignment flanks up to hundreds of bp/amino acids in length