Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
Discovery and hypothesis generation through bioinformatics.
PMID 16522224 · PMC1431734 · Genome biology · 2006 · 8 claims · 8 setups
Bioinformatics should be used as a tool for discovery and hypothesis generation, not merely to manage biological data
-
Has reproduction · 74
ChIP-seq guidelines and practices of the ENCODE and modENCODE consortia.
PMID 22955991 · PMC3431496 · Genome research · 2012 · 8 claims · 8 setups
ENCODE/modENCODE define a set of working standards and guidelines for ChIP-seq covering antibody validation, experimental replication, sequencing depth, data/metadata reporting, and data quality assessment.
-
Full-text index only
A model-based approach to selection of tag SNPs.
PMID 16776821 · PMC1525207 · BMC bioinformatics · 2006 · 7 claims · 5 setups
The Li and Stephens hidden Markov model outperforms other tested models (simple Markov, two-state HMM, HMM-4D, greedy GR-1/GR-2) in description code-length, tag set information content, and prediction of tagged SNPs.
-
Full-text index only
Ensembl 2007.
PMID 17148474 · PMC1761443 · Nucleic acids research · 2007 · 8 claims · 7 setups
Ensembl added 18 new chordate genomes this year, increasing total genomes available from 15 to 33, the largest yearly increase to date.
-
Full-text index only
The vertebrate genome annotation (Vega) database.
PMID 18003653 · PMC2238886 · Nucleic acids research · 2008 · 8 claims · 8 setups
Vega is a database for viewing manual genome annotation of human, mouse and zebrafish genomic sequences produced at the Wellcome Trust Sanger Institute.
-
Has reproduction · 77
Accurate chromatin marks peak calling with Omnipeak.
PMID 41521664 · PMC12784980 · Nucleic acids research · 2026 · 8 claims · 6 setups
Omnipeak is a universal unsupervised peak-calling algorithm based on a constrained three-state hidden Markov model (zero, noise, signal states)
-
Full-text index only
Automatic annotation of eukaryotic genes, pseudogenes and promoters.
PMID 16925832 · PMC1810547 · Genome biology · 2006 · 8 claims · 6 setups
Fgenesh++ gene prediction pipeline identifies 91% of coding nucleotides with 90% specificity
-
Full-text index only
An evaluation of the performance of tag SNPs derived from HapMap in a Caucasian population.
PMID 16532062 · PMC1391920 · PLoS genetics · 2006 · 8 claims · 5 setups
CEU HapMap-derived tSNPs capture most of the genetic variation observed in the Estonian (EGP) population sample
-
Has reproduction · 69
A comparison across non-model animals suggests an optimal sequencing depth for de novo transcriptome assembly.
PMID 23496952 · PMC3655071 · BMC genomics · 2013 · 8 claims · 8 setups
Representative de novo transcriptome assemblies are generated with as few as ~20 million reads for single-tissue samples and ~30 million reads for whole animals at the mRNA-coverage level.
-
Full-text index only
JIGSAW, GeneZilla, and GlimmerHMM: puzzling out the features of human genes in the ENCODE regions.
PMID 16925843 · PMC1810558 · Genome biology · 2006 · 8 claims · 4 setups
Adding model states for specific biological features (signal peptides, CpG islands, etc.) to non-comparative GHMM gene finders did little or nothing to enhance predictive accuracy, sometimes reducing it.
-
Full-text index only
Software for tag single nucleotide polymorphism selection.
PMID 16004730 · PMC3525260 · Human genomics · 2005 · 8 claims · 3 setups
Pairwise R2 methods tend to pick more tagging SNPs than strictly needed because they miss redundancy where two or more tag SNPs jointly predict an untagged SNP with no single direct surrogate.
-
Has reproduction · 82
Landscape of allele-specific transcription factor binding in the human genome.
PMID 33980847 · PMC8115691 · Nature communications · 2021 · 8 claims · 6 setups
A novel statistical framework (ADASTRA) calls allele-specific TF binding from existing ChIP-Seq alignments by jointly correcting for background allelic dosage (BAD, from aneuploidy/CNVs) and reference mapping bias.
-
Full-text index only
CONTRAST: a discriminative, phylogeny-free approach to multiple informant de novo gene prediction.
PMID 18096039 · PMC2246271 · Genome biology · 2007 · 8 claims · 5 setups
CONTRAST predicts exact coding region structures for 65% more human genes than the previous state-of-the-art de novo predictor (N-SCAN)
-
Full-text index only
Vertebrate gene finding from multiple-species alignments using a two-level strategy.
PMID 16925840 · PMC1810555 · Genome biology · 2006 · 8 claims · 5 setups
DOGFISH cleanly separates a multi-species alignment classifier (RVM cascade) from an HMM-based structure predictor, avoiding tight coupling of alignment complexity with HMM formalism
-
Full-text index only
Discovery of protein-protein interactions using a combination of linguistic, statistical and graphical information.
PMID 15941473 · PMC1164402 · BMC bioinformatics · 2005 · 8 claims · 5 setups
A combined linguistic+statistical+rule-based method achieves precision 0.61 and recall 0.97 (f=0.74) detecting yeast protein-protein interactions across 12,300 Medline abstracts.
-
Full-text index only
miRGen: a database for the study of animal microRNA genomic organization and function.
PMID 17108354 · PMC1669779 · Nucleic acids research · 2007 · 8 claims · 6 setups
miRGen is an integrated database combining Genomics, Targets, and Clusters interfaces to study miRNA genomic organization and function across 11 animal genomes
-
Full-text index only
Power analysis for genome-wide association studies.
PMID 17725844 · PMC2042984 · BMC genetics · 2007 · 8 claims · 6 setups
Developed a method to compute genome-wide association study power using tag SNPs and representative population genotype data (HapMap), equivalent to the cumulative r2-adjusted power of Jorgenson and Witte.
-
Full-text index only
Sequence polymorphisms cause many false cis eQTLs.
PMID 17637838 · PMC1906859 · PloS one · 2007 · 8 claims · 7 setups
Many reported local/cis eQTLs are false positives caused by probe-region sequence polymorphisms affecting hybridization rather than true cis-regulatory expression differences.
-
Full-text index only
Ensembl 2009.
PMID 19033362 · PMC2686571 · Nucleic acids research · 2009 · 8 claims · 6 setups
Ensembl provides comprehensive, consistently annotated genome information for chordate genomes with automatically generated genesets and comparative genomics data