Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
Completing the map of human genetic variation.
PMID 17495918 · PMC2685471 · Nature · 2007 · 8 claims · 5 setups
A community resource initiative will sequence fosmid and BAC clone libraries from 62 HapMap individuals to systematically discover and resolve structural genetic variants at nucleotide resolution
-
Full-text index only
Exploiting noise in array CGH data to improve detection of DNA copy number change.
PMID 17272296 · PMC1994778 · Nucleic acids research · 2007 · 7 claims · 4 setups
When aberrations are present, noise in BAC, 19k oligo, and 385k oligo array-CGH data is highly non-Gaussian and shows long-range spatial correlations.
-
Full-text index only
TPRpred: a tool for prediction of TPR-, PPR- and SEL1-like repeats from protein sequences.
PMID 17199898 · PMC1774580 · BMC bioinformatics · 2007 · 7 claims · 8 setups
TPRpred detects divergent/remote-homolog TPR repeat units that existing resources (Pfam, SMART, REP) fail to detect
-
Full-text index only
Evolutionary modeling of rate shifts reveals specificity determinants in HIV-1 subtypes.
PMID 18989394 · PMC2566816 · PLoS computational biology · 2008 · 7 claims · 4 setups
A novel Bayesian method, RASER, can detect site-specific evolutionary rate shifts and the lineages in which they occurred without pre-specifying candidate lineages.
-
Full-text index only
High-throughput chromatin information enables accurate tissue-specific prediction of transcription factor binding sites.
PMID 18988630 · PMC2662491 · Nucleic acids research · 2009 · 8 claims · 8 setups
Incorporating H3K4me3 chromatin modification estimates greatly improves the accuracy of in silico prediction of in vivo TF binding for a wide range of TFs in human and mouse
-
Full-text index only
DNA sequencing of a cytogenetically normal acute myeloid leukaemia genome.
PMID 18987736 · PMC2603574 · Nature · 2008 · 8 claims · 8 setups
Whole genome sequencing can identify unbiased, novel somatic mutations in a cytogenetically normal AML genome that would not have been found by candidate-gene resequencing.
-
Full-text index only
Functional annotation and identification of candidate disease genes by computational analysis of normal tissue gene expression data.
PMID 18560577 · PMC2409962 · PloS one · 2008 · 7 claims · 5 setups
Ranked Coexpression Groups (RCG) built from k=6 nearest coexpressed genes, combined with a majority-rule functional characterization, integrate multiple datasets/coexpression measures to generate high-confidence functional annotation predictions
-
Full-text index only
Human and mouse introns are linked to the same processes and functions through each genome's most frequent non-conserved motifs.
PMID 18450818 · PMC2425492 · Nucleic acids research · 2008 · 8 claims · 5 setups
Pyknons (recurrent, genome-specific, ≥16nt motifs with ≥30 intact intergenic/intronic copies and ≥1 exonic copy) span a substantial fraction of previously uncharacterized intronic space (7.4% human, 4.4% mouse)
-
Full-text index only
Minisequencing mitochondrial DNA pathogenic mutations.
PMID 18402672 · PMC2377236 · BMC medical genetics · 2008 · 7 claims · 7 setups
A minisequencing multiplex assay can interrogate 25 pathogenic mtDNA mutations across the whole mtDNA genome in a single reaction using 13 amplicons.
-
Has reproduction · 50
Polymorphism identification and improved genome annotation of Brassica rapa through Deep RNA sequencing.
PMID 25122667 · PMC4232532 · G3 (Bethesda, Md.) · 2014 · 8 claims · 8 setups
330,995 SNPs were identified in transcribed regions between B. rapa genotypes R500 and IMB211, at an average frequency of one SNP per 200 bases.
-
Full-text index only
Microdroplet-based PCR enrichment for large-scale targeted sequencing.
PMID 19881494 · PMC2779736 · Nature biotechnology · 2009 · 7 claims · 5 setups
Microdroplet PCR enables massively parallel singleplex amplification (up to ~1.5 million reactions, up to 4,000 targets) for targeted sequencing enrichment
-
Full-text index only
Targeted capture and massively parallel sequencing of 12 human exomes.
PMID 19684571 · PMC2844771 · Nature · 2009 · 8 claims · 8 setups
Targeted exome capture combined with massively parallel sequencing sensitively and specifically identifies rare and common variants across >300 Mb of coding sequence
-
Full-text index only
A machine learning approach uncovers principles and determinants of eukaryotic ribosome pausing.
PMID 39423268 · PMC11488575 · Science advances · 2024 · 8 claims · 5 setups
An unsupervised ML pipeline using the extended isolation forest (EIF) algorithm can reliably detect ribosome pausing sites from noisy, coverage-biased RiboSeq data across expression levels
-
Full-text index only
Prediction of candidate primary immunodeficiency disease genes using a support vector machine learning approach.
PMID 19801557 · PMC2780952 · DNA research : an international journal for rapid publication of reports on genes and genomes · 2009 · 6 claims · 3 setups
An SVM trained on 69 binary features of known PID genes can accurately classify PID vs non-PID genes and predict novel candidate PID genes
-
Full-text index only
CGMIM: automated text-mining of Online Mendelian Inheritance in Man (OMIM) to identify genetically-associated cancers and candidate genes.
PMID 15796777 · PMC1274267 · BMC bioinformatics · 2005 · 8 claims · 2 setups
CGMIM is a Perl program that text-mines OMIM entries to identify cancer-gene associations and genetically-related cancer type pairs.
-
Full-text index only
AutoCSA, an algorithm for high throughput DNA sequence variant detection in cancer genomes.
PMID 17485433 · PMC5947781 · Bioinformatics (Oxford, England) · 2007 · 7 claims · 2 setups
AutoCSA is an automated algorithm, extended from the CSA protocol, that detects DNA sequence variants in cancer genomes with minimal manual intervention
-
Has reproduction · 87
CoINcIDE: A framework for discovery of patient subtypes across multiple datasets.
PMID 26961683 · PMC4784276 · Genome medicine · 2016 · 8 claims · 4 setups
CoINcIDE is a novel framework for discovering patient subtypes across multiple datasets that requires no between-dataset transformations (e.g., batch correction)
-
Has reproduction · 85
Digital sorting of complex tissues for cell type-specific gene expression profiles.
PMID 23497278 · PMC3626856 · BMC bioinformatics · 2013 · 8 claims · 8 setups
The Digital Sorting Algorithm (DSA) deconvolves mixed tissue expression into cell type-specific profiles using only marker genes, without requiring prior knowledge of cell type frequencies or in vitro pure-cell profiles.
-
Has reproduction · 91
Insights into the evolution of cotton diploids and polyploids from whole-genome re-sequencing.
PMID 23979935 · PMC3789805 · G3 (Bethesda, Md.) · 2013 · 8 claims · 8 setups
An index of 23,859,893 (~24 million) homoeo-SNPs distinguishing A-genome from D-genome cotton was constructed at a density of one SNP per 32.3 bases of the D5 reference.
-
Full-text index only
BABELOMICS: a systems biology perspective in the functional annotation of genome-scale experiments.
PMID 16845052 · PMC1538844 · Nucleic acids research · 2006 · 8 claims · 8 setups
Babelomics is presented as an updated, complete suite of web tools for functional analysis of genome-scale experiments with new and improved modules