Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
Motif discovery in promoters of genes co-localized and co-expressed during myeloid cells differentiation.
PMID 19059999 · PMC2632922 · Nucleic acids research · 2009 · 6 claims · 8 setups
A novel multi-step computational method (built on approximate pattern enumeration, binomial over-representation scoring with FDR correction, and k-medoids clustering) can identify over-represented motifs in a selected set of promoters relative to a background promoter set.
-
Has reproduction · 50
Ancient gene duplicates in Gossypium (cotton) exhibit near-complete expression divergence.
PMID 24558256 · PMC3971588 · Genome biology and evolution · 2014 · 8 claims · 8 setups
Nearly all (99.4%) ancient paralog pairs in Gossypium raimondii are differentially expressed in at least one of three tissues (petal, leaf, seed), indicating massive, near-complete expression-level divergence.
-
Has reproduction · 49
oPOSSUM-3: advanced analysis of regulatory motif over-representation across genes or ChIP-Seq datasets.
PMID 22973536 · PMC3429929 · G3 (Bethesda, Md.) · 2012 · 8 claims · 6 setups
oPOSSUM-3 is a web-accessible system that identifies over-represented TFBS and TFBS families in DNA sequences of co-expressed genes or in sequences from high-throughput methods such as ChIP-Seq.
-
Full-text index only
DG-CST (Disease Gene Conserved Sequence Tags), a database of human-mouse conserved elements associated to disease genes.
PMID 15608249 · PMC539965 · Nucleic acids research · 2005 · 5 claims · 8 setups
Comparative human-mouse genome analysis identifies conserved sequence tags (CSTs, >=70% identity over >=100bp) that frequently correspond to non-coding elements with putative regulatory or structural roles
-
Full-text index only
Applications for protein sequence-function evolution data: mRNA/protein expression analysis and coding SNP scoring tools.
PMID 16912992 · PMC1538848 · Nucleic acids research · 2006 · 7 claims · 8 setups
PANTHER HMMs built from family/subfamily multiple sequence alignments can classify novel protein sequences into functional groups based on statistically significant HMM match scores
-
Full-text index only
Human Lsg1 defines a family of essential GTPases that correlates with the evolution of compartmentalization.
PMID 16209721 · PMC1262696 · BMC biology · 2005 · 8 claims · 9 setups
hLsg1 is the human orthologue of yeast Lsg1p and defines a family of circularly permuted GTPases named YRG (YlqF Related GTPases)
-
Full-text index only
Dog Y chromosomal DNA sequence: identification, sequencing and SNP discovery.
PMID 17026745 · PMC1630699 · BMC genetics · 2006 · 8 claims · 6 setups
Identified 32 male-specific Y-chromosome sequences totaling 24159 bp via combined Blast (human Y chromosome match, absence from female dog genome) and PCR male-specificity screening of a male poodle shotgun genome.
-
Has reproduction · 43
StatsDB: platform-agnostic storage and understanding of next generation sequencing run metrics.
PMID 24627795 · PMC3938176 · F1000Research · 2013 · 8 claims · 6 setups
StatsDB is an open-source software package for storage and analysis of next generation sequencing run metrics, backed by an SQL (MySQL) database with Perl and Java APIs.
-
Full-text index only
HIV-1 sequence evolution in vivo after superinfection with three viral strains.
PMID 17716368 · PMC2020475 · Retrovirology · 2007 · 8 claims · 8 setups
gag and env-V3 nucleotide evolution follows a similar pattern in all three strains: low substitution rate in the first 2-3 years of infection, then an increase driven mainly by synonymous substitutions
-
Full-text index only
Sequence occurrence and structural uniqueness of a G-quadruplex in the human c-kit promoter.
PMID 17720713 · PMC2034477 · Nucleic acids research · 2007 · 8 claims · 4 setups
The native 22-nt c-kit87 sequence occurs only once in the entire human genome.
-
Full-text index only
Toward the use of genomics to study microevolutionary change in bacteria.
PMID 19855823 · PMC2756242 · PLoS genetics · 2009 · 7 claims · 5 setups
The clonal population structure of bacteria, combined with occasional DNA import, provides a powerful context for identifying genetic bases of adaptive phenotypes via association studies.
-
Full-text index only
SUPERFAMILY--sophisticated comparative genomics, data mining, visualization and phylogeny.
PMID 19036790 · PMC2686452 · Nucleic acids research · 2009 · 7 claims · 6 setups
SUPERFAMILY provides structural, functional and evolutionary annotation for proteins from all completely sequenced genomes using SCOP-based hidden Markov models
-
Full-text index only
GenBank.
PMID 18940867 · PMC2686462 · Nucleic acids research · 2009 · 8 claims · 4 setups
GenBank is a comprehensive public database of nucleotide sequences with bibliographic and biological annotation, growing exponentially with a current doubling time of ~30 months.
-
Full-text index only
Species-specific protein sequence and fold optimizations.
PMID 12487631 · PMC139977 · BMC bioinformatics · 2002 · 7 claims · 7 setups
Environmental niche is a significant factor explaining variability in amino acid composition across 100 complete genomes
-
Has reproduction · 86
RNASEQR--a streamlined and accurate RNA-seq sequence analysis program.
PMID 22199257 · PMC3315322 · Nucleic acids research · 2012 · 8 claims · 7 setups
RNASEQR is a new RNA-seq mapper/aligner that combines a BWT-based (Bowtie) transcriptomic/genomic alignment with hash-based BLAT local alignment in three sequential steps: transcriptome mapping, novel exon detection, and anchor-and-align novel splice junction identification.
-
Full-text index only
The gentle art of gene arrangement: the meaning of gene clusters.
PMID 11897017 · PMC139018 · Genome biology · 2002 · 8 claims · 7 setups
Gene order in eukaryotic genomes is likely optimized by natural selection rather than arising purely by chance reshuffling.
-
Full-text index only
A novel polymorphism in the 1A promoter region of the vitamin D receptor is associated with altered susceptibilty and prognosis in malignant melanoma.
PMID 15238985 · PMC2364794 · British journal of cancer · 2004 · 7 claims · 6 setups
A novel A-1012G (adenine-guanine) polymorphism exists in the VDR exon 1a promoter region, identified by SSCP screening and sequencing
-
Full-text index only
Human-zebrafish non-coding conserved elements act in vivo to regulate transcription.
PMID 16179648 · PMC1236720 · Nucleic acids research · 2005 · 8 claims · 4 setups
Deeply conserved human-zebrafish non-coding elements are enriched for in vivo cis-acting transcriptional regulatory activity.
-
Full-text index only
ARED 3.0: the large and diverse AU-rich transcriptome.
PMID 16381826 · PMC1347415 · Nucleic acids research · 2006 · 7 claims · 6 setups
ARED 3.0 computationally mapped more than 4000 ARE-mRNAs to the human genome, representing 5-8% of human genes.
-
Full-text index only
The specificity and polymorphism of the MHC class I prevents the global adaptation of HIV-1 to the monomorphic proteasome and TAP.
PMID 18949050 · PMC2569417 · PloS one · 2008 · 6 claims · 5 setups
Within individual hosts, proteasome and TAP escape mutations in HIV-1 occur frequently