Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
A general definition and nomenclature for alternative splicing events.
PMID 18688268 · PMC2467475 · PLoS computational biology · 2008 · 6 claims · 4 setups
Existing AS nomenclatures (Malko et al.'s 5-letter strings, Nagasaki et al.'s bit matrices, and the ASD/ATD/AEdb system) are redundant, ambiguous, or incapable of representing complex or large splicing variations.
-
Full-text index only
Non-EST based prediction of exon skipping and intron retention events using Pfam information.
PMID 16204458 · PMC1243800 · Nucleic acids research · 2005 · 7 claims · 5 setups
A novel ab initio method predicts exon skipping and intron retention events using only Pfam domain annotation, via a Viterbi-like dynamic programming algorithm applied to the Pfam alignment.
-
Full-text index only
Computational analysis of splicing errors and mutations in human transcripts.
PMID 18194514 · PMC2234086 · BMC genomics · 2008 · 8 claims · 4 setups
Retained introns are significantly shorter than constitutively spliced introns
-
Full-text index only
Integrative analysis of the human cis-antisense gene pairs, miRNAs and their transcription regulation patterns.
PMID 19906709 · PMC2811022 · Nucleic acids research · 2010 · 8 claims · 5 setups
A genome-wide catalog of up to ~9000 overlapping antisense loci (23,782 non-redundant SAT pairs, clustered into 8894) was compiled and stored in the USAGP database
-
Full-text index only
A unique, consistent identifier for alternatively spliced transcript variants.
PMID 19865484 · PMC2765725 · PloS one · 2009 · 6 claims · 1 setups
Existing transcript identifiers (NM_ accessions, ENST identifiers) are unsuitable for uniquely identifying isoform structure across databases, methods, or organisms
-
Full-text index only
NCBI Reference Sequence (RefSeq): a curated non-redundant sequence database of genomes, transcripts and proteins.
PMID 15608248 · PMC539979 · Nucleic acids research · 2005 · 7 claims · 5 setups
RefSeq provides a curated, non-redundant, explicitly linked collection of genomic, transcript and protein sequences spanning prokaryotes, eukaryotes and viruses.
-
Full-text index only
ARED 3.0: the large and diverse AU-rich transcriptome.
PMID 16381826 · PMC1347415 · Nucleic acids research · 2006 · 7 claims · 6 setups
ARED 3.0 computationally mapped more than 4000 ARE-mRNAs to the human genome, representing 5-8% of human genes.
-
Full-text index only
Automatic annotation of eukaryotic genes, pseudogenes and promoters.
PMID 16925832 · PMC1810547 · Genome biology · 2006 · 8 claims · 6 setups
Fgenesh++ gene prediction pipeline identifies 91% of coding nucleotides with 90% specificity
-
Full-text index only
The biological function of some human transcription factor binding motifs varies with position relative to the transcription start site.
PMID 18367472 · PMC2377430 · Nucleic acids research · 2008 · 8 claims · 5 setups
1226 eight-letter DNA words show statistically significant positional preferences relative to the TSS across 7914 human promoter regions
-
Full-text index only
The distribution of SNPs in human gene regulatory regions.
PMID 16209714 · PMC1260019 · BMC genomics · 2005 · 8 claims · 6 setups
SNPs occur with higher density closer to the transcriptional start site within gene promoter regions than in further upstream regions
-
Full-text index only
Computational disease gene identification: a concert of methods prioritizes type 2 diabetes and obesity candidate genes.
PMID 16757574 · PMC1475747 · Nucleic acids research · 2006 · 6 claims · 8 setups
Applying seven independent computational disease-gene prioritization methods in concert to 9556 positional candidate genes identifies a prioritized set of likely T2D and obesity candidate genes
-
Full-text index only
Improving the specificity of exon prediction using comparative genomics.
PMID 18831778 · PMC2559877 · BMC genomics · 2008 · 8 claims · 6 setups
A log-odds ratio scoring method based on codon conservation across human-mouse/human-dog alignments and adjacent-codon dependency can classify putative exons as coding vs non-coding.
-
Full-text index only
Evola: Ortholog database of all human genes in H-InvDB with manual curation of phylogenetic trees.
PMID 17982176 · PMC2238928 · Nucleic acids research · 2008 · 6 claims · 7 setups
Evola combines genome synteny-based computational ortholog detection with manual curation of phylogenetic trees by experts to yield more reliable orthologs than automated pairwise methods
-
Full-text index only
Non-EST-based prediction of novel alternatively spliced cassette exons with cell signaling function in Caenorhabditis elegans and human.
PMID 17452356 · PMC1904267 · Nucleic acids research · 2007 · 8 claims · 7 setups
PASE (Prediction of Alternative Signaling Exons) is a computational algorithm combining Markov splice-site models, a Bayesian classifier, species conservation, and Scansite motif scoring to identify novel alternative cassette exons involved in cell signaling.
-
Full-text index only
PrimerStation: a highly specific multiplex genomic PCR primer design server for the human genome.
PMID 16845094 · PMC1538814 · Nucleic acids research · 2006 · 7 claims · 2 setups
Selecting primers using the stringent hybridization ratio (requiring an 'executable temperature' where target hybridization ratio >0.99 and off-target ratio <0.05) yields more specific genomic primers than the conventional melting-temperature-based approach, which only guarantees >0.5 vs <0.5
-
Full-text index only
ARED Organism: expansion of ARED reveals AU-rich element cluster variations between human and mouse.
PMID 17984078 · PMC2238997 · Nucleic acids research · 2008 · 6 claims · 4 setups
ARED Organism and ARED-Integrated are new/updated public databases cataloguing ARE-containing mRNAs/genes in human, mouse and rat
-
Full-text index only
miRGen: a database for the study of animal microRNA genomic organization and function.
PMID 17108354 · PMC1669779 · Nucleic acids research · 2007 · 8 claims · 6 setups
miRGen is an integrated database combining Genomics, Targets, and Clusters interfaces to study miRNA genomic organization and function across 11 animal genomes
-
Full-text index only
Sequence determinants of human microsatellite variability.
PMID 20015383 · PMC2806349 · BMC genomics · 2009 · 6 claims · 4 setups
Mean and maximum number of repeats across individuals are positively correlated with heterozygosity
-
Full-text index only
A re-annotation pipeline for Illumina BeadArrays: improving the interpretation of gene expression data.
PMID 19923232 · PMC2817484 · Nucleic acids research · 2010 · 8 claims · 7 setups
A Perl-based pipeline that BLASTs/BLATs Illumina probe sequences against genomes and transcript databases (RefSeq, UCSC Known Genes, UniGene/GenBank, Ensembl) can classify probes by quality grade (Perfect/Good/Bad/No match) and is applicable across 8 BeadArray platforms and other array types
-
Full-text index only
ABS: a database of Annotated regulatory Binding Sites from orthologous promoters.
PMID 16381947 · PMC1347478 · Nucleic acids research · 2006 · 7 claims · 6 setups
ABS is a public database of experimentally identified TF binding sites conserved in orthologous vertebrate gene promoters, manually curated from the literature.