Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
Large-scale trends in the evolution of gene structures within 11 animal genomes.
PMID 16518452 · PMC1386723 · PLoS computational biology · 2006 · 8 claims · 5 setups
Change in intron–exon gene structure is gradual, clock-like, and largely independent of coding-sequence (protein) evolution
-
Full-text index only
JIGSAW, GeneZilla, and GlimmerHMM: puzzling out the features of human genes in the ENCODE regions.
PMID 16925843 · PMC1810558 · Genome biology · 2006 · 8 claims · 4 setups
Adding model states for specific biological features (signal peptides, CpG islands, etc.) to non-comparative GHMM gene finders did little or nothing to enhance predictive accuracy, sometimes reducing it.
-
Full-text index only
Functional importance of different patterns of correlation between adjacent cassette exons in human and mouse.
PMID 18439302 · PMC2432081 · BMC genomics · 2008 · 8 claims · 7 setups
Adjacent cassette exon pairs can be categorized by EST-derived correlation coefficient into three groups: mutually exclusive (ME, r<=-0.7), independent (IND, -0.2<=r<=0.2), and linked (LNK, r>=0.7)
-
Full-text index only
Vertebrate gene finding from multiple-species alignments using a two-level strategy.
PMID 16925840 · PMC1810555 · Genome biology · 2006 · 8 claims · 5 setups
DOGFISH cleanly separates a multi-species alignment classifier (RVM cascade) from an HMM-based structure predictor, avoiding tight coupling of alignment complexity with HMM formalism
-
Full-text index only
NCBI Reference Sequences: current status, policy and new initiatives.
PMID 18927115 · PMC2686572 · Nucleic acids research · 2009 · 7 claims · 5 setups
RefSeq is a curated, non-redundant, explicitly linked database of nucleotide and protein sequences spanning genomes, transcripts and proteins across prokaryotes, eukaryotes and viruses
-
Has reproduction · 86
RNASEQR--a streamlined and accurate RNA-seq sequence analysis program.
PMID 22199257 · PMC3315322 · Nucleic acids research · 2012 · 8 claims · 7 setups
RNASEQR is a new RNA-seq mapper/aligner that combines a BWT-based (Bowtie) transcriptomic/genomic alignment with hash-based BLAT local alignment in three sequential steps: transcriptome mapping, novel exon detection, and anchor-and-align novel splice junction identification.
-
Full-text index only
Human SNPs resulting in premature stop codons and protein truncation.
PMID 16595072 · PMC3500177 · Human genomics · 2006 · 8 claims · 6 setups
Genome-wide screening of dbSNP identified 28 validated X-SNPs from 28 genes with known minor allele frequencies.
-
Full-text index only
Low conservation and species-specific evolution of alternative splicing in humans and mice: comparative genomics analysis using well-annotated full-length cDNAs.
PMID 18838389 · PMC2582632 · Nucleic acids research · 2008 · 7 claims · 8 setups
Although 86% of individual human exons are conserved in the mouse genome, only a small fraction (431/20392, ~2%) of human AS variants are perfectly conserved AS variants in mice.
-
Full-text index only
SVC: structured visualization of evolutionary sequence conservation.
PMID 15991338 · PMC1160265 · Nucleic acids research · 2005 · 7 claims · 5 setups
SVC aligns protein-coding sequences of orthologous gene pairs and maps them back onto their encoding exons/introns to generate a scaffold of conserved gene structure.
-
Full-text index only
Systematic identification of pseudogenes through whole genome expression evidence profiling.
PMID 16945953 · PMC1636364 · Nucleic acids research · 2006 · 8 claims · 8 setups
Developed a novel bioinformatics method that identifies pseudogenes by profiling whole-genome transcript and protein expression evidence
-
Full-text index only
GENCODE: producing a reference annotation for ENCODE.
PMID 16925838 · PMC1810553 · Genome biology · 2006 · 8 claims · 8 setups
GENCODE annotation combines initial manual annotation by HAVANA, experimental validation, and refinement based on results to identify protein-coding genes in ENCODE regions
-
Full-text index only
EGASP: Introduction.
PMID 16925831 · PMC1810546 · Genome biology · 2006 · 8 claims · 5 setups
Computational gene finding methods, when compared to the GENCODE golden standard annotation, show that the human genome annotation is nearly complete in terms of novel protein-coding loci.
-
Full-text index only
Identification and evolutionary analysis of novel exons and alternative splicing events using cross-species EST-to-genome comparisons in human, mouse and rat.
PMID 16536879 · PMC1479377 · BMC bioinformatics · 2006 · 8 claims · 6 setups
ENACE, a cross-species EST-to-genome comparison algorithm, can identify novel cassette-on exons and retained introns for EST-scanty species and distinguish conserved vs lineage-specific exons
-
Full-text index only
miRBase: tools for microRNA genomics.
PMID 17991681 · PMC2238936 · Nucleic acids research · 2008 · 8 claims · 6 setups
miRBase release 10.0 contains 5071 miRNA hairpin loci from 58 species, expressing 5922 distinct mature miRNA sequences, a growth of over 2000 sequences in 2 years
-
Full-text index only
Non-EST-based prediction of novel alternatively spliced cassette exons with cell signaling function in Caenorhabditis elegans and human.
PMID 17452356 · PMC1904267 · Nucleic acids research · 2007 · 8 claims · 7 setups
PASE (Prediction of Alternative Signaling Exons) is a computational algorithm combining Markov splice-site models, a Bayesian classifier, species conservation, and Scansite motif scoring to identify novel alternative cassette exons involved in cell signaling.
-
Full-text index only
Distribution and effects of nonsense polymorphisms in human genes.
PMID 18852891 · PMC2561068 · PloS one · 2008 · 8 claims · 8 setups
Nonsense SNPs occur at a lower density than nonsynonymous SNPs, indicating stronger purifying selection against premature stop codons than amino acid changes.
-
Full-text index only
PLANdbAffy: probe-level annotation database for Affymetrix expression microarrays.
PMID 19906711 · PMC2808952 · Nucleic acids research · 2010 · 6 claims · 4 setups
PLANdbAffy is a database of Affymetrix probe alignments to the human genome for five widely used arrays (HG-U133A, HG-U133B, HG-U133 Plus 2.0, Human Exon 1.0, Human Gene 1.0)
-
Full-text index only
SelenoDB 1.0 : a database of selenoprotein genes, proteins and SECIS elements.
PMID 18174224 · PMC2238826 · Nucleic acids research · 2008 · 6 claims · 5 setups
Standard genome annotation pipelines misannotate selenoprotein genes because they rely on UGA as a universal stop codon, failing to recognize its dual role as the selenocysteine-recoding codon.
-
Has reproduction · 66
RNAseq analysis of the parasitic nematode Strongyloides stercoralis reveals divergent regulation of canonical dauer pathways.
PMID 23145190 · PMC3493385 · PLoS neglected tropical diseases · 2012 · 8 claims · 8 setups
S. stercoralis possesses homologs of nearly all C. elegans dauer genes, but with significant differences in protein structure, developmental regulation, and gene family expansion.
-
Full-text index only
AceView: a comprehensive cDNA-supported gene and transcripts annotation.
PMID 16925834 · PMC1810549 · Genome biology · 2006 · 8 claims · 4 setups
At the mRNA level, AceView transcripts are the closest match to Gencode transcripts among all evaluated methods, including alternative splice variants