Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
AUGUSTUS at EGASP: using EST, protein and genomic alignments for improved gene prediction in the human genome.
PMID 16925833 · PMC1810548 · Genome biology · 2006 · 8 claims · 5 setups
AUGUSTUS predicted significantly more genes correctly than any other ab initio program in EGASP
-
Has reproduction · 84
Foster thy young: enhanced prediction of orphan genes in assembled genomes.
PMID 34928390 · PMC9023268 · Nucleic acids research · 2022 · 7 claims · 8 setups
Each of the five gene-prediction pipelines under-predicts orphan genes, with detection as low as 11% under one prediction scenario.
-
Full-text index only
JIGSAW, GeneZilla, and GlimmerHMM: puzzling out the features of human genes in the ENCODE regions.
PMID 16925843 · PMC1810558 · Genome biology · 2006 · 8 claims · 4 setups
Adding model states for specific biological features (signal peptides, CpG islands, etc.) to non-comparative GHMM gene finders did little or nothing to enhance predictive accuracy, sometimes reducing it.
-
Full-text index only
Pegasys: software for executing and integrating analyses of biological sequences.
PMID 15096276 · PMC406494 · BMC bioinformatics · 2004 · 8 claims · 7 setups
Pegasys is a flexible, modular, customizable software system for executing and integrating heterogeneous biological sequence analysis tools
-
Full-text index only
SelenoDB 1.0 : a database of selenoprotein genes, proteins and SECIS elements.
PMID 18174224 · PMC2238826 · Nucleic acids research · 2008 · 6 claims · 5 setups
Standard genome annotation pipelines misannotate selenoprotein genes because they rely on UGA as a universal stop codon, failing to recognize its dual role as the selenocysteine-recoding codon.
-
Full-text index only
The gene guessing game.
PMID 11025532 · PMC2448377 · Yeast (Chichester, England) · 2000 · 8 claims · 6 setups
Published methods for estimating human gene number diverge widely, from ~30,000 to over 140,000 genes.
-
Full-text index only
Analysis of protein sequence and interaction data for candidate disease gene prediction.
PMID 17020920 · PMC1636487 · Nucleic acids research · 2006 · 8 claims · 7 setups
Combining CPS and CMP using known disease genes as input achieves sensitivity 0.52 and specificity 0.97, reducing candidate lists 13-fold
-
Full-text index only
Automatic annotation of eukaryotic genes, pseudogenes and promoters.
PMID 16925832 · PMC1810547 · Genome biology · 2006 · 8 claims · 6 setups
Fgenesh++ gene prediction pipeline identifies 91% of coding nucleotides with 90% specificity
-
Has reproduction · 76
The genome and development-dependent transcriptomes of Pyronema confluens: a window into fungal evolution.
PMID 24068976 · PMC3778014 · PLoS genetics · 2013 · 8 claims · 8 setups
The 50 Mb P. confluens genome with 13,369 predicted protein-coding genes is more characteristic of higher filamentous ascomycetes than of the large, repeat-rich Tuber melanosporum genome, showing that the truffle's expanded genome is not typical of the Pezizales.
-
Full-text index only
EGASP: Introduction.
PMID 16925831 · PMC1810546 · Genome biology · 2006 · 8 claims · 5 setups
Computational gene finding methods, when compared to the GENCODE golden standard annotation, show that the human genome annotation is nearly complete in terms of novel protein-coding loci.
-
Has reproduction · 58
Revised annotations, sex-biased expression, and lineage-specific genes in the Drosophila melanogaster group.
PMID 25273863 · PMC4267930 · G3 (Bethesda, Md.) · 2014 · 8 claims · 6 setups
Revised RNA-seq-based gene models for D. ananassae, D. yakuba, and D. simulans include UTRs, empirically verified intron-exon boundaries, and previously unannotated novel exons, improving on r1.3 comparative-genomics annotations that lack UTRs.
-
Has reproduction · 92
Telomere-to-telomere reference genome for Panax ginseng highlights the evolution of saponin biosynthesis.
PMID 38883331 · PMC11179851 · Horticulture research · 2024 · 8 claims · 8 setups
A telomere-to-telomere reference genome of P. ginseng was assembled (3.45 Gb, 24 chromosomes, 77266 protein-coding genes)
-
Full-text index only
Integrating alternative splicing detection into gene prediction.
PMID 15705189 · PMC550657 · BMC bioinformatics · 2005 · 8 claims · 4 setups
An integrative intrinsic/extrinsic method was implemented in the gene finder EuGÈNE (as EuGÈNE-M) to detect AS evidence from aligned transcripts and generate alternative optimal gene predictions consistent with each detected AS event.
-
Full-text index only
Genome reannotation of Escherichia coli CFT073 with new insights into virulence.
PMID 19930606 · PMC2785843 · BMC genomics · 2009 · 8 claims · 7 setups
Reannotation excluded 608 CDSs from the original RefSeq annotation, mostly unfunctional 'hypothetical'/'putative' genes
-
Full-text index only
Non-EST-based prediction of novel alternatively spliced cassette exons with cell signaling function in Caenorhabditis elegans and human.
PMID 17452356 · PMC1904267 · Nucleic acids research · 2007 · 8 claims · 7 setups
PASE (Prediction of Alternative Signaling Exons) is a computational algorithm combining Markov splice-site models, a Bayesian classifier, species conservation, and Scansite motif scoring to identify novel alternative cassette exons involved in cell signaling.
-
Full-text index only
The sequence and de novo assembly of the giant panda genome.
PMID 20010809 · PMC3951497 · Nature · 2010 · 8 claims · 8 setups
A draft giant panda genome was successfully generated and assembled de novo using only Illumina Genome Analyser short-read sequencing