Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
Exogean: a framework for annotating protein-coding genes in eukaryotic genomic DNA.
PMID 16925841 · PMC1810556 · Genome biology · 2006 · 8 claims · 5 setups
Exogean is a framework using directed acyclic coloured multigraphs (DACMs) to represent biological objects (mRNA, ESTs, protein alignments, exons) and iteratively combine them into complex protein-coding transcript models.
-
Full-text index only
A general definition and nomenclature for alternative splicing events.
PMID 18688268 · PMC2467475 · PLoS computational biology · 2008 · 6 claims · 4 setups
Existing AS nomenclatures (Malko et al.'s 5-letter strings, Nagasaki et al.'s bit matrices, and the ASD/ATD/AEdb system) are redundant, ambiguous, or incapable of representing complex or large splicing variations.
-
Has reproduction · 78
annotate_my_genomes: an easy-to-use pipeline to improve genome annotation and uncover neglected genes by hybrid RNA sequencing.
PMID 36472574 · PMC9724561 · GigaScience · 2022 · 7 claims · 8 setups
annotate_my_genomes is an easy-to-use genome-guided pipeline that uses hybrid (PacBio+Illumina) assembled transcripts to distinguish coding genes from long non-coding RNAs and reconcile them with prior annotations.
-
Full-text index only
EGASP: the human ENCODE Genome Annotation Assessment Project.
PMID 16925836 · PMC1810551 · Genome biology · 2006 · 8 claims · 6 setups
Best-performing computational gene prediction methods correctly predict at least one transcript for close to 70% of annotated genes in the ENCODE regions.
-
Has reproduction · 50
Polymorphism identification and improved genome annotation of Brassica rapa through Deep RNA sequencing.
PMID 25122667 · PMC4232532 · G3 (Bethesda, Md.) · 2014 · 8 claims · 8 setups
330,995 SNPs were identified in transcribed regions between B. rapa genotypes R500 and IMB211, at an average frequency of one SNP per 200 bases.
-
Full-text index only
Pairagon+N-SCAN_EST: a model-based gene annotation pipeline.
PMID 16925839 · PMC1810554 · Genome biology · 2006 · 7 claims · 5 setups
Pairagon+N-SCAN_EST, using only native alignments, was as accurate as ENSEMBL and ExoGean in the EGASP mRNA/EST evidence assessment
-
Has reproduction · 88
nf-core/isoseq: simple gene and isoform annotation with PacBio Iso-Seq long-read sequencing.
PMID 36961337 · PMC10199315 · Bioinformatics (Oxford, England) · 2023 · 7 claims · 4 setups
nf-core/isoseq is a new automated Nextflow-based pipeline that processes raw Iso-Seq subreads through to genome annotation (BED format) without requiring transcriptome assembly.
-
Has reproduction · 65
FusionQ: a novel approach for gene fusion detection and quantification from paired-end RNA-Seq.
PMID 23768108 · PMC3691734 · BMC bioinformatics · 2013 · 8 claims · 8 setups
FusionQ is a novel tool that detects gene fusions, constructs chimerical transcript structures, and estimates their abundances from paired-end RNA-Seq data.
-
Full-text index only
Reference based annotation with GeneMapper.
PMID 16600017 · PMC1557983 · Genome biology · 2006 · 7 claims · 6 setups
GeneMapper transfers reference gene annotations to target genomes with higher accuracy than GeneWise and Projector
-
Full-text index only
Ensembl 2007.
PMID 17148474 · PMC1761443 · Nucleic acids research · 2007 · 8 claims · 7 setups
Ensembl added 18 new chordate genomes this year, increasing total genomes available from 15 to 33, the largest yearly increase to date.
-
Has reproduction · 68
Cell-type annotation with accurate unseen cell-type identification using multiple references.
PMID 37379341 · PMC10335708 · PLoS computational biology · 2023 · 8 claims · 4 setups
mtANN integrates multiple reference datasets and eight gene selection methods via ensemble learning (multiple deep classification models + majority voting) to improve cell-type annotation accuracy
-
Full-text index only
Pathway projector: web-based zoomable pathway browser using KEGG atlas and Google Maps API.
PMID 19907644 · PMC2770834 · PloS one · 2009 · 8 claims · 6 setups
Existing pathway databases and tools do not satisfy all requirements for a generic, comprehensive pathway browser (integrated maps, data access, mapping/editing, export, installation-free availability).
-
Full-text index only
NCBI Reference Sequences: current status, policy and new initiatives.
PMID 18927115 · PMC2686572 · Nucleic acids research · 2009 · 7 claims · 5 setups
RefSeq is a curated, non-redundant, explicitly linked database of nucleotide and protein sequences spanning genomes, transcripts and proteins across prokaryotes, eukaryotes and viruses
-
Full-text index only
Recent additions and improvements to the Onto-Tools.
PMID 15980579 · PMC1160233 · Nucleic acids research · 2005 · 7 claims · 3 setups
The Onto-Tools back-end database was redesigned around the Entrez Gene data model after NCBI phased out LocusLink in February 2005.
-
Full-text index only
Ensembl 2006.
PMID 16381931 · PMC1347495 · Nucleic acids research · 2006 · 8 claims · 5 setups
Ensembl now provides annotation for 19 genomes, up from 4 the previous year, including new mammalian (Rhesus macaque, Opossum), chordate (Ciona intestinalis), and yeast genomes.
-
Has reproduction · 89
DFAST and DAGA: web-based integrated genome annotation tools and resources.
PMID 27867804 · PMC5107635 · Bioscience of microbiota, food and health · 2016 · 8 claims · 7 setups
DFAST is a web-based bacterial genome annotation and DDBJ submission pipeline with integrated CheckM quality assessment and ANI taxonomic assessment.
-
Full-text index only
Variation analysis and gene annotation of eight MHC haplotypes: the MHC Haplotype Project.
PMID 18193213 · PMC2206249 · Immunogenetics · 2008 · 8 claims · 6 setups
Comparison of eight HLA-homozygous MHC haplotype sequences identified >44,000 variations (substitutions and indels), submitted to dbSNP
-
Has reproduction · 92
Telomere-to-telomere reference genome for Panax ginseng highlights the evolution of saponin biosynthesis.
PMID 38883331 · PMC11179851 · Horticulture research · 2024 · 8 claims · 8 setups
A telomere-to-telomere reference genome of P. ginseng was assembled (3.45 Gb, 24 chromosomes, 77266 protein-coding genes)
-
Full-text index only
BioGPS: an extensible and customizable portal for querying and organizing gene annotation resources.
PMID 19919682 · PMC3091323 · Genome biology · 2009 · 8 claims · 4 setups
BioGPS aggregates distributed, third-party gene annotation resources into a single customizable portal for human, mouse, and rat genes.
-
Has reproduction · 60
TRAPID 2.0: a web application for taxonomic and functional analysis of de novo transcriptomes.
PMID 34197621 · PMC8464036 · Nucleic acids research · 2021 · 8 claims · 8 setups
TRAPID 2.0 is a web application performing global characterization of de novo transcriptomes via structural, functional, and taxonomic annotation in an initial processing phase, followed by an exploratory phase of downstream analyses.