Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
Annotation and analysis of 10,000 expressed sequence tags from developing mouse eye and adult retina.
PMID 14519200 · PMC328454 · Genome biology · 2003 · 8 claims · 5 setups
Annotation of 8,633 high-quality non-mitochondrial/non-ribosomal ESTs shows 57% represent known genes and 43% are unknown or novel, with M15E having the highest proportion of novel ESTs
-
Has reproduction · 89
DFAST and DAGA: web-based integrated genome annotation tools and resources.
PMID 27867804 · PMC5107635 · Bioscience of microbiota, food and health · 2016 · 8 claims · 7 setups
DFAST is a web-based bacterial genome annotation and DDBJ submission pipeline with integrated CheckM quality assessment and ANI taxonomic assessment.
-
Has reproduction · 100
FA-nf: A Functional Annotation Pipeline for Proteins from Non-Model Organisms Implemented in Nextflow.
PMID 34681040 · PMC8535801 · Genes · 2021 · 8 claims · 4 setups
FA-nf, implemented in Nextflow with Docker/Singularity containerization, integrates NCBI BLAST+, DIAMOND, InterProScan, and KEGG (KAAS/KofamKOALA) into a single functional annotation pipeline.
-
Has reproduction · 84
Pharokka: a fast scalable bacteriophage annotation tool.
PMID 36453861 · PMC9805569 · Bioinformatics (Oxford, England) · 2023 · 8 claims · 5 setups
Pharokka is a one-line, fast, scalable bacteriophage annotation tool producing standards-compliant outputs, installable via a two-line bioconda command
-
Full-text index only
Molecular archeology of L1 insertions in the human genome.
PMID 12372140 · PMC134481 · Genome biology · 2002 · 8 claims · 4 setups
TSDfinder, a new algorithm, refines RepeatMasker-identified L1 boundaries by locating poly(A) tails, TSDs, and inversion breakpoints
-
Full-text index only
POCUS: mining genomic sequence annotation to predict disease genes.
PMID 14611661 · PMC329128 · Genome biology · 2003 · 8 claims · 6 setups
Genes predisposing to the same disease tend to share functional annotation IDs (GO/InterPro) more than expected by chance
-
Full-text index only
Integrative annotation of 21,037 human genes validated by full-length cDNA clones.
PMID 15103394 · PMC393292 · PLoS biology · 2004 · 8 claims · 5 setups
41,118 full-length human cDNAs from six high-throughput sequencing projects were exhaustively integratively characterized
-
Full-text index only
Genome annotation errors in pathway databases due to semantic ambiguity in partial EC numbers.
PMID 16034025 · PMC1179732 · Nucleic acids research · 2005 · 7 claims · 4 setups
Partial EC numbers are semantically ambiguous, and databases that assign a gene to all reactions sharing the same partial EC number make a faulty inference, causing systematic misannotation.
-
Full-text index only
Ontological visualization of protein-protein interactions.
PMID 15707487 · PMC550656 · BMC bioinformatics · 2005 · 8 claims · 8 setups
Aggregating independently made GO 'protein binding' (IPI) annotations reveals larger, previously undescribed mouse protein-protein interaction networks
-
Full-text index only
PeroxisomeDB: a database for the peroxisomal proteome, functional genomics and disease.
PMID 17135190 · PMC1747181 · Nucleic acids research · 2007 · 8 claims · 6 setups
PeroxisomeDB integrates the complete peroxisomal proteome of Homo sapiens and Saccharomyces cerevisiae into interrelated 'Genes', 'Functions', 'Metabolic pathways' and 'Diseases' sections with links to NCBI, ENSEMBL and UCSC
-
Full-text index only
GENCODE: producing a reference annotation for ENCODE.
PMID 16925838 · PMC1810553 · Genome biology · 2006 · 8 claims · 8 setups
GENCODE annotation combines initial manual annotation by HAVANA, experimental validation, and refinement based on results to identify protein-coding genes in ENCODE regions
-
Full-text index only
Automatic annotation of eukaryotic genes, pseudogenes and promoters.
PMID 16925832 · PMC1810547 · Genome biology · 2006 · 8 claims · 6 setups
Fgenesh++ gene prediction pipeline identifies 91% of coding nucleotides with 90% specificity
-
Full-text index only
The Edinburgh human metabolic network reconstruction and its functional analysis.
PMID 17882155 · PMC2013923 · Molecular systems biology · 2007 · 8 claims · 7 setups
EHMN is a high-quality, manually curated human metabolic network combining genome-based and literature-based (EMP) reconstruction, containing nearly 3000 reactions and over 2000 metabolic genes.
-
Full-text index only
Comprehensive annotation of bidirectional promoters identifies co-regulation among breast and ovarian cancer genes.
PMID 17447839 · PMC1853124 · PLoS computational biology · 2007 · 8 claims · 8 setups
A new algorithm using spliced ESTs (cross-validated against Known Genes and GenBank mRNA) comprehensively maps bidirectional promoters in the human genome
-
Full-text index only
Functional annotation and identification of candidate disease genes by computational analysis of normal tissue gene expression data.
PMID 18560577 · PMC2409962 · PloS one · 2008 · 7 claims · 5 setups
Ranked Coexpression Groups (RCG) built from k=6 nearest coexpressed genes, combined with a majority-rule functional characterization, integrate multiple datasets/coexpression measures to generate high-confidence functional annotation predictions
-
Full-text index only
Manual annotation and analysis of the defensin gene cluster in the C57BL/6J mouse reference genome.
PMID 20003482 · PMC2807441 · BMC genomics · 2009 · 8 claims · 6 setups
Manual annotation of the mouse Chromosome 8 defensin region identifies 98 gene loci: 54 in the alpha-defensin cluster and 44 in the beta-defensin cluster
-
Full-text index only
A metadata approach for clinical data management in translational genomics studies in breast cancer.
PMID 19948017 · PMC3225860 · BMC medical genomics · 2009 · 8 claims · 5 setups
A metadata/CDE-based approach using CancerGrid's semantic web tools enables automatic integration of heterogeneous clinical datasets without loss of original detail
-
Full-text index only
DAVID Knowledgebase: a gene-centered database integrating heterogeneous gene annotation resources to facilitate high-throughput gene functional analysis.
PMID 17980028 · PMC2186358 · BMC bioinformatics · 2007 · 7 claims · 3 setups
The DAVID Gene Concept, a single-linkage algorithm, merges gene clusters from Entrez Gene, UniRef100, and PIR-NREF100 that share protein IDs and species into unified DAVID gene clusters, improving cross-referencing between NCBI and UniProt systems
-
Full-text index only
Functional classification using phylogenomic inference.
PMID 16846248 · PMC1484587 · PLoS computational biology · 2006 · 8 claims · 1 setups
Functional annotation via top-hit database search transfer is used far more often in practice than phylogenomic inference, despite phylogenomic inference being more accurate.
-
Has reproduction · 91
De Novo Assembly and Annotation of the Larval Transcriptome of Two Spadefoot Toads Widely Divergent in Developmental Rate.
PMID 31217263 · PMC6686947 · G3 (Bethesda, Md.) · 2019 · 8 claims · 8 setups
De novo transcriptome assemblies were generated for larval P. cultripes and S. couchii, providing new genomic resources for spadefoot toads