Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
BIPASS: BioInformatics Pipeline Alternative Splicing Services.
PMID 17584795 · PMC1933140 · Nucleic acids research · 2007 · 8 claims · 4 setups
BIPASS offers two complementary services for alternative splicing (AS) research: BIPAS-SpliceDB, a queryable pre-computed AS data warehouse, and BIPAS-Align&Splice, an online pipeline for user-submitted sequences.
-
Full-text index only
Genome annotation errors in pathway databases due to semantic ambiguity in partial EC numbers.
PMID 16034025 · PMC1179732 · Nucleic acids research · 2005 · 7 claims · 4 setups
Partial EC numbers are semantically ambiguous, and databases that assign a gene to all reactions sharing the same partial EC number make a faulty inference, causing systematic misannotation.
-
Has reproduction · 68
Cell-type annotation with accurate unseen cell-type identification using multiple references.
PMID 37379341 · PMC10335708 · PLoS computational biology · 2023 · 8 claims · 4 setups
mtANN integrates multiple reference datasets and eight gene selection methods via ensemble learning (multiple deep classification models + majority voting) to improve cell-type annotation accuracy
-
Full-text index only
PeroxisomeDB: a database for the peroxisomal proteome, functional genomics and disease.
PMID 17135190 · PMC1747181 · Nucleic acids research · 2007 · 8 claims · 6 setups
PeroxisomeDB integrates the complete peroxisomal proteome of Homo sapiens and Saccharomyces cerevisiae into interrelated 'Genes', 'Functions', 'Metabolic pathways' and 'Diseases' sections with links to NCBI, ENSEMBL and UCSC
-
Full-text index only
Gene Prospector: an evidence gateway for evaluating potential susceptibility genes and interacting risk factors for human diseases.
PMID 19063745 · PMC2613935 · BMC bioinformatics · 2008 · 8 claims · 5 setups
Gene Prospector is a Web-based application that selects and prioritizes potential disease-related genes using a curated, updated literature database of genetic association studies
-
Full-text index only
Columba: an integrated database of proteins, structures, and annotations.
PMID 15801979 · PMC1087474 · BMC bioinformatics · 2005 · 8 claims · 6 setups
COLUMBA physically integrates data from twelve protein structure-related databases (PDB, KEGG, Swiss-Prot, CATH, SCOP, Gene Ontology, ENZYME, etc.) into a single PostgreSQL data warehouse.
-
Full-text index only
L1Base: from functional annotation to prediction of active LINE-1 elements.
PMID 15608246 · PMC539998 · Nucleic acids research · 2005 · 7 claims · 6 setups
L1Base is a database of putatively active LINE-1 insertions in human, mouse and rat genomes, containing FLI-L1s (intact in both ORFs), ORF2-L1s (intact ORF2, disrupted ORF1), and FLnI-L1s (full-length, >6000 bp, non-intact)
-
Full-text index only
TRED: a Transcriptional Regulatory Element Database and a platform for in silico gene regulation studies.
PMID 15608156 · PMC539958 · Nucleic acids research · 2005 · 8 claims · 5 setups
TRED is a database collecting both cis-regulatory elements (promoters) and trans-regulatory elements (transcription factor binding/regulation data) with linked access.
-
Full-text index only
RAId_DbS: mass-spectrometry based peptide identification web server with knowledge integration.
PMID 18954448 · PMC2605478 · BMC genomics · 2008 · 7 claims · 4 setups
Constructed enhanced protein databases integrating annotated SAPs, PTMs, and disease associations for 17 organisms.
-
Has reproduction · 60
TRAPID 2.0: a web application for taxonomic and functional analysis of de novo transcriptomes.
PMID 34197621 · PMC8464036 · Nucleic acids research · 2021 · 8 claims · 8 setups
TRAPID 2.0 is a web application performing global characterization of de novo transcriptomes via structural, functional, and taxonomic annotation in an initial processing phase, followed by an exploratory phase of downstream analyses.
-
Has reproduction · 94
Systematic assessment of pathway databases, based on a diverse collection of user-submitted experiments.
PMID 36088548 · PMC9487593 · Briefings in bioinformatics · 2022 · 8 claims · 6 setups
Well-established, hierarchically organized pathway annotation systems (e.g. GO, Reactome, KEGG) yield the best overall enrichment performance despite covering much of the human genome only in general terms.
-
Full-text index only
SUPERFAMILY--sophisticated comparative genomics, data mining, visualization and phylogeny.
PMID 19036790 · PMC2686452 · Nucleic acids research · 2009 · 7 claims · 6 setups
SUPERFAMILY provides structural, functional and evolutionary annotation for proteins from all completely sequenced genomes using SCOP-based hidden Markov models
-
Has reproduction · 78
A case study for large-scale human microbiome analysis using JCVI's metagenomics reports (METAREP).
PMID 22719821 · PMC3374610 · PloS one · 2012 · 8 claims · 7 setups
METAREP version 1.3.1 is an open-source, scalable tool for querying, browsing and comparing extremely large volumes of metagenomic annotations, with an extended data model, dynamic weighting, distributed searches and advanced clustering.
-
Full-text index only
Discovery of protein-protein interactions using a combination of linguistic, statistical and graphical information.
PMID 15941473 · PMC1164402 · BMC bioinformatics · 2005 · 8 claims · 5 setups
A combined linguistic+statistical+rule-based method achieves precision 0.61 and recall 0.97 (f=0.74) detecting yeast protein-protein interactions across 12,300 Medline abstracts.
-
Has reproduction · 87
R2DT is a framework for predicting and visualising RNA secondary structure using templates.
PMID 34108470 · PMC8190129 · Nature communications · 2021 · 8 claims · 6 setups
R2DT is a template-based computational framework/pipeline that predicts and visualises RNA 2D structure in standardised, community-accepted layouts
-
Full-text index only
MACSIMS: multiple alignment of complete sequences information management system.
PMID 16792820 · PMC1539025 · BMC bioinformatics · 2006 · 8 claims · 5 setups
MACSIMS is a multiple alignment-based information management system combining knowledge-based database mining with ab initio sequence predictions
-
Full-text index only
An integrated database-pipeline system for studying single nucleotide polymorphisms and diseases.
PMID 19091018 · PMC2638159 · BMC bioinformatics · 2008 · 6 claims · 5 setups
Existing SNP/disease databases are fragmented; no combined resource widely supports gene-, SNP-, and disease-related information together
-
Full-text index only
Gene-disease relationship discovery based on model-driven data integration and database view definition.
PMID 19042916 · PMC2639000 · Bioinformatics (Oxford, England) · 2009 · 8 claims · 4 setups
Explicit gene–disease relationships can be formulated as candidate gene definitions (e.g., co-localization, dysregulation, functional similarity) that may include intermediary orthologous or interacting genes
-
Full-text index only
A genome-wide survey of segmental duplications that mediate common human genetic variation of chromosomal architecture.
PMID 15588494 · PMC3525102 · Human genomics · 2004 · 8 claims · 5 setups
PSD-mediated genomic architecture analogous to the 8p23/4p16 inversion regions is not unique to those loci but recurs genome-wide.
-
Full-text index only
Satellog: a database for the identification and prioritization of satellite repeats in disease association studies.
PMID 15949044 · PMC1181805 · BMC bioinformatics · 2005 · 7 claims · 6 setups
Satellog is a database cataloging all pure 1-16 unit satellite repeats in the human genome with supplementary polymorphism, gene-location, and expression data for prioritizing repeats in disease-association studies.