Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
Genome-wide analysis of the human Alu Yb-lineage.
PMID 15588477 · PMC3525081 · Human genomics · 2004 · 8 claims · 6 setups
1,733 Alu Yb-lineage elements are present on human autosomal chromosomes
-
Full-text index only
Columba: an integrated database of proteins, structures, and annotations.
PMID 15801979 · PMC1087474 · BMC bioinformatics · 2005 · 8 claims · 6 setups
COLUMBA physically integrates data from twelve protein structure-related databases (PDB, KEGG, Swiss-Prot, CATH, SCOP, Gene Ontology, ENZYME, etc.) into a single PostgreSQL data warehouse.
-
Full-text index only
Towards alignment independent quantitative assessment of homology detection.
PMID 17205117 · PMC1762415 · PloS one · 2006 · 8 claims · 6 setups
The Fhom Estimator uses the prevalence of a conserved protein feature (X) in two protein sets to estimate the fraction of true homologs among paired proteins, independent of alignment quality.
-
Full-text index only
EGASP: the human ENCODE Genome Annotation Assessment Project.
PMID 16925836 · PMC1810551 · Genome biology · 2006 · 8 claims · 6 setups
Best-performing computational gene prediction methods correctly predict at least one transcript for close to 70% of annotated genes in the ENCODE regions.
-
Full-text index only
Automatic discovery of cross-family sequence features associated with protein function.
PMID 16409628 · PMC1395344 · BMC bioinformatics · 2006 · 8 claims · 6 setups
A self-supervised data mining approach can find relationships between sequence features and functional annotations without preconceived functional categories.
-
Full-text index only
miRBase: tools for microRNA genomics.
PMID 17991681 · PMC2238936 · Nucleic acids research · 2008 · 8 claims · 6 setups
miRBase release 10.0 contains 5071 miRNA hairpin loci from 58 species, expressing 5922 distinct mature miRNA sequences, a growth of over 2000 sequences in 2 years
-
Full-text index only
The DAVID Gene Functional Classification Tool: a novel biological module-centric algorithm to functionally analyze large gene lists.
PMID 17784955 · PMC2375021 · Genome biology · 2007 · 8 claims · 6 setups
Gene-gene functional similarity can be measured using kappa statistics applied to a binary gene-annotation-term matrix built from 14 annotation categories.
-
Has reproduction · 50
Polymorphism identification and improved genome annotation of Brassica rapa through Deep RNA sequencing.
PMID 25122667 · PMC4232532 · G3 (Bethesda, Md.) · 2014 · 8 claims · 8 setups
330,995 SNPs were identified in transcribed regions between B. rapa genotypes R500 and IMB211, at an average frequency of one SNP per 200 bases.
-
Full-text index only
Pathway projector: web-based zoomable pathway browser using KEGG atlas and Google Maps API.
PMID 19907644 · PMC2770834 · PloS one · 2009 · 8 claims · 6 setups
Existing pathway databases and tools do not satisfy all requirements for a generic, comprehensive pathway browser (integrated maps, data access, mapping/editing, export, installation-free availability).
-
Full-text index only
NeMeSys: a biological resource for narrowing the gap between sequence and function in the human pathogen Neisseria meningitidis.
PMID 19818133 · PMC2784325 · Genome biology · 2009 · 8 claims · 5 setups
Determined and manually annotated the complete genome sequence of N. meningitidis clinical isolate strain 8013
-
Full-text index only
Assessing the genomic evidence for conserved transcribed pseudogenes under selection.
PMID 19754956 · PMC2753554 · BMC genomics · 2009 · 8 claims · 8 setups
1750 transcribed pseudogene annotations (TPAs) were identified in the human genome, ~11.5% of all human pseudogene annotations.
-
Full-text index only
Pathway analysis for intracellular Porphyromonas gingivalis using a strain ATCC 33277 specific database.
PMID 19723305 · PMC2753363 · BMC microbiology · 2009 · 8 claims · 5 setups
Using the ATCC 33277-specific genome annotation improves proteome coverage (more proteins identified and more abundance ratios calculated) compared to the W83 annotation
-
Full-text index only
Genome sequences and great expectations.
PMID 11178275 · PMC150431 · Genome biology · 2001 · 8 claims · 3 setups
Function is known or can be predicted for an average of 62% of proteins across 31 analyzed genomes.
-
Full-text index only
TassDB: a database of alternative tandem splice sites.
PMID 17142241 · PMC1669710 · Nucleic acids research · 2007 · 7 claims · 3 setups
TassDB is a relational database storing GYNGYN donor and NAGNAG acceptor tandem splice sites across eight species
-
Has reproduction · 88
nf-core/isoseq: simple gene and isoform annotation with PacBio Iso-Seq long-read sequencing.
PMID 36961337 · PMC10199315 · Bioinformatics (Oxford, England) · 2023 · 7 claims · 4 setups
nf-core/isoseq is a new automated Nextflow-based pipeline that processes raw Iso-Seq subreads through to genome annotation (BED format) without requiring transcriptome assembly.
-
Full-text index only
NCBI Reference Sequences: current status, policy and new initiatives.
PMID 18927115 · PMC2686572 · Nucleic acids research · 2009 · 7 claims · 5 setups
RefSeq is a curated, non-redundant, explicitly linked database of nucleotide and protein sequences spanning genomes, transcripts and proteins across prokaryotes, eukaryotes and viruses
-
Full-text index only
MutScreener: primer design tool for PCR-direct sequencing.
PMID 16845093 · PMC1538803 · Nucleic acids research · 2006 · 8 claims · 4 setups
MutScreener is a web-based application that automates PCR-direct sequencing assay design by annotating gene structure and designing PCR and sequencing primers.
-
Has reproduction · 94
SEAseq: a portable and cloud-based chromatin occupancy analysis suite.
PMID 35193506 · PMC8864840 · BMC bioinformatics · 2022 · 7 claims · 2 setups
SEAseq is a comprehensive, infrastructure-independent pipeline that performs the major analyses needed to process ChIP-Seq/CUT&RUN chromatin binding datasets in a single execution
-
Full-text index only
BIPASS: BioInformatics Pipeline Alternative Splicing Services.
PMID 17584795 · PMC1933140 · Nucleic acids research · 2007 · 8 claims · 4 setups
BIPASS offers two complementary services for alternative splicing (AS) research: BIPAS-SpliceDB, a queryable pre-computed AS data warehouse, and BIPAS-Align&Splice, an online pipeline for user-submitted sequences.
-
Has reproduction · 75
FEM: mining biological meaning from cell level in single-cell RNA sequencing data.
PMID 34909283 · PMC8641482 · PeerJ · 2021 · 8 claims · 5 setups
The FEM algorithm converts a single-cell gene expression matrix into a functional expression matrix using multi-module gene set enrichment analysis, utilizing information from all expressed genes rather than discarding filtered genes.