Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
SNAP: predict effect of non-synonymous polymorphisms on function.
PMID 17526529 · PMC1920242 · Nucleic acids research · 2007 · 7 claims · 8 setups
SNAP, a neural network-based method using sequence-derived information, predicts whether a non-synonymous SNP is neutral or non-neutral for protein function
-
Full-text index only
Functional coverage of the human genome by existing structures, structural genomics targets, and homology models.
PMID 16118666 · PMC1188274 · PLoS computational biology · 2005 · 8 claims · 5 setups
Existing PDB structures provide single-domain coverage for 37% of functional classes in the human genome and complete (whole-protein) structure coverage for 25%.
-
Full-text index only
Human genome research in China.
PMID 15168679 · PMC7079922 · Journal of molecular medicine (Berlin, Germany) · 2004 · 8 claims · 8 setups
China completed its assigned 1% share of the international Human Genome Project sequencing effort and contributed ~10% of the HapMap effort
-
Full-text index only
LMPD: LIPID MAPS proteome database.
PMID 16381922 · PMC1347484 · Nucleic acids research · 2006 · 8 claims · 5 setups
LMPD is an object-relational database of lipid-associated protein sequences and annotations, publicly available from the LIPID MAPS Consortium website.
-
Full-text index only
GeneSeer: a sage for gene names and genomic resources.
PMID 16176584 · PMC1266031 · BMC genomics · 2005 · 7 claims · 4 setups
GeneSeer aggregates gene name synonyms from GenBank, FlyBase, ExPASy, HUGO, ENSEMBL, UCSC and Gene Ontology into a name-translation database that maps any familiar name to a reference (SOFAR) identifier.
-
Full-text index only
An SVM-based system for predicting protein subnuclear localizations.
PMID 16336650 · PMC1325059 · BMC bioinformatics · 2005 · 7 claims · 3 setups
New kernels defined on k-peptide vectors mapped by BLOSUM62-based high-scored pair matrices (D1, D2, D3) improve SVM discrimination of protein subnuclear localization compared to conventional k-peptide encodings.
-
Full-text index only
Database resources of the National Center for Biotechnology Information.
PMID 17170002 · PMC1781113 · Nucleic acids research · 2007 · 8 claims · 8 setups
NCBI maintains an integrated suite of database resources (Entrez, PubMed, RefSeq, dbSNP, BLAST, etc.) for molecular biology data retrieval and analysis
-
Full-text index only
Molecular archeology of L1 insertions in the human genome.
PMID 12372140 · PMC134481 · Genome biology · 2002 · 8 claims · 4 setups
TSDfinder, a new algorithm, refines RepeatMasker-identified L1 boundaries by locating poly(A) tails, TSDs, and inversion breakpoints
-
Full-text index only
Automatic discovery of cross-family sequence features associated with protein function.
PMID 16409628 · PMC1395344 · BMC bioinformatics · 2006 · 8 claims · 6 setups
A self-supervised data mining approach can find relationships between sequence features and functional annotations without preconceived functional categories.
-
Full-text index only
Predicting deleterious nsSNPs: an analysis of sequence and structural attributes.
PMID 16630345 · PMC1489951 · BMC bioinformatics · 2006 · 8 claims · 7 setups
Sequence conservation (PSIC score difference) at the nsSNP position is the single most useful attribute for predicting deleterious vs neutral status.
-
Full-text index only
Random amino acid mutations and protein misfolding lead to Shannon limit in sequence-structure communication.
PMID 18769673 · PMC2518838 · PloS one · 2008 · 8 claims · 6 setups
The protein sequence-structure map behaves as a noisy digital communication channel whose capacity C exceeds the transmission rate R for native structures, satisfying Shannon's noisy channel theorem
-
Full-text index only
Dyneins across eukaryotes: a comparative genomic analysis.
PMID 17897317 · PMC2239267 · Traffic (Copenhagen, Denmark) · 2007 · 8 claims · 6 setups
Phylogenetic inference identified nine DHC families (two cytoplasmic, seven axonemal) and six IC families (one cytoplasmic)
-
Has reproduction · 24
MiGPC: a comprehensive catalog of enzybiotics from environmental metagenomes.
PMID 41888223 · PMC13172421 · Scientific reports · 2026 · 8 claims · 8 setups
MiGPC is the first genome-resolved metagenomic gene and protein catalog specifically targeted to enzybiotics
-
Full-text index only
ChimerDB--a knowledgebase for fusion sequences.
PMID 16381848 · PMC1347382 · Nucleic acids research · 2006 · 8 claims · 6 setups
ChimerDB integrates bioinformatics analysis of mRNA/EST sequences, manually collected literature data, and OMIM translocation data into a single fusion sequence knowledgebase
-
Full-text index only
The global landscape of sequence diversity.
PMID 17996061 · PMC2258180 · Genome biology · 2007 · 7 claims · 5 setups
Eukaryotic sequence datasets show substantially greater genetic diversity (higher sequence/gene family discovery rates) than bacterial datasets, likely related to differences in modes of genetic inheritance.
-
Full-text index only
Genome comparison without alignment using shortest unique substrings.
PMID 15910684 · PMC1166540 · BMC bioinformatics · 2005 · 8 claims · 8 setups
A number of sequence comparison tasks, including detection of unique genomic regions, can be accomplished efficiently without an alignment step using shortest unique substrings.
-
Full-text index only
piRNABank: a web resource on classified and clustered Piwi-interacting RNAs.
PMID 17881367 · PMC2238943 · Nucleic acids research · 2008 · 6 claims · 4 setups
piRNABank is a web-accessible database storing empirically known piRNA sequences and annotations for human, mouse and rat.
-
Full-text index only
A compatible exon-exon junction database for the identification of exon skipping events using tandem mass spectrum data.
PMID 19087293 · PMC2636810 · BMC bioinformatics · 2008 · 6 claims · 6 setups
A theoretical exon-exon junction protein database accounting for all in-phase (frame-preserving) exon combinations can be built from the Ensembl Core Database using Perl/Bioperl/MySQL/Ensembl API.
-
Full-text index only
SysPIMP: the web-based systematical platform for identifying human disease-related mutated sequences from mass spectrometry.
PMID 19036792 · PMC2686442 · Nucleic acids research · 2009 · 8 claims · 7 setups
SysPIMP is a web-based platform integrating disease mutation databases with X!Tandem and BLAST to identify disease-related mutated proteins from MS results
-
Full-text index only
Classification of real and pseudo microRNA precursors using local structure-sequence features and support vector machine.
PMID 16381612 · PMC1360673 · BMC bioinformatics · 2005 · 7 claims · 7 setups
A 32-dimensional triplet structure-sequence feature vector combined with SVM (triplet-SVM) can distinguish real human pre-miRNAs from pseudo pre-miRNA hairpins with ~90% accuracy.