Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
RotaC: a web-based tool for the complete genome classification of group A rotaviruses.
PMID 19930627 · PMC2785824 · BMC microbiology · 2009 · 7 claims · 4 setups
RotaC is a freely available web-based tool for complete genome classification of group A rotaviruses across all 11 gene segments.
-
Full-text index only
SNPmasker: automatic masking of SNPs and repeats across eukaryotic genomes.
PMID 16845091 · PMC1538889 · Nucleic acids research · 2006 · 8 claims · 4 setups
SNPmasker is a web service combining SNP masking and repeat masking, supporting both coordinate-defined and homology-search-defined input regions, a combination not offered by prior tools
-
Full-text index only
CorGen--measuring and generating long-range correlations for DNA sequence analysis.
PMID 16845099 · PMC1538783 · Nucleic acids research · 2006 · 8 claims · 3 setups
CorGen is a web server that measures long-range correlations in DNA sequences and generates random sequences with the same (or user-specified) correlation and composition parameters
-
Full-text index only
Using multiple alignments to improve seeded local alignment algorithms.
PMID 16100379 · PMC1185574 · Nucleic acids research · 2005 · 8 claims · 2 setups
Using information implicit in a multiple alignment to dynamically build a spaced-seed index weighted toward promising regions increases sensitivity of local alignment search compared to indexing a sequence alone
-
Full-text index only
Having a BLAST with bioinformatics (and avoiding BLASTphemy).
PMID 11597340 · PMC138974 · Genome biology · 2001 · 8 claims · 4 setups
BLAST is the most widely used tool for searching biological sequences for regions of local similarity
-
Full-text index only
L1Base: from functional annotation to prediction of active LINE-1 elements.
PMID 15608246 · PMC539998 · Nucleic acids research · 2005 · 7 claims · 6 setups
L1Base is a database of putatively active LINE-1 insertions in human, mouse and rat genomes, containing FLI-L1s (intact in both ORFs), ORF2-L1s (intact ORF2, disrupted ORF1), and FLnI-L1s (full-length, >6000 bp, non-intact)
-
Full-text index only
PA-GOSUB: a searchable database of model organism protein sequences with their predicted Gene Ontology molecular function and subcellular localization.
PMID 15608166 · PMC540074 · Nucleic acids research · 2005 · 7 claims · 4 setups
PA-GOSUB significantly extends the coverage of GO molecular function and subcellular localization annotations for 10 model organism proteomes compared with existing databases (GOA, Swiss-Prot).
-
Full-text index only
Natural history of S-adenosylmethionine-binding proteins.
PMID 16225687 · PMC1282579 · BMC structural biology · 2005 · 8 claims · 6 setups
The last universal common ancestor (LUCA) of cellular life had between 10 and 20 SAM-binding proteins from at least 5 fold classes
-
Full-text index only
BIPASS: BioInformatics Pipeline Alternative Splicing Services.
PMID 17584795 · PMC1933140 · Nucleic acids research · 2007 · 8 claims · 4 setups
BIPASS offers two complementary services for alternative splicing (AS) research: BIPAS-SpliceDB, a queryable pre-computed AS data warehouse, and BIPAS-Align&Splice, an online pipeline for user-submitted sequences.
-
Full-text index only
Cataloging coding sequence variations in human genome databases.
PMID 18974781 · PMC2570488 · PloS one · 2008 · 8 claims · 7 setups
A significant proportion of CVs overlap between HGMD and dbSNP (4.36% of HGMD CVs registered in dbSNP; 8.11% of dbSNP CVs registered in HGMD), warranting caution when interpreting phenotypic relevance of concurrent CVs.
-
Has reproduction · 87
R2DT is a framework for predicting and visualising RNA secondary structure using templates.
PMID 34108470 · PMC8190129 · Nature communications · 2021 · 8 claims · 6 setups
R2DT is a template-based computational framework/pipeline that predicts and visualises RNA 2D structure in standardised, community-accepted layouts
-
Full-text index only
A genome-wide survey of segmental duplications that mediate common human genetic variation of chromosomal architecture.
PMID 15588494 · PMC3525102 · Human genomics · 2004 · 8 claims · 5 setups
PSD-mediated genomic architecture analogous to the 8p23/4p16 inversion regions is not unique to those loci but recurs genome-wide.
-
Full-text index only
Analysis of the human Alu Ye lineage.
PMID 15725352 · PMC554112 · BMC evolutionary biology · 2005 · 8 claims · 6 setups
Two new Alu subfamilies, Ye4 and Ye6, were discovered, complementing the previously described Ye5 subfamily.
-
Has reproduction · 43
StatsDB: platform-agnostic storage and understanding of next generation sequencing run metrics.
PMID 24627795 · PMC3938176 · F1000Research · 2013 · 8 claims · 6 setups
StatsDB is an open-source software package for storage and analysis of next generation sequencing run metrics, backed by an SQL (MySQL) database with Perl and Java APIs.
-
Full-text index only
NetworKIN: a resource for exploring cellular phosphorylation networks.
PMID 17981841 · PMC2238868 · Nucleic acids research · 2008 · 8 claims · 4 setups
NetworKIN integrates consensus substrate motifs with probabilistic network context modelling to predict cellular kinase-substrate relations.
-
Full-text index only
SUPERFAMILY--sophisticated comparative genomics, data mining, visualization and phylogeny.
PMID 19036790 · PMC2686452 · Nucleic acids research · 2009 · 7 claims · 6 setups
SUPERFAMILY provides structural, functional and evolutionary annotation for proteins from all completely sequenced genomes using SCOP-based hidden Markov models
-
Full-text index only
EGenBio: a data management system for evolutionary genomics and biodiversity.
PMID 17118150 · PMC1683573 · BMC bioinformatics · 2006 · 7 claims · 7 setups
EGenBio is a web-based system for integrated management, filtering, curation, and visualization of large-scale genomic sequences, alignments, and phylogenetic trees for evolutionary genomics and biodiversity research.
-
Has reproduction · 71
RNAmountAlign: Efficient software for local, global, semiglobal pairwise and multiple RNA sequence/structure alignment.
PMID 31978147 · PMC6980424 · PloS one · 2020 · 8 claims · 6 setups
RNAmountAlign is the first RNA sequence/structure pairwise alignment algorithm based on incremental ensemble mountain distance, running in O(n^3) time and O(n^2) space for two sequences of length n.
-
Full-text index only
MACSIMS: multiple alignment of complete sequences information management system.
PMID 16792820 · PMC1539025 · BMC bioinformatics · 2006 · 8 claims · 5 setups
MACSIMS is a multiple alignment-based information management system combining knowledge-based database mining with ab initio sequence predictions
-
Full-text index only
RAId_DbS: mass-spectrometry based peptide identification web server with knowledge integration.
PMID 18954448 · PMC2605478 · BMC genomics · 2008 · 7 claims · 4 setups
Constructed enhanced protein databases integrating annotated SAPs, PTMs, and disease associations for 17 organisms.