Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
Multi locus sequence typing of Chlamydiales: clonal groupings within the obligate intracellular bacteria Chlamydia trachomatis.
PMID 18307777 · PMC2268939 · BMC microbiology · 2008 · 8 claims · 7 setups
MLST of 26 C. trachomatis strains reveals three non-overlapping clonal complexes (groups I, II, III)
-
Full-text index only
Genomic sequencing of the severe acute respiratory syndrome-coronavirus.
PMID 16916263 · PMC7121524 · Methods in molecular biology (Clifton, N.J.) · 2006 · 7 claims · 7 setups
PCR-based amplification and direct sequencing of SARS-CoV genome fragments is feasible from uncultured clinical specimens (serum, nasopharyngeal aspirate, stool), avoiding culture-derived artifacts and biohazard risk.
-
Full-text index only
Completing the map of human genetic variation.
PMID 17495918 · PMC2685471 · Nature · 2007 · 8 claims · 5 setups
A community resource initiative will sequence fosmid and BAC clone libraries from 62 HapMap individuals to systematically discover and resolve structural genetic variants at nucleotide resolution
-
Full-text index only
The global landscape of sequence diversity.
PMID 17996061 · PMC2258180 · Genome biology · 2007 · 7 claims · 5 setups
Eukaryotic sequence datasets show substantially greater genetic diversity (higher sequence/gene family discovery rates) than bacterial datasets, likely related to differences in modes of genetic inheritance.
-
Full-text index only
Comparative sequence analysis of leucine-rich repeats (LRRs) within vertebrate toll-like receptors.
PMID 17517123 · PMC1899181 · BMC genomics · 2007 · 8 claims · 4 setups
A new method combining known LRR structures, multiple sequence alignment, and secondary structure prediction identifies and aligns LRRs in TLRs more accurately than PFAM/InterPro/SMART
-
Full-text index only
The TIGR Gene Indices: clustering and assembling EST and known genes and integration with eukaryotic genomes.
PMID 15608288 · PMC540018 · Nucleic acids research · 2005 · 8 claims · 8 setups
The TIGR Gene Indices (TGI) are a collection of 77 species-specific databases that cluster and assemble EST and known gene sequences into tentative consensus (TC) sequences to identify and characterize expressed transcripts.
-
Full-text index only
The specificity and polymorphism of the MHC class I prevents the global adaptation of HIV-1 to the monomorphic proteasome and TAP.
PMID 18949050 · PMC2569417 · PloS one · 2008 · 6 claims · 5 setups
Within individual hosts, proteasome and TAP escape mutations in HIV-1 occur frequently
-
Full-text index only
Columba: an integrated database of proteins, structures, and annotations.
PMID 15801979 · PMC1087474 · BMC bioinformatics · 2005 · 8 claims · 6 setups
COLUMBA physically integrates data from twelve protein structure-related databases (PDB, KEGG, Swiss-Prot, CATH, SCOP, Gene Ontology, ENZYME, etc.) into a single PostgreSQL data warehouse.
-
Has reproduction · 100
Genomic approaches used to investigate an atypical outbreak of Salmonella Adjame.
PMID 30648934 · PMC6412060 · Microbial genomics · 2019 · 7 claims · 7 setups
The S. Adjame outbreak produced a heterogeneous phylogeny with multiple temporally/geographically linked sub-clusters, atypical of a point-source Salmonella outbreak and consistent with contamination from an endemic mixed-strain source (imported South Asian herbs/spices).
-
Full-text index only
Automated recognition of retroviral sequences in genomic data--RetroTector.
PMID 17636050 · PMC1976444 · Nucleic acids research · 2007 · 8 claims · 8 setups
RetroTector uses 'fragment threading' (detection of chains of conserved retroviral motifs satisfying distance constraints) combined with LTR detection and protein reconstruction to identify ERVs in genomic sequences
-
Full-text index only
Large genomic rearrangements in the CFTR gene contribute to CBAVD.
PMID 17448246 · PMC1876208 · BMC medical genetics · 2007 · 7 claims · 6 setups
Large genomic rearrangements in CFTR contribute to CBAVD and should be systematically investigated alongside point mutation screening
-
Full-text index only
DDBJ in collaboration with mass-sequencing teams on annotation.
PMID 15608189 · PMC539974 · Nucleic acids research · 2005 · 7 claims · 5 setups
DDBJ collected and released 1,066,084 entries (718,072,425 bases) in the past year, including the complete chimpanzee chromosome 22 sequence and silkworm whole-genome shotgun data
-
Full-text index only
Determination of glycosylation sites and site-specific heterogeneity in glycoproteins.
PMID 19700364 · PMC2749913 · Current opinion in chemical biology · 2009 · 8 claims · 8 setups
Mass spectrometry has emerged as the premier tool for structural determination of oligosaccharides/glycans and glycopeptides
-
Has reproduction · 45
Accurate sequence variant genotyping in cattle using variation-aware genome graphs.
PMID 31092189 · PMC6521551 · Genetics, selection, evolution : GSE · 2019 · 8 claims · 7 setups
Graphtyper outperformed GATK and SAMtools in genotype concordance, non-reference sensitivity, and non-reference discrepancy compared to microarray genotypes
-
Has reproduction · 44
Detecting DNA modifications from SMRT sequencing data by modeling sequence context dependence of polymerase kinetic.
PMID 23516341 · PMC3597545 · PLoS computational biology · 2013 · 8 claims · 7 setups
Local sequence context strongly determines position-specific polymerase kinetic rate: roughly 80% of IPD variation is explained by a 10 bp context (7 bases upstream, 2 bases downstream of the incorporation site), saturating at 7 bases upstream.
-
Full-text index only
Predicting deleterious nsSNPs: an analysis of sequence and structural attributes.
PMID 16630345 · PMC1489951 · BMC bioinformatics · 2006 · 8 claims · 7 setups
Sequence conservation (PSIC score difference) at the nsSNP position is the single most useful attribute for predicting deleterious vs neutral status.
-
Full-text index only
Sequence occurrence and structural uniqueness of a G-quadruplex in the human c-kit promoter.
PMID 17720713 · PMC2034477 · Nucleic acids research · 2007 · 8 claims · 4 setups
The native 22-nt c-kit87 sequence occurs only once in the entire human genome.
-
Full-text index only
Target SNP selection in complex disease association studies.
PMID 15248903 · PMC487897 · BMC bioinformatics · 2004 · 7 claims · 3 setups
A computational pipeline can retrieve gene sequence, collect SNP variation data, and annotate SNPs falling in functional motifs (promoter, exon-intron structure, AU-rich elements, TF binding sites, splice sites) with expression in target tissue
-
Full-text index only
Identification of "pathologs" (disease-related genes) from the RIKEN mouse cDNA dataset using human curation plus FACTS, a new biological information extraction system.
PMID 15115540 · PMC420239 · BMC genomics · 2004 · 6 claims · 3 setups
Bioinformatic sequence comparison of 60,770 RIKEN FANTOM2 mouse cDNA clones identified 2,578 sequences with 70-85% identity to known human disease genes/proteins
-
Full-text index only
Slider--maximum use of probability information for alignment of short sequence reads and SNP detection.
PMID 18974170 · PMC2638935 · Bioinformatics (Oxford, England) · 2009 · 7 claims · 3 setups
Slider aligns reads using all bases above a probability threshold (baseMinPrb) from prb files, generating all possible read sequences above a read probability threshold (read_0_MinPrb), rather than only the most probable sequence