Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
Challenges and standards in integrating surveys of structural variation.
PMID 17597783 · PMC2698291 · Nature genetics · 2007 · 7 claims · 5 setups
There is no standard approach to collecting, assessing the quality of, or describing structural variants, risking the entire genome eventually being labeled 'structurally variant' based on uncurated nondisease-sample data.
-
Full-text index only
Metagenomic analysis of human diarrhea: viral detection and discovery.
PMID 18398449 · PMC2290972 · PLoS pathogens · 2008 · 8 claims · 7 setups
Micro-mass sequencing (minimal stool input, minimal purification, ~384 reads/sample) can detect known enteric viruses in diarrhea specimens
-
Full-text index only
Conservation, variability and the modeling of active protein kinases.
PMID 17912359 · PMC1989141 · PloS one · 2007 · 7 claims · 5 setups
A novel sequence-order independent (fold-independent) structural alignment algorithm was developed that maximizes side-chain similarity to produce a consensus kinase structure.
-
Full-text index only
Speeding disease gene discovery by sequence based candidate prioritization.
PMID 15766383 · PMC1274252 · BMC bioinformatics · 2005 · 7 claims · 8 setups
Disease genes (OMIM) differ significantly from non-disease genes in sequence-based features including gene/cDNA/protein size, exon number, homolog conservation, secretion signal, 3' UTR length, CpG islands, and distance to nearest gene.
-
Full-text index only
Genomic views of distant-acting enhancers.
PMID 19741700 · PMC2923221 · Nature · 2009 · 8 claims · 8 setups
Meta-analysis of ~1200 top GWAS SNPs found that in 40% of cases (472/1170) no known exons overlap the linked SNP or its haplotype block, implying noncoding variation causally contributes to many traits.
-
Full-text index only
Analysis of the human Alu Ye lineage.
PMID 15725352 · PMC554112 · BMC evolutionary biology · 2005 · 8 claims · 6 setups
Two new Alu subfamilies, Ye4 and Ye6, were discovered, complementing the previously described Ye5 subfamily.
-
Full-text index only
MiPred: classification of real and pseudo microRNA precursors using random forest prediction model with combined features.
PMID 17553836 · PMC1933124 · Nucleic acids research · 2007 · 8 claims · 8 setups
A hybrid feature combining local contiguous triplet structure-sequence composition, MFE of the secondary structure, and P-value of a randomization test improves classification of real vs pseudo pre-miRNAs
-
Full-text index only
Prediction of specificity-determining residues for small-molecule kinase inhibitors.
PMID 19032760 · PMC2655090 · BMC bioinformatics · 2008 · 8 claims · 5 setups
S-Filter is a novel method combining sequence and structural information (within PFAAT) to predict specificity-determining residues and selectivity profiles for small-molecule kinase inhibitors
-
Full-text index only
Distinctive pattern of sequence polymorphism in the NS3 protein of hepatitis C virus type 1b reflects conflicting evolutionary pressures.
PMID 18632963 · PMC2577380 · The Journal of general virology · 2008 · 7 claims · 6 setups
NS3 shows less evidence of purifying selection acting on its CTL epitopes than the other 9 HCV proteins, while outside the CTL epitopes NS3 is more conserved than the other proteins.
-
Full-text index only
Complete genome sequence of Treponema pallidum ssp. pallidum strain SS14 determined with oligonucleotide arrays.
PMID 18482458 · PMC2408589 · BMC microbiology · 2008 · 8 claims · 6 setups
CGS combined with targeted DDT sequencing and whole genome fingerprinting (WGF) can accurately determine a treponemal genome sequence using only three arrays, at accuracy comparable to or better than finished DDT sequencing
-
Full-text index only
Bayesian coestimation of phylogeny and sequence alignment.
PMID 15804354 · PMC1087833 · BMC bioinformatics · 2005 · 7 claims · 3 setups
Alignment and phylogenetic inference are mutually dependent, and treating them as separate sequential steps (align then infer tree) is fundamentally flawed and produces biased, overconfident estimates.
-
Full-text index only
Optimized mixed Markov models for motif identification.
PMID 16749929 · PMC1534070 · BMC bioinformatics · 2006 · 8 claims · 4 setups
OMiMa can incorporate more than NNSplice's pairwise dependencies
-
Full-text index only
Characterization of in vivo somatic mutations at the hypoxanthine phosphoribosyltransferase gene of a human control population.
PMID 8513767 · PMC1519656 · Environmental health perspectives · 1993 · 7 claims · 4 setups
In vivo hprt mutants from 63 independent donors were molecularly characterized by cDNA and genomic DNA PCR/sequencing.
-
Full-text index only
Recent segmental and gene duplications in the mouse genome.
PMID 12914656 · PMC193640 · Genome biology · 2003 · 8 claims · 8 setups
33.6 Mb (1.2%) of the February 2003 mouse genome assembly (2,695 Mb) is involved in recent segmental duplications
-
Full-text index only
Classification of real and pseudo microRNA precursors using local structure-sequence features and support vector machine.
PMID 16381612 · PMC1360673 · BMC bioinformatics · 2005 · 7 claims · 7 setups
A 32-dimensional triplet structure-sequence feature vector combined with SVM (triplet-SVM) can distinguish real human pre-miRNAs from pseudo pre-miRNA hairpins with ~90% accuracy.
-
Has reproduction · 68
LaSSO, a strategy for genome-wide mapping of intronic lariats and branch points using RNA-seq.
PMID 24709818 · PMC4079972 · Genome research · 2014 · 8 claims · 8 setups
LaSSO (Lariat Sequence Site Origin) identifies intronic lariat reads and pinpoints branch points genome-wide from RNA-seq data by considering every intronic base as a potential branch point and including all possible exon-skipping lariats.
-
Full-text index only
Simultaneous analysis of all SNPs in genome-wide and re-sequencing association studies.
PMID 18654633 · PMC2464715 · PLoS genetics · 2008 · 8 claims · 5 setups
A Bayesian-inspired penalised maximum likelihood stochastic search method can simultaneously analyse all SNPs (up to 500K) from a GWA study in a few hours on a desktop workstation
-
Full-text index only
Identification of deleterious non-synonymous single nucleotide polymorphisms using sequence-derived information.
PMID 18588693 · PMC2446391 · BMC bioinformatics · 2008 · 8 claims · 5 setups
A decision tree built on 10 selected sequence-derived features classifies SAPs as Disease or Polymorphism with 82.6% accuracy and 0.607 MCC in cross-validation.
-
Full-text index only
Conserved elements with potential to form polymorphic G-quadruplex structures in the first intron of human genes.
PMID 18187510 · PMC2275096 · Nucleic acids research · 2008 · 8 claims · 6 setups
G-richness downstream of the TSS is strand-biased, concentrated on the nontemplate strand, with a peak at +200 to +300 bp
-
Full-text index only
The TIGR Gene Indices: clustering and assembling EST and known genes and integration with eukaryotic genomes.
PMID 15608288 · PMC540018 · Nucleic acids research · 2005 · 8 claims · 8 setups
The TIGR Gene Indices (TGI) are a collection of 77 species-specific databases that cluster and assemble EST and known gene sequences into tentative consensus (TC) sequences to identify and characterize expressed transcripts.