Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
Intrinsic structural disorder confers cellular viability on oncogenic fusion proteins.
PMID 19888473 · PMC2768585 · PLoS computational biology · 2009 · 8 claims · 5 setups
Translocation-related human proteins are significantly enriched in intrinsic structural disorder compared to all human proteins
-
Full-text index only
Analysis of protein sequence and interaction data for candidate disease gene prediction.
PMID 17020920 · PMC1636487 · Nucleic acids research · 2006 · 8 claims · 7 setups
Combining CPS and CMP using known disease genes as input achieves sensitivity 0.52 and specificity 0.97, reducing candidate lists 13-fold
-
Full-text index only
Comparative sequence analysis of leucine-rich repeats (LRRs) within vertebrate toll-like receptors.
PMID 17517123 · PMC1899181 · BMC genomics · 2007 · 8 claims · 4 setups
A new method combining known LRR structures, multiple sequence alignment, and secondary structure prediction identifies and aligns LRRs in TLRs more accurately than PFAM/InterPro/SMART
-
Full-text index only
Genomic activation of the EGFR and HER2-neu genes in a significant proportion of invasive epithelial ovarian cancers.
PMID 18182111 · PMC2266762 · BMC cancer · 2008 · 8 claims · 4 setups
No somatic mutations were found in the entire tyrosine kinase domain (exons 18-24) of EGFR or HER2-neu in 68 tissue samples from 52 ovarian cancer patients.
-
Full-text index only
Prediction by graph theoretic measures of structural effects in proteins arising from non-synonymous single nucleotide polymorphisms.
PMID 18654622 · PMC2447880 · PLoS computational biology · 2008 · 8 claims · 5 setups
Bongo identifies mutations causing local and global structural effects with a remarkably low false positive rate
-
Full-text index only
Benchmarking ortholog identification methods using functional genomics data.
PMID 16613613 · PMC1557999 · Genome biology · 2006 · 8 claims · 7 setups
InParanoid is the best overall ortholog identification method for identifying functionally equivalent proteins when sensitivity and selectivity are combined into an overall score.
-
Full-text index only
Non-EST based prediction of exon skipping and intron retention events using Pfam information.
PMID 16204458 · PMC1243800 · Nucleic acids research · 2005 · 7 claims · 5 setups
A novel ab initio method predicts exon skipping and intron retention events using only Pfam domain annotation, via a Viterbi-like dynamic programming algorithm applied to the Pfam alignment.
-
Full-text index only
Sequence similarity network reveals common ancestry of multidomain proteins.
PMID 18475320 · PMC2377100 · PLoS computational biology · 2008 · 8 claims · 6 setups
Traditional homology definitions do not capture multidomain evolution; the authors extend the definition to include domain insertion via a common ancestral locus model.
-
Full-text index only
MEROPS: the peptidase database.
PMID 19892822 · PMC2808883 · Nucleic acids research · 2010 · 8 claims · 5 setups
MEROPS is a manually curated hierarchical classification of peptidases and protein inhibitors organized into protein species, families, and clans based on sequence and structural homology.
-
Full-text index only
Generation of a restriction minus enteropathogenic Escherichia coli E2348/69 strain that is efficiently transformed with large, low copy plasmids.
PMID 18681975 · PMC2518929 · BMC microbiology · 2008 · 8 claims · 7 setups
E2348/69 possesses a type I restriction-modification system encoded by an hsdMSR-like operon identified by homology to known Hsd proteins.
-
Full-text index only
The UCSC Proteome Browser.
PMID 15608236 · PMC540054 · Nucleic acids research · 2005 · 8 claims · 5 setups
The UCSC Proteome Browser is tightly integrated with the UCSC Genome Browser, giving users simultaneous access to genome and proteome data.
-
Has reproduction · 71
Systematic and computational identification of Androctonus crassicauda long non-coding RNAs.
PMID 33633149 · PMC7907363 · Scientific reports · 2021 · 7 claims · 7 setups
A custom ECF pipeline identified 13,401 lncRNAs in the A. crassicauda transcriptome (12,642 novel, 759 known).
-
Full-text index only
Identification and functional analyses of 11,769 full-length human cDNAs focused on alternative splicing.
PMID 19880432 · PMC2780955 · DNA research : an international journal for rapid publication of reports on genes and genomes · 2009 · 8 claims · 5 setups
Identified 23,241 human genes transcribed into protein-coding mRNAs using full-length cDNA and 5'-EST sequence data
-
Has reproduction · 67
Satellitome Analysis and Transposable Elements Comparison in Geographically Distant Populations of Spodoptera frugiperda.
PMID 35455012 · PMC9026859 · Life (Basel, Switzerland) · 2022 · 8 claims · 5 setups
Most transposable elements are commonly shared across all eight geographically distant S. frugiperda samples, except Maverick and PIF/Harbinger elements which show divergent repeat copies
-
Full-text index only
The cohesin complex: sequence homologies, interaction networks and shared motifs.
PMID 11276426 · PMC30708 · Genome biology · 2001 · 8 claims · 8 setups
Mouse Mmip1 and Smc3 (SMCD) share 99% sequence identity and are products of the same gene
-
Full-text index only
SARS--beginning to understand a new virus.
PMID 15035025 · PMC7097337 · Nature reviews. Microbiology · 2003 · 8 claims · 8 setups
A previously unknown coronavirus (SARS-CoV) was isolated from FRhK-4 and Vero E6 cells inoculated with clinical specimens from SARS patients and identified as the causative agent of SARS
-
Full-text index only
CanPredict: a computational tool for predicting cancer-associated missense mutations.
PMID 17537827 · PMC1933186 · Nucleic acids research · 2007 · 8 claims · 7 setups
CanPredict is a web application providing public access to a random forest classifier that combines SIFT, LogR.E-value, and GOSS scores to predict whether a missense mutation is cancer-associated
-
Has reproduction · 49
Interplay between Non-Coding RNA Transcription, Stringent/Relaxed Phenotype and Antibiotic Production in Streptomyces ambofaciens.
PMID 34438997 · PMC8388888 · Antibiotics (Basel, Switzerland) · 2021 · 8 claims · 5 setups
The S. ambofaciens ATCC 23877 transcriptome was redefined from RNAseq data into 5587 transcriptional units (4433 monocistronic, 1154 polycistronic) covering 90.8% of the linear chromosome
-
Full-text index only
How to find soluble proteins: a comprehensive analysis of alpha/beta hydrolases for recombinant expression in E. coli.
PMID 15804363 · PMC1079826 · BMC genomics · 2005 · 7 claims · 7 setups
Predicted solubility in E. coli (via CV-CV') depends on hydrolase size, phylogenetic origin, homologous family, and superfamily
-
Full-text index only
'Genome design' model and multicellular complexity: golden middle.
PMID 17062620 · PMC1635334 · Nucleic acids research · 2006 · 8 claims · 8 setups
Intermediately expressed human genes are the longest genes genome-wide, in both coding and intronic sequence, longer than housekeeping or tissue-specific genes.