Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
nsSNPAnalyzer: identifying disease-associated nonsynonymous single nucleotide polymorphisms.
PMID 15980516 · PMC1160133 · Nucleic acids research · 2005 · 6 claims · 4 setups
nsSNPAnalyzer is a web server that predicts whether a query nsSNP is disease-associated or functionally neutral using a Random Forest classifier combining structural and evolutionary information
-
Full-text index only
The association of Alu repeats with the generation of potential AU-rich elements (ARE) at 3' untranslated regions.
PMID 15610565 · PMC544599 · BMC genomics · 2004 · 6 claims · 4 setups
Alu repeats are a source of AREs at 3' UTRs of human mRNA, via poly-A regions of Alu generating complementary poly-T/poly-U regions that acquire regular adenine insertions to form ARE motifs.
-
Full-text index only
Cancer-specific high-throughput annotation of somatic mutations: computational prediction of driver missense mutations.
PMID 19654296 · PMC2763410 · Cancer research · 2009 · 7 claims · 7 setups
CHASM, a Random Forest-based computational method, was developed to identify and prioritize missense mutations likely to be functional drivers of tumor cell proliferation.
-
Full-text index only
Expansion of the Bactericidal/Permeability Increasing-like (BPI-like) protein locus in cattle.
PMID 17362520 · PMC1839098 · BMC genomics · 2007 · 8 claims · 8 setups
The bovine BPI-like locus spans 470 kbp and contains 14 contiguous genes (13 intact + 1 pseudogene); 9 are orthologous to human/mouse BPI-like genes and 4 (named BSP30A, BSP30B, BSP30C, BSP30D) arose through cattle-specific duplication of the PSP gene
-
Full-text index only
Complete genome sequence and comparative analysis of the wild-type commensal Escherichia coli strain SE11 isolated from a healthy adult.
PMID 18931093 · PMC2608844 · DNA research : an international journal for rapid publication of reports on genes and genomes · 2008 · 8 claims · 6 setups
The SE11 genome comprises a 4.8 Mb chromosome encoding 4679 protein-coding genes and six plasmids encoding 323 protein-coding genes
-
Full-text index only
Random amino acid mutations and protein misfolding lead to Shannon limit in sequence-structure communication.
PMID 18769673 · PMC2518838 · PloS one · 2008 · 8 claims · 6 setups
The protein sequence-structure map behaves as a noisy digital communication channel whose capacity C exceeds the transmission rate R for native structures, satisfying Shannon's noisy channel theorem
-
Full-text index only
A surrogate-based approach for post-genomic partner identification.
PMID 11602024 · PMC57814 · BMC biotechnology · 2001 · 8 claims · 5 setups
Peptide surrogates derived from random phage display libraries contain amino acid sequence information that identifies the natural biological partner of the panned target via database searching.
-
Full-text index only
A genome-wide survey demonstrates widespread non-linear mRNA in expressed sequences from multiple species.
PMID 16237125 · PMC1258171 · Nucleic acids research · 2005 · 8 claims · 6 setups
A genome-wide computational survey identifies 245 genes in mammals (264 across six species) that produce RREO events in expressed sequences
-
Full-text index only
Analysis of protein sequence and interaction data for candidate disease gene prediction.
PMID 17020920 · PMC1636487 · Nucleic acids research · 2006 · 8 claims · 7 setups
Combining CPS and CMP using known disease genes as input achieves sensitivity 0.52 and specificity 0.97, reducing candidate lists 13-fold
-
Full-text index only
CanPredict: a computational tool for predicting cancer-associated missense mutations.
PMID 17537827 · PMC1933186 · Nucleic acids research · 2007 · 8 claims · 7 setups
CanPredict is a web application providing public access to a random forest classifier that combines SIFT, LogR.E-value, and GOSS scores to predict whether a missense mutation is cancer-associated
-
Full-text index only
Integration of text- and data-mining using ontologies successfully selects disease gene candidates.
PMID 15767279 · PMC1065256 · Nucleic acids research · 2005 · 7 claims · 6 setups
Integrating eVOC anatomical ontology-based text-mining of PubMed abstracts with data-mining of gene expression annotation successfully selects and prioritizes candidate disease genes
-
Has reproduction · 50
Exploiting convergent phenotypes to derive a pan-cancer cisplatin response gene expression signature.
PMID 37076665 · PMC10115855 · NPJ precision oncology · 2023 · 8 claims · 8 setups
A convergent-phenotype-based seed gene/co-expression method can extract consensus gene expression signatures predictive of response to chemotherapeutic drugs in the GDSC database
-
Full-text index only
Spontaneous symmetry breaking in genome evolution.
PMID 18367477 · PMC2377439 · Nucleic acids research · 2008 · 6 claims · 3 setups
Exon size distributions in sequenced genomes follow a lognormal pattern typical of a random Kolmogoroff fractioning process
-
Full-text index only
A clustering property of highly-degenerate transcription factor binding sites in the mammalian genome.
PMID 16670430 · PMC1456330 · Nucleic acids research · 2006 · 8 claims · 7 setups
Highly-degenerate RE1 sites are significantly enriched in promoters of validated and putative REST target genes compared to control promoters
-
Full-text index only
The impact of peptide abundance and dynamic range on stable-isotope-based quantitative proteomic analyses.
PMID 18798661 · PMC2746028 · Journal of proteome research · 2008 · 8 claims · 7 setups
Over half of confidently identified peptides in complex mixtures have S/N ratios below 10 on both FT-ICR and Orbitrap instruments
-
Full-text index only
Metagenomic analysis of respiratory tract DNA viral communities in cystic fibrosis and non-cystic fibrosis individuals.
PMID 19816605 · PMC2756586 · PloS one · 2009 · 8 claims · 8 setups
CF phage communities are highly similar to each other, whereas Non-CF individuals have more distinct, variable phage communities reflecting transient environmental sampling
-
Full-text index only
Antibody binding loop insertions as diversity elements.
PMID 17023486 · PMC1635297 · Nucleic acids research · 2006 · 7 claims · 8 setups
A lysozyme-binding VHH CDR3 loop can be grafted into two surface-exposed loops of superfolder GFP, conferring lysozyme-binding activity while the protein remains fluorescent.
-
Full-text index only
Stability analysis of mixtures of mutagenetic trees.
PMID 18366778 · PMC2335279 · BMC bioinformatics · 2008 · 7 claims · 5 setups
Mutagenetic trees mixture models capture multiple alternative pathways of ordered accumulation of genetic events (e.g., HIV resistance mutations, cancer chromosomal aberrations).
-
Full-text index only
Extraction of human kinase mutations from literature, databases and genotyping studies.
PMID 19758464 · PMC2745582 · BMC bioinformatics · 2009 · 7 claims · 6 setups
A literature mining pipeline combining MutationFinder, false-positive filtering, and SVM-based classification can extract and disambiguate single-point mutation mentions from abstracts and full text
-
Full-text index only
Using ESTs to improve the accuracy of de novo gene prediction.
PMID 16817966 · PMC1534067 · BMC bioinformatics · 2006 · 8 claims · 8 setups
TWINSCAN_EST combines EST alignments with TWINSCAN via a trainable 'ESTseq' representation and improves exact gene structure prediction accuracy on the whole C. elegans genome