Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
Prediction of missed cleavage sites in tryptic peptides aids protein identification in proteomics.
PMID 17203985 · PMC2664920 · Journal of proteome research · 2007 · 8 claims · 4 setups
An information-theoretic log-likelihood scoring method can predict experimentally observed missed cleavage sites from amino acid sequence alone with up to 90% accuracy.
-
Has reproduction · 78
Detecting tipping points of complex diseases by network information entropy.
PMID 38960408 · PMC11221888 · Briefings in bioinformatics · 2024 · 8 claims · 4 setups
NIEE can detect critical states or tipping points in diverse data types, including bulk and single-sample expression data
-
Full-text index only
'Genome design' model and multicellular complexity: golden middle.
PMID 17062620 · PMC1635334 · Nucleic acids research · 2006 · 8 claims · 8 setups
Intermediately expressed human genes are the longest genes genome-wide, in both coding and intronic sequence, longer than housekeeping or tissue-specific genes.
-
Full-text index only
SysPIMP: the web-based systematical platform for identifying human disease-related mutated sequences from mass spectrometry.
PMID 19036792 · PMC2686442 · Nucleic acids research · 2009 · 8 claims · 7 setups
SysPIMP is a web-based platform integrating disease mutation databases with X!Tandem and BLAST to identify disease-related mutated proteins from MS results
-
Full-text index only
Random amino acid mutations and protein misfolding lead to Shannon limit in sequence-structure communication.
PMID 18769673 · PMC2518838 · PloS one · 2008 · 8 claims · 6 setups
The protein sequence-structure map behaves as a noisy digital communication channel whose capacity C exceeds the transmission rate R for native structures, satisfying Shannon's noisy channel theorem
-
Full-text index only
Computational identification of transcriptional regulatory elements in DNA sequence.
PMID 16855295 · PMC1524905 · Nucleic acids research · 2006 · 8 claims · 3 setups
Weight matrix (PWM/PSSM) models of TF binding sites are grounded in biophysical theory of protein-DNA interactions, with position weights corresponding to log-odds contributions to binding free energy
-
Full-text index only
Slider--maximum use of probability information for alignment of short sequence reads and SNP detection.
PMID 18974170 · PMC2638935 · Bioinformatics (Oxford, England) · 2009 · 7 claims · 3 setups
Slider aligns reads using all bases above a probability threshold (baseMinPrb) from prb files, generating all possible read sequences above a read probability threshold (read_0_MinPrb), rather than only the most probable sequence
-
Full-text index only
EGenBio: a data management system for evolutionary genomics and biodiversity.
PMID 17118150 · PMC1683573 · BMC bioinformatics · 2006 · 7 claims · 7 setups
EGenBio is a web-based system for integrated management, filtering, curation, and visualization of large-scale genomic sequences, alignments, and phylogenetic trees for evolutionary genomics and biodiversity research.
-
Full-text index only
GenBank.
PMID 16381837 · PMC1347519 · Nucleic acids research · 2006 · 8 claims · 8 setups
GenBank is a comprehensive public database of nucleotide sequences with supporting bibliographic and biological annotation, built and distributed by NCBI.
-
Full-text index only
Coverage and characteristics of the Affymetrix GeneChip Human Mapping 100K SNP set.
PMID 16680197 · PMC1456318 · PLoS genetics · 2006 · 7 claims · 7 setups
SNPs in the Affymetrix 100K set are undersampled from coding regions (both synonymous and nonsynonymous) and oversampled from regions outside genes, relative to HapMap SNPs
-
Full-text index only
Comprehensive splice-site analysis using comparative genomics.
PMID 16914448 · PMC1557818 · Nucleic acids research · 2006 · 8 claims · 6 setups
Over half a million splice sites were collected from five species (H. sapiens, M. musculus, D. melanogaster, C. elegans, A. thaliana) and classified into four main subtypes: U2-type GT-AG and GC-AG, and U12-type GT-AG and AT-AC.
-
Full-text index only
Predicting deleterious nsSNPs: an analysis of sequence and structural attributes.
PMID 16630345 · PMC1489951 · BMC bioinformatics · 2006 · 8 claims · 7 setups
Sequence conservation (PSIC score difference) at the nsSNP position is the single most useful attribute for predicting deleterious vs neutral status.
-
Full-text index only
In vitro and in silico analysis reveals an efficient algorithm to predict the splicing consequences of mutations at the 5' splice sites.
PMID 17726045 · PMC2094079 · Nucleic acids research · 2007 · 8 claims · 6 setups
Two exonic mutations, PINK1 E417G and PARK7 E64D, disrupt binding to U1 snRNA and cause skipping of the mutation-harboring exon
-
Has reproduction · 50
BiRNA-BERT allows efficient RNA language modeling with adaptive tokenization.
PMID 41266599 · PMC12635123 · Communications biology · 2025 · 8 claims · 8 setups
BiRNA-BERT uses adaptive dual-tokenization that dynamically selects nucleotide-level (NUC) or byte-pair encoding (BPE) tokens based on input sequence length
-
Full-text index only
CanPredict: a computational tool for predicting cancer-associated missense mutations.
PMID 17537827 · PMC1933186 · Nucleic acids research · 2007 · 8 claims · 7 setups
CanPredict is a web application providing public access to a random forest classifier that combines SIFT, LogR.E-value, and GOSS scores to predict whether a missense mutation is cancer-associated