Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Has reproduction · 99
getSequenceInfo: a suite of tools allowing to get genome sequence information from public repositories.
PMID 35804320 · PMC9264741 · BMC bioinformatics · 2022 · 8 claims · 8 setups
getSequenceInfo (gSeqI) allows programmatic (CLI) or GUI-based retrieval of sequence data and metadata from GenBank, RefSeq, and ENA across Linux, MacOS, and Windows.
-
Full-text index only
Classification of real and pseudo microRNA precursors using local structure-sequence features and support vector machine.
PMID 16381612 · PMC1360673 · BMC bioinformatics · 2005 · 7 claims · 7 setups
A 32-dimensional triplet structure-sequence feature vector combined with SVM (triplet-SVM) can distinguish real human pre-miRNAs from pseudo pre-miRNA hairpins with ~90% accuracy.
-
Full-text index only
Proteomic approaches for studying alcoholism and alcohol-induced organ damage.
PMID 23584750 · PMC3860448 · Alcohol research & health : the journal of the National Institute on Alcohol Abuse and Alcoholism · 2008 · 8 claims · 7 setups
Chronic alcohol abuse produces persistent changes in brain function (tolerance, dependence, craving) likely resulting from alterations in gene and protein expression
-
Full-text index only
Determination of glycosylation sites and site-specific heterogeneity in glycoproteins.
PMID 19700364 · PMC2749913 · Current opinion in chemical biology · 2009 · 8 claims · 8 setups
Mass spectrometry has emerged as the premier tool for structural determination of oligosaccharides/glycans and glycopeptides
-
Full-text index only
Prediction of missed cleavage sites in tryptic peptides aids protein identification in proteomics.
PMID 17203985 · PMC2664920 · Journal of proteome research · 2007 · 8 claims · 4 setups
An information-theoretic log-likelihood scoring method can predict experimentally observed missed cleavage sites from amino acid sequence alone with up to 90% accuracy.
-
Full-text index only
Random amino acid mutations and protein misfolding lead to Shannon limit in sequence-structure communication.
PMID 18769673 · PMC2518838 · PloS one · 2008 · 8 claims · 6 setups
The protein sequence-structure map behaves as a noisy digital communication channel whose capacity C exceeds the transmission rate R for native structures, satisfying Shannon's noisy channel theorem
-
Full-text index only
Slider--maximum use of probability information for alignment of short sequence reads and SNP detection.
PMID 18974170 · PMC2638935 · Bioinformatics (Oxford, England) · 2009 · 7 claims · 3 setups
Slider aligns reads using all bases above a probability threshold (baseMinPrb) from prb files, generating all possible read sequences above a read probability threshold (read_0_MinPrb), rather than only the most probable sequence
-
Full-text index only
'Genome design' model and multicellular complexity: golden middle.
PMID 17062620 · PMC1635334 · Nucleic acids research · 2006 · 8 claims · 8 setups
Intermediately expressed human genes are the longest genes genome-wide, in both coding and intronic sequence, longer than housekeeping or tissue-specific genes.
-
Full-text index only
Computational identification of transcriptional regulatory elements in DNA sequence.
PMID 16855295 · PMC1524905 · Nucleic acids research · 2006 · 8 claims · 3 setups
Weight matrix (PWM/PSSM) models of TF binding sites are grounded in biophysical theory of protein-DNA interactions, with position weights corresponding to log-odds contributions to binding free energy
-
Full-text index only
Identification of "pathologs" (disease-related genes) from the RIKEN mouse cDNA dataset using human curation plus FACTS, a new biological information extraction system.
PMID 15115540 · PMC420239 · BMC genomics · 2004 · 6 claims · 3 setups
Bioinformatic sequence comparison of 60,770 RIKEN FANTOM2 mouse cDNA clones identified 2,578 sequences with 70-85% identity to known human disease genes/proteins
-
Full-text index only
Predicting deleterious nsSNPs: an analysis of sequence and structural attributes.
PMID 16630345 · PMC1489951 · BMC bioinformatics · 2006 · 8 claims · 7 setups
Sequence conservation (PSIC score difference) at the nsSNP position is the single most useful attribute for predicting deleterious vs neutral status.
-
Full-text index only
MitoVariome: a variome database of human mitochondrial DNA.
PMID 19958475 · PMC2788364 · BMC genomics · 2009 · 8 claims · 5 setups
MitoVariome is a web-based, integrated variome database for human mitochondrial DNA that unifies sequence variation, haplogroup, and disease annotation information not jointly available in prior databases (MITOMAP, mtDB, Mitome, MitoRes).
-
Full-text index only
Genome sequences and great expectations.
PMID 11178275 · PMC150431 · Genome biology · 2001 · 8 claims · 3 setups
Function is known or can be predicted for an average of 62% of proteins across 31 analyzed genomes.
-
Full-text index only
EGenBio: a data management system for evolutionary genomics and biodiversity.
PMID 17118150 · PMC1683573 · BMC bioinformatics · 2006 · 7 claims · 7 setups
EGenBio is a web-based system for integrated management, filtering, curation, and visualization of large-scale genomic sequences, alignments, and phylogenetic trees for evolutionary genomics and biodiversity research.
-
Full-text index only
Analysis of concordance of different haplotype block partitioning algorithms.
PMID 16356172 · PMC1343594 · BMC bioinformatics · 2005 · 7 claims · 7 setups
Each block partitioning algorithm infers blocks differing in number, size, and coverage under different SNP density and allele frequency conditions.
-
Full-text index only
Improved mutation tagging with gene identifiers applied to membrane protein stability prediction.
PMID 19758467 · PMC2745585 · BMC bioinformatics · 2009 · 8 claims · 4 setups
MutationTagger achieves 87% F-measure for the mutation retrieval task on a benchmark dataset
-
Full-text index only
GenBank.
PMID 16381837 · PMC1347519 · Nucleic acids research · 2006 · 8 claims · 8 setups
GenBank is a comprehensive public database of nucleotide sequences with supporting bibliographic and biological annotation, built and distributed by NCBI.
-
Has reproduction · 50
BiRNA-BERT allows efficient RNA language modeling with adaptive tokenization.
PMID 41266599 · PMC12635123 · Communications biology · 2025 · 8 claims · 8 setups
BiRNA-BERT uses adaptive dual-tokenization that dynamically selects nucleotide-level (NUC) or byte-pair encoding (BPE) tokens based on input sequence length
-
Full-text index only
CONTRAST: a discriminative, phylogeny-free approach to multiple informant de novo gene prediction.
PMID 18096039 · PMC2246271 · Genome biology · 2007 · 8 claims · 5 setups
CONTRAST predicts exact coding region structures for 65% more human genes than the previous state-of-the-art de novo predictor (N-SCAN)
-
Full-text index only
In vitro and in silico analysis reveals an efficient algorithm to predict the splicing consequences of mutations at the 5' splice sites.
PMID 17726045 · PMC2094079 · Nucleic acids research · 2007 · 8 claims · 6 setups
Two exonic mutations, PINK1 E417G and PARK7 E64D, disrupt binding to U1 snRNA and cause skipping of the mutation-harboring exon