Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
Gene loss rate: a probabilistic measure for the conservation of eukaryotic genes.
PMID 17158152 · PMC1802574 · Nucleic acids research · 2007 · 8 claims · 8 setups
GLR is a novel maximum-likelihood measure of gene loss rate that probabilistically weighs all possible ancestral phyletic patterns rather than relying on a single parsimonious reconstruction.
-
Has reproduction · 44
Detecting DNA modifications from SMRT sequencing data by modeling sequence context dependence of polymerase kinetic.
PMID 23516341 · PMC3597545 · PLoS computational biology · 2013 · 8 claims · 7 setups
Local sequence context strongly determines position-specific polymerase kinetic rate: roughly 80% of IPD variation is explained by a 10 bp context (7 bases upstream, 2 bases downstream of the incorporation site), saturating at 7 bases upstream.
-
Full-text index only
FeatureScan: revealing property-dependent similarity of nucleotide sequences.
PMID 16845077 · PMC1538849 · Nucleic acids research · 2006 · 6 claims · 5 setups
FeatureScan transforms nucleotide sequences into numerical signals of physico-chemical/conformational properties and compares them via a convolution/correlation (Fourier transform) method rather than comparing letters
-
Full-text index only
Similarities and differences in genome-wide expression data of six organisms.
PMID 14737187 · PMC300882 · PLoS biology · 2004 · 8 claims · 8 setups
Coexpression of functionally related genes is frequently conserved across evolutionarily distant organisms
-
Full-text index only
Automatic discovery of cross-family sequence features associated with protein function.
PMID 16409628 · PMC1395344 · BMC bioinformatics · 2006 · 8 claims · 6 setups
A self-supervised data mining approach can find relationships between sequence features and functional annotations without preconceived functional categories.
-
Full-text index only
CorGen--measuring and generating long-range correlations for DNA sequence analysis.
PMID 16845099 · PMC1538783 · Nucleic acids research · 2006 · 8 claims · 3 setups
CorGen is a web server that measures long-range correlations in DNA sequences and generates random sequences with the same (or user-specified) correlation and composition parameters
-
Has reproduction · 96
GC-biased gene conversion conceals the prediction of the nearly neutral theory in avian genomes.
PMID 30616647 · PMC6322265 · Genome biology · 2019 · 8 claims · 6 setups
gBGC conceals the correlation between life-history traits and dN/dS in birds; accounting for it reveals correlations consistent with nearly neutral theory
-
Full-text index only
Large-scale analysis of human alternative protein isoforms: pattern classification and correlation with subcellular localization signals.
PMID 15860772 · PMC1087780 · Nucleic acids research · 2005 · 8 claims · 8 setups
Constructed a large-scale dataset of 6876 human alternative protein isoforms from 2624 genes by combining H-Invitational full-length cDNA data and SwissProt VARSPLIC entries
-
Full-text index only
Correlation of microsynteny conservation and disease gene distribution in mammalian genomes.
PMID 19909546 · PMC2779822 · BMC genomics · 2009 · 7 claims · 8 setups
Density of mouse orthologs of human disease genes correlates with regions of conserved microsynteny in the mouse genome
-
Full-text index only
Identification and characterization of HLA-A*0301 epitopes in HIV-1 gag proteins using a novel approach.
PMID 19903485 · PMC2836169 · Journal of immunological methods · 2010 · 7 claims · 7 setups
PS mutations V7I and I34L (p17) and K403R (p7) in HIV-1 gag significantly correlate with HLA-A*0301
-
Full-text index only
Predicting deleterious nsSNPs: an analysis of sequence and structural attributes.
PMID 16630345 · PMC1489951 · BMC bioinformatics · 2006 · 8 claims · 7 setups
Sequence conservation (PSIC score difference) at the nsSNP position is the single most useful attribute for predicting deleterious vs neutral status.
-
Full-text index only
Aggregation propensity of the human proteome.
PMID 18927604 · PMC2557143 · PLoS computational biology · 2008 · 8 claims · 7 setups
Long proteins have, on average, less intense/pronounced aggregation peaks than short proteins
-
Full-text index only
Sequence similarity network reveals common ancestry of multidomain proteins.
PMID 18475320 · PMC2377100 · PLoS computational biology · 2008 · 8 claims · 6 setups
Traditional homology definitions do not capture multidomain evolution; the authors extend the definition to include domain insertion via a common ancestral locus model.
-
Full-text index only
Benchmarking ortholog identification methods using functional genomics data.
PMID 16613613 · PMC1557999 · Genome biology · 2006 · 8 claims · 7 setups
InParanoid is the best overall ortholog identification method for identifying functionally equivalent proteins when sensitivity and selectivity are combined into an overall score.
-
Full-text index only
Diversity of preferred nucleotide sequences around the translation initiation codon in eukaryote genomes.
PMID 18086709 · PMC2241899 · Nucleic acids research · 2008 · 8 claims · 5 setups
Preferred nucleotide sequences around the initiation codon are diverse among eukaryote species, but differences roughly reflect evolutionary relationships between species
-
Full-text index only
EPGD: a comprehensive web resource for integrating and displaying eukaryotic paralog/paralogon information.
PMID 17984073 · PMC2238967 · Nucleic acids research · 2008 · 8 claims · 8 setups
EPGD is a gene-centered, internet-accessible database integrating paralog family and paralogon information for 26 eukaryotic genomes.
-
Full-text index only
Functional importance of different patterns of correlation between adjacent cassette exons in human and mouse.
PMID 18439302 · PMC2432081 · BMC genomics · 2008 · 8 claims · 7 setups
Adjacent cassette exon pairs can be categorized by EST-derived correlation coefficient into three groups: mutually exclusive (ME, r<=-0.7), independent (IND, -0.2<=r<=0.2), and linked (LNK, r>=0.7)
-
Full-text index only
Sequence determinants of human microsatellite variability.
PMID 20015383 · PMC2806349 · BMC genomics · 2009 · 6 claims · 4 setups
Mean and maximum number of repeats across individuals are positively correlated with heterozygosity
-
Full-text index only
The Origin at 150: is a new evolutionary synthesis in sight?
PMID 19836100 · PMC2784144 · Trends in genetics : TIG · 2009 · 8 claims · 4 setups
The Modern Synthesis (neo-Darwinism) has crumbled and its central tenets are overturned or radically revised in the post-genomic era
-
Full-text index only
Proteomics in alcohol research.
PMID 12875051 · PMC6683837 · Alcohol research & health : the journal of the National Institute on Alcohol Abuse and Alcoholism · 2002 · 7 claims · 8 setups
The proteome is larger and more complex than the genome due to differential splicing, post-translational modifications (PTMs), and protein-protein interactions.