Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
CpG_MI: a novel approach for identifying functional CpG islands in mammalian genomes.
PMID 19854943 · PMC2800233 · Nucleic acids research · 2010 · 8 claims · 6 setups
Functional ('bona fide') CGIs show distinct average/cumulative mutual information (AMI/CMI) distributions of neighboring CpG distances compared to non-functional CGIs and random genome segments
-
Full-text index only
How to find soluble proteins: a comprehensive analysis of alpha/beta hydrolases for recombinant expression in E. coli.
PMID 15804363 · PMC1079826 · BMC genomics · 2005 · 7 claims · 7 setups
Predicted solubility in E. coli (via CV-CV') depends on hydrolase size, phylogenetic origin, homologous family, and superfamily
-
Has reproduction · 80
Differential analysis of RNA structure probing experiments at nucleotide resolution: uncovering regulatory functions of RNA structure.
PMID 35869080 · PMC9307511 · Nature communications · 2022 · 7 claims · 4 setups
DiffScan is a computational framework combining a Normalization module and a Scan module to identify SVRs at nucleotide resolution from SP data.
-
Full-text index only
Exhaustive prediction of disease susceptibility to coding base changes in the human genome.
PMID 18793467 · PMC2537574 · BMC bioinformatics · 2008 · 8 claims · 7 setups
Inter-species conservation is the strongest single predictor of disease-associated coding mutations among the factors tested.
-
Full-text index only
Analysis of recent segmental duplications in the bovine genome.
PMID 19951423 · PMC2796684 · BMC genomics · 2009 · 8 claims · 6 setups
Recently duplicated sequence (≥1 kb, ≥90% identity) comprises 3.11% (94.4 Mb) of the bovine genome assembly (Btau_4.0)
-
Full-text index only
Global distribution of rubella virus genotypes.
PMID 14720390 · PMC3034328 · Emerging infectious diseases · 2003 · 8 claims · 6 setups
Phylogenetic analysis of 103 E1 gene sequences from 17 countries confirms at least two rubella virus genotypes, RGI and RGII
-
Full-text index only
Diversity of preferred nucleotide sequences around the translation initiation codon in eukaryote genomes.
PMID 18086709 · PMC2241899 · Nucleic acids research · 2008 · 8 claims · 5 setups
Preferred nucleotide sequences around the initiation codon are diverse among eukaryote species, but differences roughly reflect evolutionary relationships between species
-
Full-text index only
Disease-aging network reveals significant roles of aging genes in connecting genetic diseases.
PMID 19779549 · PMC2739292 · PLoS computational biology · 2009 · 8 claims · 8 setups
Human disease genes are much closer to aging genes in the PPI network than expected by chance
-
Full-text index only
Testing groups of genomic locations for enrichment in disease loci using linkage scan data: a method for hypothesis testing.
PMID 16848972 · PMC3525155 · Human genomics · 2006 · 8 claims · 2 setups
A method testing enrichment of a group of genomic locations for disease loci by comparing the average NPL score of the group to a null distribution from randomly drawn groups of equal size
-
Has reproduction · 95
MetaMap: an atlas of metatranscriptomic reads in human disease-related RNA-seq data.
PMID 29901703 · PMC6025204 · GigaScience · 2018 · 6 claims · 7 setups
A two-step 'omni' RNA-seq pipeline (MetaMap) combining STAR human alignment with CLARK-S metagenomic classification can quantify archaeal, bacterial, and viral reads from the non-human read fraction of human RNA-seq data
-
Has reproduction · 96
GC-biased gene conversion conceals the prediction of the nearly neutral theory in avian genomes.
PMID 30616647 · PMC6322265 · Genome biology · 2019 · 8 claims · 6 setups
gBGC conceals the correlation between life-history traits and dN/dS in birds; accounting for it reveals correlations consistent with nearly neutral theory
-
Full-text index only
POCUS: mining genomic sequence annotation to predict disease genes.
PMID 14611661 · PMC329128 · Genome biology · 2003 · 8 claims · 6 setups
Genes predisposing to the same disease tend to share functional annotation IDs (GO/InterPro) more than expected by chance
-
Full-text index only
Direct inference of SNP heterozygosity rates and resolution of LOH detection.
PMID 18052545 · PMC2098867 · PLoS computational biology · 2007 · 6 claims · 7 setups
A large proportion of SNPs in dbSNP have high-variance HET rate estimates, limiting their reliability for LOH study design.
-
Full-text index only
Does distance matter? Variations in alternative 3' splicing regulation.
PMID 17704130 · PMC2018619 · Nucleic acids research · 2007 · 8 claims · 7 setups
Alternative 3' splice sites can be distinguished from constitutive splice sites by a combination of sequence/conservation properties that vary depending on the distance between the splice sites.
-
Full-text index only
Genome-wide analysis of human disease alleles reveals that their locations are correlated in paralogous proteins.
PMID 18989397 · PMC2565504 · PLoS computational biology · 2008 · 7 claims · 5 setups
The locations of sequence variants are correlated between paralogous human proteins more than expected by chance.
-
Full-text index only
The biological function of some human transcription factor binding motifs varies with position relative to the transcription start site.
PMID 18367472 · PMC2377430 · Nucleic acids research · 2008 · 8 claims · 5 setups
1226 eight-letter DNA words show statistically significant positional preferences relative to the TSS across 7914 human promoter regions
-
Full-text index only
Network properties of complex human disease genes identified through genome-wide association studies.
PMID 19956617 · PMC2779513 · PloS one · 2009 · 7 claims · 6 setups
Complex disease genes are significantly less central (lower degree/closeness, higher eccentricity) in the human interactome than essential and monogenic disease genes, occupying an intermediate niche between monogenic disease genes and non-disease genes
-
Full-text index only
Structure of protein interaction networks and their implications on drug design.
PMID 19876376 · PMC2760708 · PLoS computational biology · 2009 · 8 claims · 6 setups
Budding yeast and human PINs are scale-rich and configured as highly optimized tolerance (HOT) networks similar to Internet router-level topology, rather than scale-free networks formed by preferential attachment.
-
Full-text index only
Discovering cancer genes by integrating network and functional properties.
PMID 19765316 · PMC2758898 · BMC medical genomics · 2009 · 8 claims · 6 setups
Cancer genes have distinct PPI network topology (higher connectivity, higher clustering coefficient, shorter path length to known cancer genes) compared to non-cancer genes
-
Full-text index only
Analyses and comparison of accuracy of different genotype imputation methods.
PMID 18958166 · PMC2569208 · PloS one · 2008 · 8 claims · 3 setups
Stronger LD produces higher imputation accuracy rates for all five methods