Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
Genomic divergences among cattle, dog and human estimated from large-scale alignments of genomic sequences.
PMID 16759380 · PMC1525190 · BMC genomics · 2006 · 8 claims · 6 setups
Overall pairwise genomic divergences among cattle, dog and human are relatively constant (0.32–0.37 change/site)
-
Full-text index only
Comprehensive splice-site analysis using comparative genomics.
PMID 16914448 · PMC1557818 · Nucleic acids research · 2006 · 8 claims · 6 setups
Over half a million splice sites were collected from five species (H. sapiens, M. musculus, D. melanogaster, C. elegans, A. thaliana) and classified into four main subtypes: U2-type GT-AG and GC-AG, and U12-type GT-AG and AT-AC.
-
Full-text index only
A general definition and nomenclature for alternative splicing events.
PMID 18688268 · PMC2467475 · PLoS computational biology · 2008 · 6 claims · 4 setups
Existing AS nomenclatures (Malko et al.'s 5-letter strings, Nagasaki et al.'s bit matrices, and the ASD/ATD/AEdb system) are redundant, ambiguous, or incapable of representing complex or large splicing variations.
-
Full-text index only
Exhaustive prediction of disease susceptibility to coding base changes in the human genome.
PMID 18793467 · PMC2537574 · BMC bioinformatics · 2008 · 8 claims · 7 setups
Inter-species conservation is the strongest single predictor of disease-associated coding mutations among the factors tested.
-
Full-text index only
CanPredict: a computational tool for predicting cancer-associated missense mutations.
PMID 17537827 · PMC1933186 · Nucleic acids research · 2007 · 8 claims · 7 setups
CanPredict is a web application providing public access to a random forest classifier that combines SIFT, LogR.E-value, and GOSS scores to predict whether a missense mutation is cancer-associated
-
Has reproduction · 87
R2DT is a framework for predicting and visualising RNA secondary structure using templates.
PMID 34108470 · PMC8190129 · Nature communications · 2021 · 8 claims · 6 setups
R2DT is a template-based computational framework/pipeline that predicts and visualises RNA 2D structure in standardised, community-accepted layouts
-
Has reproduction · 50
MEDUSA: A Pipeline for Sensitive Taxonomic Classification and Flexible Functional Annotation of Metagenomic Shotgun Sequences.
PMID 35330728 · PMC8940201 · Frontiers in genetics · 2022 · 6 claims · 6 setups
MEDUSA is an automated, Conda-installable and Snakemake-managed pipeline performing preprocessing, assembly, alignment, taxonomic classification, and functional annotation on shotgun data.
-
Full-text index only
Large-scale analysis of human alternative protein isoforms: pattern classification and correlation with subcellular localization signals.
PMID 15860772 · PMC1087780 · Nucleic acids research · 2005 · 8 claims · 8 setups
Constructed a large-scale dataset of 6876 human alternative protein isoforms from 2624 genes by combining H-Invitational full-length cDNA data and SwissProt VARSPLIC entries
-
Full-text index only
Systematic identification of pseudogenes through whole genome expression evidence profiling.
PMID 16945953 · PMC1636364 · Nucleic acids research · 2006 · 8 claims · 8 setups
Developed a novel bioinformatics method that identifies pseudogenes by profiling whole-genome transcript and protein expression evidence
-
Full-text index only
ARED 3.0: the large and diverse AU-rich transcriptome.
PMID 16381826 · PMC1347415 · Nucleic acids research · 2006 · 7 claims · 6 setups
ARED 3.0 computationally mapped more than 4000 ARE-mRNAs to the human genome, representing 5-8% of human genes.
-
Full-text index only
Genome mapping and expression analyses of human intronic noncoding RNAs reveal tissue-specific patterns and enrichment in genes related to regulation of transcription.
PMID 17386095 · PMC1868932 · Genome biology · 2007 · 8 claims · 4 setups
More than 55,000 totally intronic noncoding (TIN) RNAs are transcribed from the introns of 74% of unique RefSeq genes.
-
Full-text index only
Recent segmental and gene duplications in the mouse genome.
PMID 12914656 · PMC193640 · Genome biology · 2003 · 8 claims · 8 setups
33.6 Mb (1.2%) of the February 2003 mouse genome assembly (2,695 Mb) is involved in recent segmental duplications
-
Full-text index only
Discovery of novel human transcript variants by analysis of intronic single-block EST with polyadenylation site.
PMID 19906316 · PMC2784480 · BMC genomics · 2009 · 8 claims · 7 setups
Intronic single-block ESTs with poly(A/T) tails reveal previously unidentified novel transcript variants missed by existing databases.
-
Has reproduction · 95
A whole genome duplication drives the genome evolution of Phytophthora betacei, a closely related species to Phytophthora infestans.
PMID 34740326 · PMC8571832 · BMC genomics · 2021 · 8 claims · 7 setups
P. betacei P8084 has the largest sequenced genome in the Phytophthora genus (270 Mb)
-
Full-text index only
Widespread ultraconservation divergence in primates.
PMID 18492662 · PMC2464743 · Molecular biology and evolution · 2008 · 8 claims · 4 setups
The number of UCEs has decreased throughout primate evolution, from ~1,000 in ancestral primates to 635 in modern humans.
-
Has reproduction · 65
Estimating biodiversity across the tree of life on Mount Everest's southern flank with environmental DNA.
PMID 36148432 · PMC9486557 · iScience · 2022 · 8 claims · 6 setups
eDNA from ten high-alpine ponds and streams (4,500-5,500 m) on Mt. Everest's southern flank revealed 187 potential orders from 36 phyla across the Tree of Life.
-
Has reproduction · 62
Metatranscriptomics of the human oral microbiome during health and disease.
PMID 24692635 · PMC3977359 · mBio · 2014 · 8 claims · 8 setups
Disease-associated periodontal communities display conserved community-level metabolic gene expression profiles between patients, whereas the metabolic gene expression of individual species is highly variable between patients.