Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
The vertebrate genome annotation (Vega) database.
PMID 18003653 · PMC2238886 · Nucleic acids research · 2008 · 8 claims · 8 setups
Vega is a database for viewing manual genome annotation of human, mouse and zebrafish genomic sequences produced at the Wellcome Trust Sanger Institute.
-
Full-text index only
Designating eukaryotic orthology via processed transcription units.
PMID 18445630 · PMC2425467 · Nucleic acids research · 2008 · 8 claims · 5 setups
Existing ortholog databases discard/ignore alternative splicing via all-against-all protein comparisons, causing ambiguous ortholog calls and misclassification of AS isoforms as in-paralogs
-
Full-text index only
Comparisons of substitution, insertion and deletion probes for resequencing and mutational analysis using oligonucleotide microarrays.
PMID 15722479 · PMC549431 · Nucleic acids research · 2005 · 7 claims · 4 setups
Two base deletion probes display the highest average hybridization specificity, followed by single base substitution, single base deletion, and single base insertion probes.
-
Has reproduction · 92
Telomere-to-telomere reference genome for Panax ginseng highlights the evolution of saponin biosynthesis.
PMID 38883331 · PMC11179851 · Horticulture research · 2024 · 8 claims · 8 setups
A telomere-to-telomere reference genome of P. ginseng was assembled (3.45 Gb, 24 chromosomes, 77266 protein-coding genes)
-
Full-text index only
Satellog: a database for the identification and prioritization of satellite repeats in disease association studies.
PMID 15949044 · PMC1181805 · BMC bioinformatics · 2005 · 7 claims · 6 setups
Satellog is a database cataloging all pure 1-16 unit satellite repeats in the human genome with supplementary polymorphism, gene-location, and expression data for prioritizing repeats in disease-association studies.
-
Full-text index only
Genetic variation in an individual human exome.
PMID 18704161 · PMC2493042 · PLoS genetics · 2008 · 8 claims · 7 setups
The ~12,500 nonsilent coding variants in the HuRef exome can be reduced ~8-fold to a set of ~1,600 variants most likely to affect protein function.
-
Has reproduction · 75
geneshot: gene-level metagenomics identifies genome islands associated with immunotherapy response.
PMID 33952321 · PMC8097837 · Genome biology · 2021 · 8 claims · 4 setups
geneshot is a gene-level metagenomic bioinformatics tool that clusters de novo assembled protein-coding genes into co-abundant gene groups (CAGs) to reduce dimensionality and generate testable hypotheses from WGS microbiome data
-
Has reproduction · 60
TRAPID 2.0: a web application for taxonomic and functional analysis of de novo transcriptomes.
PMID 34197621 · PMC8464036 · Nucleic acids research · 2021 · 8 claims · 8 setups
TRAPID 2.0 is a web application performing global characterization of de novo transcriptomes via structural, functional, and taxonomic annotation in an initial processing phase, followed by an exploratory phase of downstream analyses.
-
Full-text index only
Genome-wide diversity and selective pressure in the human rhinovirus.
PMID 17477878 · PMC1892812 · Virology journal · 2007 · 7 claims · 6 setups
Whole genome and subgenomic phylogenies of HRV are essentially identical at every locus, indicating consistent phylogenetic patterns across the genome.
-
Full-text index only
A potentially deleterious new CYP2C9 polymorphism identified in an African American patient with major hemorrhage on warfarin therapy.
PMID 19083245 · PMC2662477 · Blood cells, molecules & diseases · 2009 · 7 claims · 4 setups
A new CYP2C9 coding polymorphism, G1078A (D360N) in exon 7, was identified in an African American patient who experienced major gastrointestinal hemorrhage on warfarin therapy
-
Full-text index only
A pharmacogenetics study of the human glucuronosyltransferase UGT1A4.
PMID 19890225 · PMC6177227 · Pharmacogenetics and genomics · 2009 · 7 claims · 6 setups
Extensive sequencing of UGT1A4 (promoter to exon 1+2000bp) identified numerous novel polymorphisms: 13 intronic, 39 promoter, and 14 exonic variants (10 causing amino acid changes)
-
Full-text index only
Variation in conserved non-coding sequences on chromosome 5q and susceptibility to asthma and atopy.
PMID 16336695 · PMC1325232 · Respiratory research · 2005 · 6 claims · 8 setups
There is overall little sequence variation in the conserved non-coding elements (CNEs) on 5q31, including none detected in CNE-B/CNS-1
-
Full-text index only
Phylogenomic approaches to common problems encountered in the analysis of low copy repeats: the sulfotransferase 1A gene family example.
PMID 15752422 · PMC555591 · BMC evolutionary biology · 2005 · 8 claims · 8 setups
A previously unidentified fourth human SULT1A gene (SULT1A4) exists on chromosome 16 and is transcriptionally active
-
Has reproduction · 73
Vespucci: a system for building annotated databases of nascent transcripts.
PMID 24304890 · PMC3936758 · Nucleic acids research · 2014 · 8 claims · 7 setups
Existing ChIP-seq and RNA-seq analysis platforms (e.g. Cufflinks, peak callers) are unsuited to GRO-seq because they assume spliced/exonic reads, uniform density and paired-end data, and cannot identify transcriptional units de novo across the whole genome.
-
Full-text index only
Exogean: a framework for annotating protein-coding genes in eukaryotic genomic DNA.
PMID 16925841 · PMC1810556 · Genome biology · 2006 · 8 claims · 5 setups
Exogean is a framework using directed acyclic coloured multigraphs (DACMs) to represent biological objects (mRNA, ESTs, protein alignments, exons) and iteratively combine them into complex protein-coding transcript models.
-
Full-text index only
Automatic annotation of eukaryotic genes, pseudogenes and promoters.
PMID 16925832 · PMC1810547 · Genome biology · 2006 · 8 claims · 6 setups
Fgenesh++ gene prediction pipeline identifies 91% of coding nucleotides with 90% specificity
-
Full-text index only
The Genographic Project public participation mitochondrial DNA database.
PMID 17604454 · PMC1904368 · PLoS genetics · 2007 · 7 claims · 4 setups
The Genographic Project created the largest standardized human mtDNA database to date, comprising 78,590 genotypes from the first 18 months of public participation.
-
Full-text index only
A SNP-centric database for the investigation of the human genome.
PMID 15046636 · PMC395999 · BMC bioinformatics · 2004 · 8 claims · 3 setups
SNPper is a web-based, integrated SNP database combining dbSNP, the Human Genome sequence (Goldenpath), LocusLink, GeneOntology, and SWISS-PROT data with querying, visualization, and export tools.
-
Full-text index only
mtDB: Human Mitochondrial Genome Database, a resource for population genetics and medical sciences.
PMID 16381973 · PMC1347373 · Nucleic acids research · 2006 · 8 claims · 3 setups
mtDB is a comprehensive, actively maintained database of published human mitochondrial genome sequences, providing a common resource for population genetics and medical research
-
Full-text index only
Analysis of nucleotide diversity of NAT2 coding region reveals homogeneity across Native American populations and high intra-population diversity.
PMID 16847467 · PMC3099416 · The pharmacogenomics journal · 2007 · 8 claims · 6 setups
NAT2 variants are homogeneously distributed across native populations of the American continent