Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
EPGD: a comprehensive web resource for integrating and displaying eukaryotic paralog/paralogon information.
PMID 17984073 · PMC2238967 · Nucleic acids research · 2008 · 8 claims · 8 setups
EPGD is a gene-centered, internet-accessible database integrating paralog family and paralogon information for 26 eukaryotic genomes.
-
Full-text index only
Comparative analysis reveals signatures of differentiation amid genomic polymorphism in Lake Malawi cichlids.
PMID 18616806 · PMC2530870 · Genome biology · 2008 · 8 claims · 8 setups
Lake Malawi cichlids are phenotypically and behaviorally diverse but appear genetically like a single subdivided population rather than distinct species
-
Has reproduction · 81
Whole-Genome Sequence of Cervid atadenovirus A from the Initial Cases of an Adenovirus Hemorrhagic Disease Epizootic of Black-Tailed Deer in Canada.
PMID 36129291 · PMC9584332 · Microbiology resource announcements · 2022 · 7 claims · 4 setups
A complete 30,616-nucleotide genome of Cervid atadenovirus A was determined from lung tissue of black-tailed deer that died of AHD in British Columbia in 2020
-
Full-text index only
Coverage of whole proteome by structural genomics observed through protein homology modeling database.
PMID 17146617 · PMC1769342 · Journal of structural and functional genomics · 2006 · 8 claims · 7 setups
FAMSBASE, a homology-modeling database of whole-genome ORFs, currently covers about 50% of predicted ORFs (368,724 of 734,193) across 276 genomes with modeled 3D structures.
-
Full-text index only
Gene-centric characteristics of genome-wide association studies.
PMID 18060058 · PMC2092383 · PloS one · 2007 · 8 claims · 5 setups
High-density SNP chips using either direct or indirect selection approaches provide very high coverage in genic regions and capture most known common disease variants under the HapMap framework.
-
Full-text index only
Evaluating the performance of commercial whole-genome marker sets for capturing common genetic variation.
PMID 17562002 · PMC1914356 · BMC genomics · 2007 · 8 claims · 5 setups
Commercial SNP panels provide levels of coverage in a non-reference Caucasian (Estonian) population similar to those seen in the HapMap CEPH (CEU) population sample
-
Full-text index only
Low conservation and species-specific evolution of alternative splicing in humans and mice: comparative genomics analysis using well-annotated full-length cDNAs.
PMID 18838389 · PMC2582632 · Nucleic acids research · 2008 · 7 claims · 8 setups
Although 86% of individual human exons are conserved in the mouse genome, only a small fraction (431/20392, ~2%) of human AS variants are perfectly conserved AS variants in mice.
-
Has reproduction · 87
Insights into the Evolution of the New World Diploid Cottons (Gossypium, Subgenus Houzingenia) Based on Genome Sequencing.
PMID 30476109 · PMC6320677 · Genome biology and evolution · 2019 · 8 claims · 8 setups
Subgenus Houzingenia originated via transoceanic dispersal from Africa ~6.6 Ma, with most biodiversity arising from rapid mid-Pleistocene (0.5–2.0 Ma) diversification plus multiple long-distance dispersals.
-
Full-text index only
Tracing the origin of functional and conserved domains in the human proteome: implications for protein evolution at the modular level.
PMID 17090320 · PMC1654190 · BMC evolutionary biology · 2006 · 8 claims · 5 setups
HHpred (HMM-HMM comparison) detects remote homologs in the human proteome with higher sensitivity than hmmpfam (HMMER), giving 10% more functional domain coverage and 20% higher residue coverage against Pfam-A families.
-
Full-text index only
Benchmarking tools for the alignment of functional noncoding DNA.
PMID 14736341 · PMC344529 · BMC bioinformatics · 2004 · 8 claims · 4 setups
Global alignment tools (Avid, ClustalW, Lagan, Needle, DiAlign-G) typically have higher sensitivity over entire noncoding sequences and within constrained blocks than local tools
-
Full-text index only
Computational tradeoffs in multiplex PCR assay design for SNP genotyping.
PMID 16042802 · PMC1190169 · BMC genomics · 2005 · 7 claims · 6 setups
Achieving high-multiplexing/high-coverage multiplex PCR designs is subject to a computational phase transition as the SNP-pair compatibility probability crosses a critical threshold
-
Full-text index only
Assessing the gene space in draft genomes.
PMID 19042974 · PMC2615622 · Nucleic acids research · 2009 · 6 claims · 7 setups
The proportion of mapped CEGs in a draft genome assembly is a useful metric for describing gene space completeness, complementing N50 and x-fold coverage.
-
Has reproduction · 89
TrEMOLO: accurate transposable element allele frequency estimation using long-read sequencing data combining assembly and mapping-based approaches.
PMID 37013657 · PMC10069131 · Genome biology · 2023 · 6 claims · 6 setups
TrEMOLO combines an assembly-based INSIDER module and a mapping-based OUTSIDER module to detect TE insertions/deletions from long-read sequencing data and estimate their allele frequency
-
Full-text index only
Using several pair-wise informant sequences for de novo prediction of alternatively spliced transcripts.
PMID 16925842 · PMC1810557 · Genome biology · 2006 · 8 claims · 4 setups
MARS, an extension of the Twinscan algorithm, uses multiple pairwise informant genomes to predict human alternatively spliced transcripts de novo without expressed sequence information.
-
Full-text index only
MODBASE, a database of annotated comparative protein structure models and associated resources.
PMID 18948282 · PMC2686492 · Nucleic acids research · 2009 · 8 claims · 8 setups
MODBASE contains 5,152,695 reliable comparative protein structure models for 1,593,209 unique protein sequences.
-
Has reproduction · 91
Genomic Description of 'Candidatus Abyssubacteria,' a Novel Subsurface Lineage Within the Candidate Phylum Hydrogenedentes.
PMID 30210471 · PMC6121073 · Frontiers in microbiology · 2018 · 8 claims · 7 setups
SURF_5 and SURF_17 are the first full genomes of a novel bacterial lineage, 'Candidatus Abyssubacteria,' within the candidate phylum Hydrogenedentes
-
Full-text index only
Characterisation of the genomic architecture of human chromosome 17q and evaluation of different methods for haplotype block definition.
PMID 15850495 · PMC1090572 · BMC genetics · 2005 · 8 claims · 6 setups
Haplotype block definitions based on LD measures (Definitions 1, 2, 3, 5) produce fewer, shorter blocks with limited sequence coverage compared to the haplotype diversity-based method (Definition 4)
-
Full-text index only
Automatic discovery of cross-family sequence features associated with protein function.
PMID 16409628 · PMC1395344 · BMC bioinformatics · 2006 · 8 claims · 6 setups
A self-supervised data mining approach can find relationships between sequence features and functional annotations without preconceived functional categories.
-
Full-text index only
Efficacy assessment of SNP sets for genome-wide disease association studies.
PMID 17726055 · PMC2034459 · Nucleic acids research · 2007 · 6 claims · 4 setups
τ, derived from Shannon entropy and swept radius ɛ, approximates the relative sample size efficiency of a marker set for mapping a causal variant at a given map position compared to a maximally polymorphic SNP
-
Full-text index only
Power analysis for genome-wide association studies.
PMID 17725844 · PMC2042984 · BMC genetics · 2007 · 8 claims · 6 setups
Developed a method to compute genome-wide association study power using tag SNPs and representative population genotype data (HapMap), equivalent to the cumulative r2-adjusted power of Jorgenson and Witte.