Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Has reproduction · 67
A consensus approach to vertebrate de novo transcriptome assembly from RNA-seq data: assembly of the duck (Anas platyrhynchos) transcriptome.
PMID 25009556 · PMC4070175 · Frontiers in genetics · 2014 · 8 claims · 8 setups
Multiple k-mer (MK) assemblies are more complete than single k-mer (SK) assemblies, showing higher reads-mapped-back-to-transcripts (RMBT) and higher CEGMA complete-gene percentages for all three tools.
-
Has reproduction · 76
The genome and development-dependent transcriptomes of Pyronema confluens: a window into fungal evolution.
PMID 24068976 · PMC3778014 · PLoS genetics · 2013 · 8 claims · 8 setups
The 50 Mb P. confluens genome with 13,369 predicted protein-coding genes is more characteristic of higher filamentous ascomycetes than of the large, repeat-rich Tuber melanosporum genome, showing that the truffle's expanded genome is not typical of the Pezizales.
-
Full-text index only
Gene loss rate: a probabilistic measure for the conservation of eukaryotic genes.
PMID 17158152 · PMC1802574 · Nucleic acids research · 2007 · 8 claims · 8 setups
GLR is a novel maximum-likelihood measure of gene loss rate that probabilistically weighs all possible ancestral phyletic patterns rather than relying on a single parsimonious reconstruction.
-
Full-text index only
Pegasys: software for executing and integrating analyses of biological sequences.
PMID 15096276 · PMC406494 · BMC bioinformatics · 2004 · 8 claims · 7 setups
Pegasys is a flexible, modular, customizable software system for executing and integrating heterogeneous biological sequence analysis tools
-
Has reproduction · 38
Genomic capacities for Reactive Oxygen Species metabolism across marine phytoplankton.
PMID 37098087 · PMC10128935 · PloS one · 2023 · 8 claims · 3 setups
Genes encoding superoxide (O2•−) scavenging are ubiquitous across phytoplankton, but their fractional gene allocation decreases with increasing cell radius, consistent with a nearly fixed core gene set.
-
Full-text index only
Diversity of tRNA genes in eukaryotes.
PMID 17088292 · PMC1693877 · Nucleic acids research · 2006 · 8 claims · 6 setups
The number of tRNA genes having the same anticodon but different sequences elsewhere (isodecoder genes) varies significantly (10–246) across 11 eukaryotes despite isoacceptor numbers being similar (41–55)
-
Full-text index only
The global landscape of sequence diversity.
PMID 17996061 · PMC2258180 · Genome biology · 2007 · 7 claims · 5 setups
Eukaryotic sequence datasets show substantially greater genetic diversity (higher sequence/gene family discovery rates) than bacterial datasets, likely related to differences in modes of genetic inheritance.
-
Full-text index only
EPGD: a comprehensive web resource for integrating and displaying eukaryotic paralog/paralogon information.
PMID 17984073 · PMC2238967 · Nucleic acids research · 2008 · 8 claims · 8 setups
EPGD is a gene-centered, internet-accessible database integrating paralog family and paralogon information for 26 eukaryotic genomes.
-
Full-text index only
The TIGR Gene Indices: clustering and assembling EST and known genes and integration with eukaryotic genomes.
PMID 15608288 · PMC540018 · Nucleic acids research · 2005 · 8 claims · 8 setups
The TIGR Gene Indices (TGI) are a collection of 77 species-specific databases that cluster and assemble EST and known gene sequences into tentative consensus (TC) sequences to identify and characterize expressed transcripts.
-
Has reproduction · 98
Sequence-based pangenomic core detection.
PMID 35663029 · PMC9160775 · iScience · 2022 · 7 claims · 3 setups
Sequence-based pangenomic core detection can be performed directly on unannotated genome sequences using a colored de Bruijn graph, avoiding bias from error-prone gene annotations
-
Full-text index only
MBGD update 2010: toward a comprehensive resource for exploring microbial genome diversity.
PMID 19906735 · PMC2808943 · Nucleic acids research · 2010 · 8 claims · 6 setups
MBGD allows users to create ortholog groups using a specified subgroup of organisms, distinguishing it from other comparative genomics resources
-
Has reproduction · 58
Mucospheres produced by a mixotrophic protist impact ocean carbon cycling.
PMID 35288549 · PMC8921327 · Nature communications · 2022 · 8 claims · 8 setups
P. cf. balticum produces carbon-rich 'mucospheres' that attract, capture and immobilise a wide range of prokaryotic and eukaryotic prey
-
Has reproduction · 69
High-resolution metagenomic reconstruction of the freshwater spring bloom.
PMID 36698172 · PMC9878933 · Microbiome · 2023 · 8 claims · 8 setups
High-frequency metagenomic sampling (57 samples, 3 filter sizes, 2 depths, 37 days) recovers thousands of prokaryotic, eukaryotic, and viral genomes tracking spring bloom succession in fine detail
-
Full-text index only
The human L-threonine 3-dehydrogenase gene is an expressed pseudogene.
PMID 12361482 · PMC131051 · BMC genetics · 2002 · 8 claims · 7 setups
The human TDH gene is located at chromosome 8p23-22, spans 10 kb, and has 8 exons that would be expected to encode a 369-residue ORF.
-
Full-text index only
The excess of 5' introns in eukaryotic genomes.
PMID 16314314 · PMC1292992 · Nucleic acids research · 2005 · 7 claims · 4 setups
All 21 eukaryotic genomes studied show a statistically significant 5′-biased distribution of introns in protein-coding genes
-
Full-text index only
Comprehensive genome analysis of 203 genomes provides structural genomics with new insights into protein family space.
PMID 16481312 · PMC1373602 · Nucleic acids research · 2006 · 8 claims · 7 setups
The number of protein families continues to expand steadily as more genomes are sequenced, showing no sign of saturation.
-
Full-text index only
Proceedings of the First International Conference on Phylogenomics. March 15-19, 2006. Quebec, Canada.
PMID 17288567 · PMC1796603 · BMC evolutionary biology · 2007 · 8 claims · 8 setups
Gene tree parsimony applied to EST data with widespread gene duplication can infer an organismal phylogeny in excellent agreement with the expected angiosperm phylogeny.
-
Full-text index only
Pseudofam: the pseudogene families database.
PMID 18957444 · PMC2686518 · Nucleic acids research · 2009 · 8 claims · 7 setups
Pseudofam is an online database of pseudogene families built by mapping pseudogenes to Pfam protein families, providing query tools, statistics, and sequence alignments
-
Full-text index only
Expansion of the human mitochondrial proteome by intra- and inter-compartmental protein duplication.
PMID 19930686 · PMC3091328 · Genome biology · 2009 · 8 claims · 6 setups
The human mitochondrial proteome expanded via two prevailing gene duplication modes: intra-mitochondrial and inter-compartmental duplication
-
Full-text index only
GeneMark: web software for gene finding in prokaryotes, eukaryotes and viruses.
PMID 15980510 · PMC1160247 · Nucleic acids research · 2005 · 8 claims · 2 setups
The GeneMark website provides web interfaces to the GeneMark family of ab initio gene-finding programs for prokaryotic, eukaryotic and viral genomic sequences