Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Has reproduction · 56
Comparative Metagenomic Analysis of Biosynthetic Diversity across Sponge Microbiomes Highlights Metabolic Novelty, Conservation, and Diversification.
PMID 35862823 · PMC9426513 · mSystems · 2022 · 8 claims · 5 setups
The vast majority of recovered gene cluster families (GCFs) in sponge microbiomes show no similarity to any characterized BGC, revealing extreme biosynthetic novelty
-
Full-text index only
Clustering by neurocognition for fine mapping of the schizophrenia susceptibility loci on chromosome 6p.
PMID 19694819 · PMC4286260 · Genes, brain, and behavior · 2009 · 6 claims · 6 setups
A family-based clustering strategy using neurocognitive test scores (CPT, WCST) can identify more homogeneous subgroups of schizophrenia families for genetic association analysis
-
Full-text index only
Comprehensive genome analysis of 203 genomes provides structural genomics with new insights into protein family space.
PMID 16481312 · PMC1373602 · Nucleic acids research · 2006 · 8 claims · 7 setups
The number of protein families continues to expand steadily as more genomes are sequenced, showing no sign of saturation.
-
Full-text index only
Polymorphism discovery and association analyses of the interferon genes in type 1 diabetes.
PMID 16504056 · PMC1402321 · BMC genetics · 2006 · 7 claims · 8 setups
No statistical evidence of a major association between T1D and any of the interferon or interferon-related genes tested (IFNA cluster, IFNB1, IFNW1, IFNG, ICSBP1)
-
Full-text index only
nf-core/proteinfamilies: a scalable pipeline for the generation of protein families.
PMID 41563008 · PMC12950615 · GigaScience · 2026 · 8 claims · 3 setups
nf-core/proteinfamilies is a scalable, parametrizable, open-source Nextflow pipeline that generates new protein families or assigns sequences to existing families using profile HMMs and MSAs
-
Full-text index only
The Princeton Protein Orthology Database (P-POD): a comparative genomics analysis tool for biologists.
PMID 17712414 · PMC1942082 · PloS one · 2007 · 8 claims · 5 setups
P-POD is the first comparative genomics database to combine results from multiple computational ortholog/homolog prediction methods with manually curated literature-derived experimental evidence of functional conservation.
-
Full-text index only
Automatic discovery of cross-family sequence features associated with protein function.
PMID 16409628 · PMC1395344 · BMC bioinformatics · 2006 · 8 claims · 6 setups
A self-supervised data mining approach can find relationships between sequence features and functional annotations without preconceived functional categories.
-
Full-text index only
An analysis of human microRNA and disease associations.
PMID 18923704 · PMC2559869 · PloS one · 2008 · 8 claims · 8 setups
MicroRNAs tend to show similar dysfunctional evidence (both up- or both down-regulated) for diseases within the same disease cluster, and different dysfunctional evidence between different disease clusters.
-
Full-text index only
Manual annotation and analysis of the defensin gene cluster in the C57BL/6J mouse reference genome.
PMID 20003482 · PMC2807441 · BMC genomics · 2009 · 8 claims · 6 setups
Manual annotation of the mouse Chromosome 8 defensin region identifies 98 gene loci: 54 in the alpha-defensin cluster and 44 in the beta-defensin cluster
-
Full-text index only
Mutation screen and association studies in the diacylglycerol O-acyltransferase homolog 2 gene (DGAT2), a positional candidate gene for early onset obesity on chromosome 11q13.
PMID 17477860 · PMC1871603 · BMC genetics · 2007 · 7 claims · 5 setups
DGAT2 is a plausible positional and functional candidate gene for obesity due to its localization at chr.11q13 (a linkage region) and its key role in triglyceride synthesis
-
Full-text index only
Hybrid sequencing reveals incompleteness of the H37Rv reference genome and highlights lineage-specific genomic divergence in Mycobacterium tuberculosis.
PMID 42224013 · PMC13225438 · Microbial genomics · 2026 · 8 claims · 6 setups
The H37Rv_ref reference genome, sequenced in 1998 with early technology, is incomplete relative to modern hybrid-sequenced assemblies
-
Has reproduction · 49
oPOSSUM-3: advanced analysis of regulatory motif over-representation across genes or ChIP-Seq datasets.
PMID 22973536 · PMC3429929 · G3 (Bethesda, Md.) · 2012 · 8 claims · 6 setups
oPOSSUM-3 is a web-accessible system that identifies over-represented TFBS and TFBS families in DNA sequences of co-expressed genes or in sequences from high-throughput methods such as ChIP-Seq.
-
Full-text index only
Coiled-coil protein composition of 22 proteomes--differences and common themes in subcellular infrastructure and traffic control.
PMID 16288662 · PMC1322226 · BMC evolutionary biology · 2005 · 7 claims · 5 setups
Proteins with extended coiled-coil domains (>250 amino acids) are largely absent from bacterial genomes but present in archaea and eukaryotes.
-
Has reproduction · 68
Mining the equine gut metagenome: poorly-characterized taxa associated with cardiovascular fitness in endurance athletes.
PMID 36192523 · PMC9529974 · Communications biology · 2022 · 8 claims · 8 setups
Built an integrated horse gut microbiome gene catalog (~25 million unique genes) and 372 metagenome-assembled genomes (MAGs) spanning 4179 genera and 95 phyla
-
Full-text index only
Ancient adaptive evolution of the primate antiviral DNA-editing enzyme APOBEC3G.
PMID 15269786 · PMC479043 · PLoS biology · 2004 · 7 claims · 6 setups
APOBEC3G has been under strong, recurrent positive selection throughout primate evolution (~33 million years)
-
Full-text index only
'Unknown' proteins and 'orphan' enzymes: the missing half of the engineering parts list--and how to find it.
PMID 20001958 · PMC3022307 · The Biochemical journal · 2009 · 8 claims · 8 setups
Comparative genomics is the single most effective strategy for predicting functions of unknown proteins and finding genes for orphan enzymes
-
Full-text index only
Tight junction-high and CDH17-positive cell population is the source of colorectal cancer liver metastases.
PMID 41484106 · PMC12881549 · Nature communications · 2026 · 8 claims · 8 setups
Loss/inhibition of IKKα unexpectedly promotes, rather than suppresses, CRC liver metastasis in patient-derived organoid (PDO) xenograft models.
-
Has reproduction · 95
A whole genome duplication drives the genome evolution of Phytophthora betacei, a closely related species to Phytophthora infestans.
PMID 34740326 · PMC8571832 · BMC genomics · 2021 · 8 claims · 7 setups
P. betacei P8084 has the largest sequenced genome in the Phytophthora genus (270 Mb)
-
Has reproduction · 67
Satellitome Analysis and Transposable Elements Comparison in Geographically Distant Populations of Spodoptera frugiperda.
PMID 35455012 · PMC9026859 · Life (Basel, Switzerland) · 2022 · 8 claims · 5 setups
Most transposable elements are commonly shared across all eight geographically distant S. frugiperda samples, except Maverick and PIF/Harbinger elements which show divergent repeat copies
-
Has reproduction · 61
TEMP: a computational method for analyzing transposable element polymorphism in populations.
PMID 24753423 · PMC4066757 · Nucleic acids research · 2014 · 8 claims · 8 setups
TEMP combines pair-end (discordant) read and split (soft-clipped) read information to identify both presence and absence of TE insertions in genomic DNA from heterogeneous/pooled samples.