Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Has reproduction · 68
Mining the equine gut metagenome: poorly-characterized taxa associated with cardiovascular fitness in endurance athletes.
PMID 36192523 · PMC9529974 · Communications biology · 2022 · 8 claims · 8 setups
Built an integrated horse gut microbiome gene catalog (~25 million unique genes) and 372 metagenome-assembled genomes (MAGs) spanning 4179 genera and 95 phyla
-
Full-text index only
Automatic discovery of cross-family sequence features associated with protein function.
PMID 16409628 · PMC1395344 · BMC bioinformatics · 2006 · 8 claims · 6 setups
A self-supervised data mining approach can find relationships between sequence features and functional annotations without preconceived functional categories.
-
Full-text index only
Oligomeric protein structure networks: insights into protein-protein interactions.
PMID 16336694 · PMC1326230 · BMC bioinformatics · 2005 · 8 claims · 6 setups
Interface amino acid clusters identified at Imin=6% correlate well with residues losing accessible surface area (δASA) upon oligomerization
-
Full-text index only
GeneSeer: a sage for gene names and genomic resources.
PMID 16176584 · PMC1266031 · BMC genomics · 2005 · 7 claims · 4 setups
GeneSeer aggregates gene name synonyms from GenBank, FlyBase, ExPASy, HUGO, ENSEMBL, UCSC and Gene Ontology into a name-translation database that maps any familiar name to a reference (SOFAR) identifier.
-
Full-text index only
A compatible exon-exon junction database for the identification of exon skipping events using tandem mass spectrum data.
PMID 19087293 · PMC2636810 · BMC bioinformatics · 2008 · 6 claims · 6 setups
A theoretical exon-exon junction protein database accounting for all in-phase (frame-preserving) exon combinations can be built from the Ensembl Core Database using Perl/Bioperl/MySQL/Ensembl API.
-
Full-text index only
SysPIMP: the web-based systematical platform for identifying human disease-related mutated sequences from mass spectrometry.
PMID 19036792 · PMC2686442 · Nucleic acids research · 2009 · 8 claims · 7 setups
SysPIMP is a web-based platform integrating disease mutation databases with X!Tandem and BLAST to identify disease-related mutated proteins from MS results
-
Full-text index only
Molecular archeology of L1 insertions in the human genome.
PMID 12372140 · PMC134481 · Genome biology · 2002 · 8 claims · 4 setups
TSDfinder, a new algorithm, refines RepeatMasker-identified L1 boundaries by locating poly(A) tails, TSDs, and inversion breakpoints
-
Full-text index only
Shaken not stirred: a global research cocktail served in Hinxton.
PMID 18036269 · PMC2258181 · Genome biology · 2007 · 8 claims · 8 setups
Network-guided reverse genetics using probabilistic functional gene networks (e.g. YeastNet, WormNet) reduces the search space for identifying genes in a given biological process
-
Has reproduction · 98
Large-scale quality assessment of prokaryotic genomes with metashot/prok-quality.
PMID 35136576 · PMC8804904 · F1000Research · 2021 · 8 claims · 6 setups
metashot/prok-quality is a container-enabled Nextflow pipeline for quality assessment and dereplication of draft prokaryotic genomes
-
Full-text index only
Functional coverage of the human genome by existing structures, structural genomics targets, and homology models.
PMID 16118666 · PMC1188274 · PLoS computational biology · 2005 · 8 claims · 5 setups
Existing PDB structures provide single-domain coverage for 37% of functional classes in the human genome and complete (whole-protein) structure coverage for 25%.
-
Full-text index only
EPD in its twentieth year: towards complete promoter coverage of selected model organisms.
PMID 16381980 · PMC1347508 · Nucleic acids research · 2006 · 7 claims · 4 setups
EPD is an annotated, non-redundant collection of experimentally defined eukaryotic POL II promoters accessed via genome position pointers.
-
Full-text index only
SpliceMiner: a high-throughput database implementation of the NCBI Evidence Viewer for microarray splice variant analysis.
PMID 17338820 · PMC1839109 · BMC bioinformatics · 2007 · 6 claims · 4 setups
EVDB is a comprehensive, non-redundant relational database of known human splice variants built from NCBI Entrez Gene and Evidence Viewer data
-
Full-text index only
MODBASE: a database of annotated comparative protein structure models and associated resources.
PMID 16381869 · PMC1347422 · Nucleic acids research · 2006 · 8 claims · 7 setups
MODBASE is a database of automatically calculated comparative protein structure models covering all UniProt sequences matchable to a known structure
-
Full-text index only
Predicting deleterious nsSNPs: an analysis of sequence and structural attributes.
PMID 16630345 · PMC1489951 · BMC bioinformatics · 2006 · 8 claims · 7 setups
Sequence conservation (PSIC score difference) at the nsSNP position is the single most useful attribute for predicting deleterious vs neutral status.
-
Full-text index only
Random amino acid mutations and protein misfolding lead to Shannon limit in sequence-structure communication.
PMID 18769673 · PMC2518838 · PloS one · 2008 · 8 claims · 6 setups
The protein sequence-structure map behaves as a noisy digital communication channel whose capacity C exceeds the transmission rate R for native structures, satisfying Shannon's noisy channel theorem
-
Full-text index only
Analysis of expressed sequence tags from Actinidia: applications of a cross species EST database for gene discovery in the areas of flavor, health, color and ripening.
PMID 18655731 · PMC2515324 · BMC genomics · 2008 · 7 claims · 6 setups
A collection of 132,577 ESTs from four Actinidia species was generated and clustered into 41,858 non-redundant clusters (18,070 TCs and 23,788 singletons)
-
Full-text index only
An SVM-based system for predicting protein subnuclear localizations.
PMID 16336650 · PMC1325059 · BMC bioinformatics · 2005 · 7 claims · 3 setups
New kernels defined on k-peptide vectors mapped by BLOSUM62-based high-scored pair matrices (D1, D2, D3) improve SVM discrimination of protein subnuclear localization compared to conventional k-peptide encodings.
-
Has reproduction · 84
Network Controllability Reveals Key Mitigation Points for Tumor-Promoting Signaling in Tumor-Educated Platelets.
PMID 41226816 · PMC12609506 · International journal of molecular sciences · 2025 · 8 claims · 6 setups
TEPs in NSCLC show 111 upregulated and 108 downregulated genes versus non-cancer control platelets, enriched in ECM interaction, cytoskeleton, immune signaling, and platelet activation pathways.
-
Has reproduction · 24
MiGPC: a comprehensive catalog of enzybiotics from environmental metagenomes.
PMID 41888223 · PMC13172421 · Scientific reports · 2026 · 8 claims · 8 setups
MiGPC is the first genome-resolved metagenomic gene and protein catalog specifically targeted to enzybiotics
-
Full-text index only
LMPD: LIPID MAPS proteome database.
PMID 16381922 · PMC1347484 · Nucleic acids research · 2006 · 8 claims · 5 setups
LMPD is an object-relational database of lipid-associated protein sequences and annotations, publicly available from the LIPID MAPS Consortium website.