Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
Molecular evolution of Cide family proteins: novel domain formation in early vertebrates and the subsequent divergence.
PMID 18500987 · PMC2426694 · BMC evolutionary biology · 2008 · 8 claims · 5 setups
Sequences homologous to the CIDE-N domain/NCD show a wide phylogenetic distribution, from hydra and sea anemone to mammals, while true Cide proteins are restricted to vertebrates.
-
Has reproduction · 42
The electrostatic profile of consecutive Cβ atoms applied to protein structure quality assessment.
PMID 25506420 · PMC4257144 · F1000Research · 2013 · 8 claims · 8 setups
The EPD between Cβ atoms of consecutive residues provides unique signatures of amino acid pair types and can discriminate native from decoy protein structures.
-
Full-text index only
How to find soluble proteins: a comprehensive analysis of alpha/beta hydrolases for recombinant expression in E. coli.
PMID 15804363 · PMC1079826 · BMC genomics · 2005 · 7 claims · 7 setups
Predicted solubility in E. coli (via CV-CV') depends on hydrolase size, phylogenetic origin, homologous family, and superfamily
-
Full-text index only
A survey of integral alpha-helical membrane proteins.
PMID 19760129 · PMC2780624 · Journal of structural and functional genomics · 2009 · 8 claims · 8 setups
An automated annotation pipeline defines the integral membrane genome and family associations for 21,379 proteins from 34 genomes, most belonging to 598 Pfam-derived membrane protein families.
-
Full-text index only
GeneSeer: a sage for gene names and genomic resources.
PMID 16176584 · PMC1266031 · BMC genomics · 2005 · 7 claims · 4 setups
GeneSeer aggregates gene name synonyms from GenBank, FlyBase, ExPASy, HUGO, ENSEMBL, UCSC and Gene Ontology into a name-translation database that maps any familiar name to a reference (SOFAR) identifier.
-
Full-text index only
TreeFam: a curated database of phylogenetic trees of animal gene families.
PMID 16381935 · PMC1347480 · Nucleic acids research · 2006 · 7 claims · 6 setups
Tree-based inference of orthologs and paralogs is more robust than BLAST-based methods because evolutionary rates (and thus pairwise BLAST scores) vary across gene family members
-
Full-text index only
EPGD: a comprehensive web resource for integrating and displaying eukaryotic paralog/paralogon information.
PMID 17984073 · PMC2238967 · Nucleic acids research · 2008 · 8 claims · 8 setups
EPGD is a gene-centered, internet-accessible database integrating paralog family and paralogon information for 26 eukaryotic genomes.
-
Full-text index only
Filtering high-throughput protein-protein interaction data using a combination of genomic features.
PMID 15833142 · PMC1127019 · BMC bioinformatics · 2005 · 8 claims · 8 setups
A combination of three genomic features (interacting Pfam domains, GO annotations, sequence homology) using naive Bayesian networks predicts true protein-protein interactions with high sensitivity and good specificity.
-
Full-text index only
Comparative gene finding in chicken indicates that we are closing in on the set of multi-exonic widely expressed human genes.
PMID 15809229 · PMC1074396 · Nucleic acids research · 2005 · 8 claims · 6 setups
Comparative gene finding (SGP2) between human and chicken, followed by RT-PCR verification, adds at most ~0.2% new genes to the multi-exonic human gene catalog
-
Has reproduction · 71
Protein structure quality assessment based on the distance profiles of consecutive backbone Cα atoms.
PMID 24555103 · PMC3892923 · F1000Research · 2013 · 8 claims · 8 setups
The distance between consecutive backbone Cα atoms in high-quality structures is normally distributed with mean 3.8 Å and standard deviation 0.04 Å, justifying a reference state in which all consecutive Cα atoms are 3.8 Å apart.
-
Full-text index only
Automatic annotation of eukaryotic genes, pseudogenes and promoters.
PMID 16925832 · PMC1810547 · Genome biology · 2006 · 8 claims · 6 setups
Fgenesh++ gene prediction pipeline identifies 91% of coding nucleotides with 90% specificity
-
Full-text index only
SNP-VISTA: an interactive SNP visualization tool.
PMID 16336665 · PMC1325058 · BMC bioinformatics · 2005 · 7 claims · 3 setups
SNP-VISTA is an interactive Java-based visualization tool with two versions, GeneSNP-VISTA and EcoSNP-VISTA, for exploring large-scale SNP datasets
-
Has reproduction · 80
Transcriptome-Proteome Profiling in Burkholderia thailandensis during the Transition from Exponential to Stationary Phase.
PMID 40680064 · PMC12322963 · Journal of proteome research · 2025 · 8 claims · 7 setups
928 differentially accumulating mRNAs (564 up, 364 down) were identified between exponential and stationary phase
-
Full-text index only
Retroposition and evolution of the DNA-binding motifs of YY1, YY2 and REX1.
PMID 17478514 · PMC1904287 · Nucleic acids research · 2007 · 8 claims · 5 setups
62 YY1-related sequences were identified across genomes ranging from flying insects to humans, with high zinc finger domain conservation
-
Full-text index only
Comparative phosphoproteomics reveals evolutionary and functional conservation of phosphorylation across eukaryotes.
PMID 18828897 · PMC2760871 · Genome biology · 2008 · 8 claims · 8 setups
The overlap between phosphoproteomes of six eukaryotes (human, mouse, fly, yeast, plant, zebrafish) is significantly greater than expected by chance.
-
Has reproduction · 24
MiGPC: a comprehensive catalog of enzybiotics from environmental metagenomes.
PMID 41888223 · PMC13172421 · Scientific reports · 2026 · 8 claims · 8 setups
MiGPC is the first genome-resolved metagenomic gene and protein catalog specifically targeted to enzybiotics
-
Full-text index only
Upgrades to StellaBase facilitate medical and genetic studies on the starlet sea anemone, Nematostella vectensis.
PMID 17982171 · PMC2238866 · Nucleic acids research · 2008 · 6 claims · 5 setups
StellaBase Disease houses homology data for 155,904 invertebrate isoforms of human disease genes across four model systems, including 14,874 predicted Nematostella genes