Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
Steps toward broad-spectrum therapeutics: discovering virulence-associated genes present in diverse human pathogens.
PMID 19874620 · PMC2774872 · BMC genomics · 2009 · 8 claims · 8 setups
Phylogenetic profiling of protein clusters across pathogen and non-pathogen genomes can identify candidate generic virulence factors
-
Full-text index only
Identification of the proliferation/differentiation switch in the cellular network of multicellular organisms.
PMID 17166053 · PMC1664705 · PLoS computational biology · 2006 · 8 claims · 8 setups
Integrating interactome and transcriptome data reveals a pair of transcriptionally anticorrelated network modules (P and D) each comprising hundreds of genes, present across individuals and species.
-
Full-text index only
Prediction of missed cleavage sites in tryptic peptides aids protein identification in proteomics.
PMID 17203985 · PMC2664920 · Journal of proteome research · 2007 · 8 claims · 4 setups
An information-theoretic log-likelihood scoring method can predict experimentally observed missed cleavage sites from amino acid sequence alone with up to 90% accuracy.
-
Full-text index only
Functional classification using phylogenomic inference.
PMID 16846248 · PMC1484587 · PLoS computational biology · 2006 · 8 claims · 1 setups
Functional annotation via top-hit database search transfer is used far more often in practice than phylogenomic inference, despite phylogenomic inference being more accurate.
-
Has reproduction · 78
Machine learning and free energy clustering reveal PAH protein binding linked to AD risk.
PMID 41953002 · PMC13053772 · iScience · 2026 · 7 claims · 8 setups
An integrated framework of bioinformatics, machine learning, and ΔG clustering can prioritize PAHs for AD-associated neurotoxicity.
-
Full-text index only
How to find soluble proteins: a comprehensive analysis of alpha/beta hydrolases for recombinant expression in E. coli.
PMID 15804363 · PMC1079826 · BMC genomics · 2005 · 7 claims · 7 setups
Predicted solubility in E. coli (via CV-CV') depends on hydrolase size, phylogenetic origin, homologous family, and superfamily
-
Full-text index only
A systematic comparative and structural analysis of protein phosphorylation sites based on the mtcPTM database.
PMID 17521420 · PMC1929158 · Genome biology · 2007 · 7 claims · 6 setups
mtcPTM is a hierarchically organized database of human and mouse phosphosites that preserves experimental context, enabling comparison of phosphorylation patterns across conditions
-
Full-text index only
Molecular phylogeny of the kelch-repeat superfamily reveals an expansion of BTB/kelch proteins in animals.
PMID 13678422 · PMC222960 · BMC bioinformatics · 2003 · 8 claims · 8 setups
The human genome encodes at least 71 kelch-repeat proteins
-
Full-text index only
LMPD: LIPID MAPS proteome database.
PMID 16381922 · PMC1347484 · Nucleic acids research · 2006 · 8 claims · 5 setups
LMPD is an object-relational database of lipid-associated protein sequences and annotations, publicly available from the LIPID MAPS Consortium website.
-
Full-text index only
Sys-BodyFluid: a systematical database for human body fluid proteome research.
PMID 18978022 · PMC2686600 · Nucleic acids research · 2009 · 6 claims · 4 setups
Sys-BodyFluid is a web-based database integrating proteomic data from 11 human body fluids (plasma/serum, urine, cerebrospinal fluid, saliva, bronchoalveolar lavage fluid, synovial fluid, nipple aspirate fluid, tear fluid, seminal fluid, milk, amniotic fluid), containing over 10,000 proteins
-
Full-text index only
A survey of integral alpha-helical membrane proteins.
PMID 19760129 · PMC2780624 · Journal of structural and functional genomics · 2009 · 8 claims · 8 setups
An automated annotation pipeline defines the integral membrane genome and family associations for 21,379 proteins from 34 genomes, most belonging to 598 Pfam-derived membrane protein families.
-
Full-text index only
A rigorous method for multigenic families' functional annotation: the peptidyl arginine deiminase (PADs) proteins family example.
PMID 16271148 · PMC1310624 · BMC genomics · 2005 · 8 claims · 5 setups
Integrating EST-based expression data with phylogenetic analysis is a valid new method for functionally annotating multigenic protein families
-
Full-text index only
Pathway analysis for intracellular Porphyromonas gingivalis using a strain ATCC 33277 specific database.
PMID 19723305 · PMC2753363 · BMC microbiology · 2009 · 8 claims · 5 setups
Using the ATCC 33277-specific genome annotation improves proteome coverage (more proteins identified and more abundance ratios calculated) compared to the W83 annotation
-
Full-text index only
NCBI Reference Sequence (RefSeq): a curated non-redundant sequence database of genomes, transcripts and proteins.
PMID 15608248 · PMC539979 · Nucleic acids research · 2005 · 7 claims · 5 setups
RefSeq provides a curated, non-redundant, explicitly linked collection of genomic, transcript and protein sequences spanning prokaryotes, eukaryotes and viruses.
-
Full-text index only
ORFer--retrieval of protein sequences and open reading frames from GenBank and storage into relational databases or text files.
PMID 12493080 · PMC139979 · BMC bioinformatics · 2002 · 6 claims · 6 setups
ORFer retrieves protein and nucleic acid sequences and annotations from NCBI GenBank using the XML sequence format
-
Full-text index only
The 'permeome' of the malaria parasite: an overview of the membrane transport proteins of Plasmodium falciparum.
PMID 15774027 · PMC1088945 · Genome biology · 2005 · 7 claims · 4 setups
P. falciparum encodes substantially more membrane transport proteins than originally annotated
-
Has reproduction · 71
Protein structure quality assessment based on the distance profiles of consecutive backbone Cα atoms.
PMID 24555103 · PMC3892923 · F1000Research · 2013 · 8 claims · 8 setups
The distance between consecutive backbone Cα atoms in high-quality structures is normally distributed with mean 3.8 Å and standard deviation 0.04 Å, justifying a reference state in which all consecutive Cα atoms are 3.8 Å apart.
-
Full-text index only
The 10 sea urchin receptor for egg jelly proteins (SpREJ) are members of the polycystic kidney disease-1 (PKD1) family.
PMID 17629917 · PMC1934368 · BMC genomics · 2007 · 8 claims · 5 setups
Sea urchins possess 10 SpREJ (PKD1 family) genes, compared to five in humans, all defined by possession of a ~600 residue REJ domain
-
Full-text index only
Building disease-specific drug-protein connectivity maps from molecular interaction networks and PubMed abstracts.
PMID 19649302 · PMC2709445 · PLoS computational biology · 2009 · 7 claims · 4 setups
A computational framework can build disease-specific drug-protein connectivity maps by integrating protein interaction networks and PubMed literature mining, without gene expression profiles from drug perturbation experiments
-
Full-text index only
SGCEdb: a flexible database and web interface integrating experimental results and analysis for structural genomics focusing on Caenorhabditis elegans.
PMID 16381914 · PMC1347399 · Nucleic acids research · 2006 · 8 claims · 8 setups
SGCEdb is a flexible, reusable database and web interface for reporting and analyzing structural genomics experiment results, focused on C. elegans