Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
Structural evolution of the protein kinase-like superfamily.
PMID 16244704 · PMC1261164 · PLoS computational biology · 2005 · 8 claims · 5 setups
All kinases in the superfamily share a 'universal core' domain consisting only of the regions required for ATP binding and the phosphotransfer reaction.
-
Full-text index only
Interaction profile-based protein classification of death domain.
PMID 15189571 · PMC459208 · BMC bioinformatics · 2004 · 7 claims · 6 setups
An SVM-based classifier using Residue Pair Interaction Profiles (RPIPs) can classify death domain superfamily members into subfamilies with 89% average cross-validation accuracy
-
Full-text index only
Protein under-wrapping causes dosage sensitivity and decreases gene duplicability.
PMID 18208334 · PMC2211539 · PLoS genetics · 2008 · 7 claims · 6 setups
Protein under-wrapping extent is negatively correlated with gene duplicability (family size) across six organisms (E. coli, yeast, worm, fly, human, thale cress)
-
Full-text index only
Comprehensive genome analysis of 203 genomes provides structural genomics with new insights into protein family space.
PMID 16481312 · PMC1373602 · Nucleic acids research · 2006 · 8 claims · 7 setups
The number of protein families continues to expand steadily as more genomes are sequenced, showing no sign of saturation.
-
Full-text index only
Genome-wide survey of allele-specific splicing in humans.
PMID 18518984 · PMC2427040 · BMC genomics · 2008 · 8 claims · 5 setups
A genome-wide computational scan identified 30,977 SNPs located within predicted splicing regulatory sequences (donor sites, acceptor sites, branch points, and ESEs)
-
Has reproduction · 42
The electrostatic profile of consecutive Cβ atoms applied to protein structure quality assessment.
PMID 25506420 · PMC4257144 · F1000Research · 2013 · 8 claims · 8 setups
The EPD between Cβ atoms of consecutive residues provides unique signatures of amino acid pair types and can discriminate native from decoy protein structures.
-
Full-text index only
TPRpred: a tool for prediction of TPR-, PPR- and SEL1-like repeats from protein sequences.
PMID 17199898 · PMC1774580 · BMC bioinformatics · 2007 · 7 claims · 8 setups
TPRpred detects divergent/remote-homolog TPR repeat units that existing resources (Pfam, SMART, REP) fail to detect
-
Full-text index only
MODBASE, a database of annotated comparative protein structure models and associated resources.
PMID 18948282 · PMC2686492 · Nucleic acids research · 2009 · 8 claims · 8 setups
MODBASE contains 5,152,695 reliable comparative protein structure models for 1,593,209 unique protein sequences.
-
Full-text index only
A survey of integral alpha-helical membrane proteins.
PMID 19760129 · PMC2780624 · Journal of structural and functional genomics · 2009 · 8 claims · 8 setups
An automated annotation pipeline defines the integral membrane genome and family associations for 21,379 proteins from 34 genomes, most belonging to 598 Pfam-derived membrane protein families.
-
Has reproduction · 68
Rfam 15: RNA families database in 2025.
PMID 39526405 · PMC11701678 · Nucleic acids research · 2025 · 8 claims · 6 setups
Rfamseq was expanded to 26,106 genomes, a 76% increase over Rfam 14.0, incorporating UniProt 2024_03 reference proteomes and additional viral genomes.
-
Has reproduction · 87
Enhanced Generalizability of RNA Secondary Structure Prediction via Convolutional Block Attention Network and Ensemble Learning.
PMID 40871599 · PMC12388828 · Molecules (Basel, Switzerland) · 2025 · 8 claims · 8 setups
TrioFold integrates base-pairing clues from thermodynamic- and DL-based methods via ensemble learning and a convolutional block attention mechanism to enhance RSS prediction generalizability.
-
Full-text index only
Mining the draft human genome.
PMID 11236999 · PMC2658632 · Nature · 2001 · 8 claims · 7 setups
Protein-coding exons account for only about 3% of the human genome DNA, with repeat sequences making up around 46%.
-
Full-text index only
Genome bioinformatic analysis of nonsynonymous SNPs.
PMID 17708757 · PMC1978506 · BMC bioinformatics · 2007 · 8 claims · 8 setups
Structure- and sequence-based prediction tools can generally distinguish disease-causing mutations from neutral ones
-
Full-text index only
The DNA sequence and analysis of human chromosome 13.
PMID 15057823 · PMC2665288 · Nature · 2004 · 8 claims · 8 setups
95.5 Mb of finished sequence from chromosome 13 was completed, containing 633 genes and 296 pseudogenes.
-
Full-text index only
Crystallin gene mutations in Indian families with inherited pediatric cataract.
PMID 18587492 · PMC2435160 · Molecular vision · 2008 · 8 claims · 5 setups
Crystallin gene mutations account for 16.6% of inherited pediatric cataract in this south Indian population
-
Full-text index only
MODBASE: a database of annotated comparative protein structure models and associated resources.
PMID 16381869 · PMC1347422 · Nucleic acids research · 2006 · 8 claims · 7 setups
MODBASE is a database of automatically calculated comparative protein structure models covering all UniProt sequences matchable to a known structure
-
Full-text index only
Web-based resources for comparative genomics.
PMID 16197736 · PMC3525128 · Human genomics · 2005 · 8 claims · 8 setups
Comparative genomics is an indispensable tool for identifying functional genome elements and exploring evolutionary genome dynamics