Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
Automatic discovery of cross-family sequence features associated with protein function.
PMID 16409628 · PMC1395344 · BMC bioinformatics · 2006 · 8 claims · 6 setups
A self-supervised data mining approach can find relationships between sequence features and functional annotations without preconceived functional categories.
-
Full-text index only
An SVD-based comparison of nine whole eukaryotic genomes supports a coelomate rather than ecdysozoan lineage.
PMID 15606920 · PMC544558 · BMC bioinformatics · 2004 · 8 claims · 7 setups
SVD-based analysis of tetrapeptide frequency vectors can compare whole eukaryotic proteomes without pre-defining orthologs or aligning homologous sites
-
Full-text index only
Coiled-coil protein composition of 22 proteomes--differences and common themes in subcellular infrastructure and traffic control.
PMID 16288662 · PMC1322226 · BMC evolutionary biology · 2005 · 7 claims · 5 setups
Proteins with extended coiled-coil domains (>250 amino acids) are largely absent from bacterial genomes but present in archaea and eukaryotes.
-
Full-text index only
Large-scale analysis of human alternative protein isoforms: pattern classification and correlation with subcellular localization signals.
PMID 15860772 · PMC1087780 · Nucleic acids research · 2005 · 8 claims · 8 setups
Constructed a large-scale dataset of 6876 human alternative protein isoforms from 2624 genes by combining H-Invitational full-length cDNA data and SwissProt VARSPLIC entries
-
Has reproduction · 87
Enhanced Generalizability of RNA Secondary Structure Prediction via Convolutional Block Attention Network and Ensemble Learning.
PMID 40871599 · PMC12388828 · Molecules (Basel, Switzerland) · 2025 · 8 claims · 8 setups
TrioFold integrates base-pairing clues from thermodynamic- and DL-based methods via ensemble learning and a convolutional block attention mechanism to enhance RSS prediction generalizability.
-
Full-text index only
Genome-wide identification of human functional DNA using a neutral indel model.
PMID 16410828 · PMC1326222 · PLoS computational biology · 2006 · 8 claims · 8 setups
A neutral indel model predicting a geometric distribution of intergap segment (IGS) lengths fits human-mouse ancestral repeat (AR) alignment data excellently
-
Full-text index only
Conservation, variability and the modeling of active protein kinases.
PMID 17912359 · PMC1989141 · PloS one · 2007 · 7 claims · 5 setups
A novel sequence-order independent (fold-independent) structural alignment algorithm was developed that maximizes side-chain similarity to produce a consensus kinase structure.
-
Full-text index only
Transduplication resulted in the incorporation of two protein-coding sequences into the turmoil-1 transposable element of C. elegans.
PMID 18842128 · PMC2572040 · Biology direct · 2008 · 8 claims · 6 setups
The Turmoil-1 transposable element in C. elegans incorporated two unrelated protein-coding sequences into its inverted terminal repeats (ITRs)
-
Full-text index only
Coverage of whole proteome by structural genomics observed through protein homology modeling database.
PMID 17146617 · PMC1769342 · Journal of structural and functional genomics · 2006 · 8 claims · 7 setups
FAMSBASE, a homology-modeling database of whole-genome ORFs, currently covers about 50% of predicted ORFs (368,724 of 734,193) across 276 genomes with modeled 3D structures.
-
Full-text index only
Low conservation and species-specific evolution of alternative splicing in humans and mice: comparative genomics analysis using well-annotated full-length cDNAs.
PMID 18838389 · PMC2582632 · Nucleic acids research · 2008 · 7 claims · 8 setups
Although 86% of individual human exons are conserved in the mouse genome, only a small fraction (431/20392, ~2%) of human AS variants are perfectly conserved AS variants in mice.
-
Full-text index only
The TIGR Gene Indices: clustering and assembling EST and known genes and integration with eukaryotic genomes.
PMID 15608288 · PMC540018 · Nucleic acids research · 2005 · 8 claims · 8 setups
The TIGR Gene Indices (TGI) are a collection of 77 species-specific databases that cluster and assemble EST and known gene sequences into tentative consensus (TC) sequences to identify and characterize expressed transcripts.
-
Full-text index only
Comparative gene finding in chicken indicates that we are closing in on the set of multi-exonic widely expressed human genes.
PMID 15809229 · PMC1074396 · Nucleic acids research · 2005 · 8 claims · 6 setups
Comparative gene finding (SGP2) between human and chicken, followed by RT-PCR verification, adds at most ~0.2% new genes to the multi-exonic human gene catalog
-
Full-text index only
Using several pair-wise informant sequences for de novo prediction of alternatively spliced transcripts.
PMID 16925842 · PMC1810557 · Genome biology · 2006 · 8 claims · 4 setups
MARS, an extension of the Twinscan algorithm, uses multiple pairwise informant genomes to predict human alternatively spliced transcripts de novo without expressed sequence information.
-
Full-text index only
MODBASE, a database of annotated comparative protein structure models and associated resources.
PMID 18948282 · PMC2686492 · Nucleic acids research · 2009 · 8 claims · 8 setups
MODBASE contains 5,152,695 reliable comparative protein structure models for 1,593,209 unique protein sequences.
-
Full-text index only
MODBASE: a database of annotated comparative protein structure models and associated resources.
PMID 16381869 · PMC1347422 · Nucleic acids research · 2006 · 8 claims · 7 setups
MODBASE is a database of automatically calculated comparative protein structure models covering all UniProt sequences matchable to a known structure
-
Full-text index only
OPTIC: orthologous and paralogous transcripts in clades.
PMID 17933761 · PMC2238935 · Nucleic acids research · 2008 · 6 claims · 7 setups
OPTIC is a database providing gene predictions and orthology assignments for three clades: amniotes (human, dog, mouse, opossum, platypus, chicken), 12 Drosophila species, and 4 Caenorhabditis nematodes.
-
Full-text index only
Genomic and bioinformatics analysis of human adenovirus type 37: new insights into corneal tropism.
PMID 18471294 · PMC2397415 · BMC genomics · 2008 · 7 claims · 7 setups
The complete genome of HAdV-37 was sequenced and annotated (35,213 bp, 56.6% GC content, 35 predicted coding sequences plus 8 hypothetical ORFs)
-
Full-text index only
Pegasys: software for executing and integrating analyses of biological sequences.
PMID 15096276 · PMC406494 · BMC bioinformatics · 2004 · 8 claims · 7 setups
Pegasys is a flexible, modular, customizable software system for executing and integrating heterogeneous biological sequence analysis tools
-
Full-text index only
The distribution of SNPs in human gene regulatory regions.
PMID 16209714 · PMC1260019 · BMC genomics · 2005 · 8 claims · 6 setups
SNPs occur with higher density closer to the transcriptional start site within gene promoter regions than in further upstream regions
-
Full-text index only
TRED: a Transcriptional Regulatory Element Database and a platform for in silico gene regulation studies.
PMID 15608156 · PMC539958 · Nucleic acids research · 2005 · 8 claims · 5 setups
TRED is a database collecting both cis-regulatory elements (promoters) and trans-regulatory elements (transcription factor binding/regulation data) with linked access.