Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
SUPERFAMILY--sophisticated comparative genomics, data mining, visualization and phylogeny.
PMID 19036790 · PMC2686452 · Nucleic acids research · 2009 · 7 claims · 6 setups
SUPERFAMILY provides structural, functional and evolutionary annotation for proteins from all completely sequenced genomes using SCOP-based hidden Markov models
-
Full-text index only
In silico segmentations of lentivirus envelope sequences.
PMID 17376229 · PMC1847453 · BMC bioinformatics · 2007 · 8 claims · 8 setups
C and V regions of lentivirus SU sequences have distinct statistical (oligonucleotide/amino-acid) compositions that HMMs can learn and use to delimit them.
-
Full-text index only
Tracing the origin of functional and conserved domains in the human proteome: implications for protein evolution at the modular level.
PMID 17090320 · PMC1654190 · BMC evolutionary biology · 2006 · 8 claims · 5 setups
HHpred (HMM-HMM comparison) detects remote homologs in the human proteome with higher sensitivity than hmmpfam (HMMER), giving 10% more functional domain coverage and 20% higher residue coverage against Pfam-A families.
-
Full-text index only
The protein-phosphatome of the human malaria parasite Plasmodium falciparum.
PMID 18793411 · PMC2559854 · BMC genomics · 2008 · 8 claims · 8 setups
P. falciparum possesses 27 putative protein phosphatase sequences across the four major PP families (PPP, PPM, PTP, NIF), plus 7 additional sequences predicted to dephosphorylate non-protein substrates, totaling 34.
-
Full-text index only
Applications for protein sequence-function evolution data: mRNA/protein expression analysis and coding SNP scoring tools.
PMID 16912992 · PMC1538848 · Nucleic acids research · 2006 · 7 claims · 8 setups
PANTHER HMMs built from family/subfamily multiple sequence alignments can classify novel protein sequences into functional groups based on statistically significant HMM match scores
-
Full-text index only
Towards a comprehensive structural coverage of completed genomes: a structural genomics viewpoint.
PMID 17349043 · PMC1829165 · BMC bioinformatics · 2007 · 8 claims · 6 setups
A combined target-selection approach — pursuing both structurally uncharacterised domain families and additional targets from large structurally characterised superfamilies — is essential for comprehensive structural coverage of the genomes.
-
Full-text index only
A computational screen for type I polyketide synthases in metagenomics shotgun data.
PMID 18953415 · PMC2568958 · PloS one · 2008 · 8 claims · 6 setups
Combining HMM domain searches with maximum-likelihood phylogenetic trees can discriminate true PKS I sequences from evolutionarily related but functionally different enzymes (e.g., FAS I) in metagenomic data.
-
Full-text index only
The MAPPER database: a multi-genome catalog of putative transcription factor binding sites.
PMID 15608292 · PMC540057 · Nucleic acids research · 2005 · 8 claims · 6 setups
Built a library of 1134 HMM models (359 matrix-derived, 718 factor-derived, 57 JASPAR-derived), corresponding to 863 distinct TF names, from TRANSFAC and JASPAR binding site data
-
Full-text index only
Dyneins across eukaryotes: a comparative genomic analysis.
PMID 17897317 · PMC2239267 · Traffic (Copenhagen, Denmark) · 2007 · 8 claims · 6 setups
Phylogenetic inference identified nine DHC families (two cytoplasmic, seven axonemal) and six IC families (one cytoplasmic)
-
Full-text index only
SysZNF: the C2H2 zinc finger gene database.
PMID 18974185 · PMC2686507 · Nucleic acids research · 2009 · 7 claims · 6 setups
SysZNF is a database that systematically catalogs C2H2-ZNF genes in human and mouse with physical location, gene models, expression probes, protein domains, homologs, and literature links
-
Full-text index only
Protein coding potential of retroviruses and other transposable elements in vertebrate genomes.
PMID 15716312 · PMC549403 · Nucleic acids research · 2005 · 8 claims · 5 setups
About 1000 genes across four vertebrate gene sets analyzed contain at least one RETRA marker protein domain
-
Full-text index only
Inventory and analysis of the protein subunits of the ribonucleases P and MRP provides further evidence of homology between the yeast and human enzymes.
PMID 16998185 · PMC1636426 · Nucleic acids research · 2006 · 8 claims · 6 setups
Fungal Pop8 is evolutionarily related to the Rpp14/Pop5 protein family, suggesting Pop8 is the fungal orthologue of Rpp14
-
Full-text index only
Comprehensive genome analysis of 203 genomes provides structural genomics with new insights into protein family space.
PMID 16481312 · PMC1373602 · Nucleic acids research · 2006 · 8 claims · 7 setups
The number of protein families continues to expand steadily as more genomes are sequenced, showing no sign of saturation.
-
Full-text index only
Reconstructing the evolution of the mitochondrial ribosomal proteome.
PMID 17604309 · PMC1950548 · Nucleic acids research · 2007 · 8 claims · 6 setups
The ancestral mitoribosome was of alpha-proteobacterial descent and more than doubled its protein content in most eukaryotic lineages.
-
Full-text index only
Local combinational variables: an approach used in DNA-binding helix-turn-helix motif prediction with sequence information.
PMID 19651875 · PMC2761287 · Nucleic acids research · 2009 · 8 claims · 7 setups
The LCV approach predicts HTH motifs with 93.29% accuracy, 93.93% sensitivity and 92.66% specificity using only primary sequence information
-
Has reproduction · 76
What the Phage: a scalable workflow for the identification and analysis of phage sequences.
PMID 36399058 · PMC9673492 · GigaScience · 2022 · 8 claims · 7 setups
WtP combines 11 tools (14 approaches) for phage prediction in a parallel, containerized Nextflow workflow
-
Full-text index only
Bioinformatic mapping of AlkB homology domains in viruses.
PMID 15627404 · PMC544882 · BMC genomics · 2005 · 8 claims · 8 setups
AlkB-like domains are found in at least 22 different single-stranded RNA positive-strand plant viruses, mainly within a subgroup of the Flexiviridae family.
-
Full-text index only
Comparison of complete nuclear receptor sets from the human, Caenorhabditis elegans and Drosophila genomes.
PMID 11532213 · PMC55326 · Genome biology · 2001 · 7 claims · 5 setups
The human genome contains fewer than 50 functional nuclear receptors, far fewer than C. elegans and about twice as many as Drosophila
-
Has reproduction · 67
binny: an automated binning algorithm to recover high-quality genomes from complex metagenomic datasets.
PMID 36239393 · PMC9677464 · Briefings in bioinformatics · 2022 · 8 claims · 8 setups
binny outperforms or is highly competitive with commonly used and state-of-the-art binning methods (MetaBAT2, MaxBin2, CONCOCT, VAMB, SemiBin, MetaDecoder)
-
Full-text index only
L1Base: from functional annotation to prediction of active LINE-1 elements.
PMID 15608246 · PMC539998 · Nucleic acids research · 2005 · 7 claims · 6 setups
L1Base is a database of putatively active LINE-1 insertions in human, mouse and rat genomes, containing FLI-L1s (intact in both ORFs), ORF2-L1s (intact ORF2, disrupted ORF1), and FLnI-L1s (full-length, >6000 bp, non-intact)