Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
Protein coding potential of retroviruses and other transposable elements in vertebrate genomes.
PMID 15716312 · PMC549403 · Nucleic acids research · 2005 · 8 claims · 5 setups
About 1000 genes across four vertebrate gene sets analyzed contain at least one RETRA marker protein domain
-
Full-text index only
Tracing the origin of functional and conserved domains in the human proteome: implications for protein evolution at the modular level.
PMID 17090320 · PMC1654190 · BMC evolutionary biology · 2006 · 8 claims · 5 setups
HHpred (HMM-HMM comparison) detects remote homologs in the human proteome with higher sensitivity than hmmpfam (HMMER), giving 10% more functional domain coverage and 20% higher residue coverage against Pfam-A families.
-
Full-text index only
Dyneins across eukaryotes: a comparative genomic analysis.
PMID 17897317 · PMC2239267 · Traffic (Copenhagen, Denmark) · 2007 · 8 claims · 6 setups
Phylogenetic inference identified nine DHC families (two cytoplasmic, seven axonemal) and six IC families (one cytoplasmic)
-
Full-text index only
In silico segmentations of lentivirus envelope sequences.
PMID 17376229 · PMC1847453 · BMC bioinformatics · 2007 · 8 claims · 8 setups
C and V regions of lentivirus SU sequences have distinct statistical (oligonucleotide/amino-acid) compositions that HMMs can learn and use to delimit them.
-
Full-text index only
Towards a comprehensive structural coverage of completed genomes: a structural genomics viewpoint.
PMID 17349043 · PMC1829165 · BMC bioinformatics · 2007 · 8 claims · 6 setups
A combined target-selection approach — pursuing both structurally uncharacterised domain families and additional targets from large structurally characterised superfamilies — is essential for comprehensive structural coverage of the genomes.
-
Full-text index only
SysZNF: the C2H2 zinc finger gene database.
PMID 18974185 · PMC2686507 · Nucleic acids research · 2009 · 7 claims · 6 setups
SysZNF is a database that systematically catalogs C2H2-ZNF genes in human and mouse with physical location, gene models, expression probes, protein domains, homologs, and literature links
-
Full-text index only
A computational screen for type I polyketide synthases in metagenomics shotgun data.
PMID 18953415 · PMC2568958 · PloS one · 2008 · 8 claims · 6 setups
Combining HMM domain searches with maximum-likelihood phylogenetic trees can discriminate true PKS I sequences from evolutionarily related but functionally different enzymes (e.g., FAS I) in metagenomic data.
-
Full-text index only
The protein-phosphatome of the human malaria parasite Plasmodium falciparum.
PMID 18793411 · PMC2559854 · BMC genomics · 2008 · 8 claims · 8 setups
P. falciparum possesses 27 putative protein phosphatase sequences across the four major PP families (PPP, PPM, PTP, NIF), plus 7 additional sequences predicted to dephosphorylate non-protein substrates, totaling 34.
-
Full-text index only
DBD--taxonomically broad transcription factor predictions: new content and functionality.
PMID 18073188 · PMC2238844 · Nucleic acids research · 2008 · 8 claims · 3 setups
DBD is a database of predicted sequence-specific DNA-binding transcription factors covering over 700 publicly available proteomes, up from 150 in the initial version.
-
Full-text index only
Modeling genetic inheritance of copy number variations.
PMID 18832372 · PMC2588508 · Nucleic acids research · 2008 · 8 claims · 4 setups
A joint HMM framework for parents-offspring trios significantly improves CNV call rates and boundary inference accuracy compared to existing methods.
-
Has reproduction · 75
Graph-Based Approaches Significantly Improve the Recovery of Antibiotic Resistance Genes From Complex Metagenomic Datasets.
PMID 34690959 · PMC8528159 · Frontiers in microbiology · 2021 · 8 claims · 6 setups
GraphAMR, a Nextflow pipeline that aligns AMR profile HMMs (or AA sequences) to metagenomic assembly graphs via PathRacer, then dereplicates and annotates hits, recovers more and more complete AMR genes than contig-based or read-based methods.
-
Full-text index only
The MAPPER database: a multi-genome catalog of putative transcription factor binding sites.
PMID 15608292 · PMC540057 · Nucleic acids research · 2005 · 8 claims · 6 setups
Built a library of 1134 HMM models (359 matrix-derived, 718 factor-derived, 57 JASPAR-derived), corresponding to 863 distinct TF names, from TRANSFAC and JASPAR binding site data
-
Full-text index only
A Hidden Markov Model to estimate population mixture and allelic copy-numbers in cancers using Affymetrix SNP arrays.
PMID 17996079 · PMC2206057 · BMC bioinformatics · 2007 · 8 claims · 7 setups
An HMM using paired germline genotype calls and tumour allelic SNP intensities can estimate allele-specific copy-numbers, distinguishing events like uniparental disomy from allelic imbalance.
-
Full-text index only
Applications for protein sequence-function evolution data: mRNA/protein expression analysis and coding SNP scoring tools.
PMID 16912992 · PMC1538848 · Nucleic acids research · 2006 · 7 claims · 8 setups
PANTHER HMMs built from family/subfamily multiple sequence alignments can classify novel protein sequences into functional groups based on statistically significant HMM match scores
-
Full-text index only
A response to Yu et al. "A forward-backward fragment assembling algorithm for the identification of genomic amplification and deletion breakpoints using high-density single nucleotide polymorphism (SNP) array", BMC Bioinformatics 2007, 8: 145.
PMID 17939873 · PMC2222656 · BMC bioinformatics · 2007 · 8 claims · 4 setups
Yu et al.'s original comparison ran RJaCGH's MCMC sampler for a severely insufficient number of iterations (50 burn-in, 500 total)
-
Full-text index only
QuantiSNP: an Objective Bayes Hidden-Markov Model to detect and accurately map copy number variation using SNP genotyping data.
PMID 17341461 · PMC1874617 · Nucleic acids research · 2007 · 8 claims · 7 setups
QuantiSNP (OB-HMM) provides probabilistic quantification of copy number states and significantly improves accuracy of segmental aneuploidy identification and breakpoint mapping relative to existing tools (BeadStudio/Illumina)
-
Full-text index only
GeneMark: web software for gene finding in prokaryotes, eukaryotes and viruses.
PMID 15980510 · PMC1160247 · Nucleic acids research · 2005 · 8 claims · 2 setups
The GeneMark website provides web interfaces to the GeneMark family of ab initio gene-finding programs for prokaryotic, eukaryotic and viral genomic sequences
-
Full-text index only
The EH1 motif in metazoan transcription factors.
PMID 16309560 · PMC1310626 · BMC genomics · 2005 · 8 claims · 5 setups
There is a statistically significant association between EH1hox motif HMM score and transcription factor function across human, Drosophila and C. elegans proteomes.
-
Full-text index only
ADaCGH: A parallelized web-based application and R package for the analysis of aCGH data.
PMID 17710137 · PMC1940324 · PloS one · 2007 · 8 claims · 4 setups
ADaCGH implements eight CNA detection methods, including the best-performing ones from recent reviews (CBS, GLAD, CGHseg, HMM)
-
Has reproduction · 89
Evolution of Highly Repetitive Silk Genes in the Luna Moth, Actias luna.
PMID 41738778 · PMC12962854 · Genome biology and evolution · 2026 · 8 claims · 8 setups
Eight sericin genes were identified in the Actias luna genome, including two clusters of closely related paralogs.