Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Has reproduction · 78
Identification of the Wheat (Triticum aestivum) IQD Gene Family and an Expression Analysis of Candidate Genes Associated with Seed Dormancy and Germination.
PMID 35456910 · PMC9025732 · International journal of molecular sciences · 2022 · 8 claims · 8 setups
73 IQD gene family members were identified in the wheat genome and classified into six phylogenetic groups
-
Has reproduction · 95
Determining virus-host interactions and glycerol metabolism profiles in geographically diverse solar salterns with metagenomics.
PMID 28097058 · PMC5228507 · PeerJ · 2017 · 8 claims · 8 setups
Similar virus-host interactions and glycerol metabolism gene associations (notably dihydroxyacetone kinase with Haloquadratum/Halorubrum) exist across geographically diverse solar salterns
-
Full-text index only
Towards a comprehensive structural coverage of completed genomes: a structural genomics viewpoint.
PMID 17349043 · PMC1829165 · BMC bioinformatics · 2007 · 8 claims · 6 setups
A combined target-selection approach — pursuing both structurally uncharacterised domain families and additional targets from large structurally characterised superfamilies — is essential for comprehensive structural coverage of the genomes.
-
Full-text index only
Applications for protein sequence-function evolution data: mRNA/protein expression analysis and coding SNP scoring tools.
PMID 16912992 · PMC1538848 · Nucleic acids research · 2006 · 7 claims · 8 setups
PANTHER HMMs built from family/subfamily multiple sequence alignments can classify novel protein sequences into functional groups based on statistically significant HMM match scores
-
Full-text index only
nf-core/proteinfamilies: a scalable pipeline for the generation of protein families.
PMID 41563008 · PMC12950615 · GigaScience · 2026 · 8 claims · 3 setups
nf-core/proteinfamilies is a scalable, parametrizable, open-source Nextflow pipeline that generates new protein families or assigns sequences to existing families using profile HMMs and MSAs
-
Full-text index only
Tracing the origin of functional and conserved domains in the human proteome: implications for protein evolution at the modular level.
PMID 17090320 · PMC1654190 · BMC evolutionary biology · 2006 · 8 claims · 5 setups
HHpred (HMM-HMM comparison) detects remote homologs in the human proteome with higher sensitivity than hmmpfam (HMMER), giving 10% more functional domain coverage and 20% higher residue coverage against Pfam-A families.
-
Full-text index only
pmid-41515103
PMID 41515103 · PMC12787715 · 8 claims · 8 setups
295 AP2/ERF transcription factor genes were identified and classified in the S. scabra genome
-
Full-text index only
DBD--taxonomically broad transcription factor predictions: new content and functionality.
PMID 18073188 · PMC2238844 · Nucleic acids research · 2008 · 8 claims · 3 setups
DBD is a database of predicted sequence-specific DNA-binding transcription factors covering over 700 publicly available proteomes, up from 150 in the initial version.
-
Full-text index only
Comprehensive genome analysis of 203 genomes provides structural genomics with new insights into protein family space.
PMID 16481312 · PMC1373602 · Nucleic acids research · 2006 · 8 claims · 7 setups
The number of protein families continues to expand steadily as more genomes are sequenced, showing no sign of saturation.
-
Full-text index only
A new family of giardial cysteine-rich non-VSP protein genes and a novel cyst protein.
PMID 17183673 · PMC1762436 · PloS one · 2006 · 8 claims · 7 setups
HCNCp is a novel invariant (non-variant) cyst protein belonging to a new family of high-cysteine membrane proteins (HCMp) abundant in the Giardia genome
-
Has reproduction · 93
A comparative study on recombination activity in cattle.
PMID 41942849 · PMC13067647 · Genetics, selection, evolution : GSE · 2026 · 8 claims · 8 setups
Genotype data with high systematic missingness across breeds and arrays can be streamlined and analysed with three complementary recombination-estimation approaches (HMM-based LINKPHASE3, deterministic hsphase, likelihood-based hsrecombi)
-
Full-text index only
Comparison of complete nuclear receptor sets from the human, Caenorhabditis elegans and Drosophila genomes.
PMID 11532213 · PMC55326 · Genome biology · 2001 · 7 claims · 5 setups
The human genome contains fewer than 50 functional nuclear receptors, far fewer than C. elegans and about twice as many as Drosophila
-
Has reproduction · 76
What the Phage: a scalable workflow for the identification and analysis of phage sequences.
PMID 36399058 · PMC9673492 · GigaScience · 2022 · 8 claims · 7 setups
WtP combines 11 tools (14 approaches) for phage prediction in a parallel, containerized Nextflow workflow
-
Full-text index only
NovelFam3000--uncharacterized human protein domains conserved across model organisms.
PMID 16533400 · PMC1440326 · BMC genomics · 2006 · 8 claims · 7 setups
NovelFam3000 is an online data centre unifying bioinformatics resource links, news, comments, and user-submitted experimental data (including a Gene Characterization Index) for ~3000 uncharacterized Pfam-B/DUF domain families conserved across worm, fly, and human
-
Full-text index only
L1Base: from functional annotation to prediction of active LINE-1 elements.
PMID 15608246 · PMC539998 · Nucleic acids research · 2005 · 7 claims · 6 setups
L1Base is a database of putatively active LINE-1 insertions in human, mouse and rat genomes, containing FLI-L1s (intact in both ORFs), ORF2-L1s (intact ORF2, disrupted ORF1), and FLnI-L1s (full-length, >6000 bp, non-intact)
-
Full-text index only
Local combinational variables: an approach used in DNA-binding helix-turn-helix motif prediction with sequence information.
PMID 19651875 · PMC2761287 · Nucleic acids research · 2009 · 8 claims · 7 setups
The LCV approach predicts HTH motifs with 93.29% accuracy, 93.93% sensitivity and 92.66% specificity using only primary sequence information