Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
Large-scale trends in the evolution of gene structures within 11 animal genomes.
PMID 16518452 · PMC1386723 · PLoS computational biology · 2006 · 8 claims · 5 setups
Change in intron–exon gene structure is gradual, clock-like, and largely independent of coding-sequence (protein) evolution
-
Full-text index only
Vertebrate gene finding from multiple-species alignments using a two-level strategy.
PMID 16925840 · PMC1810555 · Genome biology · 2006 · 8 claims · 5 setups
DOGFISH cleanly separates a multi-species alignment classifier (RVM cascade) from an HMM-based structure predictor, avoiding tight coupling of alignment complexity with HMM formalism
-
Full-text index only
Structural evolution of the protein kinase-like superfamily.
PMID 16244704 · PMC1261164 · PLoS computational biology · 2005 · 8 claims · 5 setups
All kinases in the superfamily share a 'universal core' domain consisting only of the regions required for ATP binding and the phosphotransfer reaction.
-
Full-text index only
SNAP predicts effect of mutations on protein function.
PMID 18757876 · PMC2562009 · Bioinformatics (Oxford, England) · 2008 · 8 claims · 3 setups
SNAP is a publicly available web-server implementation predicting functional effects (neutral/non-neutral) of single amino acid substitutions.
-
Has reproduction · 80
Differential analysis of RNA structure probing experiments at nucleotide resolution: uncovering regulatory functions of RNA structure.
PMID 35869080 · PMC9307511 · Nature communications · 2022 · 7 claims · 4 setups
DiffScan is a computational framework combining a Normalization module and a Scan module to identify SVRs at nucleotide resolution from SP data.
-
Full-text index only
SVC: structured visualization of evolutionary sequence conservation.
PMID 15991338 · PMC1160265 · Nucleic acids research · 2005 · 7 claims · 5 setups
SVC aligns protein-coding sequences of orthologous gene pairs and maps them back onto their encoding exons/introns to generate a scaffold of conserved gene structure.
-
Full-text index only
The vertebrate genome annotation (Vega) database.
PMID 18003653 · PMC2238886 · Nucleic acids research · 2008 · 8 claims · 8 setups
Vega is a database for viewing manual genome annotation of human, mouse and zebrafish genomic sequences produced at the Wellcome Trust Sanger Institute.
-
Full-text index only
Genome sequences and great expectations.
PMID 11178275 · PMC150431 · Genome biology · 2001 · 8 claims · 3 setups
Function is known or can be predicted for an average of 62% of proteins across 31 analyzed genomes.
-
Full-text index only
The effect of diet on the human gut microbiome: a metagenomic analysis in humanized gnotobiotic mice.
PMID 20368178 · PMC2894525 · Science translational medicine · 2009 · 7 claims · 6 setups
Germ-free C57BL/6J mice can be stably and heritably colonized with a human fecal microbiota, reproducing much of the donor's bacterial diversity.
-
Full-text index only
EGASP: Introduction.
PMID 16925831 · PMC1810546 · Genome biology · 2006 · 8 claims · 5 setups
Computational gene finding methods, when compared to the GENCODE golden standard annotation, show that the human genome annotation is nearly complete in terms of novel protein-coding loci.
-
Full-text index only
Performance assessment of promoter predictions on ENCODE regions in the EGASP experiment.
PMID 16925837 · PMC1810552 · Genome biology · 2006 · 6 claims · 3 setups
Promoter predictors that combine promoter prediction with gene prediction (N-SCAN, Fprom) achieve better performance than pure ab initio promoter predictors, mainly by reducing the promoter search space and false positives
-
Has reproduction · 68
Rfam 15: RNA families database in 2025.
PMID 39526405 · PMC11701678 · Nucleic acids research · 2025 · 8 claims · 6 setups
Rfamseq was expanded to 26 106 genomes, a 76% increase, by incorporating the latest UniProt reference proteomes and additional viral genomes
-
Full-text index only
Human SNPs resulting in premature stop codons and protein truncation.
PMID 16595072 · PMC3500177 · Human genomics · 2006 · 8 claims · 6 setups
Genome-wide screening of dbSNP identified 28 validated X-SNPs from 28 genes with known minor allele frequencies.
-
Has reproduction · 79
Symbiosis genes show a unique pattern of introgression and selection within a Rhizobium leguminosarum species complex.
PMID 32176601 · PMC7276703 · Microbial genomics · 2020 · 8 claims · 8 setups
The 196 R. leguminosarum sv. trifolii strains constitute a five-species complex (genospecies gsA-gsE) that occur in sympatry but show little recent between-species gene transfer in core or accessory genomes, except for a few highly mobile regions.
-
Full-text index only
Systematic identification of pseudogenes through whole genome expression evidence profiling.
PMID 16945953 · PMC1636364 · Nucleic acids research · 2006 · 8 claims · 8 setups
Developed a novel bioinformatics method that identifies pseudogenes by profiling whole-genome transcript and protein expression evidence
-
Full-text index only
In silico analysis of missense substitutions using sequence-alignment based methods.
PMID 18951440 · PMC3431198 · Human mutation · 2008 · 8 claims · 7 setups
Carefully validated PMSA-based computational algorithms can achieve predictive values of ~75-95% for classifying missense substitutions as pathogenic or neutral.
-
Full-text index only
Distribution and effects of nonsense polymorphisms in human genes.
PMID 18852891 · PMC2561068 · PloS one · 2008 · 8 claims · 8 setups
Nonsense SNPs occur at a lower density than nonsynonymous SNPs, indicating stronger purifying selection against premature stop codons than amino acid changes.
-
Full-text index only
SelenoDB 1.0 : a database of selenoprotein genes, proteins and SECIS elements.
PMID 18174224 · PMC2238826 · Nucleic acids research · 2008 · 6 claims · 5 setups
Standard genome annotation pipelines misannotate selenoprotein genes because they rely on UGA as a universal stop codon, failing to recognize its dual role as the selenocysteine-recoding codon.
-
Full-text index only
AceView: a comprehensive cDNA-supported gene and transcripts annotation.
PMID 16925834 · PMC1810549 · Genome biology · 2006 · 8 claims · 4 setups
At the mRNA level, AceView transcripts are the closest match to Gencode transcripts among all evaluated methods, including alternative splice variants
-
Full-text index only
GENCODE: producing a reference annotation for ENCODE.
PMID 16925838 · PMC1810553 · Genome biology · 2006 · 8 claims · 8 setups
GENCODE annotation combines initial manual annotation by HAVANA, experimental validation, and refinement based on results to identify protein-coding genes in ENCODE regions