Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
POCUS: mining genomic sequence annotation to predict disease genes.
PMID 14611661 · PMC329128 · Genome biology · 2003 · 8 claims · 6 setups
Genes predisposing to the same disease tend to share functional annotation IDs (GO/InterPro) more than expected by chance
-
Full-text index only
SNP@Evolution: a hierarchical database of positive selection on the human genome.
PMID 19732458 · PMC2755008 · BMC evolutionary biology · 2009 · 7 claims · 6 setups
SNP@Evolution is a hierarchical database integrating HET, FST, and iHS from HapMap Phase II and III to identify genome-wide positive selection signals
-
Full-text index only
Array-based profiling of reference-independent methylation status (aPRIMES) identifies frequent promoter methylation and consecutive downregulation of ZIC2 in pediatric medulloblastoma.
PMID 17344319 · PMC1874664 · Nucleic acids research · 2007 · 7 claims · 7 setups
aPRIMES is a novel array-based method that detects direct (absolute) methylation status of CGIs via competitive hybridization of McrBC-digested (methylated) versus HpaII/BstUI-digested (unmethylated) DNA from the same genome, avoiding reference-tissue and copy-number biases
-
Full-text index only
Speeding disease gene discovery by sequence based candidate prioritization.
PMID 15766383 · PMC1274252 · BMC bioinformatics · 2005 · 7 claims · 8 setups
Disease genes (OMIM) differ significantly from non-disease genes in sequence-based features including gene/cDNA/protein size, exon number, homolog conservation, secretion signal, 3' UTR length, CpG islands, and distance to nearest gene.
-
Full-text index only
Comprehensive genome analysis of 203 genomes provides structural genomics with new insights into protein family space.
PMID 16481312 · PMC1373602 · Nucleic acids research · 2006 · 8 claims · 7 setups
The number of protein families continues to expand steadily as more genomes are sequenced, showing no sign of saturation.
-
Full-text index only
The UCSC Genome Browser Database: update 2009.
PMID 18996895 · PMC2686463 · Nucleic acids research · 2009 · 8 claims · 6 setups
The UCSC Genome Browser Database (GBD) is a publicly available, integrated collection of genome assembly sequences and annotations across many organisms, including extensive comparative-genomic resources.
-
Has reproduction
Systematic analysis of CNGCs in cotton and the positive role of GhCNGC32 and GhCNGC35 in salt tolerance.
PMID 35931984 · PMC9356423 · BMC genomics · 2022 · 8 claims · 8 setups
114 CNGC genes were identified across the genomes of four cotton species (G. arboreum, G. raimondii, G. barbadense, G. hirsutum)
-
Full-text index only
Prediction of candidate primary immunodeficiency disease genes using a support vector machine learning approach.
PMID 19801557 · PMC2780952 · DNA research : an international journal for rapid publication of reports on genes and genomes · 2009 · 6 claims · 3 setups
An SVM trained on 69 binary features of known PID genes can accurately classify PID vs non-PID genes and predict novel candidate PID genes
-
Full-text index only
Four genomic islands that mark post-1995 pandemic Vibrio parahaemolyticus isolates.
PMID 16672049 · PMC1464126 · BMC genomics · 2006 · 8 claims · 7 setups
Seven genomic islands (VPaI-1 to VPaI-7, 10-81 kb) were identified in V. parahaemolyticus RIMD2210633 by aberrant GC content, presence of integrases/transposases, flanking direct repeats, and absence from related Vibrionaceae genomes.
-
Full-text index only
Assessment of algorithms for high throughput detection of genomic copy number variation in oligonucleotide microarray data.
PMID 17910767 · PMC2148068 · BMC bioinformatics · 2007 · 8 claims · 4 setups
Different CNV analysis software packages produce highly variable numbers and types of candidate CNVs from the same data
-
Has reproduction
Methylation patterns of the nasal epigenome of hospitalized SARS-CoV-2 positive patients reveal insights into molecular mechanisms of COVID-19.
PMID 40170038 · PMC11963311 · BMC medical genomics · 2025 · 7 claims · 7 setups
Differential DNA methylation occurs predominantly in intergenic regions and low methylated regions (LMRs), highlighting the role of distal regulatory elements in COVID-19 severity.
-
Has reproduction · 89
TrEMOLO: accurate transposable element allele frequency estimation using long-read sequencing data combining assembly and mapping-based approaches.
PMID 37013657 · PMC10069131 · Genome biology · 2023 · 6 claims · 6 setups
TrEMOLO combines an assembly-based INSIDER module and a mapping-based OUTSIDER module to detect TE insertions/deletions from long-read sequencing data and estimate their allele frequency
-
Full-text index only
Commonality of functional annotation: a method for prioritization of candidate genes from genome-wide linkage studies.
PMID 18263617 · PMC2275105 · Nucleic acids research · 2008 · 8 claims · 7 setups
Genes correlated with a common complex trait are more likely to share GO functional annotations than genes not correlated with that trait
-
Full-text index only
Automatic annotation of eukaryotic genes, pseudogenes and promoters.
PMID 16925832 · PMC1810547 · Genome biology · 2006 · 8 claims · 6 setups
Fgenesh++ gene prediction pipeline identifies 91% of coding nucleotides with 90% specificity
-
Full-text index only
Genetic linkage study of high-grade myopia in a Hutterite population from South Dakota.
PMID 17327828 · PMC2633468 · Molecular vision · 2007 · 6 claims · 5 setups
AD non-syndromic high-grade myopia in the Hutterite family MYO-101 shows significant linkage to a locus on chromosome 10q21.1
-
Full-text index only
Clustering of phosphorylation site recognition motifs can be exploited to predict the targets of cyclin-dependent kinase.
PMID 17316440 · PMC1852407 · Genome biology · 2007 · 8 claims · 6 setups
CDK consensus motifs are frequently clustered (closely spaced) in known CDK substrate proteins rather than uniformly distributed
-
Full-text index only
Genomics and biology come together to fight HIV.
PMID 18366259 · PMC2270331 · PLoS biology · 2008 · 8 claims · 6 setups
Genome-wide association studies have identified ~100 genetic polymorphisms robustly (genome-wide significant) linked to common human traits/diseases, but the biological mechanisms behind most remain unknown.
-
Full-text index only
Characterizing natural variation using next-generation sequencing technologies.
PMID 19801172 · PMC3994700 · Trends in genetics : TIG · 2009 · 8 claims · 8 setups
Next-generation sequencing enables complete, genome-wide surveys of genetic variation at unprecedented resolution, overcoming limitations of genotyping panels and microarrays.
-
Full-text index only
Single-molecule sequencing of an individual human genome.
PMID 19668243 · PMC4117198 · Nature biotechnology · 2009 · 8 claims · 7 setups
Single-molecule sequencing without cloning, amplification or ligation can sequence an individual human genome on one instrument by a single operator in four runs
-
Full-text index only
Performance assessment of promoter predictions on ENCODE regions in the EGASP experiment.
PMID 16925837 · PMC1810552 · Genome biology · 2006 · 6 claims · 3 setups
Promoter predictors that combine promoter prediction with gene prediction (N-SCAN, Fprom) achieve better performance than pure ab initio promoter predictors, mainly by reducing the promoter search space and false positives