Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
POCUS: mining genomic sequence annotation to predict disease genes.
PMID 14611661 · PMC329128 · Genome biology · 2003 · 8 claims · 6 setups
Genes predisposing to the same disease tend to share functional annotation IDs (GO/InterPro) more than expected by chance
-
Full-text index only
Network of Cancer Genes: a web resource to analyze duplicability, orthology and network properties of cancer genes.
PMID 19906700 · PMC2808873 · Nucleic acids research · 2010 · 7 claims · 4 setups
NCG is a web database integrating duplicability, orthology, evolutionary appearance, and network topology data for 736 human cancer genes
-
Full-text index only
The MAPPER database: a multi-genome catalog of putative transcription factor binding sites.
PMID 15608292 · PMC540057 · Nucleic acids research · 2005 · 8 claims · 6 setups
Built a library of 1134 HMM models (359 matrix-derived, 718 factor-derived, 57 JASPAR-derived), corresponding to 863 distinct TF names, from TRANSFAC and JASPAR binding site data
-
Full-text index only
A parsimony approach to biological pathway reconstruction/inference for genomes and metagenomes.
PMID 19680427 · PMC2714467 · PLoS computational biology · 2009 · 8 claims · 6 setups
The naïve mapping approach (present if ≥1 associated function is found) leads to an inflated estimate of biological pathways and overestimates functional diversity of a sample.
-
Has reproduction · 100
Genomic insight into the influence of selection, crossbreeding, and geography on population structure in poultry.
PMID 36670351 · PMC9854048 · Genetics, selection, evolution : GSE · 2023 · 6 claims · 8 setups
Dutch traditional chicken breeds display a complex, admixed, subdivided population structure that broadly matches historical management-based clustering (past-productive, ornamental, country fowl, Lakenvelder).
-
Full-text index only
Integrating alternative splicing detection into gene prediction.
PMID 15705189 · PMC550657 · BMC bioinformatics · 2005 · 8 claims · 4 setups
An integrative intrinsic/extrinsic method was implemented in the gene finder EuGÈNE (as EuGÈNE-M) to detect AS evidence from aligned transcripts and generate alternative optimal gene predictions consistent with each detected AS event.
-
Full-text index only
MitoVariome: a variome database of human mitochondrial DNA.
PMID 19958475 · PMC2788364 · BMC genomics · 2009 · 8 claims · 5 setups
MitoVariome is a web-based, integrated variome database for human mitochondrial DNA that unifies sequence variation, haplogroup, and disease annotation information not jointly available in prior databases (MITOMAP, mtDB, Mitome, MitoRes).
-
Full-text index only
DAVID Knowledgebase: a gene-centered database integrating heterogeneous gene annotation resources to facilitate high-throughput gene functional analysis.
PMID 17980028 · PMC2186358 · BMC bioinformatics · 2007 · 7 claims · 3 setups
The DAVID Gene Concept, a single-linkage algorithm, merges gene clusters from Entrez Gene, UniRef100, and PIR-NREF100 that share protein IDs and species into unified DAVID gene clusters, improving cross-referencing between NCBI and UniProt systems
-
Has reproduction · 71
RNAmountAlign: Efficient software for local, global, semiglobal pairwise and multiple RNA sequence/structure alignment.
PMID 31978147 · PMC6980424 · PloS one · 2020 · 8 claims · 6 setups
RNAmountAlign is the first RNA sequence/structure pairwise alignment algorithm based on incremental ensemble mountain distance, running in O(n^3) time and O(n^2) space for two sequences of length n.
-
Has reproduction · 73
Vespucci: a system for building annotated databases of nascent transcripts.
PMID 24304890 · PMC3936758 · Nucleic acids research · 2014 · 8 claims · 7 setups
Existing ChIP-seq and RNA-seq analysis platforms (e.g. Cufflinks, peak callers) are unsuited to GRO-seq because they assume spliced/exonic reads, uniform density and paired-end data, and cannot identify transcriptional units de novo across the whole genome.