Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
PeroxisomeDB: a database for the peroxisomal proteome, functional genomics and disease.
PMID 17135190 · PMC1747181 · Nucleic acids research · 2007 · 8 claims · 6 setups
PeroxisomeDB integrates the complete peroxisomal proteome of Homo sapiens and Saccharomyces cerevisiae into interrelated 'Genes', 'Functions', 'Metabolic pathways' and 'Diseases' sections with links to NCBI, ENSEMBL and UCSC
-
Full-text index only
Large-scale trends in the evolution of gene structures within 11 animal genomes.
PMID 16518452 · PMC1386723 · PLoS computational biology · 2006 · 8 claims · 5 setups
Change in intron–exon gene structure is gradual, clock-like, and largely independent of coding-sequence (protein) evolution
-
Full-text index only
Recent additions and improvements to the Onto-Tools.
PMID 15980579 · PMC1160233 · Nucleic acids research · 2005 · 7 claims · 3 setups
The Onto-Tools back-end database was redesigned around the Entrez Gene data model after NCBI phased out LocusLink in February 2005.
-
Full-text index only
L1Base: from functional annotation to prediction of active LINE-1 elements.
PMID 15608246 · PMC539998 · Nucleic acids research · 2005 · 7 claims · 6 setups
L1Base is a database of putatively active LINE-1 insertions in human, mouse and rat genomes, containing FLI-L1s (intact in both ORFs), ORF2-L1s (intact ORF2, disrupted ORF1), and FLnI-L1s (full-length, >6000 bp, non-intact)
-
Full-text index only
TassDB: a database of alternative tandem splice sites.
PMID 17142241 · PMC1669710 · Nucleic acids research · 2007 · 7 claims · 3 setups
TassDB is a relational database storing GYNGYN donor and NAGNAG acceptor tandem splice sites across eight species
-
Full-text index only
pTARGET: a web server for predicting protein subcellular localization.
PMID 16844995 · PMC1538910 · Nucleic acids research · 2006 · 7 claims · 3 setups
pTARGET web server predicts nine distinct subcellular localizations in eukaryotic non-plant proteins using an algorithm based on location-specific Pfam domain occurrence patterns and amino acid composition (AAC)
-
Full-text index only
DDBJ in collaboration with mass-sequencing teams on annotation.
PMID 15608189 · PMC539974 · Nucleic acids research · 2005 · 7 claims · 5 setups
DDBJ collected and released 1,066,084 entries (718,072,425 bases) in the past year, including the complete chimpanzee chromosome 22 sequence and silkworm whole-genome shotgun data
-
Has reproduction · 78
annotate_my_genomes: an easy-to-use pipeline to improve genome annotation and uncover neglected genes by hybrid RNA sequencing.
PMID 36472574 · PMC9724561 · GigaScience · 2022 · 7 claims · 8 setups
annotate_my_genomes is an easy-to-use genome-guided pipeline that uses hybrid (PacBio+Illumina) assembled transcripts to distinguish coding genes from long non-coding RNAs and reconcile them with prior annotations.
-
Full-text index only
A map of human protein interactions derived from co-expression of human mRNAs and their orthologs.
PMID 18414481 · PMC2387231 · Molecular systems biology · 2008 · 8 claims · 6 setups
Comparing human mRNA co-expression with co-expression of orthologous gene pairs in five other organisms identifies proteins that physically associate
-
Full-text index only
The truth about mouse, human, worms and yeast.
PMID 15601543 · PMC3525071 · Human genomics · 2004 · 8 claims · 8 setups
Comparing genomes in pairs or larger sets (mouse-human, C. elegans-C. briggsae, multiple Saccharomyces, human-pufferfish, etc.) reveals unsuspected genes and helps eliminate false-positive gene predictions
-
Full-text index only
Improved reconstruction of transcripts and coding sequences from RNA-seq data.
PMID 41700087 · PMC12910111 · Nucleic acids research · 2026 · 7 claims · 3 setups
GeMoSeq combines combinatorial enumeration of candidate transcripts, splitting heuristics, and likelihood-based (EM) quantification for transcript reconstruction from RNA-seq data
-
Full-text index only
Phylogenetic profiling of the Arabidopsis thaliana proteome: what proteins distinguish plants from other organisms?
PMID 15287975 · PMC507878 · Genome biology · 2004 · 8 claims · 6 setups
3,848 Arabidopsis proteins were identified as likely plant-specific based on phylogenetic profiling and EST confirmation in multiple plant species
-
Full-text index only
Benchmarking tools for the alignment of functional noncoding DNA.
PMID 14736341 · PMC344529 · BMC bioinformatics · 2004 · 8 claims · 4 setups
Global alignment tools (Avid, ClustalW, Lagan, Needle, DiAlign-G) typically have higher sensitivity over entire noncoding sequences and within constrained blocks than local tools
-
Full-text index only
GeneExt: a gene model extension tool for enhanced single-cell RNA-seq analysis.
PMID 41769841 · PMC12970594 · Bioinformatics (Oxford, England) · 2026 · 8 claims · 8 setups
Incomplete/inaccurate gene annotations, especially missing or truncated 3' UTRs, cause reads to map to non-genic regions and genes to be under-quantified or missing from scRNA-seq expression matrices in non-model species