Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
MODBASE, a database of annotated comparative protein structure models and associated resources.
PMID 18948282 · PMC2686492 · Nucleic acids research · 2009 · 8 claims · 8 setups
MODBASE contains 5,152,695 reliable comparative protein structure models for 1,593,209 unique protein sequences.
-
Has reproduction · 100
FA-nf: A Functional Annotation Pipeline for Proteins from Non-Model Organisms Implemented in Nextflow.
PMID 34681040 · PMC8535801 · Genes · 2021 · 8 claims · 4 setups
FA-nf, implemented in Nextflow with Docker/Singularity containerization, integrates NCBI BLAST+, DIAMOND, InterProScan, and KEGG (KAAS/KofamKOALA) into a single functional annotation pipeline.
-
Full-text index only
Comparative phosphoproteomics reveals evolutionary and functional conservation of phosphorylation across eukaryotes.
PMID 18828897 · PMC2760871 · Genome biology · 2008 · 8 claims · 8 setups
The overlap between phosphoproteomes of six eukaryotes (human, mouse, fly, yeast, plant, zebrafish) is significantly greater than expected by chance.
-
Full-text index only
Sequence variation in G-protein-coupled receptors: analysis of single nucleotide polymorphisms.
PMID 15784611 · PMC1069129 · Nucleic acids research · 2005 · 7 claims · 8 setups
Position-specific phylogenetic features describing evolutionary conservation at a site (e.g. SIFT score, normalized site entropy, residue frequency change) are the best individual discriminators of disease-causing versus neutral GPCR mutations.
-
Full-text index only
Similarities and differences in genome-wide expression data of six organisms.
PMID 14737187 · PMC300882 · PLoS biology · 2004 · 8 claims · 8 setups
Coexpression of functionally related genes is frequently conserved across evolutionarily distant organisms
-
Has reproduction · 65
FusionQ: a novel approach for gene fusion detection and quantification from paired-end RNA-Seq.
PMID 23768108 · PMC3691734 · BMC bioinformatics · 2013 · 8 claims · 8 setups
FusionQ is a novel tool that detects gene fusions, constructs chimerical transcript structures, and estimates their abundances from paired-end RNA-Seq data.
-
Full-text index only
Sys-BodyFluid: a systematical database for human body fluid proteome research.
PMID 18978022 · PMC2686600 · Nucleic acids research · 2009 · 6 claims · 4 setups
Sys-BodyFluid is a web-based database integrating proteomic data from 11 human body fluids (plasma/serum, urine, cerebrospinal fluid, saliva, bronchoalveolar lavage fluid, synovial fluid, nipple aspirate fluid, tear fluid, seminal fluid, milk, amniotic fluid), containing over 10,000 proteins
-
Full-text index only
GOLD.db: genomics of lipid-associated disorders database.
PMID 15588328 · PMC544894 · BMC genomics · 2004 · 8 claims · 4 setups
GOLD.db integrates annotated pathways, gene/protein reference information, and curated gene expression datasets for lipid-associated disorders research
-
Full-text index only
TreeFam: a curated database of phylogenetic trees of animal gene families.
PMID 16381935 · PMC1347480 · Nucleic acids research · 2006 · 7 claims · 6 setups
Tree-based inference of orthologs and paralogs is more robust than BLAST-based methods because evolutionary rates (and thus pairwise BLAST scores) vary across gene family members
-
Full-text index only
Assessing the genomic evidence for conserved transcribed pseudogenes under selection.
PMID 19754956 · PMC2753554 · BMC genomics · 2009 · 8 claims · 8 setups
1750 transcribed pseudogene annotations (TPAs) were identified in the human genome, ~11.5% of all human pseudogene annotations.
-
Full-text index only
Discovery and hypothesis generation through bioinformatics.
PMID 16522224 · PMC1431734 · Genome biology · 2006 · 8 claims · 8 setups
Bioinformatics should be used as a tool for discovery and hypothesis generation, not merely to manage biological data
-
Has reproduction · 71
Systematic and computational identification of Androctonus crassicauda long non-coding RNAs.
PMID 33633149 · PMC7907363 · Scientific reports · 2021 · 8 claims · 6 setups
13,401 lncRNAs were identified in the A. crassicauda transcriptome using the ECF pipeline
-
Full-text index only
Investigating hookworm genomes by comparative analysis of two Ancylostoma species.
PMID 15854223 · PMC1112591 · BMC genomics · 2005 · 8 claims · 8 setups
Nearly 20,000 ESTs from 7 cDNA libraries define nearly 7,000 hookworm genes across A. caninum and A. ceylanicum
-
Has reproduction · 91
Chromosome-level genome assembly of agar-producing red seaweed Gracilaria vermiculophylla.
PMID 41629338 · PMC12966425 · Scientific data · 2026 · 8 claims · 8 setups
Assembled the first chromosome-level genome of G. vermiculophylla: 77.5 Mb, 22 pseudochromosomes, contig N50 2.61 Mb, scaffold N50 3.16 Mb
-
Full-text index only
Characterization of the Schistosoma transcriptome opens up the world of helminth genomics.
PMID 14709167 · PMC395727 · Genome biology · 2003 · 8 claims · 5 setups
Near-complete transcriptome complements have now been described for S. japonicum and S. mansoni