Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
EGenBio: a data management system for evolutionary genomics and biodiversity.
PMID 17118150 · PMC1683573 · BMC bioinformatics · 2006 · 7 claims · 7 setups
EGenBio is a web-based system for integrated management, filtering, curation, and visualization of large-scale genomic sequences, alignments, and phylogenetic trees for evolutionary genomics and biodiversity research.
-
Has reproduction · 83
Macrel: antimicrobial peptide screening in genomes and metagenomes.
PMID 33384902 · PMC7751412 · PeerJ · 2020 · 8 claims · 8 setups
Macrel is an end-to-end pipeline that predicts high-quality AMP candidates from peptides, contigs, or reads of (meta)genomes
-
Full-text index only
SelTarbase, a database of human mononucleotide-microsatellite mutations and their potential impact to tumorigenesis and immunology.
PMID 19820113 · PMC2808963 · Nucleic acids research · 2010 · 7 claims · 6 setups
SelTarbase is a curated relational database of published mononucleotide-repeat mutation data from MSI-H human colorectal, gastric, endometrial tumors and colon cancer cell lines.
-
Full-text index only
nsSNPAnalyzer: identifying disease-associated nonsynonymous single nucleotide polymorphisms.
PMID 15980516 · PMC1160133 · Nucleic acids research · 2005 · 6 claims · 4 setups
nsSNPAnalyzer is a web server that predicts whether a query nsSNP is disease-associated or functionally neutral using a Random Forest classifier combining structural and evolutionary information
-
Full-text index only
Sequence similarity network reveals common ancestry of multidomain proteins.
PMID 18475320 · PMC2377100 · PLoS computational biology · 2008 · 8 claims · 6 setups
Traditional homology definitions do not capture multidomain evolution; the authors extend the definition to include domain insertion via a common ancestral locus model.
-
Full-text index only
Towards alignment independent quantitative assessment of homology detection.
PMID 17205117 · PMC1762415 · PloS one · 2006 · 8 claims · 6 setups
The Fhom Estimator uses the prevalence of a conserved protein feature (X) in two protein sets to estimate the fraction of true homologs among paired proteins, independent of alignment quality.
-
Full-text index only
TreeFam: a curated database of phylogenetic trees of animal gene families.
PMID 16381935 · PMC1347480 · Nucleic acids research · 2006 · 7 claims · 6 setups
Tree-based inference of orthologs and paralogs is more robust than BLAST-based methods because evolutionary rates (and thus pairwise BLAST scores) vary across gene family members
-
Full-text index only
GOLD.db: genomics of lipid-associated disorders database.
PMID 15588328 · PMC544894 · BMC genomics · 2004 · 8 claims · 4 setups
GOLD.db integrates annotated pathways, gene/protein reference information, and curated gene expression datasets for lipid-associated disorders research
-
Full-text index only
AnimalQTLdb: a livestock QTL database tool set for positional QTL information mining and beyond.
PMID 17135205 · PMC1781224 · Nucleic acids research · 2007 · 7 claims · 3 setups
AnimalQTLdb houses publicly available QTL data for multiple livestock species (pig, cattle, chicken) in one comparable database.
-
Has reproduction · 24
MiGPC: a comprehensive catalog of enzybiotics from environmental metagenomes.
PMID 41888223 · PMC13172421 · Scientific reports · 2026 · 8 claims · 8 setups
MiGPC is the first genome-resolved metagenomic gene and protein catalog specifically targeted to enzybiotics
-
Full-text index only
Integrating alternative splicing detection into gene prediction.
PMID 15705189 · PMC550657 · BMC bioinformatics · 2005 · 8 claims · 4 setups
An integrative intrinsic/extrinsic method was implemented in the gene finder EuGÈNE (as EuGÈNE-M) to detect AS evidence from aligned transcripts and generate alternative optimal gene predictions consistent with each detected AS event.
-
Has reproduction · 45
Identifying and classifying trait linked polymorphisms in non-reference species by walking coloured de bruijn graphs.
PMID 23536903 · PMC3607606 · PloS one · 2013 · 8 claims · 9 setups
Bubbleparse detects sequence variants directly from NGS reads without a reference genome, using the coloured de Bruijn graph implementation of Cortex plus a new depth-first bubble-finding module.
-
Full-text index only
Sys-BodyFluid: a systematical database for human body fluid proteome research.
PMID 18978022 · PMC2686600 · Nucleic acids research · 2009 · 6 claims · 4 setups
Sys-BodyFluid is a web-based database integrating proteomic data from 11 human body fluids (plasma/serum, urine, cerebrospinal fluid, saliva, bronchoalveolar lavage fluid, synovial fluid, nipple aspirate fluid, tear fluid, seminal fluid, milk, amniotic fluid), containing over 10,000 proteins
-
Has reproduction · 80
Curation of over 10 000 transcriptomic studies to enable data reuse.
PMID 33599246 · PMC7904053 · Database : the journal of biological databases and curation · 2021 · 8 claims · 6 setups
Gemma is a curated database and bioinformatics system that addresses metadata, probe annotation, and expression data inconsistencies in GEO to enable transcriptomic data reuse
-
Full-text index only
MEROPS: the peptidase database.
PMID 19892822 · PMC2808883 · Nucleic acids research · 2010 · 8 claims · 5 setups
MEROPS is a manually curated hierarchical classification of peptidases and protein inhibitors organized into protein species, families, and clans based on sequence and structural homology.
-
Full-text index only
Undergraduate research. Genomics Education Partnership.
PMID 18974335 · PMC2953277 · Science (New York, N.Y.) · 2008 · 6 claims · 4 setups
A course-embedded, multi-institution undergraduate research model (the Genomics Education Partnership) can deliver authentic research experiences during the academic year rather than only in summer programs.
-
Full-text index only
In silico analysis of missense substitutions using sequence-alignment based methods.
PMID 18951440 · PMC3431198 · Human mutation · 2008 · 8 claims · 7 setups
Carefully validated PMSA-based computational algorithms can achieve predictive values of ~75-95% for classifying missense substitutions as pathogenic or neutral.
-
Full-text index only
A space-efficient and accurate method for mapping and aligning cDNA sequences onto genomic sequence.
PMID 18344523 · PMC2377433 · Nucleic acids research · 2008 · 7 claims · 6 setups
Spaln maps and aligns large cDNA sequence sets onto whole mammalian genomes using substantially less memory than comparable existing tools
-
Full-text index only
pSTIING: a 'systems' approach towards integrating signalling pathways, interaction and transcriptional regulatory networks in inflammation and cancer.
PMID 16381926 · PMC1347407 · Nucleic acids research · 2006 · 8 claims · 3 setups
pSTIING is a publicly accessible web-based knowledgebase integrating protein-protein, protein-lipid, protein-small molecule interactions, transcriptional regulatory associations, ligand-receptor-cell type information, and signal transduction modules, with a focus on inflammation, cell migration and cancer.
-
Has reproduction · 83
Gene Expression Atlas update--a value-added database of microarray and sequencing-based functional genomics experiments.
PMID 22064864 · PMC3245177 · Nucleic acids research · 2012 · 8 claims · 5 setups
Gene Expression Atlas is an added-value database providing curated, re-annotated and statistically analysed gene expression data across cell types, organism parts, developmental stages, disease states and other biological/experimental conditions, derived from ArrayExpress Archive and the European Nucleotide Archive.