Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
Gene Prospector: an evidence gateway for evaluating potential susceptibility genes and interacting risk factors for human diseases.
PMID 19063745 · PMC2613935 · BMC bioinformatics · 2008 · 8 claims · 5 setups
Gene Prospector is a Web-based application that selects and prioritizes potential disease-related genes using a curated, updated literature database of genetic association studies
-
Full-text index only
Comparative Toxicogenomics Database: a knowledgebase and discovery tool for chemical-gene-disease networks.
PMID 18782832 · PMC2686584 · Nucleic acids research · 2009 · 8 claims · 5 setups
CTD is a manually curated knowledgebase that integrates chemical-gene interactions, chemical-disease relationships, and gene-disease relationships into a chemical-gene-disease triad
-
Has reproduction · 53
PulmonDB: a curated lung disease gene expression database.
PMID 31949184 · PMC6965635 · Scientific reports · 2020 · 6 claims · 6 setups
PulmonDB is a curated, web-based gene expression database and R package integrating microarray and RNA-seq data for COPD and IPF with manually curated controlled-vocabulary annotation.
-
Full-text index only
Evola: Ortholog database of all human genes in H-InvDB with manual curation of phylogenetic trees.
PMID 17982176 · PMC2238928 · Nucleic acids research · 2008 · 6 claims · 7 setups
Evola combines genome synteny-based computational ortholog detection with manual curation of phylogenetic trees by experts to yield more reliable orthologs than automated pairwise methods
-
Full-text index only
A comprehensive modular map of molecular interactions in RB/E2F pathway.
PMID 18319725 · PMC2290939 · Molecular systems biology · 2008 · 8 claims · 4 setups
A comprehensive, curated map of RB/E2F pathway molecular interactions was built using SBGN notation in CellDesigner and converted to BioPAX 2.0 format
-
Full-text index only
Systematic identification of pseudogenes through whole genome expression evidence profiling.
PMID 16945953 · PMC1636364 · Nucleic acids research · 2006 · 8 claims · 8 setups
Developed a novel bioinformatics method that identifies pseudogenes by profiling whole-genome transcript and protein expression evidence
-
Full-text index only
The DAVID Gene Functional Classification Tool: a novel biological module-centric algorithm to functionally analyze large gene lists.
PMID 17784955 · PMC2375021 · Genome biology · 2007 · 8 claims · 6 setups
Gene-gene functional similarity can be measured using kappa statistics applied to a binary gene-annotation-term matrix built from 14 annotation categories.
-
Full-text index only
How many human genes can be defined as housekeeping with current expression data?
PMID 18416810 · PMC2396180 · BMC genomics · 2008 · 8 claims · 4 setups
Current EST and microarray transcriptome sampling is far from saturated, limiting gene detectability and understanding of tissue-specific expression
-
Full-text index only
SelTarbase, a database of human mononucleotide-microsatellite mutations and their potential impact to tumorigenesis and immunology.
PMID 19820113 · PMC2808963 · Nucleic acids research · 2010 · 7 claims · 6 setups
SelTarbase is a curated relational database of published mononucleotide-repeat mutation data from MSI-H human colorectal, gastric, endometrial tumors and colon cancer cell lines.
-
Full-text index only
Ensembl 2006.
PMID 16381931 · PMC1347495 · Nucleic acids research · 2006 · 8 claims · 5 setups
Ensembl now provides annotation for 19 genomes, up from 4 the previous year, including new mammalian (Rhesus macaque, Opossum), chordate (Ciona intestinalis), and yeast genomes.
-
Has reproduction · 80
Curation of over 10 000 transcriptomic studies to enable data reuse.
PMID 33599246 · PMC7904053 · Database : the journal of biological databases and curation · 2021 · 8 claims · 6 setups
Gemma is a curated database and bioinformatics system that addresses metadata, probe annotation, and expression data inconsistencies in GEO to enable transcriptomic data reuse
-
Has reproduction · 100
Intratumoral heterogeneity in microsatellite instability status at single-cell resolution.
PMID 41767255 · PMC12936829 · iScience · 2026 · 8 claims · 8 setups
MSI status can be heterogeneous at the single-cell level within a tumor, challenging its use as a binary biomarker
-
Has reproduction · 100
Computational modeling demonstrates that glioblastoma cells can survive spatial environmental challenges through exploratory adaptation.
PMID 31836713 · PMC6911112 · Nature communications · 2019 · 8 claims · 6 setups
Stochastic exploration of the gene-regulatory network structure confers enhanced adaptive capacity, enabling GBM cells to converge to new target phenotypes in novel environments.
-
Has reproduction · 19
Bioinformatics Strategies to Identify Shared Molecular Biomarkers That Link Ischemic Stroke and Moyamoya Disease with Glioblastoma.
PMID 36015199 · PMC9413912 · Pharmaceutics · 2022 · 8 claims · 8 setups
Shared differentially expressed genes link glioblastoma with ischemic stroke and with moyamoya disease, revealing molecular associations among the diseases.
-
Has reproduction · 76
The genome and development-dependent transcriptomes of Pyronema confluens: a window into fungal evolution.
PMID 24068976 · PMC3778014 · PLoS genetics · 2013 · 8 claims · 8 setups
The 50 Mb P. confluens genome with 13,369 predicted protein-coding genes is more characteristic of higher filamentous ascomycetes than of the large, repeat-rich Tuber melanosporum genome, showing that the truffle's expanded genome is not typical of the Pezizales.
-
Full-text index only
In silico analysis of missense substitutions using sequence-alignment based methods.
PMID 18951440 · PMC3431198 · Human mutation · 2008 · 8 claims · 7 setups
Carefully validated PMSA-based computational algorithms can achieve predictive values of ~75-95% for classifying missense substitutions as pathogenic or neutral.
-
Full-text index only
The MAPPER database: a multi-genome catalog of putative transcription factor binding sites.
PMID 15608292 · PMC540057 · Nucleic acids research · 2005 · 8 claims · 6 setups
Built a library of 1134 HMM models (359 matrix-derived, 718 factor-derived, 57 JASPAR-derived), corresponding to 863 distinct TF names, from TRANSFAC and JASPAR binding site data
-
Full-text index only
Human disease classification in the postgenomic era: a complex systems approach to human pathobiology.
PMID 17625512 · PMC1948102 · Molecular systems biology · 2007 · 8 claims · 5 setups
Current syndromic disease classification lacks specificity despite historically serving clinicians well
-
Full-text index only
Integrating proteomic, transcriptional, and interactome data reveals hidden components of signaling and regulatory networks.
PMID 19638617 · PMC2889494 · Science signaling · 2009 · 8 claims · 6 setups
Pathway reconstruction can be modeled as a prize-collecting Steiner tree problem, balancing penalties for excluding terminal nodes against costs for including edges, controlled by a parameter β.
-
Full-text index only
EGenBio: a data management system for evolutionary genomics and biodiversity.
PMID 17118150 · PMC1683573 · BMC bioinformatics · 2006 · 7 claims · 7 setups
EGenBio is a web-based system for integrated management, filtering, curation, and visualization of large-scale genomic sequences, alignments, and phylogenetic trees for evolutionary genomics and biodiversity research.