Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
PubMatrix: a tool for multiplex literature mining.
PMID 14667255 · PMC317283 · BMC bioinformatics · 2003 · 8 claims · 3 setups
PubMatrix is a web-based CGI tool that queries PubMed with two lists of terms (search terms vs modifier terms) and returns a matrix of pairwise co-occurrence frequency counts
-
Has reproduction · 74
Wide-Open: Accelerating public data release by automating detection of overdue datasets.
PMID 28594819 · PMC5464523 · PLoS biology · 2017 · 7 claims · 5 setups
Wide-Open is a general text-mining approach that automatically detects overdue datasets by scanning PubMed articles for dataset accession identifiers and querying repositories to determine if the datasets remain private.
-
Has reproduction · 95
The archives are half-empty: an assessment of the availability of microbial community sequencing data.
PMID 32859925 · PMC7455719 · Communications biology · 2020 · 8 claims · 5 setups
More than half of surveyed amplicon sequencing studies were affected by lack of data deposition, improper file formatting, or inconsistent labeling that impede reuse.
-
Has reproduction · 75
An informatics research platform to make public gene expression time-course datasets reusable for more scientific discoveries.
PMID 33247935 · PMC7698665 · Database : the journal of biological databases and curation · 2020 · 8 claims · 6 setups
GETc enables discovery and visualization of time-course gene expression data and analytical results from GEO
-
Full-text index only
CGMIM: automated text-mining of Online Mendelian Inheritance in Man (OMIM) to identify genetically-associated cancers and candidate genes.
PMID 15796777 · PMC1274267 · BMC bioinformatics · 2005 · 8 claims · 2 setups
CGMIM is a Perl program that text-mines OMIM entries to identify cancer-gene associations and genetically-related cancer type pairs.
-
Has reproduction · 57
KARAJ: An Efficient Adaptive Multi-Processor Tool to Streamline Genomic and Transcriptomic Sequence Data Acquisition.
PMID 36430895 · PMC9694301 · International journal of molecular sciences · 2022 · 8 claims · 6 setups
KARAJ automates end-to-end querying and downloading of genomic/transcriptomic sequence data from a list of PMCIDs, URLs, or accession numbers
-
Full-text index only
Integration of text- and data-mining using ontologies successfully selects disease gene candidates.
PMID 15767279 · PMC1065256 · Nucleic acids research · 2005 · 7 claims · 6 setups
Integrating eVOC anatomical ontology-based text-mining of PubMed abstracts with data-mining of gene expression annotation successfully selects and prioritizes candidate disease genes
-
Full-text index only
Computational disease gene identification: a concert of methods prioritizes type 2 diabetes and obesity candidate genes.
PMID 16757574 · PMC1475747 · Nucleic acids research · 2006 · 6 claims · 8 setups
Applying seven independent computational disease-gene prioritization methods in concert to 9556 positional candidate genes identifies a prioritized set of likely T2D and obesity candidate genes
-
Full-text index only
Linking disease-associated genes to regulatory networks via promoter organization.
PMID 15701758 · PMC549397 · Nucleic acids research · 2005 · 8 claims · 7 setups
Pairs of TFBSs conserved both vertically (orthologous genes) and horizontally (co-regulated genes) can serve as seeds to build promoter models representing potential co-regulation networks
-
Full-text index only
G2Cdb: the Genes to Cognition database.
PMID 18984621 · PMC2686544 · Nucleic acids research · 2009 · 7 claims · 7 setups
G2Cdb integrates experimentally validated synapse proteome datasets with mouse/human genomic annotation, phenotype, and human disease data in a gene-centric database.
-
Full-text index only
Extraction of human kinase mutations from literature, databases and genotyping studies.
PMID 19758464 · PMC2745582 · BMC bioinformatics · 2009 · 7 claims · 6 setups
A literature mining pipeline combining MutationFinder, false-positive filtering, and SVM-based classification can extract and disambiguate single-point mutation mentions from abstracts and full text
-
Full-text index only
Proteomic-based identification of maternal proteins in mature mouse oocytes.
PMID 19646285 · PMC2730056 · BMC genomics · 2009 · 8 claims · 6 setups
625 different proteins were identified from 2700 zona pellucida-free mature mouse MII oocytes, the largest oocyte proteome catalog to date
-
Full-text index only
Babelomics: advanced functional profiling of transcriptomics, proteomics and genomics experiments.
PMID 18515841 · PMC2447758 · Nucleic acids research · 2008 · 8 claims · 5 setups
Babelomics is a web suite offering both conventional functional enrichment methods and more advanced gene set analysis (GSA) methods, a combination offered by only one other tool (FuncAssociate) among competitors.