Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Has reproduction · 74
Wide-Open: Accelerating public data release by automating detection of overdue datasets.
PMID 28594819 · PMC5464523 · PLoS biology · 2017 · 7 claims · 5 setups
Wide-Open is a general text-mining approach that automatically detects overdue datasets by scanning PubMed articles for dataset accession identifiers and querying repositories to determine if the datasets remain private.
-
Full-text index only
Integration of text- and data-mining using ontologies successfully selects disease gene candidates.
PMID 15767279 · PMC1065256 · Nucleic acids research · 2005 · 7 claims · 6 setups
Integrating eVOC anatomical ontology-based text-mining of PubMed abstracts with data-mining of gene expression annotation successfully selects and prioritizes candidate disease genes
-
Full-text index only
PolySearch: a web-based text mining system for extracting relationships between human diseases, genes, mutations, drugs and metabolites.
PMID 18487273 · PMC2447794 · Nucleic acids research · 2008 · 8 claims · 7 setups
PolySearch supports more than 50 different classes of queries against nearly a dozen types of text, abstract, or bioinformatic databases
-
Full-text index only
SuperCYP: a comprehensive database on Cytochrome P450 enzymes including a tool for analysis of CYP-drug interactions.
PMID 19934256 · PMC2808967 · Nucleic acids research · 2010 · 8 claims · 6 setups
SuperCYP is a comprehensive relational database aggregating CYP enzyme, drug metabolism, SNP/mutation, and structural information from literature and web resources.
-
Full-text index only
Inherited disorder phenotypes: controlled annotation and statistical analysis for knowledge mining from gene lists.
PMID 16351744 · PMC1866390 · BMC bioinformatics · 2005 · 5 claims · 3 setups
OMIM Clinical Synopsis free-text phenotype and location names can be normalized and hierarchically structured into a controlled vocabulary suitable for computational analysis
-
Has reproduction · 33
Tracing truth: dynamic temporal networks for multi-modal fake news detection.
PMID 40989433 · PMC12453775 · PeerJ. Computer science · 2025 · 8 claims · 5 setups
The proposed dynamic temporal network (DTN) model improves multi-modal fake news detection accuracy by capturing temporal dynamics of propagation nodes and dynamically fusing multi-modal information.
-
Full-text index only
BRENDA, AMENDA and FRENDA: the enzyme information system in 2007.
PMID 17202167 · PMC1899097 · Nucleic acids research · 2007 · 7 claims · 6 setups
BRENDA is the largest publicly available enzyme information system worldwide, manually curated from primary literature and covering all identified enzymes regardless of source.
-
Has reproduction · 95
The archives are half-empty: an assessment of the availability of microbial community sequencing data.
PMID 32859925 · PMC7455719 · Communications biology · 2020 · 8 claims · 5 setups
More than half of surveyed amplicon sequencing studies were affected by lack of data deposition, improper file formatting, or inconsistent labeling that impede reuse.
-
Full-text index only
Genome annotation errors in pathway databases due to semantic ambiguity in partial EC numbers.
PMID 16034025 · PMC1179732 · Nucleic acids research · 2005 · 7 claims · 4 setups
Partial EC numbers are semantically ambiguous, and databases that assign a gene to all reactions sharing the same partial EC number make a faulty inference, causing systematic misannotation.
-
Full-text index only
CGMIM: automated text-mining of Online Mendelian Inheritance in Man (OMIM) to identify genetically-associated cancers and candidate genes.
PMID 15796777 · PMC1274267 · BMC bioinformatics · 2005 · 8 claims · 2 setups
CGMIM is a Perl program that text-mines OMIM entries to identify cancer-gene associations and genetically-related cancer type pairs.
-
Full-text index only
The Mammalian Phenotype Ontology as a tool for annotating, analyzing and comparing phenotypic information.
PMID 15642099 · PMC549068 · Genome biology · 2005 · 7 claims · 1 setups
The MP Ontology enables robust, standardized annotation of mammalian phenotypes for mutations, QTLs, and strains used as models of human biology and disease.
-
Has reproduction · 92
Analytical code sharing practices in biomedical research.
PMID 38983240 · PMC11232620 · PeerJ. Computer science · 2024 · 8 claims · 4 setups
Nearly half (49.9%) of 453 examined biomedical manuscripts failed to share the analytical code used to generate their results
-
Has reproduction · 75
An informatics research platform to make public gene expression time-course datasets reusable for more scientific discoveries.
PMID 33247935 · PMC7698665 · Database : the journal of biological databases and curation · 2020 · 8 claims · 6 setups
GETc enables discovery and visualization of time-course gene expression data and analytical results from GEO
-
Full-text index only
Information extraction from full text scientific articles: where are the keywords?
PMID 12775220 · PMC166134 · BMC bioinformatics · 2003 · 8 claims · 5 setups
The keyword content of the five article sections (A, I, M, R, D) is heterogeneous, i.e., different sections carry different kinds of information.
-
Full-text index only
QTL MatchMaker: a multi-species quantitative trait loci (QTL) database and query system for annotation of genes and QTL.
PMID 16381937 · PMC1347390 · Nucleic acids research · 2006 · 8 claims · 5 setups
QTL MatchMaker integrates QTL information with physical, genetic and cytogenetic maps across human, mouse and rat genomes
-
Full-text index only
SysPIMP: the web-based systematical platform for identifying human disease-related mutated sequences from mass spectrometry.
PMID 19036792 · PMC2686442 · Nucleic acids research · 2009 · 8 claims · 7 setups
SysPIMP is a web-based platform integrating disease mutation databases with X!Tandem and BLAST to identify disease-related mutated proteins from MS results
-
Full-text index only
A biomedically enriched collection of 7000 human ORF clones.
PMID 18231609 · PMC2211400 · PloS one · 2008 · 8 claims · 4 setups
Produced and made available over 7000 fully sequence-verified plasmid ORF clones representing over 3400 unique human genes, in both closed (stop codon) and fusion (no stop codon) formats.
-
Full-text index only
Improved mutation tagging with gene identifiers applied to membrane protein stability prediction.
PMID 19758467 · PMC2745585 · BMC bioinformatics · 2009 · 8 claims · 4 setups
MutationTagger achieves 87% F-measure for the mutation retrieval task on a benchmark dataset
-
Full-text index only
PubMatrix: a tool for multiplex literature mining.
PMID 14667255 · PMC317283 · BMC bioinformatics · 2003 · 8 claims · 3 setups
PubMatrix is a web-based CGI tool that queries PubMed with two lists of terms (search terms vs modifier terms) and returns a matrix of pairwise co-occurrence frequency counts
-
Has reproduction · 93
aCLImatise: automated generation of tool definitions for bioinformatics workflows.
PMID 33325479 · PMC8016486 · Bioinformatics (Oxford, England) · 2021 · 6 claims · 3 setups
aCLImatise automatically generates workflow-language tool definitions by parsing a command-line tool's help output