Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Has reproduction · 95
The archives are half-empty: an assessment of the availability of microbial community sequencing data.
PMID 32859925 · PMC7455719 · Communications biology · 2020 · 8 claims · 5 setups
More than half of surveyed amplicon sequencing studies were affected by lack of data deposition, improper file formatting, or inconsistent labeling that impede reuse.
-
Full-text index only
Information extraction from full text scientific articles: where are the keywords?
PMID 12775220 · PMC166134 · BMC bioinformatics · 2003 · 8 claims · 5 setups
The keyword content of the five article sections (A, I, M, R, D) is heterogeneous, i.e., different sections carry different kinds of information.
-
Has reproduction · 74
Wide-Open: Accelerating public data release by automating detection of overdue datasets.
PMID 28594819 · PMC5464523 · PLoS biology · 2017 · 7 claims · 5 setups
Wide-Open is a general text-mining approach that automatically detects overdue datasets by scanning PubMed articles for dataset accession identifiers and querying repositories to determine if the datasets remain private.
-
Full-text index only
SuperCYP: a comprehensive database on Cytochrome P450 enzymes including a tool for analysis of CYP-drug interactions.
PMID 19934256 · PMC2808967 · Nucleic acids research · 2010 · 8 claims · 6 setups
SuperCYP is a comprehensive relational database aggregating CYP enzyme, drug metabolism, SNP/mutation, and structural information from literature and web resources.
-
Full-text index only
Extraction of human kinase mutations from literature, databases and genotyping studies.
PMID 19758464 · PMC2745582 · BMC bioinformatics · 2009 · 7 claims · 6 setups
A literature mining pipeline combining MutationFinder, false-positive filtering, and SVM-based classification can extract and disambiguate single-point mutation mentions from abstracts and full text
-
Has reproduction · 75
An informatics research platform to make public gene expression time-course datasets reusable for more scientific discoveries.
PMID 33247935 · PMC7698665 · Database : the journal of biological databases and curation · 2020 · 8 claims · 6 setups
GETc enables discovery and visualization of time-course gene expression data and analytical results from GEO
-
Has reproduction · 57
KARAJ: An Efficient Adaptive Multi-Processor Tool to Streamline Genomic and Transcriptomic Sequence Data Acquisition.
PMID 36430895 · PMC9694301 · International journal of molecular sciences · 2022 · 8 claims · 6 setups
KARAJ automates end-to-end querying and downloading of genomic/transcriptomic sequence data from a list of PMCIDs, URLs, or accession numbers