Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
Information extraction from full text scientific articles: where are the keywords?
PMID 12775220 · PMC166134 · BMC bioinformatics · 2003 · 8 claims · 5 setups
The keyword content of the five article sections (A, I, M, R, D) is heterogeneous, i.e., different sections carry different kinds of information.
-
Has reproduction · 43
StatsDB: platform-agnostic storage and understanding of next generation sequencing run metrics.
PMID 24627795 · PMC3938176 · F1000Research · 2013 · 8 claims · 6 setups
StatsDB is an open-source software package for storage and analysis of next generation sequencing run metrics, backed by an SQL (MySQL) database with Perl and Java APIs.
-
Has reproduction · 94
Systematic assessment of pathway databases, based on a diverse collection of user-submitted experiments.
PMID 36088548 · PMC9487593 · Briefings in bioinformatics · 2022 · 8 claims · 6 setups
Well-established, hierarchically organized pathway annotation systems (e.g. GO, Reactome, KEGG) yield the best overall enrichment performance despite covering much of the human genome only in general terms.
-
Full-text index only
Inherited disorder phenotypes: controlled annotation and statistical analysis for knowledge mining from gene lists.
PMID 16351744 · PMC1866390 · BMC bioinformatics · 2005 · 5 claims · 3 setups
OMIM Clinical Synopsis free-text phenotype and location names can be normalized and hierarchically structured into a controlled vocabulary suitable for computational analysis
-
Has reproduction · 75
An informatics research platform to make public gene expression time-course datasets reusable for more scientific discoveries.
PMID 33247935 · PMC7698665 · Database : the journal of biological databases and curation · 2020 · 8 claims · 6 setups
GETc enables discovery and visualization of time-course gene expression data and analytical results from GEO
-
Full-text index only
Integration of text- and data-mining using ontologies successfully selects disease gene candidates.
PMID 15767279 · PMC1065256 · Nucleic acids research · 2005 · 7 claims · 6 setups
Integrating eVOC anatomical ontology-based text-mining of PubMed abstracts with data-mining of gene expression annotation successfully selects and prioritizes candidate disease genes
-
Full-text index only
Gene Prospector: an evidence gateway for evaluating potential susceptibility genes and interacting risk factors for human diseases.
PMID 19063745 · PMC2613935 · BMC bioinformatics · 2008 · 8 claims · 5 setups
Gene Prospector is a Web-based application that selects and prioritizes potential disease-related genes using a curated, updated literature database of genetic association studies
-
Full-text index only
A biomedically enriched collection of 7000 human ORF clones.
PMID 18231609 · PMC2211400 · PloS one · 2008 · 8 claims · 4 setups
Produced and made available over 7000 fully sequence-verified plasmid ORF clones representing over 3400 unique human genes, in both closed (stop codon) and fusion (no stop codon) formats.
-
Full-text index only
Babelomics: advanced functional profiling of transcriptomics, proteomics and genomics experiments.
PMID 18515841 · PMC2447758 · Nucleic acids research · 2008 · 8 claims · 5 setups
Babelomics is a web suite offering both conventional functional enrichment methods and more advanced gene set analysis (GSA) methods, a combination offered by only one other tool (FuncAssociate) among competitors.
-
Full-text index only
Extraction of human kinase mutations from literature, databases and genotyping studies.
PMID 19758464 · PMC2745582 · BMC bioinformatics · 2009 · 7 claims · 6 setups
A literature mining pipeline combining MutationFinder, false-positive filtering, and SVM-based classification can extract and disambiguate single-point mutation mentions from abstracts and full text
-
Has reproduction · 99
The systematic assessment of completeness of public metadata accompanying omics studies in the Gene Expression Omnibus data repository.
PMID 40926267 · PMC12421755 · Genome biology · 2025 · 8 claims · 3 setups
Over 25% of critical metadata are omitted, with only 74.8% of relevant phenotypes available in publications or public repositories.
-
Full-text index only
Application of two machine learning algorithms to genetic association studies in the presence of covariates.
PMID 19014573 · PMC2620353 · BMC genetics · 2008 · 8 claims · 3 setups
The relative performance of RF and MARS for detecting genotype-trait associations depends on both the strategy used to handle covariates and the true underlying model of association (e.g., confounding vs. mediation vs. interaction).
-
Full-text index only
Hit selection with false discovery rate control in genome-scale RNAi screens.
PMID 18628291 · PMC2504311 · Nucleic acids research · 2008 · 8 claims · 3 setups
A Bayesian FDR-controlling methodology for hit selection in genome-scale RNAi HTS is proposed, using a direct posterior probability approach analogous to Newton et al.
-
Has reproduction · 57
KARAJ: An Efficient Adaptive Multi-Processor Tool to Streamline Genomic and Transcriptomic Sequence Data Acquisition.
PMID 36430895 · PMC9694301 · International journal of molecular sciences · 2022 · 8 claims · 6 setups
KARAJ automates end-to-end querying and downloading of genomic/transcriptomic sequence data from a list of PMCIDs, URLs, or accession numbers
-
Has reproduction · 84
COXPRESdb v8: an animal gene coexpression database navigating from a global view to detailed investigations.
PMID 36350658 · PMC9825429 · Nucleic acids research · 2023 · 8 claims · 6 setups
COXPRESdb version 8 adds CoexMap (UMAP-based genome-scale coexpression visualization), KEGG pathway enrichment summaries, and CoexPub (literature-linking tool) as new analysis features.
-
Full-text index only
Building disease-specific drug-protein connectivity maps from molecular interaction networks and PubMed abstracts.
PMID 19649302 · PMC2709445 · PLoS computational biology · 2009 · 7 claims · 4 setups
A computational framework can build disease-specific drug-protein connectivity maps by integrating protein interaction networks and PubMed literature mining, without gene expression profiles from drug perturbation experiments
-
Has reproduction · 86
Prediction, syntax and semantic grounding in the brain and large language models.
PMID 41807493 · PMC12979642 · Scientific reports · 2026 · 8 claims · 6 setups
Nouns show significant pre-onset neural activity, suggesting enhanced anticipatory processing of this word class.
-
Full-text index only
Columba: an integrated database of proteins, structures, and annotations.
PMID 15801979 · PMC1087474 · BMC bioinformatics · 2005 · 8 claims · 6 setups
COLUMBA physically integrates data from twelve protein structure-related databases (PDB, KEGG, Swiss-Prot, CATH, SCOP, Gene Ontology, ENZYME, etc.) into a single PostgreSQL data warehouse.