Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Has reproduction · 75
An informatics research platform to make public gene expression time-course datasets reusable for more scientific discoveries.
PMID 33247935 · PMC7698665 · Database : the journal of biological databases and curation · 2020 · 8 claims · 6 setups
GETc enables discovery and visualization of time-course gene expression data and analytical results from GEO
-
Has reproduction · 76
Bayesian prediction of microbial oxygen requirement.
PMID 26913185 · PMC4743139 · F1000Research · 2013 · 7 claims · 8 setups
A naive Bayesian classifier based on presence/absence of class-associated Pfam-A domains can distinguish three oxygen requirement classes (aerobe, anaerobe, facultative anaerobe) from genome sequence, unlike prior studies that only made pairwise distinctions.
-
Full-text index only
PubMatrix: a tool for multiplex literature mining.
PMID 14667255 · PMC317283 · BMC bioinformatics · 2003 · 8 claims · 3 setups
PubMatrix is a web-based CGI tool that queries PubMed with two lists of terms (search terms vs modifier terms) and returns a matrix of pairwise co-occurrence frequency counts
-
Full-text index only
PeroxisomeDB: a database for the peroxisomal proteome, functional genomics and disease.
PMID 17135190 · PMC1747181 · Nucleic acids research · 2007 · 8 claims · 6 setups
PeroxisomeDB integrates the complete peroxisomal proteome of Homo sapiens and Saccharomyces cerevisiae into interrelated 'Genes', 'Functions', 'Metabolic pathways' and 'Diseases' sections with links to NCBI, ENSEMBL and UCSC
-
Full-text index only
CancerGenes: a gene selection resource for cancer genome projects.
PMID 17088289 · PMC1781153 · Nucleic acids research · 2007 · 6 claims · 4 setups
CancerGenes is a gene list-centric web resource that combines expert-annotated gene lists with data from public databases (Entrez Gene, Ensembl BioMart, Kim et al. promoter data, Sanger COSMIC) to support gene selection for cancer re-sequencing projects.
-
Full-text index only
Extraction of human kinase mutations from literature, databases and genotyping studies.
PMID 19758464 · PMC2745582 · BMC bioinformatics · 2009 · 7 claims · 6 setups
A literature mining pipeline combining MutationFinder, false-positive filtering, and SVM-based classification can extract and disambiguate single-point mutation mentions from abstracts and full text
-
Full-text index only
The Universal Protein Resource (UniProt) in 2010.
PMID 19843607 · PMC2808944 · Nucleic acids research · 2010 · 8 claims · 5 setups
UniProt is a centralized, freely accessible, comprehensive knowledgebase of protein sequence and functional annotation maintained by the EBI, SIB and PIR consortium.
-
Has reproduction · 95
The archives are half-empty: an assessment of the availability of microbial community sequencing data.
PMID 32859925 · PMC7455719 · Communications biology · 2020 · 8 claims · 5 setups
More than half of surveyed amplicon sequencing studies were affected by lack of data deposition, improper file formatting, or inconsistent labeling that impede reuse.
-
Has reproduction · 50
SMAC, a computational system to link literature, biomedical and expression data.
PMID 31324861 · PMC6642118 · Scientific reports · 2019 · 8 claims · 8 setups
SMAC is a tool that extracts, prioritises, integrates and analyses biomedical and molecular data according to user-defined terms
-
Has reproduction · 78
Fungal metabarcoding data integration framework for the MycoDiversity DataBase (MDDB).
PMID 32463383 · PMC7734503 · Journal of integrative bioinformatics · 2020 · 7 claims · 4 setups
Public fungal metabarcoding raw DNA data and their associated environmental metadata in sequence archives are heterogeneously annotated and lack a uniform processing pipeline, preventing large-scale biodiversity/distribution assessments.
-
Full-text index only
Quasimonomorphic mononucleotide repeats for high-level microsatellite instability analysis.
PMID 15528790 · PMC3888729 · Disease markers · 2004 · 8 claims · 8 setups
Mononucleotide repeats are more sensitive, specific, and easier to use than dinucleotide repeats for detecting MSI-H tumors
-
Full-text index only
Integration of text- and data-mining using ontologies successfully selects disease gene candidates.
PMID 15767279 · PMC1065256 · Nucleic acids research · 2005 · 7 claims · 6 setups
Integrating eVOC anatomical ontology-based text-mining of PubMed abstracts with data-mining of gene expression annotation successfully selects and prioritizes candidate disease genes
-
Full-text index only
TRED: a Transcriptional Regulatory Element Database and a platform for in silico gene regulation studies.
PMID 15608156 · PMC539958 · Nucleic acids research · 2005 · 8 claims · 5 setups
TRED is a database collecting both cis-regulatory elements (promoters) and trans-regulatory elements (transcription factor binding/regulation data) with linked access.
-
Full-text index only
ABS: a database of Annotated regulatory Binding Sites from orthologous promoters.
PMID 16381947 · PMC1347478 · Nucleic acids research · 2006 · 7 claims · 6 setups
ABS is a public database of experimentally identified TF binding sites conserved in orthologous vertebrate gene promoters, manually curated from the literature.
-
Full-text index only
Epidemiology of doublet/multiplet mutations in lung cancers: evidence that a subset arises by chronocoordinate events.
PMID 19005564 · PMC2579325 · PloS one · 2008 · 8 claims · 7 setups
Doublet mutations are significantly more frequent in EGFR (6.0%) and TP53 (2.3%) in human lung cancer than spontaneous doublets in mouse lacI (0.7%), about 8-fold and 3-fold higher respectively.
-
Full-text index only
In silico promoters: modelling of cis-regulatory context facilitates target predictio.
PMID 18505473 · PMC3823354 · Journal of cellular and molecular medicine · 2009 · 8 claims · 8 setups
An integrated 'profiling of transcriptional targets' (PTT) strategy by Freebern et al. identified IGF-1 as a co-modulator of immune cell function genes in mitogen/drug-activated T cells.
-
Full-text index only
What can genome-wide association studies tell us about the genetics of common disease?
PMID 18454206 · PMC2323402 · PLoS genetics · 2008 · 8 claims · 4 setups
Apparent patterns of common, low-effect disease-associated alleles largely reflect statistical power of studies rather than the true underlying distribution of disease variants
-
Has reproduction · 57
KARAJ: An Efficient Adaptive Multi-Processor Tool to Streamline Genomic and Transcriptomic Sequence Data Acquisition.
PMID 36430895 · PMC9694301 · International journal of molecular sciences · 2022 · 8 claims · 6 setups
KARAJ automates end-to-end querying and downloading of genomic/transcriptomic sequence data from a list of PMCIDs, URLs, or accession numbers
-
Full-text index only
Predicting candidate genes for human deafness disorders: a bioinformatics approach.
PMID 16854223 · PMC1564145 · BMC genomics · 2006 · 8 claims · 4 setups
A bioinformatic approach combining expression databases and protein interaction data narrows ~2400 candidate genes across deafness loci to a manageable set of candidates.
-
Full-text index only
Reconstruction of pathways associated with amino acid metabolism in human mitochondria.
PMID 18267298 · PMC5054205 · Genomics, proteomics & bioinformatics · 2007 · 8 claims · 5 setups
Out of 20 amino acids, the metabolic pathways of 17 utilize mitochondrial enzymes, and dysfunction of these enzymes causes over 40 known human mitochondrial diseases/disorders