Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
Discovery of protein-protein interactions using a combination of linguistic, statistical and graphical information.
PMID 15941473 · PMC1164402 · BMC bioinformatics · 2005 · 8 claims · 5 setups
A combined linguistic+statistical+rule-based method achieves precision 0.61 and recall 0.97 (f=0.74) detecting yeast protein-protein interactions across 12,300 Medline abstracts.
-
Full-text index only
Exogean: a framework for annotating protein-coding genes in eukaryotic genomic DNA.
PMID 16925841 · PMC1810556 · Genome biology · 2006 · 8 claims · 5 setups
Exogean is a framework using directed acyclic coloured multigraphs (DACMs) to represent biological objects (mRNA, ESTs, protein alignments, exons) and iteratively combine them into complex protein-coding transcript models.
-
Full-text index only
Large-scale identification and characterization of alternative splicing variants of human gene transcripts using 56,419 completely sequenced and manually annotated full-length cDNAs.
PMID 16914452 · PMC1557807 · Nucleic acids research · 2006 · 8 claims · 8 setups
Analysis of 56,419 full-length cDNAs identified 6877 alternative splicing genes encoding 18,297 alternative splicing variants made of 37,670 exons.
-
Has reproduction · 83
Public Omics Explorer (POE): Enabling integrative semantic search across GEO omics datasets based on PubMed publications.
PMID 41282419 · PMC12636342 · Computational and structural biotechnology journal · 2025 · 7 claims · 3 setups
POE performs literature-informed dataset retrieval by semantically linking GEO datasets and ENA records through associated PubMed publications
-
Full-text index only
Structural evolution of the protein kinase-like superfamily.
PMID 16244704 · PMC1261164 · PLoS computational biology · 2005 · 8 claims · 5 setups
All kinases in the superfamily share a 'universal core' domain consisting only of the regions required for ATP binding and the phosphotransfer reaction.
-
Has reproduction · 84
Genome of the Asian longhorned beetle (Anoplophora glabripennis), a globally significant invasive species, reveals key functional and evolutionary innovations at the beetle-plant interface.
PMID 27832824 · PMC5105290 · Genome biology · 2016 · 8 claims · 7 setups
The A. glabripennis genome encodes a uniquely diverse arsenal of enzymes that degrade plant cell wall polysaccharide networks (cellulose, hemicellulose, pectin) and detoxify plant allelochemicals.
-
Has reproduction · 89
Evolution of Highly Repetitive Silk Genes in the Luna Moth, Actias luna.
PMID 41738778 · PMC12962854 · Genome biology and evolution · 2026 · 8 claims · 6 setups
Eight sericin genes were identified in the A. luna genome, including two clusters of closely related paralogs (serB-D and serE-G)
-
Full-text index only
The homeobox gene CDX2 in colorectal carcinoma: a genetic analysis.
PMID 11161380 · PMC2363702 · British journal of cancer · 2001 · 7 claims · 7 setups
No CDX2 mutations predisposing to sporadic colorectal cancer were identified
-
Full-text index only
The DNA sequence and analysis of human chromosome 13.
PMID 15057823 · PMC2665288 · Nature · 2004 · 8 claims · 8 setups
95.5 Mb of finished sequence from chromosome 13 was completed, containing 633 genes and 296 pseudogenes.
-
Full-text index only
EGASP: the human ENCODE Genome Annotation Assessment Project.
PMID 16925836 · PMC1810551 · Genome biology · 2006 · 8 claims · 6 setups
Best-performing computational gene prediction methods correctly predict at least one transcript for close to 70% of annotated genes in the ENCODE regions.
-
Full-text index only
EGASP: Introduction.
PMID 16925831 · PMC1810546 · Genome biology · 2006 · 8 claims · 5 setups
Computational gene finding methods, when compared to the GENCODE golden standard annotation, show that the human genome annotation is nearly complete in terms of novel protein-coding loci.
-
Full-text index only
Variation analysis and gene annotation of eight MHC haplotypes: the MHC Haplotype Project.
PMID 18193213 · PMC2206249 · Immunogenetics · 2008 · 8 claims · 6 setups
Comparison of eight HLA-homozygous MHC haplotype sequences identified >44,000 variations (substitutions and indels), submitted to dbSNP
-
Full-text index only
Genome sequences and great expectations.
PMID 11178275 · PMC150431 · Genome biology · 2001 · 8 claims · 3 setups
Function is known or can be predicted for an average of 62% of proteins across 31 analyzed genomes.
-
Full-text index only
CGMIM: automated text-mining of Online Mendelian Inheritance in Man (OMIM) to identify genetically-associated cancers and candidate genes.
PMID 15796777 · PMC1274267 · BMC bioinformatics · 2005 · 8 claims · 2 setups
CGMIM is a Perl program that text-mines OMIM entries to identify cancer-gene associations and genetically-related cancer type pairs.
-
Full-text index only
VectorBase: a home for invertebrate vectors of human pathogens.
PMID 17145709 · PMC1751530 · Nucleic acids research · 2007 · 8 claims · 5 setups
VectorBase is a web-accessible data repository for information about invertebrate vectors of human pathogens
-
Has reproduction · 93
aCLImatise: automated generation of tool definitions for bioinformatics workflows.
PMID 33325479 · PMC8016486 · Bioinformatics (Oxford, England) · 2021 · 6 claims · 3 setups
aCLImatise automatically generates workflow-language tool definitions by parsing a command-line tool's help output
-
Has reproduction · 70
GAUGE-Annotated Microbial Transcriptomic Data Facilitate Parallel Mining and High-Throughput Reanalysis To Form Data-Driven Hypotheses.
PMID 33758032 · PMC8547006 · mSystems · 2021 · 8 claims · 7 setups
GAUGE automatically annotates GEO microbial transcriptomic data sets (microarray and RNA-seq), increasing the proportion of annotatable studies from 4% to 33%
-
Has reproduction · 74
Wide-Open: Accelerating public data release by automating detection of overdue datasets.
PMID 28594819 · PMC5464523 · PLoS biology · 2017 · 6 claims · 5 setups
A general text-mining + API-query approach (Wide-Open) can automatically identify datasets that are overdue for public release in a repository
-
Has reproduction · 100
Smart spatial omics (S2-omics) optimizes region of interest selection to capture molecular heterogeneity in diverse tissues.
PMID 41298871 · PMC12662399 · Nature cell biology · 2025 · 7 claims · 6 setups
S2-omics is an end-to-end workflow that automatically selects ROIs from H&E histology images to maximize molecular information content for spatial omics profiling.
-
Has reproduction · 92
Analytical code sharing practices in biomedical research.
PMID 38983240 · PMC11232620 · PeerJ. Computer science · 2024 · 8 claims · 4 setups
Nearly half (49.9%) of 453 examined biomedical manuscripts failed to share the analytical code used to generate their results