Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
Metappuccino: large language model-driven reconstruction of sequence read archive metadata for cancer research.
PMID 42057294 · PMC13148957 · Bioinformatics (Oxford, England) · 2026 · 8 claims · 3 setups
Metappuccino reconstructs 19 metadata classes by combining deterministic rule-based extraction/normalization (for explicit context) with LoRA-specialized Mistral-7B-Instruct completion (for missing/ambiguous fields)
-
Has reproduction · 64
GeMI: interactive interface for transformer-based Genomic Metadata Integration.
PMID 35657113 · PMC9216561 · Database : the journal of biological databases and curation · 2022 · 8 claims · 5 setups
GeMI is a web tool that uses a fine-tuned GPT2 model to extract 15 structured key-value attributes from free-text GEO sample metadata.
-
Has reproduction · 75
An informatics research platform to make public gene expression time-course datasets reusable for more scientific discoveries.
PMID 33247935 · PMC7698665 · Database : the journal of biological databases and curation · 2020 · 8 claims · 6 setups
GETc enables discovery and visualization of time-course gene expression data and analytical results from GEO
-
Has reproduction · 88
Human methylome variation across Infinium 450K data on the Gene Expression Omnibus.
PMID 33937763 · PMC8061458 · NAR genomics and bioinformatics · 2021 · 8 claims · 8 setups
Among annotated HM450K GEO samples, about two-thirds were from blood, one-quarter from brain, and about one-third were from cancer patients.
-
Full-text index only
Human-scATAC-Corpus: a comprehensive database of scATAC-seq data.
PMID 41296545 · PMC12807747 · Nucleic acids research · 2026 · 8 claims · 6 setups
Human-scATAC-Corpus is a comprehensive database of human scATAC-seq data comprising 5,407,621 cells from 35 datasets across 37 tissues or cell lines
-
Has reproduction · 67
Generative and integrative modeling for transcriptomics with formalin fixed paraffin embedded material.
PMID 41029822 · PMC12486589 · Journal of translational medicine · 2025 · 8 claims · 6 setups
The negative binomial distribution best fits fRNA-seq transcript counts, with little evidence supporting zero-inflated extensions
-
Full-text index only
Water mass specific genes dominate the Southern Ocean microbiome.
PMID 41803086 · PMC12972064 · Nature communications · 2026 · 8 claims · 8 setups
The Southern Ocean microbial gene catalog is highly original and largely distinct from existing marine gene catalogs