Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
Common pathogenic mechanisms in the hippocampus across neurodegenerative dementias: Alzheimer's disease, Down syndrome, and Parkinson's disease.
PMID 42078116 · PMC13128439 · NPJ dementia · 2026 · 8 claims · 6 setups
Chronological age does not align with biological (transcriptional) age in hippocampal tissue from PDD, AD, and DSD dementia patients
-
Has reproduction · 90
pysradb: A Python package to query next-generation sequencing metadata and data from NCBI Sequence Read Archive.
PMID 31114675 · PMC6505635 · F1000Research · 2019 · 7 claims · 4 setups
pysradb provides a command-line interface to query metadata and download raw sequencing data from NCBI SRA using the SRAdb SQLite database.
-
Has reproduction · 64
GeMI: interactive interface for transformer-based Genomic Metadata Integration.
PMID 35657113 · PMC9216561 · Database : the journal of biological databases and curation · 2022 · 8 claims · 5 setups
GeMI is a web tool that uses a fine-tuned GPT2 model to extract 15 structured key-value attributes from free-text GEO sample metadata.
-
Has reproduction · 64
Celline: a flexible tool for one-step retrieval and integrative analysis of public single-cell RNA sequencing data.
PMID 41458999 · PMC12738925 · Frontiers in bioinformatics · 2025 · 8 claims · 6 setups
Celline is a Python package that automates the full scRNA-seq workflow (retrieval, metadata extraction, preprocessing, cell-type annotation, batch correction, trajectory inference) via single-line commands.
-
Has reproduction · 83
Current status of use of high throughput nucleotide sequencing in rheumatology.
PMID 33408124 · PMC7789458 · RMD open · 2021 · 8 claims · 8 setups
RNA-Seq is the most represented HTS assay used in rheumatology research, primarily for biomarker identification in blood or synovial tissue.
-
Full-text index only
Metappuccino: large language model-driven reconstruction of sequence read archive metadata for cancer research.
PMID 42057294 · PMC13148957 · Bioinformatics (Oxford, England) · 2026 · 8 claims · 3 setups
Metappuccino reconstructs 19 metadata classes by combining deterministic rule-based extraction/normalization (for explicit context) with LoRA-specialized Mistral-7B-Instruct completion (for missing/ambiguous fields)
-
Full-text index only
Human-scATAC-Corpus: a comprehensive database of scATAC-seq data.
PMID 41296545 · PMC12807747 · Nucleic acids research · 2026 · 8 claims · 6 setups
Human-scATAC-Corpus is a comprehensive database of human scATAC-seq data comprising 5,407,621 cells from 35 datasets across 37 tissues or cell lines
-
Has reproduction · 67
Generative and integrative modeling for transcriptomics with formalin fixed paraffin embedded material.
PMID 41029822 · PMC12486589 · Journal of translational medicine · 2025 · 8 claims · 6 setups
The negative binomial distribution best fits fRNA-seq transcript counts, with little evidence supporting zero-inflated extensions
-
Full-text index only
Human endogenous retrovirus profiling reveals heterogenous expression in cutaneous melanoma.
PMID 41971443 · PMC13061668 · Frontiers in oncology · 2026 · 7 claims · 8 setups
HERV expression differs between primary and metastatic cutaneous melanoma
-
Has reproduction · 69
COVID-19 vaccination atlas using an integrative systems vaccinology approach.
PMID 40456760 · PMC12130191 · NPJ vaccines · 2025 · 8 claims · 6 setups
mRNA vaccines induce transient but strong immune responses after booster doses
-
Has reproduction · 80
Curation of over 10 000 transcriptomic studies to enable data reuse.
PMID 33599246 · PMC7904053 · Database : the journal of biological databases and curation · 2021 · 8 claims · 6 setups
Gemma is a curated database and bioinformatics system that addresses metadata, probe annotation, and expression data inconsistencies in GEO to enable transcriptomic data reuse