Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Has reproduction · 90
pysradb: A Python package to query next-generation sequencing metadata and data from NCBI Sequence Read Archive.
PMID 31114675 · PMC6505635 · F1000Research · 2019 · 7 claims · 4 setups
pysradb provides a command-line interface to query metadata and download raw sequencing data from NCBI SRA using the SRAdb SQLite database.
-
Full-text index only
A metadata approach for clinical data management in translational genomics studies in breast cancer.
PMID 19948017 · PMC3225860 · BMC medical genomics · 2009 · 8 claims · 5 setups
A metadata/CDE-based approach using CancerGrid's semantic web tools enables automatic integration of heterogeneous clinical datasets without loss of original detail
-
Has reproduction · 64
GeMI: interactive interface for transformer-based Genomic Metadata Integration.
PMID 35657113 · PMC9216561 · Database : the journal of biological databases and curation · 2022 · 8 claims · 5 setups
GeMI is a web tool that uses a fine-tuned GPT2 model to extract 15 structured key-value attributes from free-text GEO sample metadata.
-
Has reproduction · 100
poreCov-An Easy to Use, Fast, and Robust Workflow for SARS-CoV-2 Genome Reconstruction via Nanopore Sequencing.
PMID 34394197 · PMC8355734 · Frontiers in genetics · 2021 · 8 claims · 8 setups
poreCov is an easy-to-use, fast, and robust Nextflow-based workflow for reference-based SARS-CoV-2 genome reconstruction and lineage determination from nanopore sequencing data
-
Has reproduction · 99
The systematic assessment of completeness of public metadata accompanying omics studies in the Gene Expression Omnibus data repository.
PMID 40926267 · PMC12421755 · Genome biology · 2025 · 8 claims · 3 setups
Across 253 manually curated studies (164,000+ samples), over 25% of critical metadata are omitted, with only 74.8% of relevant phenotypes available overall.
-
Has reproduction
All of gene expression (AOE): An integrated index for public gene expression databases.
PMID 31978081 · PMC6980531 · PloS one · 2020 · 6 claims · 4 setups
ArrayExpress (AE) discontinued importing data from GEO in 2017, causing new GEO entries to become unavailable from AE
-
Has reproduction · 75
An informatics research platform to make public gene expression time-course datasets reusable for more scientific discoveries.
PMID 33247935 · PMC7698665 · Database : the journal of biological databases and curation · 2020 · 8 claims · 6 setups
GETc enables discovery and visualization of time-course gene expression data and analytical results from GEO
-
Has reproduction · 64
Celline: a flexible tool for one-step retrieval and integrative analysis of public single-cell RNA sequencing data.
PMID 41458999 · PMC12738925 · Frontiers in bioinformatics · 2025 · 8 claims · 6 setups
Celline is a Python package that automates the full scRNA-seq workflow (retrieval, metadata extraction, preprocessing, cell-type annotation, batch correction, trajectory inference) via single-line commands.
-
Has reproduction · 80
Curation of over 10 000 transcriptomic studies to enable data reuse.
PMID 33599246 · PMC7904053 · Database : the journal of biological databases and curation · 2021 · 8 claims · 6 setups
Gemma is a curated database and bioinformatics system that addresses metadata, probe annotation, and expression data inconsistencies in GEO to enable transcriptomic data reuse
-
Has reproduction · 78
Fungal metabarcoding data integration framework for the MycoDiversity DataBase (MDDB).
PMID 32463383 · PMC7734503 · Journal of integrative bioinformatics · 2020 · 7 claims · 4 setups
Public fungal metabarcoding raw DNA data and their associated environmental metadata in sequence archives are heterogeneously annotated and lack a uniform processing pipeline, preventing large-scale biodiversity/distribution assessments.
-
Has reproduction · 83
Current status of use of high throughput nucleotide sequencing in rheumatology.
PMID 33408124 · PMC7789458 · RMD open · 2021 · 8 claims · 8 setups
RNA-Seq is the most represented HTS assay used in rheumatology research, primarily for biomarker identification in blood or synovial tissue.
-
Has reproduction · 99
getSequenceInfo: a suite of tools allowing to get genome sequence information from public repositories.
PMID 35804320 · PMC9264741 · BMC bioinformatics · 2022 · 8 claims · 8 setups
getSequenceInfo (gSeqI) allows programmatic (CLI) or GUI-based retrieval of sequence data and metadata from GenBank, RefSeq, and ENA across Linux, MacOS, and Windows.
-
Has reproduction · 100
Shiny-Calorie: a context-aware application for indirect calorimetry data analysis and visualization using R.
PMID 41640623 · PMC12867577 · Bioinformatics advances · 2026 · 8 claims · 8 setups
Shiny-Calorie is an open-source interactive application for transparent data and metadata integration, statistical analysis, and visualization of indirect calorimetry datasets.
-
Full-text index only
The Genomes On Line Database (GOLD) in 2009: status of genomic and metagenomic projects and their associated metadata.
PMID 19914934 · PMC2808860 · Nucleic acids research · 2010 · 8 claims · 5 setups
GOLD is a comprehensive, centralized resource for tracking genome and metagenome sequencing projects and their associated metadata worldwide.
-
Has reproduction · 95
The archives are half-empty: an assessment of the availability of microbial community sequencing data.
PMID 32859925 · PMC7455719 · Communications biology · 2020 · 8 claims · 6 setups
A large proportion of 16S rRNA amplicon sequencing studies contain data that is not available or not reusable despite being reported as deposited.
-
Has reproduction · 88
Human methylome variation across Infinium 450K data on the Gene Expression Omnibus.
PMID 33937763 · PMC8061458 · NAR genomics and bioinformatics · 2021 · 8 claims · 8 setups
Among annotated HM450K GEO samples, about two-thirds were from blood, one-quarter from brain, and about one-third were from cancer patients.
-
Has reproduction · 57
KARAJ: An Efficient Adaptive Multi-Processor Tool to Streamline Genomic and Transcriptomic Sequence Data Acquisition.
PMID 36430895 · PMC9694301 · International journal of molecular sciences · 2022 · 8 claims · 6 setups
KARAJ automates end-to-end querying and downloading of genomic/transcriptomic sequence data from a list of PMCIDs, URLs, or accession numbers
-
Full-text index only
Human-scATAC-Corpus: a comprehensive database of scATAC-seq data.
PMID 41296545 · PMC12807747 · Nucleic acids research · 2026 · 8 claims · 6 setups
Human-scATAC-Corpus is a comprehensive database of human scATAC-seq data comprising 5,407,621 cells from 35 datasets across 37 tissues or cell lines
-
Has reproduction · 38
Genomic capacities for Reactive Oxygen Species metabolism across marine phytoplankton.
PMID 37098087 · PMC10128935 · PloS one · 2023 · 8 claims · 3 setups
Genes encoding superoxide (O2•−) scavenging are ubiquitous across phytoplankton, but their fractional gene allocation decreases with increasing cell radius, consistent with a nearly fixed core gene set.
-
Has reproduction · 75
PHA4GE quality control contextual data tags: standardized annotations for sharing public health sequence datasets with known quality issues to facilitate testing and training.
PMID 38860884 · PMC11261899 · Microbial genomics · 2024 · 7 claims · 4 setups
No standardized attributes or mechanisms existed for tagging poor-quality or purpose-specific pathogen sequence datasets prior to this work