Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Has reproduction · 89
DFAST and DAGA: web-based integrated genome annotation tools and resources.
PMID 27867804 · PMC5107635 · Bioscience of microbiota, food and health · 2016 · 8 claims · 7 setups
DFAST is a web-based genome annotation pipeline with integrated quality assessment (CheckM) and taxonomic assessment (ANI) that produces DDBJ submission-ready files
-
Has reproduction · 50
MEDUSA: A Pipeline for Sensitive Taxonomic Classification and Flexible Functional Annotation of Metagenomic Shotgun Sequences.
PMID 35330728 · PMC8940201 · Frontiers in genetics · 2022 · 7 claims · 6 setups
MEDUSA correctly identifies more species than MEGAN 6 CE, especially less abundant species.
-
Has reproduction · 67
snpQT: flexible, reproducible, and comprehensive quality control and imputation of genomic data.
PMID 34900230 · PMC8637247 · F1000Research · 2021 · 8 claims · 4 setups
snpQT is a scalable, stand-alone software pipeline using nextflow and BioContainers for comprehensive, reproducible, interactive QC of human genomic data.
-
Full-text index only
NCBI Reference Sequence (RefSeq): a curated non-redundant sequence database of genomes, transcripts and proteins.
PMID 15608248 · PMC539979 · Nucleic acids research · 2005 · 7 claims · 5 setups
RefSeq provides a curated, non-redundant, explicitly linked collection of genomic, transcript and protein sequences spanning prokaryotes, eukaryotes and viruses.
-
Full-text index only
MoGAAAP: a modular Snakemake workflow for automated genome assembly and annotation with quality assessment.
PMID 41585413 · PMC12824462 · NAR genomics and bioinformatics · 2026 · 8 claims · 8 setups
MoGAAAP is a modular Snakemake pipeline that automates assembly, provisional annotation, and quality assessment (QA) for any diploid eukaryotic organism
-
Full-text index only
nf-core/viralmetagenome: A novel pipeline for untargeted viral genome reconstruction.
PMID 42057295 · PMC13141149 · Bioinformatics (Oxford, England) · 2026 · 8 claims · 5 setups
nf-core/viralmetagenome is a Nextflow pipeline that automates untargeted reconstruction and variant analysis of eukaryotic DNA and RNA viruses from short-read metagenomic or hybridisation-capture data.
-
Has reproduction · 95
nf-rnaSeqCount: A Nextflow pipeline for obtaining raw read counts from RNA-seq data.
PMID 35574063 · PMC9097006 · South African computer journal = Suid-Afrikaanse rekenaartydskrif · 2021 · 7 claims · 5 setups
nf-rnaSeqCount is a portable, reproducible Nextflow pipeline that maps RNA-seq reads to a reference genome and quantifies gene abundance for differential expression analysis
-
Has reproduction · 66
RiboTaxa: combined approaches for rRNA genes taxonomic resolution down to the species level from metagenomics data revealing novelties.
PMID 36159175 · PMC9492272 · NAR genomics and bioinformatics · 2022 · 8 claims · 6 setups
RiboTaxa, combining BBTools, FastQC, SortMeRNA, MetaRib, EMIRGE, VSEARCH, BBMap and QIIME 2's Sklearn classifier, was built as a pipeline for SSU rRNA-based taxonomic profiling of metagenomics data.
-
Full-text index only
umite: fast quantification of Smart-seq3 libraries with improved UMI retrieval.
PMID 41692984 · PMC12989134 · Bioinformatics (Oxford, England) · 2026 · 8 claims · 6 setups
umite offers efficient mismatch-tolerant (fuzzy) UMI detection that boosts UMI retrieval by 5%-15% compared to standard position-based matching
-
Full-text index only
TOFU-MAaPO: fast, scalable and reproducible analysis of large metagenome sequence data from the Sequence Read Archive.
PMID 42277027 · PMC13260335 · Nature communications · 2026 · 8 claims · 5 setups
TOFU-MAaPO yields significantly more high-quality MAGs than metaFun, nf-core/mag, and ATLAS due to integration of multiple complementary binning tools with unified MAGScoT refinement
-
Has reproduction · 59
WASP: a versatile, web-accessible single cell RNA-Seq processing platform.
PMID 33736596 · PMC7977290 · BMC genomics · 2021 · 7 claims · 7 setups
WASP is a software platform for processing Drop-Seq-based scRNA-seq data generated with ddSEQ or 10x protocols, combining a Snakemake pre-processing pipeline with an R Shiny post-processing application.
-
Full-text index only
BIPASS: BioInformatics Pipeline Alternative Splicing Services.
PMID 17584795 · PMC1933140 · Nucleic acids research · 2007 · 8 claims · 4 setups
BIPASS offers two complementary services for alternative splicing (AS) research: BIPAS-SpliceDB, a queryable pre-computed AS data warehouse, and BIPAS-Align&Splice, an online pipeline for user-submitted sequences.
-
Has reproduction · 85
An extensive evaluation of read trimming effects on Illumina NGS data analysis.
PMID 24376861 · PMC3871669 · PloS one · 2013 · 8 claims · 8 setups
Read trimming increases the quality and reliability of downstream NGS analyses (RNA-Seq mapping, SNP identification, genome assembly) while reducing execution time and computational resources.
-
Has reproduction · 80
VGEA: an RNA viral assembly toolkit.
PMID 34567846 · PMC8428259 · PeerJ · 2021 · 8 claims · 5 setups
VGEA is a Snakemake workflow that chains existing tools (fastp, BWA, SAMtools, IVA, shiver, SeqKit, QUAST, MultiQC) into an all-in-one RNA viral genome assembly pipeline
-
Full-text index only
Optimizing data-driven excellence: Canada's approach to using pathogen test datasets for quality control, pipeline development and training initiatives.
PMID 41591806 · PMC12847982 · Microbial genomics · 2026 · 8 claims · 5 setups
Standardized SARS-CoV-2 test datasets (Illumina and Nanopore) were developed as benchmarks for validating sequencing/bioinformatics pipelines across Canadian public health labs
-
Full-text index only
Metapipeline-DNA: A comprehensive germline and somatic genomics Nextflow pipeline.
PMID 41850291 · PMC13030954 · Cell reports methods · 2026 · 8 claims · 7 setups
Metapipeline-DNA automates germline and somatic DNA sequencing analysis end-to-end, from raw reads through preprocessing, feature detection, QC, and visualization.
-
Has reproduction · 100
A Bioinformatics Workflow to Identify eccDNA Using ECCFP From Long-Read Nanopore Sequencing Data.
PMID 41924242 · PMC13037781 · Bio-protocol · 2026 · 7 claims · 5 setups
ECCFP significantly improves eccDNA detection sensitivity, accuracy, and runtime efficiency compared to other pipelines
-
Has reproduction · 88
Human methylome variation across Infinium 450K data on the Gene Expression Omnibus.
PMID 33937763 · PMC8061458 · NAR genomics and bioinformatics · 2021 · 8 claims · 8 setups
Among annotated HM450K GEO samples, about two-thirds were from blood, one-quarter from brain, and about one-third were from cancer patients.
-
Has reproduction · 51
SGCP: a spectral self-learning method for clustering genes in co-expression networks.
PMID 38956463 · PMC11221046 · BMC bioinformatics · 2024 · 7 claims · 4 setups
SGCP, a spectral self-learning method, yields gene co-expression modules with higher GO enrichment than WGCNA, CoExpNets, and CEMiTool across 12 real gene expression datasets.
-
Has reproduction · 80
Curation of over 10 000 transcriptomic studies to enable data reuse.
PMID 33599246 · PMC7904053 · Database : the journal of biological databases and curation · 2021 · 8 claims · 6 setups
Gemma is a curated database and bioinformatics system that addresses metadata, probe annotation, and expression data inconsistencies in GEO to enable transcriptomic data reuse