Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
Umi-pipeline-nf: a modular and scalable workflow for UMI-tagged nanopore amplicon analysis with real-time sequencing integration and GPU-acceleration.
PMID 41923360 · PMC13070649 · Bioinformatics (Oxford, England) · 2026 · 8 claims · 6 setups
umi-pipeline-nf is a portable, fully containerized, modular Nextflow DSL2 workflow that generates single-molecule consensus sequences from UMI-tagged nanopore amplicon data and scales linearly from single samples to large cohorts.
-
Has reproduction · 76
Organelle Genomes and Transcriptomes of Nymphaea Reveal the Interplay between Intron Splicing and RNA Editing.
PMID 34576004 · PMC8466565 · International journal of molecular sciences · 2021 · 8 claims · 8 setups
Both cis- and trans-splicing group II introns in Nymphaea organelle genomes are spliced in random order, generating diverse co-existing intermediates rather than following a fixed splicing sequence.
-
Full-text index only
Optimizing data-driven excellence: Canada's approach to using pathogen test datasets for quality control, pipeline development and training initiatives.
PMID 41591806 · PMC12847982 · Microbial genomics · 2026 · 8 claims · 5 setups
Standardized SARS-CoV-2 test datasets (Illumina and Nanopore) were developed as benchmarks for validating sequencing/bioinformatics pipelines across Canadian public health labs
-
Has reproduction · 100
A workflow reproducibility scale for automatic validation of biological interpretation results.
PMID 37150537 · PMC10164546 · GigaScience · 2022 · 8 claims · 4 setups
Comparing output files by checksum alone is insufficient to verify reproducibility, since checksums can differ even when the underlying biological interpretation is unchanged
-
Full-text index only
Variation analysis and gene annotation of eight MHC haplotypes: the MHC Haplotype Project.
PMID 18193213 · PMC2206249 · Immunogenetics · 2008 · 8 claims · 6 setups
Comparison of eight HLA-homozygous MHC haplotype sequences identified >44,000 variations (substitutions and indels), submitted to dbSNP
-
Full-text index only
A comparison of random sequence reads versus 16S rDNA sequences for estimating the biodiversity of a metagenomic library.
PMID 18682527 · PMC2532719 · Nucleic acids research · 2008 · 8 claims · 7 setups
Biodiversity observed by RSR analysis is consistent with that obtained by 16S rDNA analysis
-
Has reproduction · 89
DFAST and DAGA: web-based integrated genome annotation tools and resources.
PMID 27867804 · PMC5107635 · Bioscience of microbiota, food and health · 2016 · 8 claims · 7 setups
DFAST is a web-based genome annotation pipeline with integrated quality assessment (CheckM) and taxonomic assessment (ANI) that produces DDBJ submission-ready files
-
Has reproduction · 50
MEDUSA: A Pipeline for Sensitive Taxonomic Classification and Flexible Functional Annotation of Metagenomic Shotgun Sequences.
PMID 35330728 · PMC8940201 · Frontiers in genetics · 2022 · 7 claims · 6 setups
MEDUSA correctly identifies more species than MEGAN 6 CE, especially less abundant species.
-
Has reproduction · 95
nf-rnaSeqCount: A Nextflow pipeline for obtaining raw read counts from RNA-seq data.
PMID 35574063 · PMC9097006 · South African computer journal = Suid-Afrikaanse rekenaartydskrif · 2021 · 7 claims · 5 setups
nf-rnaSeqCount is a portable, reproducible Nextflow pipeline that maps RNA-seq reads to a reference genome and quantifies gene abundance for differential expression analysis
-
Full-text index only
rMAP 2.0: a modular, reproducible, and scalable WDL-Cromwell-Docker workflow for genomic analysis of ESKAPEE pathogens.
PMID 41782684 · PMC12955837 · Bioinformatics advances · 2026 · 8 claims · 8 setups
rMAP 2.0 standardizes end-to-end bacterial WGS analysis (QC, trimming, assembly, annotation, AMR/virulence/mobile-element profiling, sequence typing, pangenome inference, phylogenetics) via containerized WDL/Cromwell execution
-
Has reproduction · 100
FA-nf: A Functional Annotation Pipeline for Proteins from Non-Model Organisms Implemented in Nextflow.
PMID 34681040 · PMC8535801 · Genes · 2021 · 8 claims · 4 setups
FA-nf, implemented in Nextflow with Docker/Singularity containerization, integrates NCBI BLAST+, DIAMOND, InterProScan, and KEGG (KAAS/KofamKOALA) into a single functional annotation pipeline.
-
Has reproduction · 63
hgtseq: A Standard Pipeline to Study Horizontal Gene Transfer.
PMID 36498841 · PMC9738810 · International journal of molecular sciences · 2022 · 8 claims · 8 setups
hgtseq is a fully automated, portable, and scalable Nextflow/nf-core pipeline for detecting horizontal gene transfer signatures from unmapped sequencing reads.
-
Full-text index only
Discovery of novel human transcript variants by analysis of intronic single-block EST with polyadenylation site.
PMID 19906316 · PMC2784480 · BMC genomics · 2009 · 8 claims · 7 setups
Intronic single-block ESTs with poly(A/T) tails reveal previously unidentified novel transcript variants missed by existing databases.
-
Full-text index only
Duplex-Indel: a Snakemake pipeline for somatic Indel calling in Tn5 transposase-based duplex sequencing data.
PMID 42046229 · PMC13171174 · Bioinformatics (Oxford, England) · 2026 · 8 claims · 8 setups
Duplex-Indel is a Snakemake pipeline for somatic Indel calling from Tn5 transposase-based duplex sequencing data that requires consensus support from both DNA strands to minimize technical artifacts.
-
Full-text index only
A compatible exon-exon junction database for the identification of exon skipping events using tandem mass spectrum data.
PMID 19087293 · PMC2636810 · BMC bioinformatics · 2008 · 6 claims · 6 setups
A theoretical exon-exon junction protein database accounting for all in-phase (frame-preserving) exon combinations can be built from the Ensembl Core Database using Perl/Bioperl/MySQL/Ensembl API.
-
Full-text index only
MobiCT: a UMI-based circulating tumor DNA analysis pipeline.
PMID 41503160 · PMC12770973 · NAR genomics and bioinformatics · 2026 · 7 claims · 7 setups
MobiCT is a Nextflow/nf-core UMI-based ctDNA pipeline (deduplication, alignment, variant calling with VarDict, annotation with VEP) achieving sensitivity, precision, and F1-score around 90% after comprehensive filtering.
-
Full-text index only
Optimized library preparation, sequencing, and data analysis protocols for the generation of orbivirus consensus sequences.
PMID 41527034 · PMC12809950 · BMC genomics · 2026 · 8 claims · 8 setups
Optimized sample and library preparation protocols achieved comparable results to established methods while requiring simpler sample preparation.
-
Full-text index only
Application of qualifying variants for genomic analysis.
PMID 41570118 · PMC12926777 · Bioinformatics (Oxford, England) · 2026 · 7 claims · 4 setups
QVs should be treated as dynamic, multifaceted elements permeating the entire analysis workflow, not as a single static filtering step
-
Full-text index only
ANOMALY: a Snakemake pipeline for identifying NuMTs from long-read sequencing data.
PMID 41647924 · PMC12869244 · NAR genomics and bioinformatics · 2026 · 8 claims · 8 setups
ANOMALY is a novel Snakemake pipeline for detecting NuMTs from long-read sequencing data
-
Has reproduction · 58
iCOMIC: a graphical interface-driven bioinformatics pipeline for analyzing cancer omics data.
PMID 35899080 · PMC9310080 · NAR genomics and bioinformatics · 2022 · 8 claims · 4 setups
iCOMIC provides a GUI-driven, Snakemake-based pipeline integrating multiple tools for DNA-Seq and RNA-Seq analysis with minimal command-line interaction.