Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Has reproduction · 92
A network-guided protocol to discover susceptibility genes in genome-wide association studies using stability selection.
PMID 36609152 · PMC9850185 · STAR protocols · 2023 · 5 claims · 5 setups
The protocol identifies genes that are both statistically associated with a phenotype and functionally interconnected in a biological network
-
Full-text index only
SpliceMiner: a high-throughput database implementation of the NCBI Evidence Viewer for microarray splice variant analysis.
PMID 17338820 · PMC1839109 · BMC bioinformatics · 2007 · 6 claims · 4 setups
EVDB is a comprehensive, non-redundant relational database of known human splice variants built from NCBI Entrez Gene and Evidence Viewer data
-
Full-text index only
An integrated database-pipeline system for studying single nucleotide polymorphisms and diseases.
PMID 19091018 · PMC2638159 · BMC bioinformatics · 2008 · 6 claims · 5 setups
Existing SNP/disease databases are fragmented; no combined resource widely supports gene-, SNP-, and disease-related information together
-
Full-text index only
A compatible exon-exon junction database for the identification of exon skipping events using tandem mass spectrum data.
PMID 19087293 · PMC2636810 · BMC bioinformatics · 2008 · 6 claims · 6 setups
A theoretical exon-exon junction protein database accounting for all in-phase (frame-preserving) exon combinations can be built from the Ensembl Core Database using Perl/Bioperl/MySQL/Ensembl API.
-
Full-text index only
ANOMALY: a Snakemake pipeline for identifying NuMTs from long-read sequencing data.
PMID 41647924 · PMC12869244 · NAR genomics and bioinformatics · 2026 · 8 claims · 8 setups
ANOMALY is a novel Snakemake pipeline for detecting NuMTs from long-read sequencing data
-
Full-text index only
rMAP 2.0: a modular, reproducible, and scalable WDL-Cromwell-Docker workflow for genomic analysis of ESKAPEE pathogens.
PMID 41782684 · PMC12955837 · Bioinformatics advances · 2026 · 8 claims · 8 setups
rMAP 2.0 standardizes end-to-end bacterial WGS analysis (QC, trimming, assembly, annotation, AMR/virulence/mobile-element profiling, sequence typing, pangenome inference, phylogenetics) via containerized WDL/Cromwell execution
-
Full-text index only
Eduomics: a Nextflow pipeline to simulate -omics data for education.
PMID 41816779 · PMC12972896 · NAR genomics and bioinformatics · 2026 · 8 claims · 4 setups
Eduomics is a Nextflow DSL2 pipeline that automates generation of validated variant-calling and RNA-seq datasets for education while abstracting away technical requirements
-
Full-text index only
MIMIC: a flexible pipeline to register and summarize IMC-MSI experiments.
PMID 41917425 · PMC13201759 · Communications biology · 2026 · 7 claims · 6 setups
MIMIC is a reproducible, semi-automated workflow that co-registers and jointly analyzes MALDI-MSI and IMC data using a chain of before/after microscopy images.
-
Full-text index only
Duplex-Indel: a Snakemake pipeline for somatic Indel calling in Tn5 transposase-based duplex sequencing data.
PMID 42046229 · PMC13171174 · Bioinformatics (Oxford, England) · 2026 · 8 claims · 8 setups
Duplex-Indel is a Snakemake pipeline for somatic Indel calling from Tn5 transposase-based duplex sequencing data that requires consensus support from both DNA strands to minimize technical artifacts.
-
Full-text index only
Target SNP selection in complex disease association studies.
PMID 15248903 · PMC487897 · BMC bioinformatics · 2004 · 7 claims · 3 setups
A computational pipeline can retrieve gene sequence, collect SNP variation data, and annotate SNPs falling in functional motifs (promoter, exon-intron structure, AU-rich elements, TF binding sites, splice sites) with expression in target tissue
-
Has reproduction · 59
WASP: a versatile, web-accessible single cell RNA-Seq processing platform.
PMID 33736596 · PMC7977290 · BMC genomics · 2021 · 7 claims · 7 setups
WASP is a software platform for processing Drop-Seq-based scRNA-seq data generated with ddSEQ or 10x protocols, combining a Snakemake pre-processing pipeline with an R Shiny post-processing application.
-
Has reproduction · 57
Design considerations for workflow management systems use in production genomics research and the clinic.
PMID 34737383 · PMC8569008 · Scientific reports · 2021 · 8 claims · 2 setups
The choice of WfMS depends both on intrinsic language/engine features and on adoption, collaboration, and technical support within bioinformatics consortia.
-
Full-text index only
Ensembl 2008.
PMID 18000006 · PMC2238821 · Nucleic acids research · 2008 · 8 claims · 6 setups
The Ensembl regulatory build integrates multiple genome-wide functional genomics datasets to automatically annotate regulatory regions and assign putative functions across the genome.
-
Full-text index only
Ensembl 2009.
PMID 19033362 · PMC2686571 · Nucleic acids research · 2009 · 8 claims · 6 setups
Ensembl provides comprehensive, consistently annotated genome information for chordate genomes with automatically generated genesets and comparative genomics data
-
Full-text index only
BIPASS: BioInformatics Pipeline Alternative Splicing Services.
PMID 17584795 · PMC1933140 · Nucleic acids research · 2007 · 8 claims · 4 setups
BIPASS offers two complementary services for alternative splicing (AS) research: BIPAS-SpliceDB, a queryable pre-computed AS data warehouse, and BIPAS-Align&Splice, an online pipeline for user-submitted sequences.
-
Has reproduction · 80
Characterization of ALTO-encoding circular RNAs expressed by Merkel cell polyomavirus and trichodysplasia spinulosa polyomavirus.
PMID 33999949 · PMC8158866 · PLoS pathogens · 2021 · 8 claims · 8 setups
MCPyV generates two circular RNAs (circALTO1, circALTO2) spanning the early region/ALTO ORF, detectable in VP-MCC cell lines and patient tumors
-
Has reproduction · 67
snpQT: flexible, reproducible, and comprehensive quality control and imputation of genomic data.
PMID 34900230 · PMC8637247 · F1000Research · 2021 · 8 claims · 4 setups
snpQT is a scalable, stand-alone software pipeline using nextflow and BioContainers for comprehensive, reproducible, interactive QC of human genomic data.
-
Has reproduction · 58
iCOMIC: a graphical interface-driven bioinformatics pipeline for analyzing cancer omics data.
PMID 35899080 · PMC9310080 · NAR genomics and bioinformatics · 2022 · 8 claims · 4 setups
iCOMIC provides a GUI-driven, Snakemake-based pipeline integrating multiple tools for DNA-Seq and RNA-Seq analysis with minimal command-line interaction.
-
Has reproduction · 67
GAVISUNK: genome assembly validation via inter-SUNK distances in Oxford Nanopore reads.
PMID 36321867 · PMC9805576 · Bioinformatics (Oxford, England) · 2023 · 7 claims · 4 setups
GAVISUNK is an open-source pipeline that validates phased diploid HiFi assemblies by assessing concordance of inter-SUNK distances against orthogonal Oxford Nanopore (ONT) reads.
-
Full-text index only
Pseudofam: the pseudogene families database.
PMID 18957444 · PMC2686518 · Nucleic acids research · 2009 · 8 claims · 7 setups
Pseudofam is an online database of pseudogene families built by mapping pseudogenes to Pfam protein families, providing query tools, statistics, and sequence alignments