Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
Phylogenomic approaches to common problems encountered in the analysis of low copy repeats: the sulfotransferase 1A gene family example.
PMID 15752422 · PMC555591 · BMC evolutionary biology · 2005 · 8 claims · 8 setups
A previously unidentified fourth human SULT1A gene (SULT1A4) exists on chromosome 16 and is transcriptionally active
-
Full-text index only
CEAS: cis-regulatory element annotation system.
PMID 16845068 · PMC1538818 · Nucleic acids research · 2006 · 7 claims · 5 setups
CEAS is the first web server to streamline genome-scale ChIP-chip downstream analyses for biologists without strong bioinformatics support
-
Has reproduction · 84
Pharokka: a fast scalable bacteriophage annotation tool.
PMID 36453861 · PMC9805569 · Bioinformatics (Oxford, England) · 2023 · 8 claims · 5 setups
Pharokka is a one-line, fast, scalable bacteriophage annotation tool producing standards-compliant outputs, installable via a two-line bioconda command
-
Full-text index only
DupyliCate: mining, classifying, and characterizing gene duplications.
PMID 42209743 · PMC13219399 · Scientific reports · 2026 · 8 claims · 8 setups
DupyliCate is a Python tool for identifying and classifying gene duplication arrays, using BUSCO-based species-specific thresholds and offering integrated expression divergence and Ka/Ks analysis.
-
Has reproduction · 67
binny: an automated binning algorithm to recover high-quality genomes from complex metagenomic datasets.
PMID 36239393 · PMC9677464 · Briefings in bioinformatics · 2022 · 8 claims · 8 setups
binny outperforms or is highly competitive with commonly used and state-of-the-art binning methods (MetaBAT2, MaxBin2, CONCOCT, VAMB, SemiBin, MetaDecoder)
-
Full-text index only
Iterative pruning PCA improves resolution of highly structured populations.
PMID 19930644 · PMC2790469 · BMC bioinformatics · 2009 · 7 claims · 7 setups
ipPCA is a novel algorithm that assigns individuals to subpopulations and infers the total number of subpopulations (K) present in genotypic data
-
Full-text index only
Evolution of two distinct phylogenetic lineages of the emerging human pathogen Mycobacterium ulcerans.
PMID 17900363 · PMC2098775 · BMC evolutionary biology · 2007 · 8 claims · 4 setups
M. ulcerans has evolved into five InDel haplotypes that separate into two distinct phylogenetic lineages: a 'classical' lineage (Africa, Australia, South East Asia) and an 'ancestral' lineage (Asia, South America, Mexico)
-
Full-text index only
Performance of methods to detect genetic variants from bisulphite sequencing data in a non-model species.
PMID 34435438 · PMC9290141 · Molecular ecology resources · 2022 · 6 claims · 6 setups
Bisulphite conversion of unmethylated cytosines to thymines violates strand-complementarity assumptions of SNP callers and confounds true C->T SNPs with unmethylated cytosines, complicating SNP calling from bisulphite sequencing data.
-
Has reproduction · 55
Genome-Wide Survey and Development of the First Microsatellite Markers Database (AnCorDB) in Anemone coronaria L.
PMID 35328546 · PMC8949970 · International journal of molecular sciences · 2022 · 8 claims · 8 setups
Generated the first draft genome assembly of A. coronaria by Illumina sequencing a haploid androgenetic plant
-
Has reproduction · 78
GenTB: A user-friendly genome-based predictor for tuberculosis resistance powered by machine learning.
PMID 34461978 · PMC8407037 · Genome medicine · 2021 · 8 claims · 6 setups
GenTB is a free, open, web-based application offering two ML predictors (Random Forest and WDNN) that predict resistance to 13 and 10 anti-TB drugs, respectively.
-
Has reproduction · 42
KAGE: fast alignment-free graph-based genotyping of SNPs and short indels.
PMID 36195962 · PMC9531401 · Genome biology · 2022 · 7 claims · 7 setups
KAGE combines population-based kmer count modeling with single-variant prior adjustment into an alignment-free genotyper that matches the accuracy of the best existing alignment-free genotypers while being an order of magnitude faster.
-
Full-text index only
Multi-context seeds enable fast and high-accuracy read mapping.
PMID 41764549 · PMC13059148 · Genome biology · 2026 · 7 claims · 5 setups
Multi-context seeds (MCS) allow storage of seeds with different lengths in the same index structure by splitting hash bits among strobes, enabling full and partial matches
-
Has reproduction · 74
MicroPIPE: validating an end-to-end workflow for high-quality complete bacterial genome construction.
PMID 34172000 · PMC8235852 · BMC genomics · 2021 · 8 claims · 8 setups
MicroPIPE, an end-to-end Nextflow/Singularity-based pipeline built from systematically validated tool choices, produces high-quality complete bacterial genome assemblies without manual intervention.
-
Has reproduction · 78
QuasiFlow: a Nextflow pipeline for analysis of NGS-based HIV-1 drug resistance data.
PMID 36699347 · PMC9722223 · Bioinformatics advances · 2022 · 6 claims · 8 setups
QuasiFlow is a Nextflow pipeline that runs entirely locally via command-line tools and a local HIVdb database copy to analyze NGS-based HIV-1 drug resistance testing data.
-
Full-text index only
Benchmarking methods for genome annotation using nanopore direct RNA in a non-model crop plant.
PMID 41800382 · PMC12967217 · Bioinformatics advances · 2026 · 6 claims · 8 setups
Annotation tools show substantial variation in isoform detection, structural completeness, splicing classification, and handling of 5' read truncation when applied to plant dRNA-seq data.
-
Full-text index only
Distinguishing benign from pathogenic duplications involving GPR101 and VGLL1-adjacent enhancers in the clinical setting with the bioinformatic tool POSTRE.
PMID 41540017 · PMC12890961 · NPJ genomic medicine · 2026 · 6 claims · 7 setups
POSTRE correctly classified all 34 GPR101-associated duplications (27 pathogenic X-LAG, 7 non-pathogenic) as pathogenic or benign
-
Full-text index only
EGenBio: a data management system for evolutionary genomics and biodiversity.
PMID 17118150 · PMC1683573 · BMC bioinformatics · 2006 · 7 claims · 7 setups
EGenBio is a web-based system for integrated management, filtering, curation, and visualization of large-scale genomic sequences, alignments, and phylogenetic trees for evolutionary genomics and biodiversity research.
-
Has reproduction · 80
VGEA: an RNA viral assembly toolkit.
PMID 34567846 · PMC8428259 · PeerJ · 2021 · 8 claims · 5 setups
VGEA is a Snakemake workflow that chains existing tools (fastp, BWA, SAMtools, IVA, shiver, SeqKit, QUAST, MultiQC) into an all-in-one RNA viral genome assembly pipeline
-
Has reproduction · 90
Optimal Dual RNA-Seq Mapping for Accurate Pathogen Detection in Complex Eukaryotic Hosts.
PMID 39959292 · PMC11825298 · Bio-protocol · 2025 · 7 claims · 6 setups
Mapping adapter-trimmed reads first to the pathogen genome recovers more pathogen reads than the traditional host-first mapping approach.
-
Has reproduction · 85
An extensive evaluation of read trimming effects on Illumina NGS data analysis.
PMID 24376861 · PMC3871669 · PloS one · 2013 · 8 claims · 8 setups
Read trimming increases the quality and reliability of downstream NGS analyses (RNA-Seq mapping, SNP identification, genome assembly) while reducing execution time and computational resources.