Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
TOFU-MAaPO: fast, scalable and reproducible analysis of large metagenome sequence data from the Sequence Read Archive.
PMID 42277027 · PMC13260335 · Nature communications · 2026 · 8 claims · 5 setups
TOFU-MAaPO yields significantly more high-quality MAGs than metaFun, nf-core/mag, and ATLAS due to integration of multiple complementary binning tools with unified MAGScoT refinement
-
Full-text index only
Frag'n'Flow: automated workflow for large-scale quantitative proteomics in high performance computing environments.
PMID 41486154 · PMC12828970 · BMC bioinformatics · 2026 · 8 claims · 8 setups
Frag'n'Flow is a Nextflow-based pipeline that encapsulates FragPipe, automating manifest/workflow generation, tool dependency management, and downstream analysis for HPC/cloud/cluster environments.
-
Full-text index only
umite: fast quantification of Smart-seq3 libraries with improved UMI retrieval.
PMID 41692984 · PMC12989134 · Bioinformatics (Oxford, England) · 2026 · 8 claims · 6 setups
umite offers efficient mismatch-tolerant (fuzzy) UMI detection that boosts UMI retrieval by 5%-15% compared to standard position-based matching
-
Has reproduction · 83
Hobbes: optimized gram-based methods for efficient read alignment.
PMID 22199254 · PMC3315303 · Nucleic acids research · 2012 · 8 claims · 4 setups
Hobbes, a gram-based short-read mapper supporting Hamming and edit distance, is faster than all other read-mapping programs tested while maintaining high mapping quality.
-
Has reproduction · 84
Pharokka: a fast scalable bacteriophage annotation tool.
PMID 36453861 · PMC9805569 · Bioinformatics (Oxford, England) · 2023 · 8 claims · 5 setups
Pharokka is a one-line, fast, scalable bacteriophage annotation tool producing standards-compliant outputs, installable via a two-line bioconda command
-
Has reproduction · 100
poreCov-An Easy to Use, Fast, and Robust Workflow for SARS-CoV-2 Genome Reconstruction via Nanopore Sequencing.
PMID 34394197 · PMC8355734 · Frontiers in genetics · 2021 · 8 claims · 8 setups
poreCov is an easy-to-use, fast, and robust Nextflow-based workflow for reference-based SARS-CoV-2 genome reconstruction and lineage determination from nanopore sequencing data
-
Full-text index only
Scalable nonparametric clustering with unified marker gene selection for single-cell RNA-seq data.
PMID 41825449 · PMC13030991 · Cell reports methods · 2026 · 7 claims · 3 setups
NCLUSION matches the performance of state-of-the-art single-cell clustering techniques with significantly reduced runtime
-
Has reproduction · 71
Hyb: a bioinformatics pipeline for the analysis of CLASH (crosslinking, ligation and sequencing of hybrids) data.
PMID 24211736 · PMC3969109 · Methods (San Diego, Calif.) · 2014 · 8 claims · 6 setups
The 'hyb' pipeline detects, calls, folds and annotates chimeric reads from CLASH high-throughput sequencing data.
-
Has reproduction · 79
RetroSnake: A modular pipeline to detect human endogenous retroviruses in genome sequencing data.
PMID 36339261 · PMC9626663 · iScience · 2022 · 8 claims · 4 setups
RetroSnake is an end-to-end, modular, computationally efficient Snakemake pipeline for detecting HERV-K insertions in short-read NGS data, from raw alignment files to an annotated interactive HTML report
-
Has reproduction · 78
QuasiFlow: a Nextflow pipeline for analysis of NGS-based HIV-1 drug resistance data.
PMID 36699347 · PMC9722223 · Bioinformatics advances · 2022 · 6 claims · 8 setups
QuasiFlow is a Nextflow pipeline that runs entirely locally via command-line tools and a local HIVdb database copy to analyze NGS-based HIV-1 drug resistance testing data.
-
Full-text index only
The use of edge-betweenness clustering to investigate biological function in protein interaction networks.
PMID 15740614 · PMC555937 · BMC bioinformatics · 2005 · 8 claims · 7 setups
Edge-Betweenness clustering separates protein interaction graphs into subgraphs whose GO term distributions show significant correlations, revealing biologically meaningful functional modules.
-
Full-text index only
Predicting the protein interaction landscape of a free-living bacterium with pooled-AlphaFold3.
PMID 41559189 · PMC13047044 · Molecular systems biology · 2026 · 8 claims · 6 setups
Pooled-AlphaFold3 prediction improves accuracy of genome-scale PPI screens compared to a paired approach while reducing inference time (~2-fold) and job count (~100-fold)
-
Full-text index only
FLASH-MM: fast and scalable single-cell differential expression analysis using linear mixed-effects models.
PMID 41644528 · PMC12982622 · Nature communications · 2026 · 8 claims · 6 setups
FLASH-MM produces LMM parameter estimates identical to lmer (lme4) up to the sixth decimal place while being 50- to 140-fold faster as sample size increases from 20,000 to 120,000 cells
-
Full-text index only
Reliable Inference of Phylogenomic Relationship via Assembly-Based Strategy Accommodating Raw Reads and Proteins.
PMID 41800729 · PMC12969758 · Molecular ecology resources · 2026 · 7 claims · 8 setups
VEHoP infers protein-coding regions from diverse input types (raw reads, draft genomes, transcriptomes, annotated genomes) and automates generation of orthologous alignments, concatenated supermatrices, and phylogenetic trees in a single pipeline run.
-
Full-text index only
SGCRNA: spectral clustering-guided co-expression network analysis without scale-free constraints for multi-omic data.
PMID 41615289 · PMC12856952 · Briefings in bioinformatics · 2026 · 8 claims · 8 setups
WGCNA's reliance on a scale-free topology assumption is problematic because real co-expression networks do not consistently exhibit scale-free properties
-
Full-text index only
iMapper: a web application for the automated analysis and mapping of insertional mutagenesis sequence data against Ensembl genomes.
PMID 18974167 · PMC2639305 · Bioinformatics (Oxford, England) · 2008 · 6 claims · 3 setups
iMapper is a web application for automated analysis and mapping of insertional mutagenesis sequence data against vertebrate and invertebrate Ensembl genomes (human, mouse, rat, zebrafish, Drosophila, S. cerevisiae).
-
Has reproduction · 73
Genetic polyploid phasing from low-depth progeny samples.
PMID 35692633 · PMC9184567 · iScience · 2022 · 8 claims · 7 setups
WH-PPG phases polyploid parental samples by scoring informative variant pairs with a Bayesian log-likelihood model of progeny allele depths, clustering alleles by co-occurrence likelihood, and assigning clusters to haplotypes via interval scheduling
-
Full-text index only
Metapipeline-DNA: A comprehensive germline and somatic genomics Nextflow pipeline.
PMID 41850291 · PMC13030954 · Cell reports methods · 2026 · 8 claims · 7 setups
Metapipeline-DNA automates germline and somatic DNA sequencing analysis end-to-end, from raw reads through preprocessing, feature detection, QC, and visualization.
-
Full-text index only
PeakPrime: a peak-guided primer design pipeline for target enrichment in 3'-end RNA-seq.
PMID 41919010 · PMC13034549 · Bioinformatics advances · 2026 · 8 claims · 7 setups
PeakPrime is a reproducible Nextflow pipeline that calls 3′ RNA-seq coverage peaks (MACS2), selects exonic windows, designs strand-appropriate primers (Primer3), and screens specificity (Bowtie2)
-
Full-text index only
Robust and efficient annotation of cell states through gene signature scoring.
PMID 41708334 · PMC12951948 · Genome research · 2026 · 8 claims · 8 setups
Established scoring methods (Seurat, SCANPY, UCell, JASMINE) fail to provide robust and comparable score distributions across diverse signatures and experimental conditions, precluding accurate unsupervised cell-state annotation.