Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Has reproduction · 90
CONSULT: accurate contamination removal using locality-sensitive hashing.
PMID 34377979 · PMC8340999 · NAR genomics and bioinformatics · 2021 · 8 claims · 4 setups
CONSULT is a k-mer read-matching tool that uses locality-sensitive hashing (LSH) to allow inexact k-mer matches (within a user-defined Hamming distance) between query reads and a reference dataset.
-
Full-text index only
Automatic annotation of eukaryotic genes, pseudogenes and promoters.
PMID 16925832 · PMC1810547 · Genome biology · 2006 · 8 claims · 6 setups
Fgenesh++ gene prediction pipeline identifies 91% of coding nucleotides with 90% specificity
-
Full-text index only
MACSIMS: multiple alignment of complete sequences information management system.
PMID 16792820 · PMC1539025 · BMC bioinformatics · 2006 · 8 claims · 5 setups
MACSIMS is a multiple alignment-based information management system combining knowledge-based database mining with ab initio sequence predictions
-
Full-text index only
Performance assessment of promoter predictions on ENCODE regions in the EGASP experiment.
PMID 16925837 · PMC1810552 · Genome biology · 2006 · 6 claims · 3 setups
Promoter predictors that combine promoter prediction with gene prediction (N-SCAN, Fprom) achieve better performance than pure ab initio promoter predictors, mainly by reducing the promoter search space and false positives
-
Has reproduction · 29
MOSAIK: a hash-based algorithm for accurate next-generation sequencing short-read mapping.
PMID 24599324 · PMC3944147 · PloS one · 2014 · 8 claims · 8 setups
MOSAIK is the only aligner that consistently aligns reads from all major sequencing platforms (Illumina, AB SOLiD, Roche 454, Ion Torrent, Pacific Biosciences SMRT) using the same algorithmic approach.
-
Has reproduction · 100
poreCov-An Easy to Use, Fast, and Robust Workflow for SARS-CoV-2 Genome Reconstruction via Nanopore Sequencing.
PMID 34394197 · PMC8355734 · Frontiers in genetics · 2021 · 8 claims · 8 setups
poreCov is an easy-to-use, fast, and robust Nextflow-based workflow for reference-based SARS-CoV-2 genome reconstruction and lineage determination from nanopore sequencing data
-
Has reproduction · 89
TrEMOLO: accurate transposable element allele frequency estimation using long-read sequencing data combining assembly and mapping-based approaches.
PMID 37013657 · PMC10069131 · Genome biology · 2023 · 6 claims · 6 setups
TrEMOLO combines an assembly-based INSIDER module and a mapping-based OUTSIDER module to detect TE insertions/deletions from long-read sequencing data and estimate their allele frequency
-
Full-text index only
From endosymbiont to host-controlled organelle: the hijacking of mitochondrial protein synthesis and metabolism.
PMID 17983265 · PMC2062474 · PLoS computational biology · 2007 · 8 claims · 7 setups
There has been a large turnover of the mitochondrial proteome during evolution: cell envelope synthesis proteins virtually disappeared, and replication, transcription, cell division, transport, regulation, and signal transduction proteins were replaced by eukaryotic proteins
-
Full-text index only
Detecting unannotated splicing events in short-read RNA-seq with SAMI, a UMI-aware Nextflow pipeline.
PMID 42166739 · PMC13242923 · Bioinformatics (Oxford, England) · 2026 · 8 claims · 5 setups
SAMI is a UMI-aware, Singularity-contained Nextflow pipeline that detects splicing events diverging from transcript annotations directly from raw FASTQ files.
-
Has reproduction · 79
RetroSnake: A modular pipeline to detect human endogenous retroviruses in genome sequencing data.
PMID 36339261 · PMC9626663 · iScience · 2022 · 8 claims · 4 setups
RetroSnake is an end-to-end, modular, computationally efficient Snakemake pipeline for detecting HERV-K insertions in short-read NGS data, from raw alignment files to an annotated interactive HTML report
-
Full-text index only
nf-core/viralmetagenome: A novel pipeline for untargeted viral genome reconstruction.
PMID 42057295 · PMC13141149 · Bioinformatics (Oxford, England) · 2026 · 8 claims · 5 setups
nf-core/viralmetagenome is a Nextflow pipeline that automates untargeted reconstruction and variant analysis of eukaryotic DNA and RNA viruses from short-read metagenomic or hybridisation-capture data.
-
Has reproduction · 71
RNAmountAlign: Efficient software for local, global, semiglobal pairwise and multiple RNA sequence/structure alignment.
PMID 31978147 · PMC6980424 · PloS one · 2020 · 7 claims · 6 setups
RNAmountAlign performs pairwise local, global, and semiglobal (query search) alignment and progressive multiple alignment (global and local) using incremental ensemble mountain height, running in O(n^3) time and O(n^2) space for two sequences of length n
-
Has reproduction · 71
Hyb: a bioinformatics pipeline for the analysis of CLASH (crosslinking, ligation and sequencing of hybrids) data.
PMID 24211736 · PMC3969109 · Methods (San Diego, Calif.) · 2014 · 8 claims · 6 setups
The 'hyb' pipeline detects, calls, folds and annotates chimeric reads from CLASH high-throughput sequencing data.
-
Has reproduction · 45
Identifying and classifying trait linked polymorphisms in non-reference species by walking coloured de bruijn graphs.
PMID 23536903 · PMC3607606 · PloS one · 2013 · 8 claims · 9 setups
Bubbleparse detects sequence variants directly from NGS reads without a reference genome, using the coloured de Bruijn graph implementation of Cortex plus a new depth-first bubble-finding module.
-
Full-text index only
Function2Gene: a gene selection tool to increase the power of genetic association studies by utilizing public databases and expert knowledge.
PMID 18631403 · PMC2500032 · BMC bioinformatics · 2008 · 6 claims · 5 setups
Function2Gene is a set of Perl programs that queries public databases (NCBI, GeneCards, Harvester, with Uniprot/Ensembl also supported) using expert-selected keywords to rank genes by prior probability of disease association.
-
Full-text index only
Identification of deleterious non-synonymous single nucleotide polymorphisms using sequence-derived information.
PMID 18588693 · PMC2446391 · BMC bioinformatics · 2008 · 8 claims · 5 setups
A decision tree built on 10 selected sequence-derived features classifies SAPs as Disease or Polymorphism with 82.6% accuracy and 0.607 MCC in cross-validation.
-
Full-text index only
Functional annotation and identification of candidate disease genes by computational analysis of normal tissue gene expression data.
PMID 18560577 · PMC2409962 · PloS one · 2008 · 7 claims · 5 setups
Ranked Coexpression Groups (RCG) built from k=6 nearest coexpressed genes, combined with a majority-rule functional characterization, integrate multiple datasets/coexpression measures to generate high-confidence functional annotation predictions
-
Full-text index only
Evolutionary trace annotation of protein function in the structural proteome.
PMID 20036248 · PMC2831211 · Journal of molecular biology · 2010 · 8 claims · 7 setups
ET-ranked residue clusters can be used to build 3D templates that predict GO function in enzymes and non-enzymes alike, without prior knowledge of functional mechanism.
-
Full-text index only
A machine learning approach uncovers principles and determinants of eukaryotic ribosome pausing.
PMID 39423268 · PMC11488575 · Science advances · 2024 · 8 claims · 5 setups
An unsupervised ML pipeline using the extended isolation forest (EIF) algorithm can reliably detect ribosome pausing sites from noisy, coverage-biased RiboSeq data across expression levels
-
Has reproduction · 85
Digital sorting of complex tissues for cell type-specific gene expression profiles.
PMID 23497278 · PMC3626856 · BMC bioinformatics · 2013 · 8 claims · 8 setups
The Digital Sorting Algorithm (DSA) deconvolves mixed tissue expression into cell type-specific profiles using only marker genes, without requiring prior knowledge of cell type frequencies or in vitro pure-cell profiles.