Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
CellMap: precision mapping of cellular landscape in spatial transcriptomics.
PMID 41505103 · PMC12781899 · Nucleic acids research · 2026 · 7 claims · 3 setups
CellMap combines co-linearity of seed genes, a random forest model, and the linear assignment algorithm to achieve optimal assignment of single cells to spatial spots
-
Has reproduction · 29
MOSAIK: a hash-based algorithm for accurate next-generation sequencing short-read mapping.
PMID 24599324 · PMC3944147 · PloS one · 2014 · 8 claims · 8 setups
MOSAIK is the only aligner that consistently aligns reads from all major sequencing platforms (Illumina, AB SOLiD, Roche 454, Ion Torrent, Pacific Biosciences SMRT) using the same algorithmic approach.
-
Full-text index only
PPC: an algorithm for accurate estimation of SNP allele frequencies in small equimolar pools of DNA using data from high density microarrays.
PMID 16199750 · PMC1240117 · Nucleic acids research · 2005 · 7 claims · 6 setups
The PPC algorithm, which applies a probe-pair-specific second-degree polynomial correction, increases the accuracy of allele frequency estimates from pooled DNA compared with previously described algorithms
-
Full-text index only
GeneKeyDB: a lightweight, gene-centric, relational database to support data mining environments.
PMID 15790402 · PMC1274265 · BMC bioinformatics · 2005 · 8 claims · 6 setups
GeneKeyDB is a lightweight, gene-centric relational database that supports data mining and integration with computational analysis tools.
-
Full-text index only
Automatic annotation of eukaryotic genes, pseudogenes and promoters.
PMID 16925832 · PMC1810547 · Genome biology · 2006 · 8 claims · 6 setups
Fgenesh++ gene prediction pipeline identifies 91% of coding nucleotides with 90% specificity
-
Full-text index only
ADaCGH: A parallelized web-based application and R package for the analysis of aCGH data.
PMID 17710137 · PMC1940324 · PloS one · 2007 · 8 claims · 4 setups
ADaCGH implements eight CNA detection methods, including the best-performing ones from recent reviews (CBS, GLAD, CGHseg, HMM)
-
Full-text index only
BFAST: an alignment tool for large scale genome resequencing.
PMID 19907642 · PMC2770639 · PloS one · 2009 · 7 claims · 4 setups
BFAST is a new algorithm and freely available software tool for aligning large-scale short-read sequencing data to large reference genomes with user-customizable speed and accuracy
-
Full-text index only
Eukan: a fully automated nuclear genome annotation pipeline for less studied and divergent eukaryotes.
PMID 41567515 · PMC12817076 · NAR genomics and bioinformatics · 2026 · 8 claims · 7 setups
Eukan automatically leverages RNA-Seq coverage to inform generalized Hidden Markov Model gene prediction and intron lengths to inform protein sequence alignments
-
Full-text index only
iMapper: a web application for the automated analysis and mapping of insertional mutagenesis sequence data against Ensembl genomes.
PMID 18974167 · PMC2639305 · Bioinformatics (Oxford, England) · 2008 · 6 claims · 3 setups
iMapper is a web application for automated analysis and mapping of insertional mutagenesis sequence data against vertebrate and invertebrate Ensembl genomes (human, mouse, rat, zebrafish, Drosophila, S. cerevisiae).
-
Full-text index only
UBD: incorporating uncertainty in cell type proportion estimates from bulk samples to infer cell-type-specific profiles.
PMID 41520227 · PMC12895075 · Briefings in bioinformatics · 2026 · 7 claims · 4 setups
Existing CTS deconvolution methods (e.g., CIBERSORTx, TCA, bMIND, CellDMC, HBI) require cell type proportions that are in practice only estimated, not known, introducing unaccounted uncertainty into CTS inference.
-
Has reproduction · 97
CellFishing.jl: an ultrafast and scalable cell search method for single-cell RNA sequencing.
PMID 30744683 · PMC6371477 · Genome biology · 2019 · 8 claims · 5 setups
CellFishing.jl achieves accuracy comparable to state-of-the-art software (scmap-cell) but is markedly faster
-
Has reproduction · 89
TrEMOLO: accurate transposable element allele frequency estimation using long-read sequencing data combining assembly and mapping-based approaches.
PMID 37013657 · PMC10069131 · Genome biology · 2023 · 6 claims · 6 setups
TrEMOLO combines an assembly-based INSIDER module and a mapping-based OUTSIDER module to detect TE insertions/deletions from long-read sequencing data and estimate their allele frequency
-
Has reproduction · 76
Tracing human genetic histories and natural selection with precise local ancestry inference.
PMID 40379651 · PMC12084304 · Nature communications · 2025 · 7 claims · 7 setups
Orchestra, a two-stage LAI method combining a recombination-distance base layer with a deep learning (convolutional + attention) smoothing module, outperforms RFmix, FLARE and Gnomix in precision and recall across simulated admixture generations.
-
Full-text index only
A space-efficient and accurate method for mapping and aligning cDNA sequences onto genomic sequence.
PMID 18344523 · PMC2377433 · Nucleic acids research · 2008 · 7 claims · 6 setups
Spaln maps and aligns large cDNA sequence sets onto whole mammalian genomes using substantially less memory than comparable existing tools
-
Has reproduction · 50
RNA-Seq alignment to individualized genomes improves transcript abundance estimates in multiparent populations.
PMID 25236449 · PMC4174954 · Genetics · 2014 · 8 claims · 7 setups
Genetic variants distinguishing an individual genome from the reference cause read misalignment and biased transcript abundance estimates, and fine-tuning of alignment algorithms does not correct this problem.
-
Full-text index only
Reconstructing single-cell resolution from spatial transcriptomics with CellRefiner.
PMID 41760664 · PMC13066420 · Nature communications · 2026 · 8 claims · 8 setups
CellRefiner is a physical/particle-based model (subcellular element method) that integrates scRNA-seq and spatial transcriptomics data to reconstruct single-cell resolution spatial data
-
Full-text index only
Genomics--from Neanderthals to high-throughput sequencing.
PMID 16934106 · PMC1779599 · Genome biology · 2006 · 8 claims · 8 setups
Next-generation sequencing platforms (GS20/454 and Solexa) can deliver the throughput and cost reductions needed for population-scale and medical resequencing.
-
Has reproduction · 89
Statistical framework for calling allelic imbalance in high-throughput sequencing data.
PMID 39966391 · PMC11836314 · Nature communications · 2025 · 8 claims · 6 setups
MIXALIME is a versatile computational framework for calling allele-specific variants (ASVs) from diverse high-throughput omics data
-
Has reproduction · 85
An extensive evaluation of read trimming effects on Illumina NGS data analysis.
PMID 24376861 · PMC3871669 · PloS one · 2013 · 8 claims · 8 setups
Read trimming increases the quality and reliability of downstream NGS analyses (RNA-Seq mapping, SNP identification, genome assembly) while reducing execution time and computational resources.
-
Full-text index only
sedimix: a workflow for the analysis of hominin nuclear DNA sequences from sediments.
PMID 41512286 · PMC12866666 · Bioinformatics (Oxford, England) · 2026 · 8 claims · 6 setups
sedimix is a snakemake workflow that processes raw sediment DNA sequencing reads (fastq) through filtering, taxonomic classification, mapping, and duplicate/quality filtering to output hominin-derived BAM files and summary statistics.