Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
Multi-context seeds enable fast and high-accuracy read mapping.
PMID 41764549 · PMC13059148 · Genome biology · 2026 · 7 claims · 5 setups
Multi-context seeds (MCS) allow storage of seeds with different lengths in the same index structure by splitting hash bits among strobes, enabling full and partial matches
-
Full-text index only
Cleanifier: contamination removal from microbial sequences using spaced seeds of a human pangenome index.
PMID 41252442 · PMC12758600 · Bioinformatics (Oxford, England) · 2026 · 8 claims · 4 setups
Cleanifier is a fast, memory-frugal alignment-free tool for detecting and removing human contamination using gapped k-mers (spaced seeds) and a human pangenome index.
-
Has reproduction · 67
binny: an automated binning algorithm to recover high-quality genomes from complex metagenomic datasets.
PMID 36239393 · PMC9677464 · Briefings in bioinformatics · 2022 · 8 claims · 8 setups
binny outperforms or is highly competitive with commonly used and state-of-the-art binning methods (MetaBAT2, MaxBin2, CONCOCT, VAMB, SemiBin, MetaDecoder)
-
Has reproduction · 67
GAVISUNK: genome assembly validation via inter-SUNK distances in Oxford Nanopore reads.
PMID 36321867 · PMC9805576 · Bioinformatics (Oxford, England) · 2023 · 7 claims · 4 setups
GAVISUNK is an open-source pipeline that validates phased diploid HiFi assemblies by assessing concordance of inter-SUNK distances against orthogonal Oxford Nanopore (ONT) reads.
-
Full-text index only
CLAMP: predicting specific protein-mediated chromatin loops in diverse species with a chromatin accessibility language model.
PMID 41555433 · PMC12903630 · Genome biology · 2026 · 8 claims · 8 setups
CLAMP, a chromatin-accessibility language model, predicts protein-mediated chromatin loops across 10 species, 18 proteins, and 24 cell types with superior performance versus existing methods.
-
Full-text index only
Retentive Network promotes efficient RNA language modeling of long sequences.
PMID 41814064 · PMC13111708 · Communications biology · 2026 · 8 claims · 6 setups
RNAret, a RetNet-based RNA language model with O(n) complexity, achieves training parallelism and low computational overhead while processing long RNA sequences
-
Full-text index only
The 1000 Chinese Pangenome empowers medical and population genetics.
PMID 41922767 · PMC13233627 · Nature · 2026 · 8 claims · 8 setups
1,116 diploid genome assemblies (55 de novo, 1,061 pangenome-informed) were generated as part of the 1KCP project