Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Has reproduction · 78
Determining the quality and complexity of next-generation sequencing data without a reference genome.
PMID 25514851 · PMC4298064 · Genome biology · 2014 · 8 claims · 8 setups
kPAL, an open-source alignment-free package, assesses sequencing data quality and complexity using k-mer frequency profiles and pairwise distances between them, without a reference sequence.
-
Has reproduction · 91
A reference profile-free deconvolution method to infer cancer cell-intrinsic subtypes and tumor-type-specific stromal profiles.
PMID 32111252 · PMC7049190 · Genome medicine · 2020 · 8 claims · 8 setups
DeClust is a reference-profile-free deconvolution method that incorporates molecular subtyping directly into the deconvolution process, outputting cohort-level cancer subtype and stromal reference profiles rather than per-individual profiles
-
Has reproduction · 93
Population genomics of the Wolbachia endosymbiont in Drosophila melanogaster.
PMID 23284297 · PMC3527207 · PLoS genetics · 2012 · 8 claims · 8 setups
Wolbachia infection status can be accurately predicted in silico from whole-genome shotgun sequence of individual host strains, showing 99% concordance with diagnostic PCR.
-
Full-text index only
CholeraSeq: a comprehensive genomic pipeline for cholera surveillance and near real-time outbreak investigation.
PMID 41400832 · PMC12790814 · Bioinformatics (Oxford, England) · 2026 · 7 claims · 6 setups
CholeraSeq is an automated, V. cholerae-specific Nextflow pipeline that processes WGS outbreak data from raw reads/assemblies to high-quality SNPs and phylogenies in near-real-time.
-
Has reproduction · 45
Accurate sequence variant genotyping in cattle using variation-aware genome graphs.
PMID 31092189 · PMC6521551 · Genetics, selection, evolution : GSE · 2019 · 8 claims · 7 setups
Graphtyper outperformed GATK and SAMtools in genotype concordance, non-reference sensitivity, and non-reference discrepancy compared to microarray genotypes
-
Full-text index only
DupyliCate: mining, classifying, and characterizing gene duplications.
PMID 42209743 · PMC13219399 · Scientific reports · 2026 · 8 claims · 8 setups
DupyliCate is a Python tool for identifying and classifying gene duplication arrays, using BUSCO-based species-specific thresholds and offering integrated expression divergence and Ka/Ks analysis.
-
Has reproduction · 83
Accurate prediction of metagenome-assembled genome completeness by MAGISTA, a random forest model built on alignment-free intra-bin statistics.
PMID 35248155 · PMC8898458 · Environmental microbiome · 2022 · 7 claims · 7 setups
MAGISTA, a random forest model built on alignment-free intra-bin distance-distribution statistics, can estimate MAG completeness and purity without relying on reference marker genes.
-
Full-text index only
Whole-genome sequencing with AVITI and NovaSeq X Plus reveals comparable performance with contextual biases.
PMID 42206012 · PMC13202175 · NAR genomics and bioinformatics · 2026 · 8 claims · 7 setups
AVITI and NovaSeq X Plus are highly comparable overall for variant-calling performance in WGS