Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
A negative binomial latent factor model for paired microbiome sequencing data.
PMID 41572173 · PMC12910815 · BMC bioinformatics · 2026 · 8 claims · 2 setups
A negative binomial model with a shared taxon-specific latent factor (JNBM) captures cross-site correlation between paired microbiome samples from two body sites.
-
Full-text index only
Extending differential gene expression testing to handle genome aneuploidy in cancer.
PMID 41894415 · PMC13061324 · PLoS computational biology · 2026 · 8 claims · 4 setups
DeConveil integrates CNV data into DGE analysis using a GLM with negative binomial distribution to correct for CN-driven gene dosage effects
-
Has reproduction · 85
Single-Cell Differential Network Analysis with Sparse Bayesian Factor Models.
PMID 35186014 · PMC8855158 · Frontiers in genetics · 2021 · 8 claims · 2 setups
A hierarchical Bayesian factor model using treatment-dependent latent factor loadings can construct gene co-expression networks from scRNA-seq data and identify differences in network structure between two (or more) biological conditions.
-
Full-text index only
Score Matching for Differential Abundance Testing of Compositional High-Throughput Sequencing Data.
PMID 41944570 · PMC13055433 · Statistics in medicine · 2026 · 8 claims · 3 setups
cosmoDA extends the a-b power interaction model by adding a linear covariate effect on the location vector, enabling differential abundance testing on compositional data with feature interactions.
-
Full-text index only
Assessment of dispersion metrics for estimating single-cell transcriptional variability.
PMID 41770747 · PMC12970974 · PLoS computational biology · 2026 · 7 claims · 4 setups
The variance-to-mean ratio (VMR/Fano factor) scales approximately linearly with increasing dispersion and is independent of dataset size.
-
Has reproduction · 73
treeclimbR pinpoints the data-dependent resolution of hierarchical hypotheses.
PMID 34001188 · PMC8127214 · Genome biology · 2021 · 7 claims · 6 setups
treeclimbR proposes multiple candidate resolutions on a tree and selects the optimal one in a data-driven manner using three criteria (FDR-controlling range of t, number of rejected leaves, fewest internal nodes)
-
Full-text index only
saseR: juggling offsets unlocks RNA-seq tools for fast and scalable differential usage, aberrant splicing and expression retrieval.
PMID 41709279 · PMC13019952 · Genome biology · 2026 · 8 claims · 5 setups
Replacing the library-size offset with the log of the total gene count in NB-based bulk RNA-seq models (edgeR/DESeq2) lets the mean-model parameters be interpreted as transcript/exon usage, unlocking these tools for differential usage and aberrant splicing without DEXSeq-style subject-specific blocking covariates.
-
Full-text index only
Bayesian estimates of linkage disequilibrium.
PMID 17592642 · PMC1924864 · BMC genetics · 2007 · 8 claims · 3 setups
The MLE of D' is biased toward disequilibrium, with the bias particularly severe in small samples (<100 subjects) and rare alleles (MAF<0.05)
-
Has reproduction · 42
KAGE: fast alignment-free graph-based genotyping of SNPs and short indels.
PMID 36195962 · PMC9531401 · Genome biology · 2022 · 7 claims · 7 setups
KAGE combines population-based kmer count modeling with single-variant prior adjustment into an alignment-free genotyper that matches the accuracy of the best existing alignment-free genotypers while being an order of magnitude faster.
-
Full-text index only
Benchmarking of methods to analyse data derived from GBS-MeDIP.
PMID 41555215 · PMC12829230 · BMC bioinformatics · 2026 · 7 claims · 4 setups
featureCounts is the most reliable tool for count matrix generation from GBS-MeDIP data, outperforming MEDIPS
-
Has reproduction · 61
TEMP: a computational method for analyzing transposable element polymorphism in populations.
PMID 24753423 · PMC4066757 · Nucleic acids research · 2014 · 8 claims · 8 setups
TEMP combines pair-end (discordant) read and split (soft-clipped) read information to identify both presence and absence of TE insertions in genomic DNA from heterogeneous/pooled samples.
-
Full-text index only
Duplication count distributions in DNA sequences.
PMID 19256873 · PMC3121164 · Physical review. E, Statistical, nonlinear, and soft matter physics · 2008 · 8 claims · 8 setups
Duplication count distributions N(c) for complex 40-mers show power-law-like decay for c roughly 3 to 50 (or higher) across human, C. elegans, A. thaliana, and D. melanogaster genomes.
-
Full-text index only
RUMINA: high-throughput deduplication of unique molecular identifiers for amplicon and whole-genome sequencing with enhanced error correction.
PMID 41734278 · PMC12975283 · Bioinformatics (Oxford, England) · 2026 · 8 claims · 4 setups
RUMINA improves detection accuracy of ultra-low frequency SNVs (0.01%-1%) compared to UMI-tools and UMICollapse
-
Full-text index only
Semi-parametric empirical bayes method for multiplet detection in snATAC-seq with probabilistic multi-omic integration.
PMID 42054434 · PMC13148828 · PLoS computational biology · 2026 · 8 claims · 5 setups
SEBULA models the singlet background directly from observed HCLC (high-coverage locus count) statistics using fragment-level snATAC-seq information, avoiding reliance on synthetic/artificial doublets.
-
Full-text index only
Modeling ChIP sequencing in silico with applications.
PMID 18725927 · PMC2507756 · PLoS computational biology · 2008 · 8 claims · 4 setups
Observed ChIP-seq tag counts follow an initial power-law distribution followed by a long right tail.
-
Full-text index only
StrainMake: reproducible hybrid metagenomics with MAG recovery and strain-level resolution.
PMID 42097292 · PMC13188985 · Bioinformatics (Oxford, England) · 2026 · 8 claims · 5 setups
StrainMake is a Snakemake-based, Conda-managed workflow for de novo metagenomic analysis from short, long, or hybrid sequencing data.
-
Full-text index only
Cell neighborhood topology directs rare cell population identification.
PMID 41912521 · PMC13199379 · Nature communications · 2026 · 8 claims · 8 setups
RareQ is a framework that quantifies neighborhood connectivity (Q), a cell-specific measure of kNN-graph cliquishness, to detect rare cell populations from single-cell and spatial omics data
-
Full-text index only
Differential expression analysis in single-cell and spatial RNA-seq without model assumptions.
PMID 41980775 · PMC13198004 · Cell reports methods · 2026 · 7 claims · 4 setups
Common DGE analysis methods (Wilcoxon test, unweighted t-test, pseudo-bulk aggregation, SCTransform-style parametrization) rely on unnecessary simplifications and assumptions that are inconsistent with experimental data and cause false findings
-
Has reproduction · 53
spliceJAC: transition genes and state-specific gene regulation from single-cell transcriptome data.
PMID 36321549 · PMC9627675 · Molecular systems biology · 2022 · 8 claims · 8 setups
spliceJAC uses unspliced and spliced mRNA count matrices to construct cell state-specific gene-gene regulatory interaction (Jacobian) matrices from scRNA-seq data
-
Has reproduction · 83
Hobbes: optimized gram-based methods for efficient read alignment.
PMID 22199254 · PMC3315303 · Nucleic acids research · 2012 · 8 claims · 4 setups
Hobbes, a gram-based short-read mapper supporting Hamming and edit distance, is faster than all other read-mapping programs tested while maintaining high mapping quality.