Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
Benchmarking tools for the alignment of functional noncoding DNA.
PMID 14736341 · PMC344529 · BMC bioinformatics · 2004 · 8 claims · 4 setups
Global alignment tools (Avid, ClustalW, Lagan, Needle, DiAlign-G) typically have higher sensitivity over entire noncoding sequences and within constrained blocks than local tools
-
Has reproduction · 62
Gbdmr: identifying differentially methylated CpG regions in the human genome via generalized beta regressions.
PMID 38443825 · PMC10916021 · BMC bioinformatics · 2024 · 8 claims · 4 setups
gbdmr models DNA methylation levels of CpG sites using a generalized beta distribution instead of assuming normality as in linear-regression-based methods
-
Has reproduction · 78
IsomiR_Window: a system for analyzing small-RNA-seq data in an integrative and user-friendly manner.
PMID 33522913 · PMC7852101 · BMC bioinformatics · 2021 · 8 claims · 2 setups
IsomiR Window is an integrated, user-friendly platform that systematically identifies, quantifies, and functionally explores isomiR expression in small-RNA-seq datasets without requiring computational skills
-
Has reproduction · 84
An accurate method for identifying recent recombinants from unaligned sequences.
PMID 35025988 · PMC8963311 · Bioinformatics (Oxford, England) · 2022 · 8 claims · 4 setups
A novel algorithm combining the JHMM (Zilversmit et al. 2013) mosaic representation with a distance-based triple comparison can identify recombinant sequences and their parents from unaligned, gene-length sequences without a reference panel.
-
Full-text index only
Computational tradeoffs in multiplex PCR assay design for SNP genotyping.
PMID 16042802 · PMC1190169 · BMC genomics · 2005 · 7 claims · 6 setups
Achieving high-multiplexing/high-coverage multiplex PCR designs is subject to a computational phase transition as the SNP-pair compatibility probability crosses a critical threshold
-
Has reproduction · 53
spliceJAC: transition genes and state-specific gene regulation from single-cell transcriptome data.
PMID 36321549 · PMC9627675 · Molecular systems biology · 2022 · 8 claims · 6 setups
spliceJAC quantifies multivariate mRNA splicing from unspliced/spliced count matrices to construct cell state-specific gene-gene (Jacobian) interaction matrices.
-
Has reproduction · 50
RNA modifications detection by comparative Nanopore direct RNA sequencing.
PMID 34893601 · PMC8664944 · Nature communications · 2021 · 7 claims · 5 setups
Nanocompore is a model-free comparative method that uses a 2-component Gaussian mixture model (GMM) and univariate statistical tests on signal intensity/dwell time to detect RNA modifications in Nanopore direct RNA sequencing data without needing a training set
-
Has reproduction · 50
MoDLE: high-performance stochastic modeling of DNA loop extrusion interactions.
PMID 36451166 · PMC9710047 · Genome biology · 2022 · 7 claims · 6 setups
MoDLE is a high-performance stochastic model that simulates DNA-DNA contacts from loop extrusion genome-wide in minutes using less than 1 GB of RAM
-
Has reproduction · 89
TrEMOLO: accurate transposable element allele frequency estimation using long-read sequencing data combining assembly and mapping-based approaches.
PMID 37013657 · PMC10069131 · Genome biology · 2023 · 6 claims · 6 setups
TrEMOLO combines an assembly-based INSIDER module and a mapping-based OUTSIDER module to detect TE insertions/deletions from long-read sequencing data and estimate their allele frequency
-
Has reproduction · 63
Community assessment of methods to deconvolve cellular composition from bulk gene expression.
PMID 39191725 · PMC11350143 · Nature communications · 2024 · 8 claims · 4 setups
Most deconvolution methods accurately predict coarse-grained immune/stromal cell populations from bulk expression.
-
Full-text index only
Decoding of superimposed traces produced by direct sequencing of heterozygous indels.
PMID 18654614 · PMC2429969 · PLoS computational biology · 2008 · 7 claims · 3 setups
A dynamic programming method (implemented as web app Indelligent) can decode superimposed allelic sequences from a single mixed trace, using only the observed string of ambiguous peak calls, without a reference sequence or reverse trace.
-
Has reproduction · 79
Computationally scalable regression modeling for ultrahigh-dimensional omics data with ParProx.
PMID 34254998 · PMC8575036 · Briefings in bioinformatics · 2021 · 6 claims · 4 setups
ParProx implements overlapping and non-overlapping (latent) group lasso regression for time-to-event (Cox) and classification (logistic) analysis with variables grouped by biological priors.
-
Has reproduction · 95
Mouse-Geneformer: A deep learning model for mouse single-cell transcriptome and its cross-species utility.
PMID 40106407 · PMC11964219 · PLoS genetics · 2025 · 7 claims · 6 setups
Mouse-Geneformer, a Transformer Encoder model pre-trained via masked-token self-supervised learning on mouse-Genecorpus-20M, was successfully constructed following the original human Geneformer architecture.
-
Has reproduction · 67
Generative and integrative modeling for transcriptomics with formalin fixed paraffin embedded material.
PMID 41029822 · PMC12486589 · Journal of translational medicine · 2025 · 8 claims · 5 setups
fRNA-seq transcript counts are best fit by the negative binomial distribution, with little evidence supporting zero-inflated extensions
-
Has reproduction · 85
Digital sorting of complex tissues for cell type-specific gene expression profiles.
PMID 23497278 · PMC3626856 · BMC bioinformatics · 2013 · 8 claims · 8 setups
The Digital Sorting Algorithm (DSA) deconvolves mixed tissue expression into cell type-specific profiles using only marker genes, without requiring prior knowledge of cell type frequencies or in vitro pure-cell profiles.
-
Full-text index only
POCUS: mining genomic sequence annotation to predict disease genes.
PMID 14611661 · PMC329128 · Genome biology · 2003 · 8 claims · 6 setups
Genes predisposing to the same disease tend to share functional annotation IDs (GO/InterPro) more than expected by chance
-
Full-text index only
Speeding disease gene discovery by sequence based candidate prioritization.
PMID 15766383 · PMC1274252 · BMC bioinformatics · 2005 · 7 claims · 8 setups
Disease genes (OMIM) differ significantly from non-disease genes in sequence-based features including gene/cDNA/protein size, exon number, homolog conservation, secretion signal, 3' UTR length, CpG islands, and distance to nearest gene.