Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
Large-scale estimation of bacterial and archaeal DNA prevalence in metagenomes reveals biome-specific patterns.
PMID 41854267 · PMC13098197 · mSystems · 2026 · 8 claims · 6 setups
SPF scalably and robustly estimates the fraction of bacterial and archaeal reads in a metagenome using detection of prokaryotic single-copy marker genes, without requiring eukaryotic or viral reference genomes
-
Full-text index only
Machine-learning approaches for classifying haplogroup from Y chromosome STR data.
PMID 18551166 · PMC2396484 · PLoS computational biology · 2008 · 8 claims · 5 setups
Y-STR allelic variability is partitioned more by differences among haplogroups than by differences among populations, suggesting Y-STRs carry haplogroup information
-
Has reproduction · 63
Community assessment of methods to deconvolve cellular composition from bulk gene expression.
PMID 39191725 · PMC11350143 · Nature communications · 2024 · 8 claims · 4 setups
Most deconvolution methods accurately predict coarse-grained immune/stromal cell populations from bulk expression.
-
Has reproduction · 87
Enhanced Generalizability of RNA Secondary Structure Prediction via Convolutional Block Attention Network and Ensemble Learning.
PMID 40871599 · PMC12388828 · Molecules (Basel, Switzerland) · 2025 · 8 claims · 8 setups
TrioFold integrates base-pairing clues from thermodynamic- and DL-based methods via ensemble learning and a convolutional block attention mechanism to enhance RSS prediction generalizability.
-
Full-text index only
Clustering of phosphorylation site recognition motifs can be exploited to predict the targets of cyclin-dependent kinase.
PMID 17316440 · PMC1852407 · Genome biology · 2007 · 8 claims · 6 setups
CDK consensus motifs are frequently clustered (closely spaced) in known CDK substrate proteins rather than uniformly distributed
-
Full-text index only
CanSig Benchmarks Methods for Reproducible Cancer Cell State Discovery from Single-Cell Transcriptomic Data.
PMID 41231245 · PMC13053056 · Cancer research · 2026 · 7 claims · 7 setups
CanSig is a comprehensive benchmarking tool for evaluating computational methods that identify shared transcriptional signatures in cancer from scRNA-seq data
-
Full-text index only
omnideconv: a unifying framework for using and benchmarking single-cell-informed deconvolution of bulk RNA-seq data.
PMID 41582216 · PMC12837286 · Genome biology · 2026 · 8 claims · 6 setups
omnideconv is an R package providing a unified interface to twelve second-generation deconvolution methods (AutoGeneS, BayesPrism, Bseq-SC, Bisque, CDseq, CIBERSORTx, CPM, DWLS, MOMF, MuSiC, SCDC, Scaden)
-
Full-text index only
Evaluating deep learning based structure prediction methods on antibody-antigen complexes.
PMID 41863324 · PMC13061134 · Bioinformatics (Oxford, England) · 2026 · 8 claims · 8 setups
Increased sampling improves the chance of generating a correct antibody-antigen model in a roughly log-linear manner with sample size
-
Full-text index only
SPrUCE: Utilizing Ultraconserved Elements of DNA for Population-Level Genetic Diversity Estimation.
PMID 42026820 · PMC13106921 · Molecular ecology resources · 2026 · 7 claims · 5 setups
Naive diversity estimators applied directly to UCE alignments underestimate nucleotide diversity due to negative selection/conservation at the UCE core.
-
Has reproduction · 53
spliceJAC: transition genes and state-specific gene regulation from single-cell transcriptome data.
PMID 36321549 · PMC9627675 · Molecular systems biology · 2022 · 8 claims · 8 setups
spliceJAC uses unspliced and spliced mRNA count matrices to construct cell state-specific gene-gene regulatory interaction (Jacobian) matrices from scRNA-seq data
-
Has reproduction · 89
Graph Random Forest: A Graph Embedded Algorithm for Identifying Highly Connected Important Features.
PMID 37509188 · PMC10377046 · Biomolecules · 2023 · 8 claims · 6 setups
GRF identifies effective features that form highly connected sub-graphs on the underlying biological network
-
Has reproduction · 87
Forseti: a mechanistic and predictive model of the splicing status of scRNA-seq reads.
PMID 38940130 · PMC11256924 · Bioinformatics (Oxford, England) · 2024 · 7 claims · 5 setups
Forseti is the first probabilistic model for resolving the splicing status of exonic scRNA-seq reads by scoring putative fragments linking read alignments to proximate priming sites
-
Has reproduction · 75
stDyer-image improves clustering analysis of spatially resolved transcriptomics and proteomics with morphological images.
PMID 41692960 · PMC12960910 · Bioinformatics (Oxford, England) · 2026 · 8 claims · 5 setups
stDyer-image directly associates the image modality with predicted cluster labels rather than using images to enhance/impute gene expression data
-
Has reproduction · 82
Landscape of allele-specific transcription factor binding in the human genome.
PMID 33980847 · PMC8115691 · Nature communications · 2021 · 8 claims · 6 setups
A novel statistical framework (ADASTRA) calls allele-specific TF binding from existing ChIP-Seq alignments by jointly correcting for background allelic dosage (BAD, from aneuploidy/CNVs) and reference mapping bias.
-
Has reproduction · 77
Accurate chromatin marks peak calling with Omnipeak.
PMID 41521664 · PMC12784980 · Nucleic acids research · 2026 · 8 claims · 6 setups
Omnipeak is a universal unsupervised peak-calling algorithm based on a constrained three-state hidden Markov model (zero, noise, signal states)
-
Has reproduction · 97
Determination of complete chromosomal haplotypes by bulk DNA sequencing.
PMID 33957932 · PMC8101039 · Genome biology · 2021 · 8 claims · 8 setups
A hierarchical computational strategy that first builds high-confidence local haplotype blocks from long-range/linked-read linkage and then concatenates them into whole-chromosome haplotypes using Hi-C contacts
-
Has reproduction · 50
MoDLE: high-performance stochastic modeling of DNA loop extrusion interactions.
PMID 36451166 · PMC9710047 · Genome biology · 2022 · 8 claims · 6 setups
MoDLE is a high-performance stochastic model/software for simulating DNA-DNA contacts generated by loop extrusion genome-wide
-
Has reproduction
Fast, accurate, and racially unbiased pan-cancer tumor-only variant calling with tabular machine learning.
PMID 36611079 · PMC9825621 · NPJ precision oncology · 2023 · 8 claims · 8 setups
Tree-based (XGBoost, LightGBM) and deep-learning (TabNet) tabular ML classifiers achieve state-of-the-art somatic vs germline classification in tumor-only WES samples, outperforming PureCN.
-
Has reproduction · 73
Genetic polyploid phasing from low-depth progeny samples.
PMID 35692633 · PMC9184567 · iScience · 2022 · 8 claims · 7 setups
WH-PPG phases polyploid parental samples by scoring informative variant pairs with a Bayesian log-likelihood model of progeny allele depths, clustering alleles by co-occurrence likelihood, and assigning clusters to haplotypes via interval scheduling
-
Full-text index only
Exome sequencing of a multigenerational human pedigree.
PMID 20011588 · PMC2788131 · PloS one · 2009 · 8 claims · 6 setups
Microarray-based exome capture combined with 454 GS FLX NGS is an efficient and reliable method to enrich for chromosomal regions of interest, validated on eight individuals from a three-generation pedigree