Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
Dynamics of intronic polyadenylation in the hematopoietic lineage and its regulation by DNA methylation.
PMID 41974581 · PMC13262950 · Genome research · 2026 · 8 claims · 5 setups
IPAseek, a dynamic programming framework combining the PELT algorithm with CROPS, enables de novo identification of IPA events from bulk RNA-seq data.
-
Full-text index only
Evaluating deconvolution methods using real bulk RNA-expression data for robust prognostic insights across cancer types.
PMID 41566530 · PMC12906006 · Genome biology · 2026 · 7 claims · 6 setups
Pseudobulk and real bulk RNA-seq deconvolution performance differ significantly, and method ranking consistency is lower between pseudobulk and real bulk than within either data type alone
-
Full-text index only
scSurv: a deep generative model for single-cell survival analysis.
PMID 41429574 · PMC12797213 · Bioinformatics (Oxford, England) · 2026 · 8 claims · 6 setups
scSurv combines a Cox proportional hazards model with a deep generative model (VAE) of single-cell transcriptomes to estimate individual cellular contributions to clinical outcomes
-
Full-text index only
Clustering of phosphorylation site recognition motifs can be exploited to predict the targets of cyclin-dependent kinase.
PMID 17316440 · PMC1852407 · Genome biology · 2007 · 8 claims · 6 setups
CDK consensus motifs are frequently clustered (closely spaced) in known CDK substrate proteins rather than uniformly distributed
-
Full-text index only
A cellular epigenetic classification system for glioblastoma.
PMID 41499453 · PMC13128495 · Neuro-oncology · 2026 · 8 claims · 8 setups
ITHresolveGBM, a hierarchical two-step NMF method, deconvolutes bulk GBM DNA methylation profiles into three non-malignant (immune, glial, neuronal) and three malignant components
-
Full-text index only
PreTSA: computationally efficient modeling of temporal and spatial gene expression patterns.
PMID 41673899 · PMC12998178 · Genome biology · 2026 · 7 claims · 8 setups
PreTSA dramatically reduces computational time and memory versus GAM (Monocle, TSCAN) and PseudotimeDE for identifying temporally variable genes (TVGs) while producing highly similar results
-
Has reproduction · 89
Graph Random Forest: A Graph Embedded Algorithm for Identifying Highly Connected Important Features.
PMID 37509188 · PMC10377046 · Biomolecules · 2023 · 8 claims · 6 setups
GRF identifies effective features that form highly connected sub-graphs on the underlying biological network
-
Full-text index only
A sequence knowledge-guided deep learning method for single-cell multi-omics translation.
PMID 41975483 · PMC13185235 · Genome biology · 2026 · 7 claims · 7 setups
scProTrans, a deep learning framework combining sequence knowledge (dna2vec gene embeddings, ProtT5 protein embeddings) with a cross-omics attention mechanism, translates single-cell transcriptome data into proteome profiles
-
Has reproduction · 87
Forseti: a mechanistic and predictive model of the splicing status of scRNA-seq reads.
PMID 38940130 · PMC11256924 · Bioinformatics (Oxford, England) · 2024 · 7 claims · 5 setups
Forseti is the first probabilistic model for resolving the splicing status of exonic scRNA-seq reads by scoring putative fragments linking read alignments to proximate priming sites
-
Has reproduction · 87
Enhanced Generalizability of RNA Secondary Structure Prediction via Convolutional Block Attention Network and Ensemble Learning.
PMID 40871599 · PMC12388828 · Molecules (Basel, Switzerland) · 2025 · 8 claims · 8 setups
TrioFold integrates base-pairing clues from thermodynamic- and DL-based methods via ensemble learning and a convolutional block attention mechanism to enhance RSS prediction generalizability.
-
Has reproduction
Fast, accurate, and racially unbiased pan-cancer tumor-only variant calling with tabular machine learning.
PMID 36611079 · PMC9825621 · NPJ precision oncology · 2023 · 8 claims · 8 setups
Tree-based (XGBoost, LightGBM) and deep-learning (TabNet) tabular ML classifiers achieve state-of-the-art somatic vs germline classification in tumor-only WES samples, outperforming PureCN.
-
Full-text index only
High-fidelity bidirectional translation between single-cell transcriptomes and DNA methylomes with scBOND.
PMID 41887797 · PMC13138010 · Genome research · 2026 · 7 claims · 6 setups
scBOND is a bidirectional dual-channel VAE framework for cross-modality translation between scRNA-seq and scDNAm that outperforms existing baseline methods (scCross, MAPLE) in both translation directions
-
Full-text index only
Integrating single-cell and single-nucleus datasets improves bulk RNA-seq deconvolution.
PMID 41895263 · PMC13106970 · Cell reports methods · 2026 · 8 claims · 5 setups
scRNA-seq references yield significantly higher Pearson correlation and lower RMSE than snRNA-seq references for deconvolution across all four tissue datasets
-
Has reproduction · 53
spliceJAC: transition genes and state-specific gene regulation from single-cell transcriptome data.
PMID 36321549 · PMC9627675 · Molecular systems biology · 2022 · 8 claims · 8 setups
spliceJAC uses unspliced and spliced mRNA count matrices to construct cell state-specific gene-gene regulatory interaction (Jacobian) matrices from scRNA-seq data
-
Full-text index only
AICellType: a large language model-based platform for accurate cell type annotation.
PMID 42001469 · PMC13092268 · Briefings in bioinformatics · 2026 · 8 claims · 8 setups
Claude 3.5 Sonnet achieved the best overall performance among 79 benchmarked LLMs for cell type annotation, balancing accuracy, robustness, speed, and cost-efficiency
-
Full-text index only
Embeddings from language models are good learners for single-cell data analysis.
PMID 41726097 · PMC12921509 · Patterns (New York, N.Y.) · 2026 · 8 claims · 8 setups
scELMo combines LLM-derived embeddings of gene and cell metadata with raw single-cell expression data via matrix operations to generate cell embeddings without pretraining a new model