Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
Geometry-aware graph attention networks to explain single-cell chromatin states and gene expression with SEAGALL.
PMID 42026624 · PMC13238118 · Genome biology · 2026 · 8 claims · 6 setups
SEAGALL combines a geometry-regularised autoencoder (GRAE) to embed cells and build a cell-cell graph with a graph attention network (GAT) classifier and GNNExplainer-based XAI to identify features driving cell type/phenotype.
-
Has reproduction · 57
Genome-wide kinetic properties of transcriptional bursting in mouse embryonic stem cells.
PMID 32596448 · PMC7299619 · Science advances · 2020 · 8 claims · 8 setups
Transcriptional bursting kinetics is regulated by a combination of promoter- and gene body-binding proteins, including the polycomb repressive complex 2 (PRC2) and transcription elongation factors
-
Full-text index only
HIF3A-mediated aberrant activation of TXNIP promotes Alzheimer's disease progression.
PMID 41807716 · PMC13096512 · Scientific reports · 2026 · 8 claims · 8 setups
OS activity is significantly elevated in AD and shows pronounced heterogeneity across brain cell types
-
Full-text index only
Early feature extraction drives model performance in high-resolution chromatin accessibility prediction.
PMID 41526189 · PMC12951969 · Genome research · 2026 · 8 claims · 6 setups
Early feature extraction (via ConvNeXt V2 blocks), rather than downstream architecture type, is the primary determinant of prediction accuracy in high-resolution chromatin accessibility prediction.
-
Full-text index only
Human-scATAC-Corpus: a comprehensive database of scATAC-seq data.
PMID 41296545 · PMC12807747 · Nucleic acids research · 2026 · 8 claims · 6 setups
Human-scATAC-Corpus is a comprehensive database of human scATAC-seq data comprising 5,407,621 cells from 35 datasets across 37 tissues or cell lines
-
Full-text index only
Benchmarking component choices for unpaired single cell RNA and epigenomic integration.
PMID 41987329 · PMC13192178 · Genome biology · 2026 · 7 claims · 8 setups
Gene activity scores (GAS) show limited correlation with actual gene expression but effectively preserve cellular neighborhood structure and support clustering.
-
Full-text index only
Visualizing stability: a sensitivity analysis framework for t-SNE embeddings.
PMID 41552665 · PMC12808344 · Frontiers in bioinformatics · 2025 · 8 claims · 5 setups
The Implicit Function Theorem, combined with automatic differentiation, can be used to efficiently compute the complete sensitivity Jacobian of a converged t-SNE embedding with respect to the input data, avoiding differentiation through the full iterative optimizer.
-
Has reproduction · 67
Adaptive learning embedding features to improve the predictive performance of SARS-CoV-2 phosphorylation sites.
PMID 37847658 · PMC10628388 · Bioinformatics (Oxford, England) · 2023 · 8 claims · 6 setups
PSPred-ALE outperforms state-of-the-art SARS-CoV-2 phosphorylation site predictors (e.g. DeepIPs) and handcrafted feature-based methods in benchmarking comparisons
-
Full-text index only
Leveraging the germ layer development patterns to predict prognosis and identify MEST as a novel therapeutic target in glioma.
PMID 41501725 · PMC12870398 · Cancer cell international · 2026 · 7 claims · 8 setups
MEST is a key oncogenic GLD-related gene and a novel therapeutic target in glioma, identified via a machine learning feature selection framework
-
Full-text index only
KRT15 identified by scRNA-Seq and machine learning as stemness regulator and prognostic biomarker in ESCC.
PMID 41907411 · PMC13018899 · iScience · 2026 · 8 claims · 8 setups
A malignant epithelial subpopulation (Cluster 0) with the highest CytoTRACE stemness potential and highest EMT activity exists in ESCC scRNA-seq data
-
Full-text index only
TDAGENE: Inference of Gene Regulatory Network Based on Topological Data Analysis and Graph Attention Network for Single-Cell RNA Sequencing Data.
PMID 42093817 · PMC13139726 · Computational and structural biotechnology journal · 2026 · 7 claims · 5 setups
TDAGENE combines TDA features with a multilayer GAT via gate-controlled fusion to improve GRN inference accuracy
-
Has reproduction · 100
A workflow reproducibility scale for automatic validation of biological interpretation results.
PMID 37150537 · PMC10164546 · GigaScience · 2022 · 8 claims · 4 setups
Comparing output files by checksum alone is insufficient to verify reproducibility, since checksums can differ even when the underlying biological interpretation is unchanged
-
Has reproduction · 70
Predicting enhancers in mammalian genomes using supervised hidden Markov models.
PMID 30917778 · PMC6437899 · BMC bioinformatics · 2019 · 8 claims · 8 setups
eHMM predicts enhancers with high precision and recall comparable to state-of-the-art methods and consistently outperforms them in accuracy and resolution
-
Has reproduction · 32
Developing prognostic gene panel of survival time in lung adenocarcinoma patients using machine learning.
PMID 35117753 · PMC8799101 · Translational cancer research · 2020 · 8 claims · 5 setups
Naïve Bayes using a 22-gene panel is the best-performing and most stable machine learning model for predicting LUAD survival time (>3 vs <3 years)
-
Full-text index only
TSniffer: unbiased de novo identification of RNA editing sites and quantification of editing activity in RNA-seq data.
PMID 41549280 · PMC12838065 · Genome biology · 2026 · 8 claims · 6 setups
TSniffer is a novel tool that uses a rolling window Fisher's exact test approach to identify RNA editing sites (TsRegions) de novo in RNA-seq data without relying on editing databases or two-sample differential comparison.
-
Full-text index only
Multi-omics feature engineering driven by biomedical foundation models improves drug response prediction for inflammatory bowel disease patients.
PMID 41844950 · PMC13129071 · Scientific reports · 2026 · 8 claims · 7 setups
FM (MAMMAL)-derived drug-target binding affinity (BA) inference can be used to rank/select biologically relevant protein targets and their associated genes/SNPs for a drug of interest without knowledge of protein structure or active sites
-
Full-text index only
GDSim: accurate simulation for single-cell transcriptomes based on the guided diffusion model.
PMID 41978379 · PMC13076945 · Briefings in bioinformatics · 2026 · 8 claims · 4 setups
GDSim, a label-guided diffusion-based deep generative network, can simulate scRNA-seq data that closely reflects the true distribution of original data
-
Has reproduction · 59
WASP: a versatile, web-accessible single cell RNA-Seq processing platform.
PMID 33736596 · PMC7977290 · BMC genomics · 2021 · 7 claims · 7 setups
WASP is a software platform for processing Drop-Seq-based scRNA-seq data generated with ddSEQ or 10x protocols, combining a Snakemake pre-processing pipeline with an R Shiny post-processing application.
-
Full-text index only
Score Matching for Differential Abundance Testing of Compositional High-Throughput Sequencing Data.
PMID 41944570 · PMC13055433 · Statistics in medicine · 2026 · 8 claims · 3 setups
cosmoDA extends the a-b power interaction model by adding a linear covariate effect on the location vector, enabling differential abundance testing on compositional data with feature interactions.
-
Has reproduction · 44
Dynamic Gene Attention Focus (DyGAF): Enhancing Biomarker Identification Through Dual-Model Attention Networks.
PMID 40160891 · PMC11951896 · Bioinformatics and biology insights · 2025 · 6 claims · 5 setups
DyGAF, a dual-model attention neural network (independent Model A + dependent Model B), identifies and ranks genes by significance for COVID-19 biomarker discovery more effectively than differential expression analysis (DEA) and random forest (RF) feature selection