Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Has reproduction · 42
KAGE: fast alignment-free graph-based genotyping of SNPs and short indels.
PMID 36195962 · PMC9531401 · Genome biology · 2022 · 7 claims · 7 setups
KAGE combines population-based kmer count modeling with single-variant prior adjustment into an alignment-free genotyper that matches the accuracy of the best existing alignment-free genotypers while being an order of magnitude faster.
-
Full-text index only
iAODE for benchmarking and continuum modeling of single-cell chromatin accessibility.
PMID 41775921 · PMC13066597 · Communications biology · 2026 · 8 claims · 5 setups
iAODE combines a ZINB-likelihood VAE, a latent Neural ODE, low-weight KL regularization, and an interpretable reconstruction (irecon) bottleneck to learn generative, temporally continuous latent spaces for scATAC-seq.
-
Has reproduction · 88
Comprehensive benchmarking of large language models for RNA secondary structure prediction.
PMID 40205851 · PMC11982019 · Briefings in bioinformatics · 2025 · 7 claims · 4 setups
Existing RNA-LLMs had not previously been evaluated for secondary structure prediction in a unified, fair experimental setup with the same datasets and prediction model.
-
Has reproduction · 77
SurvConvMixer: robust and interpretable cancer survival prediction based on ConvMixer using pathway-level gene expression images.
PMID 38539106 · PMC10967213 · BMC bioinformatics · 2024 · 6 claims · 5 setups
SurvConvMixer, using pathway-level gene expression images and ConvMixer, achieves strong internal validation AUC for overall survival prediction, especially on larger datasets like LUAD
-
Has reproduction · 98
maxATAC: Genome-scale transcription-factor binding prediction from ATAC-seq with deep neural networks.
PMID 36719906 · PMC9917285 · PLoS computational biology · 2023 · 8 claims · 6 setups
maxATAC is a suite of deep neural network models enabling state-of-the-art, genome-scale TFBS prediction from ATAC-seq, with models for 127 human transcription factors
-
Full-text index only
Metappuccino: large language model-driven reconstruction of sequence read archive metadata for cancer research.
PMID 42057294 · PMC13148957 · Bioinformatics (Oxford, England) · 2026 · 8 claims · 3 setups
Metappuccino reconstructs 19 metadata classes by combining deterministic rule-based extraction/normalization (for explicit context) with LoRA-specialized Mistral-7B-Instruct completion (for missing/ambiguous fields)
-
Has reproduction · 58
iCOMIC: a graphical interface-driven bioinformatics pipeline for analyzing cancer omics data.
PMID 35899080 · PMC9310080 · NAR genomics and bioinformatics · 2022 · 8 claims · 4 setups
iCOMIC provides a GUI-driven, Snakemake-based pipeline integrating multiple tools for DNA-Seq and RNA-Seq analysis with minimal command-line interaction.
-
Has reproduction · 67
Evaluating native-like structures of RNA-protein complexes through the deep learning method.
PMID 36828844 · PMC9958188 · Nature communications · 2023 · 8 claims · 7 setups
DRPScore identifies native-like RNA-protein structures with higher success rates than ITScore-PR, DARS-RNP, and 3dRPC across bound and unbound testing sets.
-
Full-text index only
A space-efficient and accurate method for mapping and aligning cDNA sequences onto genomic sequence.
PMID 18344523 · PMC2377433 · Nucleic acids research · 2008 · 7 claims · 6 setups
Spaln maps and aligns large cDNA sequence sets onto whole mammalian genomes using substantially less memory than comparable existing tools
-
Full-text index only
Identifying drug effects via pathway alterations using an integer linear programming optimization formulation on phosphoproteomic data.
PMID 19997482 · PMC2776985 · PLoS computational biology · 2009 · 7 claims · 4 setups
An ILP formulation of the Boolean pathway optimization problem fits phosphoproteomic data faster and more efficiently than the previously used genetic algorithm (GA) approach.
-
Full-text index only
Improved mutation tagging with gene identifiers applied to membrane protein stability prediction.
PMID 19758467 · PMC2745585 · BMC bioinformatics · 2009 · 8 claims · 4 setups
MutationTagger achieves 87% F-measure for the mutation retrieval task on a benchmark dataset
-
Full-text index only
CLUES A Comprehensive Workflow for Integrating Geospatial Data in Biomedical Research.
PMID 42128886 · PMC13172076 · Nature communications · 2026 · 8 claims · 5 setups
CLUES is an open-source, end-to-end workflow that automates selection, download, harmonization, and linkage of open-access geospatial environmental data to individual-level biomedical data without requiring geospatial expertise.
-
Full-text index only
VISTA uncovers missing gene expression and spatial-induced information for spatial transcriptomic data analysis.
PMID 41507434 · PMC12891734 · Communications biology · 2026 · 8 claims · 6 setups
VISTA predicts unmeasured gene expression in subcellular spatial transcriptomic data by integrating scRNA-seq and SST through variational inference and geometric deep learning with built-in uncertainty quantification
-
Has reproduction · 75
Inference of RNA polymerase II transcription dynamics from chromatin immunoprecipitation time course data.
PMID 24830797 · PMC4022483 · PLoS computational biology · 2014 · 8 claims · 8 setups
A convolved Gaussian process model of pol-II occupancy across gene segments captures the transcription wave and yields estimates of transcription speed and promoter-proximal pol-II activity.
-
Full-text index only
rMAP 2.0: a modular, reproducible, and scalable WDL-Cromwell-Docker workflow for genomic analysis of ESKAPEE pathogens.
PMID 41782684 · PMC12955837 · Bioinformatics advances · 2026 · 8 claims · 8 setups
rMAP 2.0 standardizes end-to-end bacterial WGS analysis (QC, trimming, assembly, annotation, AMR/virulence/mobile-element profiling, sequence typing, pangenome inference, phylogenetics) via containerized WDL/Cromwell execution
-
Full-text index only
EpiXFormer: a cross-attention neural network for predicting cell type-specific transcription factor binding sites.
PMID 41527854 · PMC12796812 · Briefings in bioinformatics · 2026 · 8 claims · 8 setups
EpiXFormer achieves high accuracy (mean AUROC ~0.99) predicting binding sites of both TFs and non-sequence-specific DBPs across 199 DBP-cell type pairs
-
Full-text index only
SMART: spatial multi-omic aggregation using graph neural networks and metric learning.
PMID 41896208 · PMC13031631 · Nature communications · 2026 · 8 claims · 5 setups
SMART accurately identifies spatial regions of anatomical structures and is compatible with spatial datasets of any type and number of omics layers
-
Has reproduction · 50
DeeReCT-APA: Prediction of Alternative Polyadenylation Site Usage Through Deep Learning.
PMID 33662629 · PMC9801043 · Genomics, proteomics & bioinformatics · 2022 · 8 claims · 8 setups
DeeReCT-APA quantitatively predicts the usage of all competing PASs of a gene simultaneously, rather than casting the problem as pairwise comparison like prior methods.
-
Has reproduction · 65
Wireless sensor network design with reliable and long network lifetime.
PMID 41981015 · PMC13083951 · Scientific reports · 2026 · 8 claims · 2 setups
Three strategies (Single Copy, Double Copy, Hybrid) jointly address the four WSN design problems (coverage, sink placement/routing, activity scheduling, data routing) together with network reliability in a unified framework.
-
Has reproduction · 85
Digital sorting of complex tissues for cell type-specific gene expression profiles.
PMID 23497278 · PMC3626856 · BMC bioinformatics · 2013 · 8 claims · 8 setups
The Digital Sorting Algorithm (DSA) deconvolves mixed tissue expression into cell type-specific profiles using only marker genes, without requiring prior knowledge of cell type frequencies or in vitro pure-cell profiles.