Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
PolySearch: a web-based text mining system for extracting relationships between human diseases, genes, mutations, drugs and metabolites.
PMID 18487273 · PMC2447794 · Nucleic acids research · 2008 · 8 claims · 7 setups
PolySearch supports more than 50 different classes of queries against nearly a dozen types of text, abstract, or bioinformatic databases
-
Has reproduction · 78
Taxonomic analysis of metagenomic data with kASA.
PMID 33784400 · PMC8266618 · Nucleic acids research · 2021 · 8 claims · 3 setups
kASA achieves high sensitivity and precision by using an amino acid-like encoding of k-mers together with a range of multiple k's
-
Full-text index only
PreTSA: computationally efficient modeling of temporal and spatial gene expression patterns.
PMID 41673899 · PMC12998178 · Genome biology · 2026 · 7 claims · 8 setups
PreTSA dramatically reduces computational time and memory versus GAM (Monocle, TSCAN) and PseudotimeDE for identifying temporally variable genes (TVGs) while producing highly similar results
-
Full-text index only
TiRank prioritizes phenotypic niches in tumor microenvironment for clinical biomarker discovery.
PMID 41689080 · PMC12910759 · Genome medicine · 2026 · 7 claims · 4 setups
TiRank is a framework that integrates scRNA-seq, ST, and bulk transcriptomes using an REO-transformation module and multitask transfer learning to align data into a unified embedding space for prioritizing clinically relevant spatial niches
-
Has reproduction · 50
MEDUSA: A Pipeline for Sensitive Taxonomic Classification and Flexible Functional Annotation of Metagenomic Shotgun Sequences.
PMID 35330728 · PMC8940201 · Frontiers in genetics · 2022 · 7 claims · 6 setups
MEDUSA correctly identifies more species than MEGAN 6 CE, especially less abundant species.
-
Has reproduction · 67
Evaluating native-like structures of RNA-protein complexes through the deep learning method.
PMID 36828844 · PMC9958188 · Nature communications · 2023 · 8 claims · 7 setups
DRPScore identifies native-like RNA-protein structures with higher success rates than ITScore-PR, DARS-RNP, and 3dRPC across bound and unbound testing sets.
-
Full-text index only
Systematic evaluation of single-cell multimodal data integration enhances cell type resolution and discovery of clinically relevant states in complex tissues.
PMID 41821037 · PMC12983708 · Genome biology · 2026 · 8 claims · 8 setups
Horizontal integration of scRNA-seq and snRNA-seq improves cell-type identification
-
Has reproduction · 42
CanCellCap: robust cancer cell capture across tissue types on single-cell RNA-seq data by multi-domain learning.
PMID 40739511 · PMC12312500 · BMC biology · 2025 · 8 claims · 7 setups
CanCellCap identifies cancer cells in scRNA-seq data across 13 tissue types, 23 cancer types, and 7 sequencing platforms with 0.977 average accuracy
-
Has reproduction · 64
Widespread allele-specific topological domains in the human genome are not confined to imprinted gene clusters.
PMID 36869353 · PMC9983196 · Genome biology · 2023 · 8 claims · 5 setups
HiCFlow, a new bioinformatic pipeline, performs de novo haplotype assembly, phasing, and visualization of allele-specific (parental) chromatin conformation directly from Hi-C data without requiring pre-phased haplotypes.
-
Has reproduction · 79
Interpretable prediction models for widespread m6A RNA modification across cell lines and tissues.
PMID 37995291 · PMC10697738 · Bioinformatics (Oxford, England) · 2023 · 8 claims · 8 setups
CLSM6A is a set of CNN-based deep learning models that predict single-nucleotide-resolution m6A RNA modification sites across eight cell lines and three tissues in H. sapiens
-
Has reproduction · 100
Intra-Host Co-Existing Strains of SARS-CoV-2 Reference Genome Uncovered by Exhaustive Computational Search.
PMID 37243151 · PMC10224212 · Viruses · 2023 · 8 claims · 7 setups
An exhaustive-search workflow can recover intra-host co-existing SARS-CoV-2 strains from the reference-genome read set (SRR11092062) that de Bruijn-graph assemblers discard.
-
Full-text index only
Analysis of protein sequence and interaction data for candidate disease gene prediction.
PMID 17020920 · PMC1636487 · Nucleic acids research · 2006 · 8 claims · 7 setups
Combining CPS and CMP using known disease genes as input achieves sensitivity 0.52 and specificity 0.97, reducing candidate lists 13-fold
-
Full-text index only
Uncovering Cas9 PAM diversity through metagenomic mining and machine learning.
PMID 41656299 · PMC12996302 · Nature communications · 2026 · 8 claims · 6 setups
CRISPR-PAMdb is a publicly accessible database compiling Cas9 protein sequences from 3.8 million bacterial/archaeal genomes and PAM profiles from 7.4 million phage/plasmid sequences
-
Full-text index only
Multimodal-based analysis of single-cell ATAC-seq data enables highly accurate delineation of clinically relevant tumor cell subpopulations.
PMID 41530870 · PMC12888741 · Genome medicine · 2026 · 8 claims · 8 setups
MAAS integrates chromatin accessibility, CNVs, and SNVs from scATAC-seq data to identify functional tumor cell subpopulations
-
Has reproduction · 52
epiGBS2: Improvements and evaluation of highly multiplexed, epiGBS-based reduced representation bisulfite sequencing.
PMID 35178872 · PMC9311447 · Molecular ecology resources · 2022 · 8 claims · 8 setups
epiGBS2 provides a laboratory protocol and revised bioinformatics pipeline for de novo cytosine methylation and SNP calling in species with or without a reference genome
-
Has reproduction · 45
Identifying and classifying trait linked polymorphisms in non-reference species by walking coloured de bruijn graphs.
PMID 23536903 · PMC3607606 · PloS one · 2013 · 8 claims · 9 setups
Bubbleparse detects sequence variants directly from NGS reads without a reference genome, using the coloured de Bruijn graph implementation of Cortex plus a new depth-first bubble-finding module.
-
Full-text index only
WeavePop: a bioinformatics workflow to explore and analyze genomic variants of eukaryotic populations.
PMID 41685638 · PMC13042275 · G3 (Bethesda, Md.) · 2026 · 8 claims · 7 setups
WeavePop is a novel Snakemake-based, reproducible, scalable workflow that performs reference-based read mapping, assembly, annotation, small variant calling/effect prediction, and CNV detection for eukaryotic haploid organisms
-
Full-text index only
Benchmarking LLM-based agents for single-cell omics analysis.
PMID 41742311 · PMC13064268 · Genome biology · 2026 · 8 claims · 8 setups
Introduces a comprehensive benchmarking evaluation system comprising an open-source agent platform, 18 evaluation metrics across four dimensions, and 50 real-world single-cell omics tasks
-
Has reproduction · 83
Macrel: antimicrobial peptide screening in genomes and metagenomes.
PMID 33384902 · PMC7751412 · PeerJ · 2020 · 8 claims · 8 setups
Macrel introduces a novel set of 22 peptide features (6 local, 16 global), including a new Free Energy Transition (FET) feature group, for AMP and hemolytic activity classification
-
Has reproduction · 42
The electrostatic profile of consecutive Cβ atoms applied to protein structure quality assessment.
PMID 25506420 · PMC4257144 · F1000Research · 2013 · 8 claims · 8 setups
The EPD between Cβ atoms of consecutive residues provides unique signatures of amino acid pair types and can discriminate native from decoy protein structures.