Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
Exogean: a framework for annotating protein-coding genes in eukaryotic genomic DNA.
PMID 16925841 · PMC1810556 · Genome biology · 2006 · 8 claims · 5 setups
Exogean is a framework using directed acyclic coloured multigraphs (DACMs) to represent biological objects (mRNA, ESTs, protein alignments, exons) and iteratively combine them into complex protein-coding transcript models.
-
Full-text index only
AUGUSTUS at EGASP: using EST, protein and genomic alignments for improved gene prediction in the human genome.
PMID 16925833 · PMC1810548 · Genome biology · 2006 · 8 claims · 5 setups
AUGUSTUS predicted significantly more genes correctly than any other ab initio program in EGASP
-
Has reproduction · 80
SLDMS: A Tool for Calculating the Overlapping Regions of Sequences.
PMID 35046988 · PMC8761809 · Frontiers in plant science · 2021 · 8 claims · 5 setups
SLDMS is a novel method for computing overlapping regions of sequencing reads using suffix array (SA), longest common prefix (LCP) array, document array (DA), and a monotonic stack.
-
Full-text index only
Benchmarking tools for the alignment of functional noncoding DNA.
PMID 14736341 · PMC344529 · BMC bioinformatics · 2004 · 8 claims · 4 setups
Global alignment tools (Avid, ClustalW, Lagan, Needle, DiAlign-G) typically have higher sensitivity over entire noncoding sequences and within constrained blocks than local tools
-
Has reproduction
Using random walks to identify cancer-associated modules in expression data.
PMID 24128261 · PMC4015830 · BioData mining · 2013 · 8 claims · 8 setups
Walktrap-GM, a random-walk community detection algorithm adapted with stopping criteria (maximum modularity, maximum size, maximum module score), identifies modules significantly enriched with cancer genes in expression-weighted interaction networks.
-
Full-text index only
Function2Gene: a gene selection tool to increase the power of genetic association studies by utilizing public databases and expert knowledge.
PMID 18631403 · PMC2500032 · BMC bioinformatics · 2008 · 6 claims · 5 setups
Function2Gene is a set of Perl programs that queries public databases (NCBI, GeneCards, Harvester, with Uniprot/Ensembl also supported) using expert-selected keywords to rank genes by prior probability of disease association.
-
Has reproduction · 69
Automatic discovery of 100-miRNA signature for cancer classification using ensemble feature selection.
PMID 31533612 · PMC6751684 · BMC bioinformatics · 2019 · 8 claims · 6 setups
An ensemble feature selection strategy using consensus of feature relevance across 8 classifier types identifies a 100-miRNA signature from a 1046-feature TCGA dataset
-
Has reproduction · 32
Developing prognostic gene panel of survival time in lung adenocarcinoma patients using machine learning.
PMID 35117753 · PMC8799101 · Translational cancer research · 2020 · 8 claims · 5 setups
Naïve Bayes using a 22-gene panel is the best-performing and most stable machine learning model for predicting LUAD survival time (>3 vs <3 years)
-
Full-text index only
A machine learning approach uncovers principles and determinants of eukaryotic ribosome pausing.
PMID 39423268 · PMC11488575 · Science advances · 2024 · 8 claims · 5 setups
An unsupervised ML pipeline using the extended isolation forest (EIF) algorithm can reliably detect ribosome pausing sites from noisy, coverage-biased RiboSeq data across expression levels
-
Full-text index only
nf-core/viralmetagenome: A novel pipeline for untargeted viral genome reconstruction.
PMID 42057295 · PMC13141149 · Bioinformatics (Oxford, England) · 2026 · 8 claims · 5 setups
nf-core/viralmetagenome is a Nextflow pipeline that automates untargeted reconstruction and variant analysis of eukaryotic DNA and RNA viruses from short-read metagenomic or hybridisation-capture data.
-
Full-text index only
EpiToolKit--a web server for computational immunomics.
PMID 18440979 · PMC2447732 · Nucleic acids research · 2008 · 7 claims · 3 setups
EpiToolKit is a web server integrating five MHC class I and two MHC class II epitope prediction methods in a unified, user-friendly interface.
-
Full-text index only
OTMODE: an optimal transport theory-based framework for identifying differential features in single-cell multi-omics data.
PMID 41335419 · PMC12766913 · Bioinformatics (Oxford, England) · 2026 · 8 claims · 8 setups
OTMODE, using an unbalanced Sinkhorn algorithm and Wald test, improves differential feature identification in single-cell multi-omics data
-
Full-text index only
PPC: an algorithm for accurate estimation of SNP allele frequencies in small equimolar pools of DNA using data from high density microarrays.
PMID 16199750 · PMC1240117 · Nucleic acids research · 2005 · 7 claims · 6 setups
The PPC algorithm, which applies a probe-pair-specific second-degree polynomial correction, increases the accuracy of allele frequency estimates from pooled DNA compared with previously described algorithms
-
Full-text index only
Computational tradeoffs in multiplex PCR assay design for SNP genotyping.
PMID 16042802 · PMC1190169 · BMC genomics · 2005 · 7 claims · 6 setups
Achieving high-multiplexing/high-coverage multiplex PCR designs is subject to a computational phase transition as the SNP-pair compatibility probability crosses a critical threshold
-
Full-text index only
Duplex-Indel: a Snakemake pipeline for somatic Indel calling in Tn5 transposase-based duplex sequencing data.
PMID 42046229 · PMC13171174 · Bioinformatics (Oxford, England) · 2026 · 8 claims · 8 setups
Duplex-Indel is a Snakemake pipeline for somatic Indel calling from Tn5 transposase-based duplex sequencing data that requires consensus support from both DNA strands to minimize technical artifacts.
-
Full-text index only
Features affecting Cas9-induced editing efficiency and patterns in tomato: evidence from a large CRISPR dataset.
PMID 41877594 · PMC13014117 · The Plant journal : for cell and molecular biology · 2026 · 8 claims · 5 setups
Chromatin accessibility significantly increases editing efficiency, with higher editing at targets in accessible versus inaccessible chromatin.
-
Full-text index only
FineST: contrastive learning integrates histology and spatial transcriptomics for nuclei-resolved ligand-receptor analysis.
PMID 41839892 · PMC13201544 · Nature communications · 2026 · 8 claims · 6 setups
FineST, a bimodal contrastive learning model integrating histology (Virchow2 ViT features) and spatial gene expression, enables nuclei-resolved high-resolution RNA imputation.
-
Has reproduction · 67
HArmonized single-cell RNA-seq Cell type Assisted Deconvolution (HASCAD).
PMID 37907883 · PMC10619225 · BMC medical genomics · 2023 · 6 claims · 4 setups
Removal of batch effects in reference scRNA-seq datasets (via Harmony-Symphony) benefits the task of cell composition deconvolution
-
Has reproduction · 92
Large-scale integration of single-cell transcriptomic data captures transitional progenitor states in mouse skeletal muscle regeneration.
PMID 34773081 · PMC8589952 · Communications biology · 2021 · 8 claims · 7 setups
Large-scale integration of 111 sc/snRNAseq datasets captures rare, transitional myogenic progenitor states (commitment and fusion) that are poorly represented in individual datasets.
-
Full-text index only
Deep-learning prediction of gene expression from personal genomes.
PMID 41495833 · PMC12869966 · Genome biology · 2026 · 8 claims · 8 setups
Fine-tuning Enformer on paired personal WGS and RNA-seq data (Variformer) corrects Enformer's failure to predict inter-individual gene expression differences across held-out people.