Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
The fragile breakage versus random breakage models of chromosome evolution.
PMID 16501665 · PMC1378107 · PLoS computational biology · 2006 · 8 claims · 6 setups
Sankoff and Trinh's synteny block identification algorithm (ST-Synteny) is flawed, producing erroneous block identifications even in small toy examples.
-
Full-text index only
Bcipep: a database of B-cell epitopes.
PMID 15921533 · PMC1173103 · BMC genomics · 2005 · 8 claims · 2 setups
Bcipep is a comprehensive database of experimentally determined linear B-cell epitopes compiled from literature and other public databases
-
Has reproduction · 98
Projecting contact matrices in 177 geographical regions: An update and comparison with empirical data for the COVID-19 era.
PMID 34310590 · PMC8354454 · PLoS computational biology · 2021 · 7 claims · 6 setups
Updated synthetic contact matrices were generated for 177 geographical locations covering 97.2% of the world's population (up from 152 locations/95.9% in 2017).
-
Full-text index only
Visualization of shared genomic regions and meiotic recombination in high-density SNP data.
PMID 19696932 · PMC2725774 · PloS one · 2009 · 8 claims · 7 setups
SNPduo is a command-line (SNPduo++) and web-accessible tool that analyzes and visualizes relatedness between two individuals using identity by state (IBS) from SNP genotypes.
-
Full-text index only
Iterative class discovery and feature selection using Minimal Spanning Trees.
PMID 15355552 · PMC520744 · BMC bioinformatics · 2004 · 7 claims · 5 setups
Iterating between MST-based clustering and t-statistic feature selection removes noise genes step-wise while sharpening the sample clustering
-
Has reproduction · 89
Graph Random Forest: A Graph Embedded Algorithm for Identifying Highly Connected Important Features.
PMID 37509188 · PMC10377046 · Biomolecules · 2023 · 8 claims · 3 setups
Graph Random Forest (GRF) embeds graph/network information directly into the decision-tree building process by splitting on features in the k-hop neighborhood of a data-driven head-splitting node.
-
Has reproduction · 78
Taxonomic analysis of metagenomic data with kASA.
PMID 33784400 · PMC8266618 · Nucleic acids research · 2021 · 8 claims · 3 setups
kASA achieves high sensitivity and precision by using an amino acid-like encoding of k-mers together with a range of multiple k's
-
Has reproduction · 80
DMN-seq enriches DNA hypomethylated regions for biomarker discovery using 5-methylcytosine glycosylase.
PMID 41673887 · PMC13097799 · Genome biology · 2026 · 8 claims · 9 setups
DMN-seq (DMN+) uses DME to nick DNA specifically at 5mC sites, enabling 5mC detection at single-base resolution via selective adaptor ligation
-
Full-text index only
DiagHunter and GenoPix2D: programs for genomic comparisons, large-scale homology discovery and visualization.
PMID 14519203 · PMC328457 · Genome biology · 2003 · 7 claims · 5 setups
DiagHunter identifies large-scale synteny blocks within or between genomes efficiently despite background noise and genomic discontinuities, without performing sequence alignment
-
Has reproduction · 73
treeclimbR pinpoints the data-dependent resolution of hierarchical hypotheses.
PMID 34001188 · PMC8127214 · Genome biology · 2021 · 7 claims · 7 setups
treeclimbR proposes multiple candidate resolutions on a hierarchical tree and selects the optimal one in a data-driven way to pinpoint signal branches/leaves of interest.
-
Full-text index only
Stability analysis of mixtures of mutagenetic trees.
PMID 18366778 · PMC2335279 · BMC bioinformatics · 2008 · 7 claims · 5 setups
Mutagenetic trees mixture models capture multiple alternative pathways of ordered accumulation of genetic events (e.g., HIV resistance mutations, cancer chromosomal aberrations).
-
Full-text index only
Grammar-based distance in progressive multiple sequence alignment.
PMID 18616828 · PMC2478692 · BMC bioinformatics · 2008 · 7 claims · 3 setups
A grammar-based (LZ complexity) distance metric can be used to determine the order in which sequences are progressively pairwise aligned
-
Has reproduction · 53
spliceJAC: transition genes and state-specific gene regulation from single-cell transcriptome data.
PMID 36321549 · PMC9627675 · Molecular systems biology · 2022 · 8 claims · 6 setups
spliceJAC quantifies multivariate mRNA splicing from unspliced/spliced count matrices to construct cell state-specific gene-gene (Jacobian) interaction matrices.
-
Full-text index only
Flanking p10 contribution and sequence bias in matrix based epitope prediction: revisiting the assumption of independent binding pockets.
PMID 18925947 · PMC2600787 · BMC structural biology · 2008 · 8 claims · 3 setups
The extended matrix PP10 (built from a proline-containing peptide library) shows significant improvement in binding prediction over the original nine-residue matrix P9
-
Full-text index only
Cancer-specific high-throughput annotation of somatic mutations: computational prediction of driver missense mutations.
PMID 19654296 · PMC2763410 · Cancer research · 2009 · 7 claims · 7 setups
CHASM, a Random Forest-based computational method, was developed to identify and prioritize missense mutations likely to be functional drivers of tumor cell proliferation.
-
Has reproduction
Human Retrotransposons and Effective Computational Detection Methods for Next-Generation Sequencing Data.
PMID 36295018 · PMC9605557 · Life (Basel, Switzerland) · 2022 · 8 claims · 7 setups
Transposable elements make up nearly 45% of the human genome, vastly exceeding the ~1.5% that is protein-coding.
-
Has reproduction · 50
MoDLE: high-performance stochastic modeling of DNA loop extrusion interactions.
PMID 36451166 · PMC9710047 · Genome biology · 2022 · 7 claims · 6 setups
MoDLE is a high-performance stochastic model that simulates DNA-DNA contacts from loop extrusion genome-wide in minutes using less than 1 GB of RAM
-
Full-text index only
Identification of gene interactions associated with disease from gene expression data using synergy networks.
PMID 18234101 · PMC2258206 · BMC systems biology · 2008 · 8 claims · 4 setups
Synergy of a gene pair with respect to disease, defined as I(G1,G2;C) - [I(G1;C)+I(G2;C)], identifies gene pairs that interact cooperatively with respect to a phenotype rather than independently.
-
Full-text index only
MEROPS: the peptidase database.
PMID 19892822 · PMC2808883 · Nucleic acids research · 2010 · 8 claims · 5 setups
MEROPS is a manually curated hierarchical classification of peptidases and protein inhibitors organized into protein species, families, and clans based on sequence and structural homology.
-
Has reproduction · 67
Generative and integrative modeling for transcriptomics with formalin fixed paraffin embedded material.
PMID 41029822 · PMC12486589 · Journal of translational medicine · 2025 · 8 claims · 6 setups
The negative binomial distribution best fits fRNA-seq transcript counts, with little evidence supporting zero-inflated extensions