Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Has reproduction · 80
SLDMS: A Tool for Calculating the Overlapping Regions of Sequences.
PMID 35046988 · PMC8761809 · Frontiers in plant science · 2021 · 8 claims · 5 setups
SLDMS is a novel method for computing overlapping regions of sequencing reads using suffix array (SA), longest common prefix (LCP) array, document array (DA), and a monotonic stack.
-
Full-text index only
Automatic annotation of eukaryotic genes, pseudogenes and promoters.
PMID 16925832 · PMC1810547 · Genome biology · 2006 · 8 claims · 6 setups
Fgenesh++ gene prediction pipeline identifies 91% of coding nucleotides with 90% specificity
-
Full-text index only
POCUS: mining genomic sequence annotation to predict disease genes.
PMID 14611661 · PMC329128 · Genome biology · 2003 · 8 claims · 6 setups
Genes predisposing to the same disease tend to share functional annotation IDs (GO/InterPro) more than expected by chance
-
Has reproduction · 83
Public Omics Explorer (POE): Enabling integrative semantic search across GEO omics datasets based on PubMed publications.
PMID 41282419 · PMC12636342 · Computational and structural biotechnology journal · 2025 · 6 claims · 4 setups
POE is a web platform that semantically links GEO datasets and ENA records through their associated PubMed publications for literature-informed dataset retrieval
-
Has reproduction · 63
RummaGEO: Automatic mining of human and mouse gene sets from GEO.
PMID 39569206 · PMC11573963 · Patterns (New York, N.Y.) · 2024 · 8 claims · 7 setups
RummaGEO is a gene expression signature search engine built from automatically mined human and mouse RNA-seq perturbation studies in GEO
-
Has reproduction · 51
SGCP: a spectral self-learning method for clustering genes in co-expression networks.
PMID 38956463 · PMC11221046 · BMC bioinformatics · 2024 · 7 claims · 4 setups
SGCP, a spectral self-learning method, yields gene co-expression modules with higher GO enrichment than WGCNA, CoExpNets, and CEMiTool across 12 real gene expression datasets.
-
Full-text index only
Gene Prospector: an evidence gateway for evaluating potential susceptibility genes and interacting risk factors for human diseases.
PMID 19063745 · PMC2613935 · BMC bioinformatics · 2008 · 8 claims · 5 setups
Gene Prospector is a Web-based application that selects and prioritizes potential disease-related genes using a curated, updated literature database of genetic association studies
-
Full-text index only
CpG_MI: a novel approach for identifying functional CpG islands in mammalian genomes.
PMID 19854943 · PMC2800233 · Nucleic acids research · 2010 · 8 claims · 6 setups
Functional ('bona fide') CGIs show distinct average/cumulative mutual information (AMI/CMI) distributions of neighboring CpG distances compared to non-functional CGIs and random genome segments
-
Has reproduction · 89
Statistical framework for calling allelic imbalance in high-throughput sequencing data.
PMID 39966391 · PMC11836314 · Nature communications · 2025 · 8 claims · 6 setups
MIXALIME is a versatile computational framework for calling allele-specific variants (ASVs) from diverse high-throughput omics data
-
Has reproduction · 53
Combining evidence of preferential gene-tissue relationships from multiple sources.
PMID 23950964 · PMC3741196 · PloS one · 2013 · 8 claims · 8 setups
A high-level integration approach combining three methods across four human microarray datasets, merged by consensus voting and a rule-based inner/total score, predicts preferentially expressed genes while reducing method- and study-specific bias.
-
Has reproduction · 68
LaSSO, a strategy for genome-wide mapping of intronic lariats and branch points using RNA-seq.
PMID 24709818 · PMC4079972 · Genome research · 2014 · 8 claims · 8 setups
LaSSO (Lariat Sequence Site Origin) identifies intronic lariat reads and pinpoints branch points genome-wide from RNA-seq data by considering every intronic base as a potential branch point and including all possible exon-skipping lariats.
-
Full-text index only
Functional annotation and identification of candidate disease genes by computational analysis of normal tissue gene expression data.
PMID 18560577 · PMC2409962 · PloS one · 2008 · 7 claims · 5 setups
Ranked Coexpression Groups (RCG) built from k=6 nearest coexpressed genes, combined with a majority-rule functional characterization, integrate multiple datasets/coexpression measures to generate high-confidence functional annotation predictions
-
Full-text index only
Designating eukaryotic orthology via processed transcription units.
PMID 18445630 · PMC2425467 · Nucleic acids research · 2008 · 8 claims · 5 setups
Existing ortholog databases discard/ignore alternative splicing via all-against-all protein comparisons, causing ambiguous ortholog calls and misclassification of AS isoforms as in-paralogs
-
Has reproduction · 100
Smart spatial omics (S2-omics) optimizes region of interest selection to capture molecular heterogeneity in diverse tissues.
PMID 41298871 · PMC12662399 · Nature cell biology · 2025 · 7 claims · 6 setups
S2-omics is an end-to-end workflow that automatically selects ROIs from H&E histology images to maximize molecular information content for spatial omics profiling.
-
Full-text index only
GENCODE: producing a reference annotation for ENCODE.
PMID 16925838 · PMC1810553 · Genome biology · 2006 · 8 claims · 8 setups
GENCODE annotation combines initial manual annotation by HAVANA, experimental validation, and refinement based on results to identify protein-coding genes in ENCODE regions
-
Has reproduction · 50
Wx: a neural network-based feature selection algorithm for transcriptomic data.
PMID 31324856 · PMC6642261 · Scientific reports · 2019 · 8 claims · 8 setups
The Wx algorithm ranks genes by a discriminative index (DI) score representing classification power for distinguishing given groups, enabling intuitive selection of optimal biomarker genes.
-
Has reproduction
Using random walks to identify cancer-associated modules in expression data.
PMID 24128261 · PMC4015830 · BioData mining · 2013 · 8 claims · 8 setups
Walktrap-GM, a random-walk community detection algorithm adapted with stopping criteria (maximum modularity, maximum size, maximum module score), identifies modules significantly enriched with cancer genes in expression-weighted interaction networks.
-
Full-text index only
Function2Gene: a gene selection tool to increase the power of genetic association studies by utilizing public databases and expert knowledge.
PMID 18631403 · PMC2500032 · BMC bioinformatics · 2008 · 6 claims · 5 setups
Function2Gene is a set of Perl programs that queries public databases (NCBI, GeneCards, Harvester, with Uniprot/Ensembl also supported) using expert-selected keywords to rank genes by prior probability of disease association.