Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Has reproduction · 58
iCOMIC: a graphical interface-driven bioinformatics pipeline for analyzing cancer omics data.
PMID 35899080 · PMC9310080 · NAR genomics and bioinformatics · 2022 · 8 claims · 4 setups
iCOMIC provides a GUI-driven, Snakemake-based pipeline integrating multiple tools for DNA-Seq and RNA-Seq analysis with minimal command-line interaction.
-
Has reproduction · 81
SEMdag: Fast learning of Directed Acyclic Graphs via node or layer ordering.
PMID 39775401 · PMC11709272 · PloS one · 2025 · 8 claims · 5 setups
SEMdag() is a two-step order-based algorithm for fast learning of high-dimensional linear SEMs, using knowledge-based (KB) or data-driven bottom-up (BU) node/layer ordering followed by penalized (L1) DAG estimation
-
Has reproduction · 50
SMAC, a computational system to link literature, biomedical and expression data.
PMID 31324861 · PMC6642118 · Scientific reports · 2019 · 7 claims · 6 setups
SMAC is a tool that extracts, prioritises, integrates and analyses biomedical literature and molecular data according to user-defined terms, linking PubMed and GEO.
-
Full-text index only
GeneKeyDB: a lightweight, gene-centric, relational database to support data mining environments.
PMID 15790402 · PMC1274265 · BMC bioinformatics · 2005 · 8 claims · 6 setups
GeneKeyDB is a lightweight, gene-centric relational database that supports data mining and integration with computational analysis tools.
-
Has reproduction · 92
Prognostic biomarker discovery in pancreatic cancer through hybrid ensemble feature selection and multi-omics data.
PMID 41957754 · PMC13188360 · BioData mining · 2026 · 7 claims · 3 setups
The hEFS framework integrates data subsampling with multiple prognostic models (embedded and wrapper-based), aggregates feature rankings via a voting-theory-inspired approach, and selects the optimal feature subset via Pareto front optimization, eliminating user-defined thresholds.
-
Has reproduction · 96
Scalable Prediction of Acute Myeloid Leukemia Using High-Dimensional Machine Learning and Blood Transcriptomics.
PMID 31918046 · PMC6992905 · iScience · 2020 · 8 claims · 8 setups
Data-driven, high-dimensional ML approaches that learn multivariate signatures directly from genome-wide transcriptomic data (no prior gene selection) yield accurate and robust AML classifiers.
-
Has reproduction · 63
RummaGEO: Automatic mining of human and mouse gene sets from GEO.
PMID 39569206 · PMC11573963 · Patterns (New York, N.Y.) · 2024 · 8 claims · 7 setups
RummaGEO is a gene expression signature search engine built from automatically mined human and mouse RNA-seq perturbation studies in GEO
-
Has reproduction · 51
SGCP: a spectral self-learning method for clustering genes in co-expression networks.
PMID 38956463 · PMC11221046 · BMC bioinformatics · 2024 · 7 claims · 4 setups
SGCP, a spectral self-learning method, yields gene co-expression modules with higher GO enrichment than WGCNA, CoExpNets, and CEMiTool across 12 real gene expression datasets.
-
Has reproduction · 85
Digital sorting of complex tissues for cell type-specific gene expression profiles.
PMID 23497278 · PMC3626856 · BMC bioinformatics · 2013 · 8 claims · 8 setups
The Digital Sorting Algorithm (DSA) deconvolves mixed tissue expression into cell type-specific profiles using only marker genes, without requiring prior knowledge of cell type frequencies or in vitro pure-cell profiles.
-
Has reproduction · 53
Combining evidence of preferential gene-tissue relationships from multiple sources.
PMID 23950964 · PMC3741196 · PloS one · 2013 · 8 claims · 8 setups
A high-level integration approach combining three methods across four human microarray datasets, merged by consensus voting and a rule-based inner/total score, predicts preferentially expressed genes while reducing method- and study-specific bias.
-
Has reproduction
Using random walks to identify cancer-associated modules in expression data.
PMID 24128261 · PMC4015830 · BioData mining · 2013 · 8 claims · 8 setups
Walktrap-GM, a random-walk community detection algorithm adapted with stopping criteria (maximum modularity, maximum size, maximum module score), identifies modules significantly enriched with cancer genes in expression-weighted interaction networks.
-
Full-text index only
AUGUSTUS at EGASP: using EST, protein and genomic alignments for improved gene prediction in the human genome.
PMID 16925833 · PMC1810548 · Genome biology · 2006 · 8 claims · 5 setups
AUGUSTUS predicted significantly more genes correctly than any other ab initio program in EGASP
-
Full-text index only
Functional annotation and identification of candidate disease genes by computational analysis of normal tissue gene expression data.
PMID 18560577 · PMC2409962 · PloS one · 2008 · 7 claims · 5 setups
Ranked Coexpression Groups (RCG) built from k=6 nearest coexpressed genes, combined with a majority-rule functional characterization, integrate multiple datasets/coexpression measures to generate high-confidence functional annotation predictions
-
Has reproduction · 67
GEMmaker: process massive RNA-seq datasets on heterogeneous computational infrastructure.
PMID 35501696 · PMC9063052 · BMC bioinformatics · 2022 · 6 claims · 3 setups
GEMmaker, an nf-core compliant Nextflow workflow, can quantify gene expression from small to massive RNA-seq datasets while remaining reproducible via versioned containerized software.
-
Has reproduction · 62
Gbdmr: identifying differentially methylated CpG regions in the human genome via generalized beta regressions.
PMID 38443825 · PMC10916021 · BMC bioinformatics · 2024 · 8 claims · 4 setups
gbdmr models DNA methylation levels of CpG sites using a generalized beta distribution instead of assuming normality as in linear-regression-based methods
-
Full-text index only
Genomics--from Neanderthals to high-throughput sequencing.
PMID 16934106 · PMC1779599 · Genome biology · 2006 · 8 claims · 8 setups
Next-generation sequencing platforms (GS20/454 and Solexa) can deliver the throughput and cost reductions needed for population-scale and medical resequencing.
-
Has reproduction
miRge3.0: a comprehensive microRNA and tRF sequencing analysis pipeline.
PMID 34308351 · PMC8294687 · NAR genomics and bioinformatics · 2021 · 6 claims · 5 setups
miRge3.0 with 12 CPUs consistently has the best execution speed compared to miRge2.0, Chimira and sRNAbench
-
Has reproduction · 50
Viewing RNA-seq data on the entire human genome.
PMID 28979763 · PMC5605993 · F1000Research · 2017 · 6 claims · 3 setups
RNA-Seq Viewer is a web application that visualizes genome-wide expression data from NCBI's SRA and GEO databases using an ideogram across the entire human genome.
-
Has reproduction · 78
IsomiR_Window: a system for analyzing small-RNA-seq data in an integrative and user-friendly manner.
PMID 33522913 · PMC7852101 · BMC bioinformatics · 2021 · 8 claims · 2 setups
IsomiR Window is an integrated, user-friendly platform that systematically identifies, quantifies, and functionally explores isomiR expression in small-RNA-seq datasets without requiring computational skills
-
Has reproduction · 97
CellFishing.jl: an ultrafast and scalable cell search method for single-cell RNA sequencing.
PMID 30744683 · PMC6371477 · Genome biology · 2019 · 8 claims · 8 setups
CellFishing.jl searches prebuilt databases for cells with similar expression patterns with high accuracy and throughput using locality-sensitive hashing.