Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Has reproduction · 63
RummaGEO: Automatic mining of human and mouse gene sets from GEO.
PMID 39569206 · PMC11573963 · Patterns (New York, N.Y.) · 2024 · 8 claims · 7 setups
RummaGEO is a gene expression signature search engine built from automatically mined human and mouse RNA-seq perturbation studies in GEO
-
Has reproduction · 50
SMAC, a computational system to link literature, biomedical and expression data.
PMID 31324861 · PMC6642118 · Scientific reports · 2019 · 7 claims · 6 setups
SMAC is a tool that extracts, prioritises, integrates and analyses biomedical literature and molecular data according to user-defined terms, linking PubMed and GEO.
-
Has reproduction · 62
Gbdmr: identifying differentially methylated CpG regions in the human genome via generalized beta regressions.
PMID 38443825 · PMC10916021 · BMC bioinformatics · 2024 · 8 claims · 4 setups
gbdmr models DNA methylation levels of CpG sites using a generalized beta distribution instead of assuming normality as in linear-regression-based methods
-
Has reproduction · 67
Generative and integrative modeling for transcriptomics with formalin fixed paraffin embedded material.
PMID 41029822 · PMC12486589 · Journal of translational medicine · 2025 · 8 claims · 5 setups
fRNA-seq transcript counts are best fit by the negative binomial distribution, with little evidence supporting zero-inflated extensions
-
Has reproduction · 83
Public Omics Explorer (POE): Enabling integrative semantic search across GEO omics datasets based on PubMed publications.
PMID 41282419 · PMC12636342 · Computational and structural biotechnology journal · 2025 · 6 claims · 4 setups
POE is a web platform that semantically links GEO datasets and ENA records through their associated PubMed publications for literature-informed dataset retrieval
-
Has reproduction · 95
nf-rnaSeqCount: A Nextflow pipeline for obtaining raw read counts from RNA-seq data.
PMID 35574063 · PMC9097006 · South African computer journal = Suid-Afrikaanse rekenaartydskrif · 2021 · 7 claims · 5 setups
nf-rnaSeqCount is a portable, reproducible Nextflow pipeline that maps RNA-seq reads to a reference genome and quantifies gene abundance for differential expression analysis
-
Has reproduction · 71
Hyb: a bioinformatics pipeline for the analysis of CLASH (crosslinking, ligation and sequencing of hybrids) data.
PMID 24211736 · PMC3969109 · Methods (San Diego, Calif.) · 2014 · 8 claims · 6 setups
The 'hyb' pipeline detects, calls, folds and annotates chimeric reads from CLASH high-throughput sequencing data.
-
Has reproduction · 95
Mouse-Geneformer: A deep learning model for mouse single-cell transcriptome and its cross-species utility.
PMID 40106407 · PMC11964219 · PLoS genetics · 2025 · 7 claims · 6 setups
Mouse-Geneformer, a Transformer Encoder model pre-trained via masked-token self-supervised learning on mouse-Genecorpus-20M, was successfully constructed following the original human Geneformer architecture.
-
Has reproduction · 75
ResnetAge: A Resnet-Based DNA Methylation Age Prediction Method.
PMID 38247911 · PMC10813502 · Bioengineering (Basel, Switzerland) · 2023 · 8 claims · 4 setups
ResnetAge, a ResNet-based neural network using 22,278 shared Illumina 27K/450K CpG sites, predicts DNA methylation age from beta values.
-
Has reproduction · 51
SGCP: a spectral self-learning method for clustering genes in co-expression networks.
PMID 38956463 · PMC11221046 · BMC bioinformatics · 2024 · 7 claims · 4 setups
SGCP, a spectral self-learning method, yields gene co-expression modules with higher GO enrichment than WGCNA, CoExpNets, and CEMiTool across 12 real gene expression datasets.
-
Has reproduction · 53
Combining evidence of preferential gene-tissue relationships from multiple sources.
PMID 23950964 · PMC3741196 · PloS one · 2013 · 8 claims · 8 setups
A high-level integration approach combining three methods across four human microarray datasets, merged by consensus voting and a rule-based inner/total score, predicts preferentially expressed genes while reducing method- and study-specific bias.
-
Has reproduction
Using random walks to identify cancer-associated modules in expression data.
PMID 24128261 · PMC4015830 · BioData mining · 2013 · 8 claims · 8 setups
Walktrap-GM, a random-walk community detection algorithm adapted with stopping criteria (maximum modularity, maximum size, maximum module score), identifies modules significantly enriched with cancer genes in expression-weighted interaction networks.
-
Has reproduction · 79
Enriched domain detector: a program for detection of wide genomic enrichment domains robust against local variations.
PMID 24782521 · PMC4066758 · Nucleic acids research · 2014 · 8 claims · 5 setups
EDD is a new algorithm that detects broad (megabase-size) enrichment domains from ChIP-seq data of widely distributed chromatin proteins such as A- and B-type lamins.