Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Has reproduction · 65
FusionQ: a novel approach for gene fusion detection and quantification from paired-end RNA-Seq.
PMID 23768108 · PMC3691734 · BMC bioinformatics · 2013 · 8 claims · 8 setups
FusionQ is a novel tool that detects gene fusions, constructs chimerical transcript structures, and estimates their abundances from paired-end RNA-Seq data.
-
Full-text index only
Testing whether genetic variation explains correlation of quantitative measures of gene expression, and application to genetic network analysis.
PMID 18444230 · PMC2729096 · Statistics in medicine · 2008 · 8 claims · 3 setups
A statistical test (delta method and Steiger-Browne optimal linear composites) is developed to test equality of the marginal correlation and the partial correlation of two gene expression traits conditional on a set of covariates.
-
Full-text index only
PedGenie: an analysis approach for genetic association testing in extended pedigrees and genealogies of arbitrary size.
PMID 16620382 · PMC1459209 · BMC bioinformatics · 2006 · 7 claims · 3 setups
PedGenie is a valid, flexible statistical tool for genetic association analysis in pedigrees of arbitrary size and structure using Monte Carlo significance testing
-
Has reproduction · 87
CoINcIDE: A framework for discovery of patient subtypes across multiple datasets.
PMID 26961683 · PMC4784276 · Genome medicine · 2016 · 8 claims · 6 setups
CoINcIDE is a methodological framework that discovers replicable patient subtypes (meta-clusters) across multiple datasets by finding consensus across dataset-specific clusterings, requiring no between-dataset transformations.
-
Has reproduction · 95
Mouse-Geneformer: A deep learning model for mouse single-cell transcriptome and its cross-species utility.
PMID 40106407 · PMC11964219 · PLoS genetics · 2025 · 7 claims · 6 setups
Mouse-Geneformer, a Transformer Encoder model pre-trained via masked-token self-supervised learning on mouse-Genecorpus-20M, was successfully constructed following the original human Geneformer architecture.
-
Has reproduction · 76
Correcting scale distortion in RNA sequencing data.
PMID 39875825 · PMC11776150 · BMC bioinformatics · 2025 · 8 claims · 8 setups
Local averaging reveals expression-level-dependent biases that differ from sample to sample across all RNA-seq datasets studied, and are not corrected by conventional normalization (TPM/FPKM)
-
Full-text index only
Identification of gene interactions associated with disease from gene expression data using synergy networks.
PMID 18234101 · PMC2258206 · BMC systems biology · 2008 · 8 claims · 4 setups
Synergy of a gene pair with respect to disease, defined as I(G1,G2;C) - [I(G1;C)+I(G2;C)], identifies gene pairs that interact cooperatively with respect to a phenotype rather than independently.
-
Full-text index only
Iterative class discovery and feature selection using Minimal Spanning Trees.
PMID 15355552 · PMC520744 · BMC bioinformatics · 2004 · 7 claims · 5 setups
Iterating between MST-based clustering and t-statistic feature selection removes noise genes step-wise while sharpening the sample clustering
-
Has reproduction · 40
DeepGSEA: explainable deep gene set enrichment analysis for single-cell transcriptomic data.
PMID 38950178 · PMC11236288 · Bioinformatics (Oxford, England) · 2024 · 8 claims · 2 setups
DeepGSEA is an explainable deep gene set enrichment analysis method built on interpretable, prototype-based neural networks.
-
Full-text index only
Bayesian survival analysis in genetic association studies.
PMID 18617538 · PMC2530885 · Bioinformatics (Oxford, England) · 2008 · 7 claims · 5 setups
A novel Bayesian method (BETA-Surv) extends prior case-control haplotype-clustering work to censored survival outcomes by clustering haplotypes via gene tree/perfect phylogeny topology and relative mutation age.
-
Full-text index only
The origins of lactase persistence in Europe.
PMID 19714206 · PMC2722739 · PLoS computational biology · 2009 · 8 claims · 5 setups
The −13,910*T allele first underwent selection among dairying farmers around 7,500 years ago in a region between the central Balkans and central Europe, possibly linked to the Linearbandkeramik culture.
-
Has reproduction · 62
Equivalent change enrichment analysis: assessing equivalent and inverse change in biological pathways between diverse experiments.
PMID 32093613 · PMC7041296 · BMC genomics · 2020 · 7 claims · 3 setups
The Equivalent Change Index (ECI), a gene-level statistic ranging from -1 to 1, quantifies whether a gene was changed to the same (1) or completely opposite (-1) degree across two experiments relative to their controls.
-
Has reproduction · 67
Comparison of Metagenomics and Metatranscriptomics Tools: A Guide to Making the Right Choice.
PMID 36553546 · PMC9777648 · Genes · 2022 · 8 claims · 1 setups
16S rRNA gene sequencing enables taxonomic identification of bacteria/archaea via hypervariable regions without amplifying human DNA, but is limited by short-read biases (GC bias, sequencing errors) and poor species-level resolution
-
Has reproduction · 75
Revealing the critical state and identifying individualized dynamic network biomarker for type 2 diabetes through advanced analysis methods on individual basis.
PMID 39890881 · PMC11785715 · Scientific reports · 2025 · 8 claims · 5 setups
sJSD, NIG, and TNFE methods can detect critical states/tipping points before disease deterioration using only a single sample
-
Full-text index only
Modeling the amplification dynamics of human Alu retrotransposons.
PMID 16201008 · PMC1239904 · PLoS computational biology · 2005 · 8 claims · 4 setups
Combining sequence diversity (π) and insertion polymorphism level (IPL) statistics can statistically exclude implausible Alu amplification scenarios and narrow the range of plausible ones for individual subfamilies.
-
Full-text index only
Operon information improves gene expression estimation for cDNA microarrays.
PMID 16630355 · PMC1513396 · BMC genomics · 2006 · 7 claims · 3 setups
A hierarchical Bayesian model that borrows expression information from other genes within the same operon improves estimation of relative transcript levels.
-
Full-text index only
Predicting survival outcomes using subsets of significant genes in prognostic marker studies with microarrays.
PMID 16549007 · PMC1544357 · BMC bioinformatics · 2006 · 7 claims · 2 setups
A methodology combining Cox proportional hazards models with a compound covariate, cross-validated log partial likelihood (ACVL) for predictive accuracy, and permutation-based significance testing can identify an optimal subset of significant genes for survival prediction
-
Has reproduction · 86
RNASEQR--a streamlined and accurate RNA-seq sequence analysis program.
PMID 22199257 · PMC3315322 · Nucleic acids research · 2012 · 8 claims · 7 setups
RNASEQR is a new RNA-seq mapper/aligner that combines a BWT-based (Bowtie) transcriptomic/genomic alignment with hash-based BLAT local alignment in three sequential steps: transcriptome mapping, novel exon detection, and anchor-and-align novel splice junction identification.
-
Has reproduction · 85
Single-Cell Differential Network Analysis with Sparse Bayesian Factor Models.
PMID 35186014 · PMC8855158 · Frontiers in genetics · 2021 · 8 claims · 2 setups
A hierarchical Bayesian factor model using treatment-dependent latent factor loadings can construct gene co-expression networks from scRNA-seq data and identify differences in network structure between two (or more) biological conditions.
-
Has reproduction · 86
Multi-INTACT: integrative analysis of the genome, transcriptome, and proteome identifies causal mechanisms of complex traits.
PMID 39901160 · PMC11789355 · Genome biology · 2025 · 8 claims · 4 setups
Multi-INTACT aggregates colocalization and TWAS evidence across diverse gene products (e.g., expression and protein) within a Bayesian/empirical Bayes framework to implicate putative causal genes and identify the relevant gene product(s).