Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Has reproduction · 94
Deep learning from phylogenies to uncover the epidemiological dynamics of outbreaks.
PMID 35794110 · PMC9258765 · Nature communications · 2022 · 8 claims · 5 setups
Deep learning (FFNN-SS and CNN-CBLV) enables accurate and fast likelihood-free estimation of epidemiological parameters and model selection from phylogenies
-
Full-text index only
On the analysis of glycomics mass spectrometry data via the regularized area under the ROC curve.
PMID 18076765 · PMC2211327 · BMC bioinformatics · 2007 · 8 claims · 4 setups
The TGDR-AUC algorithm regularizes the empirical AUC by replacing the non-differentiable 0-1 loss with a smooth sigmoid surrogate function and applies constrained threshold gradient descent regularization
-
Full-text index only
Effect of the assignment of ancestral CpG state on the estimation of nucleotide substitution rates in mammals.
PMID 18826599 · PMC2576242 · BMC evolutionary biology · 2008 · 7 claims · 4 setups
CpG/non-CpG assignment based on presence/absence of a CpG dinucleotide seriously biases substitution rate estimates, overestimating CpG changes and underestimating non-CpG changes.
-
Full-text index only
Stability analysis of mixtures of mutagenetic trees.
PMID 18366778 · PMC2335279 · BMC bioinformatics · 2008 · 7 claims · 5 setups
Mutagenetic trees mixture models capture multiple alternative pathways of ordered accumulation of genetic events (e.g., HIV resistance mutations, cancer chromosomal aberrations).
-
Full-text index only
A statistical change point model approach for the detection of DNA copy number variations in array CGH data.
PMID 19875853 · PMC4154476 · IEEE/ACM transactions on computational biology and bioinformatics · 2009 · 7 claims · 4 setups
A novel mean and variance change point model (MVCM) is proposed to detect CNVs/breakpoints in aCGH data.
-
Has reproduction · 89
Graph Random Forest: A Graph Embedded Algorithm for Identifying Highly Connected Important Features.
PMID 37509188 · PMC10377046 · Biomolecules · 2023 · 8 claims · 3 setups
Graph Random Forest (GRF) embeds graph/network information directly into the decision-tree building process by splitting on features in the k-hop neighborhood of a data-driven head-splitting node.
-
Has reproduction · 62
Equivalent change enrichment analysis: assessing equivalent and inverse change in biological pathways between diverse experiments.
PMID 32093613 · PMC7041296 · BMC genomics · 2020 · 7 claims · 3 setups
The Equivalent Change Index (ECI), a gene-level statistic ranging from -1 to 1, quantifies whether a gene was changed to the same (1) or completely opposite (-1) degree across two experiments relative to their controls.
-
Has reproduction · 98
Uncertainty in the mating strategy of honeybees causes bias and unreliability in the estimates of genetic parameters.
PMID 38632535 · PMC11022492 · Genetics, selection, evolution : GSE · 2024 · 7 claims · 3 setups
The most precise estimates of genetic parameters and genetic trends are obtained when breeding queens are mated with drones of a single DPQ that is correctly assigned in the pedigree (SS mating).
-
Has reproduction · 83
ConNIS and labeling instability: New statistical methods for improving the detection of essential genes in TraDIS libraries.
PMID 41790830 · PMC12991369 · PLoS computational biology · 2026 · 8 claims · 3 setups
ConNIS provides an analytic solution for the probability of observing the longest insertion-free sequence within a gene given its length and number of insertion sites under non-essentiality.
-
Full-text index only
Screening large-scale association study data: exploiting interactions using random forests.
PMID 15588316 · PMC545646 · BMC genetics · 2004 · 7 claims · 3 setups
Random forest importance measure significantly outperforms the Fisher Exact test as a screening tool when risk SNPs interact.
-
Full-text index only
Application of two machine learning algorithms to genetic association studies in the presence of covariates.
PMID 19014573 · PMC2620353 · BMC genetics · 2008 · 8 claims · 3 setups
The relative performance of RF and MARS for detecting genotype-trait associations depends on both the strategy used to handle covariates and the true underlying model of association (e.g., confounding vs. mediation vs. interaction).
-
Has reproduction · 62
Gbdmr: identifying differentially methylated CpG regions in the human genome via generalized beta regressions.
PMID 38443825 · PMC10916021 · BMC bioinformatics · 2024 · 8 claims · 4 setups
gbdmr models DNA methylation levels of CpG sites using a generalized beta distribution instead of assuming normality as in linear-regression-based methods
-
Full-text index only
Predicting survival outcomes using subsets of significant genes in prognostic marker studies with microarrays.
PMID 16549007 · PMC1544357 · BMC bioinformatics · 2006 · 7 claims · 2 setups
A methodology combining Cox proportional hazards models with a compound covariate, cross-validated log partial likelihood (ACVL) for predictive accuracy, and permutation-based significance testing can identify an optimal subset of significant genes for survival prediction
-
Full-text index only
Size matters: just how big is BIG?: Quantifying realistic sample size requirements for human genome epidemiology.
PMID 18676414 · PMC2639365 · International journal of epidemiology · 2009 · 7 claims · 2 setups
Conventional power calculations for case-control studies disregard analytic complexity (e.g. clinical assessment errors, unmeasured aetiological determinants) and can seriously underestimate true sample size requirements
-
Has reproduction · 40
DeepGSEA: explainable deep gene set enrichment analysis for single-cell transcriptomic data.
PMID 38950178 · PMC11236288 · Bioinformatics (Oxford, England) · 2024 · 8 claims · 2 setups
DeepGSEA is an explainable deep gene set enrichment analysis method built on interpretable, prototype-based neural networks.
-
Has reproduction · 60
UNMF: a unified nonnegative matrix factorization for multi-dimensional omics data.
PMID 37478378 · PMC10516365 · Briefings in bioinformatics · 2023 · 5 claims · 3 setups
UNMF is designed for tidy data format and structure, allowing it to handle a wide range of data structures and formats in a unified manner without requiring format-specific preprocessing.
-
Has reproduction · 78
Emergent dynamics of underlying regulatory network links EMT and androgen receptor-dependent resistance in prostate cancer.
PMID 36851919 · PMC9957767 · Computational and structural biotechnology journal · 2023 · 8 claims · 7 setups
Simulations of the EMT-AR crosstalk network reveal four possible phenotypes: epithelial-sensitive (ES), epithelial-resistant (ER), mesenchymal-resistant (MR), and mesenchymal-sensitive (MS), with MS occurring rarely
-
Has reproduction · 75
Revealing the critical state and identifying individualized dynamic network biomarker for type 2 diabetes through advanced analysis methods on individual basis.
PMID 39890881 · PMC11785715 · Scientific reports · 2025 · 8 claims · 5 setups
sJSD, NIG, and TNFE methods can detect critical states/tipping points before disease deterioration using only a single sample
-
Has reproduction · 85
A mechanistic model captures the emergence and implications of non-genetic heterogeneity and reversible drug resistance in ER+ breast cancer cells.
PMID 34316714 · PMC8271219 · NAR cancer · 2021 · 7 claims · 8 setups
EMT and tamoxifen-resistance (TamR) regulatory axes can drive one another, enabling non-genetic heterogeneity via six co-existing phenotypes (ES, ER, HS, HR, MS, MR)
-
Has reproduction · 84
An accurate method for identifying recent recombinants from unaligned sequences.
PMID 35025988 · PMC8963311 · Bioinformatics (Oxford, England) · 2022 · 8 claims · 4 setups
A novel algorithm combining the JHMM (Zilversmit et al. 2013) mosaic representation with a distance-based triple comparison can identify recombinant sequences and their parents from unaligned, gene-length sequences without a reference panel.