Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
Lightweight genome viewer: portable software for browsing genomics data in its chromosomal context.
PMID 17877794 · PMC2238324 · BMC bioinformatics · 2007 · 7 claims · 7 setups
lwgv provides a lightweight alternative to large genome browsers for visualizing biological annotations and dynamic analyses without requiring a database or complex software infrastructure
-
Has reproduction · 92
Chromosome-scale genome sequencing, assembly and annotation of six genomes from subfamily Leishmaniinae.
PMID 34489462 · PMC8421402 · Scientific data · 2021 · 8 claims · 8 setups
Chromosome-scale genomes of six Leishmaniinae species (five L. (Mundinia) species and one Porcisia species) were sequenced, assembled and annotated, providing genome, proteome, transcriptome and GFF outputs for taxa previously lacking public reference genomes
-
Has reproduction · 100
ChIP-seq Data Processing and Relative and Quantitative Signal Normalization for Saccharomyces cerevisiae.
PMID 40364978 · PMC12067309 · Bio-protocol · 2025 · 8 claims · 6 setups
siQ-ChIP measures absolute protein–DNA interaction (IP efficiency) genome-wide without relying on exogenous spike-in chromatin, overcoming limitations of spike-in normalization.
-
Has reproduction · 44
An OMICs-based meta-analysis to support infection state stratification.
PMID 33560295 · PMC8388022 · Bioinformatics (Oxford, England) · 2021 · 7 claims · 6 setups
Multi-class machine learning models built from cross-platform microarray meta-analysis can distinguish bacterial, viral and no-infection states with high accuracy (best model: 93% bacterial, 89% viral correct).
-
Has reproduction · 62
scATD: a high-throughput and interpretable framework for single-cell cancer drug resistance prediction and biomarker identification.
PMID 40501071 · PMC12159290 · Briefings in bioinformatics · 2025 · 8 claims · 6 setups
scATD enables high-throughput single-cell drug sensitivity prediction for new patients without model parameter retraining via bidirectional Bi-AdaIN style transfer
-
Has reproduction · 48
Comparative analysis of molecular signatures reveals a hybrid approach in breast cancer: Combining the Nottingham Prognostic Index with gene expressions into a hybrid signature.
PMID 35143511 · PMC8830616 · PloS one · 2022 · 8 claims · 6 setups
A hybrid signature combining the Nottingham Prognostic Index with SIS-selected gene expressions can be built in a data-driven fashion (NPI treated as a gene expression during feature selection).
-
Has reproduction · 86
Molecular Classification Models for Triple Negative Breast Cancer Subtype Using Machine Learning.
PMID 34575658 · PMC8472680 · Journal of personalized medicine · 2021 · 6 claims · 4 setups
A training gene set of 719 unique upregulated DEGs (subtype-specific) can be used to build ML models that classify TNBC into BLIA, BLIS, MES, and LAR subtypes.
-
Full-text index only
Speeding disease gene discovery by sequence based candidate prioritization.
PMID 15766383 · PMC1274252 · BMC bioinformatics · 2005 · 7 claims · 8 setups
Disease genes (OMIM) differ significantly from non-disease genes in sequence-based features including gene/cDNA/protein size, exon number, homolog conservation, secretion signal, 3' UTR length, CpG islands, and distance to nearest gene.
-
Has reproduction · 83
Gene-expression patterns in peripheral blood classify familial breast cancer susceptibility.
PMID 26538066 · PMC4634735 · BMC medical genomics · 2015 · 8 claims · 5 setups
A multigene peripheral-blood gene-expression biomarker accurately classifies which women from high-risk families develop familial breast cancer.
-
Has reproduction · 59
Refining breast cancer biomarker discovery and drug targeting through an advanced data-driven approach.
PMID 38253993 · PMC10810249 · BMC bioinformatics · 2024 · 8 claims · 8 setups
The BGWO_SA_Ens algorithm (hybrid BGWO + simulated annealing with an ensemble classifier objective function) selects predictive breast cancer biomarker genes with high classification performance
-
Has reproduction · 89
Graph Random Forest: A Graph Embedded Algorithm for Identifying Highly Connected Important Features.
PMID 37509188 · PMC10377046 · Biomolecules · 2023 · 8 claims · 3 setups
Graph Random Forest (GRF) embeds graph/network information directly into the decision-tree building process by splitting on features in the k-hop neighborhood of a data-driven head-splitting node.
-
Has reproduction · 83
Analyzing biomarker discovery: Estimating the reproducibility of biomarker sets.
PMID 35901020 · PMC9333302 · PloS one · 2022 · 7 claims · 3 setups
A Reproducibility Score, RS(D,BD), defined as the average Jaccard overlap between biomarker sets found by the same discovery process on comparable datasets from the same distribution, quantifies biomarker reproducibility on a 0-1 scale
-
Full-text index only
Atlas - a data warehouse for integrative bioinformatics.
PMID 15723693 · PMC554782 · BMC bioinformatics · 2005 · 8 claims · 3 setups
Atlas is a biological data warehouse that locally stores and integrates sequences, molecular interactions, homology information, functional annotations, and ontologies
-
Has reproduction · 55
Identification and verification of diagnostic biomarkers in recurrent pregnancy loss via machine learning algorithm and WGCNA.
PMID 37691920 · PMC10485775 · Frontiers in immunology · 2023 · 8 claims · 8 setups
352 DEGs (198 up-regulated, 154 down-regulated) were identified between RPL and control endometrial samples
-
Has reproduction · 80
Specific signature biomarkers highlight the potential mechanisms of circulating neutrophils in aneurysmal subarachnoid hemorrhage.
PMID 36438795 · PMC9685413 · Frontiers in pharmacology · 2022 · 7 claims · 8 setups
Six genes (CST7, HSP90AB1, PADI4, PLBD1, RAB32, SLAMF6) are signature diagnostic biomarkers for aSAH identified by LASSO and SVM-RFE.
-
Has reproduction · 44
Dynamic Gene Attention Focus (DyGAF): Enhancing Biomarker Identification Through Dual-Model Attention Networks.
PMID 40160891 · PMC11951896 · Bioinformatics and biology insights · 2025 · 6 claims · 5 setups
DyGAF, a dual-model attention neural network (independent Model A + dependent Model B), identifies and ranks genes by significance for COVID-19 biomarker discovery more effectively than differential expression analysis (DEA) and random forest (RF) feature selection
-
Has reproduction · 66
Integrative bioinformatics and artificial intelligence analyses of transcriptomics data identified genes associated with major depressive disorders including NRG1.
PMID 37583471 · PMC10423927 · Neurobiology of stress · 2023 · 7 claims · 5 setups
Differentially expressed genes in MDD patients are enriched in immune response, inflammatory response, neurodegeneration, and cerebellar atrophy pathways.
-
Has reproduction · 90
A Decentralized Kidney Transplant Biopsy Classifier for Transplant Rejection Developed Using Genes of the Banff-Human Organ Transplant Panel.
PMID 35619722 · PMC9128066 · Frontiers in immunology · 2022 · 6 claims · 6 setups
A random forest model trained solely on B-HOT panel genes (B-HOT Model) accurately classifies kidney transplant biopsies as NR, ABMR, or TCMR.
-
Full-text index only
Towards the identification of essential genes using targeted genome sequencing and comparative analysis.
PMID 17052348 · PMC1624830 · BMC genomics · 2006 · 8 claims · 8 setups
Phyletic retention (ortholog presence across organisms) is the single most predictive feature of gene essentiality in both E. coli and S. cerevisiae.
-
Has reproduction · 92
Prognostic biomarker discovery in pancreatic cancer through hybrid ensemble feature selection and multi-omics data.
PMID 41957754 · PMC13188360 · BioData mining · 2026 · 7 claims · 3 setups
The hEFS framework integrates data subsampling with multiple prognostic models (embedded and wrapper-based), aggregates feature rankings via a voting-theory-inspired approach, and selects the optimal feature subset via Pareto front optimization, eliminating user-defined thresholds.