Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
Early feature extraction drives model performance in high-resolution chromatin accessibility prediction.
PMID 41526189 · PMC12951969 · Genome research · 2026 · 8 claims · 6 setups
Early feature extraction (via ConvNeXt V2 blocks), rather than downstream architecture type, is the primary determinant of prediction accuracy in high-resolution chromatin accessibility prediction.
-
Full-text index only
A comprehensive toolkit for analyzing cell-free DNA genomic sequencing data in liquid biopsy.
PMID 42111187 · PMC13157187 · iScience · 2026 · 8 claims · 8 setups
cfDNAanalyzer integrates feature extraction, feature processing/selection, and machine learning model building into a single one-command-line toolkit for cfDNA genomic sequencing data
-
Full-text index only
Statistical challenges in preprocessing in microarray experiments in cancer.
PMID 18829474 · PMC3529914 · Clinical cancer research : an official journal of the American Association for Cancer Research · 2008 · 8 claims · 7 setups
Choice of pre-processing method materially changes which features are found significantly associated with survival in the Beer et al. lung cancer microarray dataset
-
Full-text index only
Multi-omics feature engineering driven by biomedical foundation models improves drug response prediction for inflammatory bowel disease patients.
PMID 41844950 · PMC13129071 · Scientific reports · 2026 · 8 claims · 7 setups
FM (MAMMAL)-derived drug-target binding affinity (BA) inference can be used to rank/select biologically relevant protein targets and their associated genes/SNPs for a drug of interest without knowledge of protein structure or active sites
-
Full-text index only
Speeding disease gene discovery by sequence based candidate prioritization.
PMID 15766383 · PMC1274252 · BMC bioinformatics · 2005 · 7 claims · 8 setups
Disease genes (OMIM) differ significantly from non-disease genes in sequence-based features including gene/cDNA/protein size, exon number, homolog conservation, secretion signal, 3' UTR length, CpG islands, and distance to nearest gene.
-
Full-text index only
Discovering cancer genes by integrating network and functional properties.
PMID 19765316 · PMC2758898 · BMC medical genomics · 2009 · 8 claims · 6 setups
Cancer genes have distinct PPI network topology (higher connectivity, higher clustering coefficient, shorter path length to known cancer genes) compared to non-cancer genes
-
Full-text index only
Leveraging the germ layer development patterns to predict prognosis and identify MEST as a novel therapeutic target in glioma.
PMID 41501725 · PMC12870398 · Cancer cell international · 2026 · 7 claims · 8 setups
MEST is a key oncogenic GLD-related gene and a novel therapeutic target in glioma, identified via a machine learning feature selection framework
-
Full-text index only
Fast-evolving noncoding sequences in the human genome.
PMID 17578567 · PMC2394770 · Genome biology · 2007 · 8 claims · 6 setups
1,356 conserved noncoding sequences show human-specific accelerated substitution rates (ANC sequences) relative to chimpanzee
-
Full-text index only
AMR-GNN: a multi-representation graph neural network framework to enable genomic antimicrobial resistance prediction.
PMID 41792137 · PMC13087051 · Nature communications · 2026 · 7 claims · 8 setups
AMR-GNN, a graph neural network integrating multiple genomic representations (unitigs, SNPs, FCGR) via low-rank multimodal fusion, improves AMR phenotype prediction in P. aeruginosa compared to single-representation baseline models.
-
Full-text index only
Single-cell transcriptomics identifies regulatory T cell heterogeneity in gestational diabetes mellitus.
PMID 41933176 · PMC13223204 · Communications medicine · 2026 · 8 claims · 8 setups
Treg cluster proportions do not significantly differ between GDM and healthy control patients
-
Full-text index only
Prioritization of candidate cancer genes--an aid to oncogenomic studies.
PMID 18710882 · PMC2566894 · Nucleic acids research · 2008 · 8 claims · 8 setups
Computational classifiers using combinations of protein conservation, gene structure, protein domains, protein interactions, and regulatory data can distinguish known cancer genes (CD/CR) from unlabelled human genes
-
Full-text index only
Detecting purely epistatic multi-locus interactions by an omnibus permutation test on ensembles of two-locus analyses.
PMID 19761607 · PMC2759961 · BMC bioinformatics · 2009 · 8 claims · 5 setups
2LOmb performs an omnibus permutation test on ensembles of two-locus analyses via a four-step algorithm (two-locus analysis, permutation test, global p-value determination, progressive ensemble search)
-
Full-text index only
Mapping Genetic Regulation of Transcription to Identify Functional Variants and Genes Associated with Pancreatic Cancer Risk.
PMID 41824785 · PMC13205582 · Advanced science (Weinheim, Baden-Wurttemberg, Germany) · 2026 · 8 claims · 8 setups
A genome-wide cis-eQTL meta-analysis of 482 pancreatic tissues (177 TCGA tumor + 305 GTEx normal) identified 1,123,483 significant SNP-gene pairs, 709,720 unique eQTLs, and 13,758 eGenes (FDR<0.05)
-
Full-text index only
Machine learning-driven transcriptomic and single-cell profiling of programed cell death patterns in colon cancer.
PMID 41854308 · PMC13009810 · Science progress · 2026 · 8 claims · 8 setups
Disulfidptosis and anoikis are consistently identified as the most robust prognostic PCD patterns in colon cancer across six independent machine learning algorithms
-
Full-text index only
A novel deep learning-driven framework for improving lncRNA comprehensive annotation with LncADeep 2.0.
PMID 41923359 · PMC13090826 · Bioinformatics (Oxford, England) · 2026 · 8 claims · 8 setups
LncADeep 2.0 outperforms LncADeep and other existing tools for lncRNA identification on both GENCODE annotated transcripts and independent RNA-seq data
-
Has reproduction · 70
Predicting enhancers in mammalian genomes using supervised hidden Markov models.
PMID 30917778 · PMC6437899 · BMC bioinformatics · 2019 · 8 claims · 8 setups
eHMM predicts enhancers with high precision and recall comparable to state-of-the-art methods and consistently outperforms them in accuracy and resolution
-
Has reproduction · 58
Identification of common genetic characteristics of rheumatoid arthritis and major depressive disorder by bioinformatics analysis and machine learning.
PMID 37415981 · PMC10320004 · Frontiers in immunology · 2023 · 7 claims · 8 setups
EAF1, SDCBP and RNF19B are common genetic characteristics (hub genes) shared by RA and MDD
-
Has reproduction · 67
binny: an automated binning algorithm to recover high-quality genomes from complex metagenomic datasets.
PMID 36239393 · PMC9677464 · Briefings in bioinformatics · 2022 · 8 claims · 8 setups
binny outperforms or is highly competitive with commonly used and state-of-the-art binning methods (MetaBAT2, MaxBin2, CONCOCT, VAMB, SemiBin, MetaDecoder)
-
Has reproduction · 52
Social complexity, life-history and lineage influence the molecular basis of castes in vespid wasps.
PMID 36828829 · PMC9958023 · Nature communications · 2023 · 8 claims · 7 setups
A shared genetic toolkit of caste-associated genes exists across vespid wasp species spanning different levels of social complexity
-
Full-text index only
A computational study of off-target effects of RNA interference.
PMID 15800213 · PMC1072799 · Nucleic acids research · 2005 · 8 claims · 5 setups
The chance of RNAi off-target effects is considerable, ranging from 5% to 80% depending on organism and parameters, when using exact sequence identity between siRNA and transcripts.