Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Has reproduction · 68
Machine learning algorithm predicts fibrosis-related blood diagnosis markers of intervertebral disc degeneration.
PMID 37915003 · PMC10619283 · BMC medical genomics · 2023 · 7 claims · 7 setups
CEP120 and SPDL1 are fibrosis-related diagnostic genes for IDD, identified via a random forest model from 29 differentially expressed fibrosis-related genes
-
Has reproduction · 69
Machine learning-based identification of an immunotherapy-related signature to enhance outcomes and immunotherapy responses in melanoma.
PMID 39355255 · PMC11442245 · Frontiers in immunology · 2024 · 8 claims · 8 setups
66 consensus immunotherapy prognostic genes (CITPGs) were identified from the intersection of WGCNA modules, immunotherapy responder-vs-non-responder DEGs, and tumor-vs-normal DEGs
-
Has reproduction · 86
Molecular Classification Models for Triple Negative Breast Cancer Subtype Using Machine Learning.
PMID 34575658 · PMC8472680 · Journal of personalized medicine · 2021 · 6 claims · 4 setups
A training gene set of 719 unique upregulated DEGs (subtype-specific) can be used to build ML models that classify TNBC into BLIA, BLIS, MES, and LAR subtypes.
-
Has reproduction · 42
Machine learning-based identification of biomarkers and drugs in immunologically cold and hot pancreatic adenocarcinomas.
PMID 39152432 · PMC11328457 · Journal of translational medicine · 2024 · 7 claims · 8 setups
PAAD tumors can be consensus-clustered into immunologically hot and cold subtypes based on CIBERSORT-derived immune cell fractions, with significantly different survival outcomes.
-
Has reproduction · 85
Predicting the pathogenicity of missense variants using features derived from AlphaFold2.
PMID 37084271 · PMC10203375 · Bioinformatics (Oxford, England) · 2023 · 6 claims · 8 setups
AlphaFold2-derived structural features (solvent accessibility, amino acid network features, physicochemical environment, pLDDT) can be used to train a random forest classifier (AlphScore) that distinguishes proxy-benign from proxy-pathogenic missense variants.
-
Has reproduction · 63
Clustering and machine learning-based integration identify cancer associated fibroblasts genes' signature in head and neck squamous cell carcinoma.
PMID 37065499 · PMC10098459 · Frontiers in genetics · 2023 · 8 claims · 8 setups
Clustering of 31 CAFs genes across 868 HNSCC samples identifies two distinct molecular patterns (C1, C2) with different survival outcomes
-
Full-text index only
Identification of diagnostic markers for tuberculosis by proteomic fingerprinting of serum.
PMID 16980117 · PMC7159276 · Lancet (London, England) · 2006 · 8 claims · 5 setups
An SVM classifier trained on serum proteomic profiles discriminated patients with active tuberculosis from controls with clinically overlapping conditions
-
Full-text index only
Cancer-specific high-throughput annotation of somatic mutations: computational prediction of driver missense mutations.
PMID 19654296 · PMC2763410 · Cancer research · 2009 · 7 claims · 7 setups
CHASM, a Random Forest-based computational method, was developed to identify and prioritize missense mutations likely to be functional drivers of tumor cell proliferation.
-
Full-text index only
Ab initio identification of human microRNAs based on structure motifs.
PMID 18088431 · PMC2238772 · BMC bioinformatics · 2007 · 8 claims · 7 setups
MiRPred predicts miRNA precursors ab initio using only predicted secondary structure motifs, ignoring nucleotide sequence
-
Has reproduction · 49
Integrative transcriptomics and single-cell transcriptomics analyses reveal potential biomarkers and mechanisms of action in papillary thyroid carcinoma.
PMID 40520228 · PMC12162626 · Frontiers in genetics · 2025 · 8 claims · 8 setups
ENTPD1, SERPINA1, and TACSTD2 are potential transcriptomic biomarkers for PTC
-
Has reproduction
Unlocking the microbial studies through computational approaches: how far have we reached?
PMID 36920617 · PMC10016191 · Environmental science and pollution research international · 2023 · 8 claims · 8 setups
Metagenomics enables culture-independent study of microbial communities directly from their natural environments, bypassing the need for clonal isolation.
-
Has reproduction · 58
iCOMIC: a graphical interface-driven bioinformatics pipeline for analyzing cancer omics data.
PMID 35899080 · PMC9310080 · NAR genomics and bioinformatics · 2022 · 8 claims · 4 setups
iCOMIC provides a GUI-driven, Snakemake-based pipeline integrating multiple tools for DNA-Seq and RNA-Seq analysis with minimal command-line interaction.
-
Has reproduction · 52
Social complexity, life-history and lineage influence the molecular basis of castes in vespid wasps.
PMID 36828829 · PMC9958023 · Nature communications · 2023 · 8 claims · 7 setups
A shared genetic toolkit of caste-associated genes exists across vespid wasp species spanning different levels of social complexity
-
Full-text index only
Integrated proteomic and transcriptomic profiling of mouse lung development and Nmyc target genes.
PMID 17486137 · PMC2673710 · Molecular systems biology · 2007 · 8 claims · 7 setups
Global MudPIT-based proteomic profiling across six mouse lung developmental time points (E13.5–P56) identifies thousands of proteins and captures developmental/cell-biological expression patterns.
-
Full-text index only
Extraction of human kinase mutations from literature, databases and genotyping studies.
PMID 19758464 · PMC2745582 · BMC bioinformatics · 2009 · 7 claims · 6 setups
A literature mining pipeline combining MutationFinder, false-positive filtering, and SVM-based classification can extract and disambiguate single-point mutation mentions from abstracts and full text
-
Full-text index only
Discovering cancer genes by integrating network and functional properties.
PMID 19765316 · PMC2758898 · BMC medical genomics · 2009 · 8 claims · 6 setups
Cancer genes have distinct PPI network topology (higher connectivity, higher clustering coefficient, shorter path length to known cancer genes) compared to non-cancer genes
-
Has reproduction · 44
Dynamic Gene Attention Focus (DyGAF): Enhancing Biomarker Identification Through Dual-Model Attention Networks.
PMID 40160891 · PMC11951896 · Bioinformatics and biology insights · 2025 · 6 claims · 5 setups
DyGAF, a dual-model attention neural network (independent Model A + dependent Model B), identifies and ranks genes by significance for COVID-19 biomarker discovery more effectively than differential expression analysis (DEA) and random forest (RF) feature selection
-
Has reproduction · 90
A Decentralized Kidney Transplant Biopsy Classifier for Transplant Rejection Developed Using Genes of the Banff-Human Organ Transplant Panel.
PMID 35619722 · PMC9128066 · Frontiers in immunology · 2022 · 6 claims · 6 setups
A random forest model trained solely on B-HOT panel genes (B-HOT Model) accurately classifies kidney transplant biopsies as NR, ABMR, or TCMR.
-
Full-text index only
Large-scale structural analysis of the core promoter in mammalian and plant genomes.
PMID 16049029 · PMC1181242 · Nucleic acids research · 2005 · 8 claims · 7 setups
DNA encodes at least two independent levels of functional information: protein/TF-binding sequence information and physical/structural properties of the molecule itself.