Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
Computational approaches for predicting the biological effect of p53 missense mutations: a comparison of three sequence analysis based methods.
PMID 16522644 · PMC1390679 · Nucleic acids research · 2006 · 7 claims · 6 setups
Align-GVGD predicts loss of transactivation activity with high specificity (~88%) but lower sensitivity (67.9-71.2%) for neutral mutants
-
Full-text index only
The DAVID Gene Functional Classification Tool: a novel biological module-centric algorithm to functionally analyze large gene lists.
PMID 17784955 · PMC2375021 · Genome biology · 2007 · 8 claims · 6 setups
Gene-gene functional similarity can be measured using kappa statistics applied to a binary gene-annotation-term matrix built from 14 annotation categories.
-
Full-text index only
Pol II promoter prediction using characteristic 4-mer motifs: a machine learning approach.
PMID 18834544 · PMC2575220 · BMC bioinformatics · 2008 · 8 claims · 8 setups
128 discriminating 4-mer motifs combined with an SVM (RBF kernel, LIBSVM) can distinguish promoter from non-promoter DNA sequences
-
Full-text index only
Exhaustive prediction of disease susceptibility to coding base changes in the human genome.
PMID 18793467 · PMC2537574 · BMC bioinformatics · 2008 · 8 claims · 7 setups
Inter-species conservation is the strongest single predictor of disease-associated coding mutations among the factors tested.
-
Full-text index only
Geometry-aware graph attention networks to explain single-cell chromatin states and gene expression with SEAGALL.
PMID 42026624 · PMC13238118 · Genome biology · 2026 · 8 claims · 6 setups
SEAGALL combines a geometry-regularised autoencoder (GRAE) to embed cells and build a cell-cell graph with a graph attention network (GAT) classifier and GNNExplainer-based XAI to identify features driving cell type/phenotype.
-
Full-text index only
Much ado about nothing: modeling amino acid replacement with predicted protein structures.
PMID 42036821 · PMC13171170 · Bioinformatics (Oxford, England) · 2026 · 8 claims · 7 setups
AFSM was constructed from over 660,000 structural alignments across ~21,000 proteins (297 InterPro families), following the BLOSUM log-odds methodology.
-
Full-text index only
Structural evolution of the protein kinase-like superfamily.
PMID 16244704 · PMC1261164 · PLoS computational biology · 2005 · 8 claims · 5 setups
All kinases in the superfamily share a 'universal core' domain consisting only of the regions required for ATP binding and the phosphotransfer reaction.
-
Full-text index only
An SVM-based system for predicting protein subnuclear localizations.
PMID 16336650 · PMC1325059 · BMC bioinformatics · 2005 · 7 claims · 3 setups
New kernels defined on k-peptide vectors mapped by BLOSUM62-based high-scored pair matrices (D1, D2, D3) improve SVM discrimination of protein subnuclear localization compared to conventional k-peptide encodings.
-
Full-text index only
Reverse polarization in amino acid and nucleotide substitution patterns between human-mouse orthologs of two compositional extrema.
PMID 17895298 · PMC2533592 · DNA research : an international journal for rapid publication of reports on genes and genomes · 2007 · 8 claims · 7 setups
Nucleotide and amino acid substitution trends between human-mouse orthologs are highly asymmetric and polarized in opposite directions for high-GC versus low-GC gene groups.
-
Has reproduction · 30
Minimal metabolic pathway structure is consistent with associated biomolecular interactions.
PMID 24987116 · PMC4299494 · Molecular systems biology · 2014 · 8 claims · 8 setups
MinSpan, a mixed-integer linear optimization algorithm, computes the shortest, linearly independent pathways (sparsest basis of the null space of the stoichiometric matrix S) for genome-scale metabolic networks, which convex approaches (extreme pathways, elementary flux modes) cannot do at genome scale.
-
Full-text index only
ChromBERT: A foundation model for learning interpretable representations for context-specific transcriptional regulatory networks.
PMID 41592570 · PMC13069865 · Cell genomics · 2026 · 8 claims · 7 setups
ChromBERT is pre-trained via masked reconstruction on the Cistrome-Human-6K dataset (6,391 cistromes, 991 transcription regulators) to learn genome-wide interaction syntax of transcription regulators
-
Full-text index only
Annotation-free prediction of immunotherapy response in melanoma using single-cell transcriptomic data.
PMID 41758825 · PMC12948085 · PloS one · 2026 · 8 claims · 6 setups
AI-based predictive models built on unannotated scRNA-seq data (cell-by-gene expression matrices) can classify melanoma patients as ICI responders vs. non-responders
-
Has reproduction · 90
Systematic clustering algorithm for chromatin accessibility data and its application to hematopoietic cells.
PMID 33253153 · PMC7728210 · PLoS computational biology · 2020 · 7 claims · 5 setups
A systematic clustering algorithm for ATAC-seq data can be built by binarizing the genome into open/closed chromatin (1/0) strings and computing Hamming distances between samples for hierarchical clustering.
-
Full-text index only
Protein expression patterns in primary carcinoma of the vagina.
PMID 15199389 · PMC2409807 · British journal of cancer · 2004 · 5 claims · 4 setups
Cluster analysis of 2-DE proteomic data allows accurate discrimination between normal vaginal mucosa, primary vaginal carcinoma and primary cervical carcinoma
-
Full-text index only
Interaction preferences across protein-protein interfaces of obligatory and non-obligatory components are different.
PMID 16105176 · PMC1201154 · BMC structural biology · 2005 · 8 claims · 5 setups
Interaction patterns across obligatory and non-obligatory interfaces are different, with obligatory contacts predominantly non-polar
-
Has reproduction · 76
Bayesian prediction of microbial oxygen requirement.
PMID 26913185 · PMC4743139 · F1000Research · 2013 · 7 claims · 8 setups
A naive Bayesian classifier based on presence/absence of class-associated Pfam-A domains can distinguish three oxygen requirement classes (aerobe, anaerobe, facultative anaerobe) from genome sequence, unlike prior studies that only made pairwise distinctions.
-
Full-text index only
BaGPipe: an automated, reproducible, and flexible pipeline for bacterial genome-wide association studies.
PMID 41896736 · PMC13147680 · BMC microbiology · 2026 · 7 claims · 8 setups
BaGPipe is an automated, reproducible Nextflow pipeline that integrates pre-processing, Pyseer-based association analysis, and downstream visualisation into a unified bacterial GWAS workflow
-
Full-text index only
RaMBat: Accurate identification of medulloblastoma subtypes from diverse data sources with severe batch effects.
PMID 41571436 · PMC13060657 · Molecular oncology · 2026 · 7 claims · 5 setups
RaMBat achieves a median accuracy of 99% across 13 independent benchmark datasets, significantly outperforming state-of-the-art MB subtyping methods and conventional ML classifiers
-
Full-text index only
Predicting FOX gene candidates for oxic nitrogen fixation using multi-omic machine learning and comparative bioinformatics.
PMID 41764348 · PMC13056922 · Scientific reports · 2026 · 8 claims · 6 setups
Random Forest, XGBoost, and logistic regression classifiers can meaningfully differentiate literature-validated FOX genes from conserved non-essential genes, with Random Forest achieving the best ROC-AUC (~0.80)
-
Full-text index only
Integrating co-expression network analysis and machine learning to reveal the regulatory landscape of GPD genes in Chlamydomonas reinhardtii under salinity stress.
PMID 42004709 · PMC13089224 · PeerJ · 2026 · 8 claims · 8 setups
GPD2 and GPD3 cluster into distinct WGCNA co-expression modules (magenta and black, respectively) with contrasting temporal expression profiles