Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Has reproduction · 78
Enhancing chemotherapy response prediction via matched colorectal tumor-organoid gene expression analysis and network-based biomarker selection.
PMID 39754813 · PMC11754497 · Translational oncology · 2025 · 6 claims · 8 setups
A consensus WGCNA approach combining matched tumor-organoid and independent organoid drug-response expression data identifies gene modules and hub genes predictive of 5-FU chemotherapy response
-
Full-text index only
Discordance of species trees with their most likely gene trees.
PMID 16733550 · PMC1464820 · PLoS genetics · 2006 · 7 claims · 2 setups
For any species tree topology with n ≥ 5 taxa, there exist branch lengths for which the most likely gene tree topology (an 'anomalous gene tree') differs from the species tree topology.
-
Has reproduction · 75
Sequencing of human genomes with nanopore technology.
PMID 31015479 · PMC6478738 · Nature communications · 2019 · 8 claims · 7 setups
A novel single-sample, reference panel-free, read-based phasing algorithm built on the STITCH model improves nanopore SNV calling from modest baseline levels.
-
Has reproduction · 84
An integrated in silico-in vitro approach for identifying therapeutic targets against osteoarthritis.
PMID 36352408 · PMC9648005 · BMC biology · 2022 · 7 claims · 5 setups
A signal transduction/gene regulatory network model of the articular chondrocyte was built combining knowledge-based curation and data-driven (machine learning) network inference
-
Full-text index only
Integrated analysis of genetic and proteomic data identifies biomarkers associated with adverse events following smallpox vaccination.
PMID 18923431 · PMC2692715 · Genes and immunity · 2009 · 7 claims · 6 setups
A two-stage strategy (Random Forest filtering followed by decision tree modeling) can integrate categorical genetic and continuous proteomic data to identify biomarkers of AE risk
-
Full-text index only
Evolutionary distance estimation and fidelity of pair wise sequence alignment.
PMID 15840174 · PMC1087827 · BMC bioinformatics · 2005 · 8 claims · 8 setups
Evolutionary distance estimation is relatively unaffected by alignment error as long as 50% or more of homologous sites remain identical between sequences
-
Has reproduction · 100
Gene signature discovery and systematic validation across diverse clinical cohorts for TB prognosis and response to treatment.
PMID 37471455 · PMC10393163 · PLoS computational biology · 2023 · 8 claims · 8 setups
A network-based meta-analysis across studies identifies a common 45-gene signature specific to active TB disease that accounts for cohort/population heterogeneity
-
Has reproduction · 85
Predicting the pathogenicity of missense variants using features derived from AlphaFold2.
PMID 37084271 · PMC10203375 · Bioinformatics (Oxford, England) · 2023 · 6 claims · 8 setups
AlphaFold2-derived structural features (solvent accessibility, amino acid network features, physicochemical environment, pLDDT) can be used to train a random forest classifier (AlphScore) that distinguishes proxy-benign from proxy-pathogenic missense variants.
-
Has reproduction · 94
Topological signatures in regulatory network enable phenotypic heterogeneity in small cell lung cancer.
PMID 33729159 · PMC8012062 · eLife · 2021 · 7 claims · 6 setups
Discrete (Boolean/Ising) and continuous (RACIPE) simulations of the SCLC regulatory network yield similar multistable phenotypic distributions, with four dominant steady states (X1-X4) that map onto experimentally observed SCLC molecular subtypes.
-
Has reproduction · 44
An OMICs-based meta-analysis to support infection state stratification.
PMID 33560295 · PMC8388022 · Bioinformatics (Oxford, England) · 2021 · 7 claims · 6 setups
Multi-class machine learning models built from cross-platform microarray meta-analysis can distinguish bacterial, viral and no-infection states with high accuracy (best model: 93% bacterial, 89% viral correct).
-
Full-text index only
Consolidating the set of known human protein-protein interactions in preparation for large-scale mapping of the human interactome.
PMID 15892868 · PMC1175952 · Genome biology · 2005 · 8 claims · 6 setups
Two quantitative benchmarks (functional-annotation-based and physical-interaction-based log likelihood ratio scores) can measure relative accuracy of human PPI datasets
-
Full-text index only
Defective splicing, disease and therapy: searching for master checkpoints in exon definition.
PMID 16855287 · PMC1524908 · Nucleic acids research · 2006 · 8 claims · 8 setups
Splicing-affecting genomic variations can account for up to 50% of mutations leading to gene dysfunction in some genes
-
Full-text index only
The fragile breakage versus random breakage models of chromosome evolution.
PMID 16501665 · PMC1378107 · PLoS computational biology · 2006 · 8 claims · 6 setups
Sankoff and Trinh's synteny block identification algorithm (ST-Synteny) is flawed, producing erroneous block identifications even in small toy examples.
-
Full-text index only
An oncogenomics-based in vivo RNAi screen identifies tumor suppressors in liver cancer.
PMID 19012953 · PMC2990916 · Cell · 2008 · 7 claims · 8 setups
shRNA pools targeting genes recurrently deleted in human HCC accelerate hepatocarcinogenesis in vivo, whereas randomly selected shRNA pools do not.
-
Has reproduction · 59
Application of Machine Learning in Predicting Hepatic Metastasis or Primary Site in Gastroenteropancreatic Neuroendocrine Tumors.
PMID 37887568 · PMC10605255 · Current oncology (Toronto, Ont.) · 2023 · 8 claims · 7 setups
Multi-gene random forest models classify primary tumor vs. liver metastasis samples with 100% accuracy in training/test cohorts and >90% accuracy in an independent validation cohort
-
Full-text index only
Expansion of the Bactericidal/Permeability Increasing-like (BPI-like) protein locus in cattle.
PMID 17362520 · PMC1839098 · BMC genomics · 2007 · 8 claims · 8 setups
The bovine BPI-like locus spans 470 kbp and contains 14 contiguous genes (13 intact + 1 pseudogene); 9 are orthologous to human/mouse BPI-like genes and 4 (named BSP30A, BSP30B, BSP30C, BSP30D) arose through cattle-specific duplication of the PSP gene
-
Full-text index only
Cancer-specific high-throughput annotation of somatic mutations: computational prediction of driver missense mutations.
PMID 19654296 · PMC2763410 · Cancer research · 2009 · 7 claims · 7 setups
CHASM, a Random Forest-based computational method, was developed to identify and prioritize missense mutations likely to be functional drivers of tumor cell proliferation.
-
Full-text index only
Application of two machine learning algorithms to genetic association studies in the presence of covariates.
PMID 19014573 · PMC2620353 · BMC genetics · 2008 · 8 claims · 3 setups
The relative performance of RF and MARS for detecting genotype-trait associations depends on both the strategy used to handle covariates and the true underlying model of association (e.g., confounding vs. mediation vs. interaction).
-
Has reproduction · 72
Prediction of prognostic signatures in triple-negative breast cancer based on the differential expression analysis via NanoString nCounter immune panel.
PMID 33138797 · PMC7607642 · BMC cancer · 2020 · 8 claims · 7 setups
edgeR identifies 9 DEGs associated with pCR and 13 DEGs associated with relapse from 579 immune genes in a small TNBC sample set (n=55)
-
Full-text index only
CanPredict: a computational tool for predicting cancer-associated missense mutations.
PMID 17537827 · PMC1933186 · Nucleic acids research · 2007 · 8 claims · 7 setups
CanPredict is a web application providing public access to a random forest classifier that combines SIFT, LogR.E-value, and GOSS scores to predict whether a missense mutation is cancer-associated