Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Has reproduction
Using random walks to identify cancer-associated modules in expression data.
PMID 24128261 · PMC4015830 · BioData mining · 2013 · 8 claims · 8 setups
Walktrap-GM, a random-walk community detection algorithm adapted with stopping criteria (maximum modularity, maximum size, maximum module score), identifies modules significantly enriched with cancer genes in expression-weighted interaction networks.
-
Full-text index only
Solving structures of protein complexes by molecular replacement with Phaser.
PMID 17164524 · PMC2483468 · Acta crystallographica. Section D, Biological crystallography · 2007 · 7 claims · 4 setups
Maximum-likelihood MR functions enable complex asymmetric units to be built up from individual components using a 'tree search with pruning' approach implemented in Phaser's automated MR mode.
-
Full-text index only
Function2Gene: a gene selection tool to increase the power of genetic association studies by utilizing public databases and expert knowledge.
PMID 18631403 · PMC2500032 · BMC bioinformatics · 2008 · 6 claims · 5 setups
Function2Gene is a set of Perl programs that queries public databases (NCBI, GeneCards, Harvester, with Uniprot/Ensembl also supported) using expert-selected keywords to rank genes by prior probability of disease association.
-
Full-text index only
sCellST predicts single-cell gene expression from H& E images.
PMID 41513659 · PMC12858858 · Nature communications · 2026 · 7 claims · 6 setups
sCellST is a weakly supervised (Multiple Instance Learning) deep learning framework that predicts single-cell gene expression from H&E images alone, trained using paired spatial transcriptomics (Visium) and H&E slides
-
Has reproduction · 69
Automatic discovery of 100-miRNA signature for cancer classification using ensemble feature selection.
PMID 31533612 · PMC6751684 · BMC bioinformatics · 2019 · 8 claims · 6 setups
An ensemble feature selection strategy using consensus of feature relevance across 8 classifier types identifies a 100-miRNA signature from a 1046-feature TCGA dataset
-
Has reproduction · 80
Colorectal Cancer Prediction Based on Weighted Gene Co-Expression Network Analysis and Variational Auto-Encoder.
PMID 32825264 · PMC7563725 · Biomolecules · 2020 · 6 claims · 7 setups
Combining WGCNA hub genes and VAE 10-dimensional representation as features for an SVM classifier achieves high accuracy in predicting CRC
-
Has reproduction · 92
Analytical code sharing practices in biomedical research.
PMID 38983240 · PMC11232620 · PeerJ. Computer science · 2024 · 8 claims · 4 setups
Nearly half (49.9%) of 453 examined biomedical manuscripts failed to share the analytical code used to generate their results
-
Has reproduction · 32
Developing prognostic gene panel of survival time in lung adenocarcinoma patients using machine learning.
PMID 35117753 · PMC8799101 · Translational cancer research · 2020 · 8 claims · 5 setups
Naïve Bayes using a 22-gene panel is the best-performing and most stable machine learning model for predicting LUAD survival time (>3 vs <3 years)
-
Has reproduction · 75
Revealing the critical state and identifying individualized dynamic network biomarker for type 2 diabetes through advanced analysis methods on individual basis.
PMID 39890881 · PMC11785715 · Scientific reports · 2025 · 8 claims · 5 setups
sJSD, NIG, and TNFE methods can detect critical states/tipping points before disease deterioration using only a single sample
-
Full-text index only
Mutational analyses of multiple target genes in histologically heterogeneous gastric cancer with microsatellite instability.
PMID 10081489 · PMC5921733 · Japanese journal of cancer research : Gann · 1998 · 7 claims · 5 setups
MSI frequency in gastric cancers with histological heterogeneity was 35% (7/20 cases) and 28% (11/40 tumor DNAs), consistent with prior gastric cancer MSI studies.
-
Full-text index only
MSI test to distinguish between HNPCC and other predisposing syndromes -- of value in tailored surveillance.
PMID 15528791 · PMC3839337 · Disease markers · 2004 · 8 claims · 5 setups
MSI (or immunohistochemistry) testing applied to a pre-selected patient with a family history or early-onset colorectal cancer can distinguish HNPCC from unknown non-HNPCC colorectal cancer syndromes.
-
Full-text index only
Novel and de novo PKD1 mutations identified by multiple restriction fragment-single strand conformation polymorphism (MRF-SSCP).
PMID 15018634 · PMC356914 · BMC medical genetics · 2004 · 6 claims · 7 setups
MRF-SSCP method (using combined restriction digestion plus SSCP with silver staining) was developed to screen PKD1 mutations in full-length cDNA fractionated into nine overlapping nested-PCR segments
-
Full-text index only
Ontological visualization of protein-protein interactions.
PMID 15707487 · PMC550656 · BMC bioinformatics · 2005 · 8 claims · 8 setups
Aggregating independently made GO 'protein binding' (IPI) annotations reveals larger, previously undescribed mouse protein-protein interaction networks
-
Full-text index only
Exogean: a framework for annotating protein-coding genes in eukaryotic genomic DNA.
PMID 16925841 · PMC1810556 · Genome biology · 2006 · 8 claims · 5 setups
Exogean is a framework using directed acyclic coloured multigraphs (DACMs) to represent biological objects (mRNA, ESTs, protein alignments, exons) and iteratively combine them into complex protein-coding transcript models.
-
Full-text index only
AUGUSTUS at EGASP: using EST, protein and genomic alignments for improved gene prediction in the human genome.
PMID 16925833 · PMC1810548 · Genome biology · 2006 · 8 claims · 5 setups
AUGUSTUS predicted significantly more genes correctly than any other ab initio program in EGASP
-
Full-text index only
Computational disease gene identification: a concert of methods prioritizes type 2 diabetes and obesity candidate genes.
PMID 16757574 · PMC1475747 · Nucleic acids research · 2006 · 6 claims · 8 setups
Applying seven independent computational disease-gene prioritization methods in concert to 9556 positional candidate genes identifies a prioritized set of likely T2D and obesity candidate genes
-
Full-text index only
Incorporation of genetic model parameters for cost-effective designs of genetic association studies using DNA pooling.
PMID 17634103 · PMC1947971 · BMC genomics · 2007 · 8 claims · 4 setups
A closed-form approximation to the F-test non-centrality parameter (NCP) incorporating genetic model parameters (disease allele frequency, marker allele frequency, prevalence, genotype relative risk, sample size, genetic model, number of pools/replicates, machine variability) can be used to compute power for DNA pooling association studies
-
Full-text index only
Adverse prognosis of epigenetic inactivation in RUNX3 gene at 1p36 in human pancreatic cancer.
PMID 18475302 · PMC2391125 · British journal of cancer · 2008 · 7 claims · 5 setups
RUNX3 promoter hypermethylation is frequent in primary pancreatic cancer tissue
-
Full-text index only
Comparing whole genomes using DNA microarrays.
PMID 18347592 · PMC7097741 · Nature reviews. Genetics · 2008 · 8 claims · 6 setups
DNA microarrays offer a relatively inexpensive and efficient alternative to genome sequencing for comparing all known classes of genomic diversity between closely related genomes.
-
Full-text index only
A machine learning approach uncovers principles and determinants of eukaryotic ribosome pausing.
PMID 39423268 · PMC11488575 · Science advances · 2024 · 8 claims · 5 setups
An unsupervised ML pipeline using the extended isolation forest (EIF) algorithm can reliably detect ribosome pausing sites from noisy, coverage-biased RiboSeq data across expression levels