Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
Genome-wide survey of allele-specific splicing in humans.
PMID 18518984 · PMC2427040 · BMC genomics · 2008 · 8 claims · 5 setups
A genome-wide computational scan identified 30,977 SNPs located within predicted splicing regulatory sequences (donor sites, acceptor sites, branch points, and ESEs)
-
Full-text index only
Optimized mixed Markov models for motif identification.
PMID 16749929 · PMC1534070 · BMC bioinformatics · 2006 · 8 claims · 4 setups
OMiMa can incorporate more than NNSplice's pairwise dependencies
-
Full-text index only
Discovery and identification of potential biomarkers of papillary thyroid carcinoma.
PMID 19785722 · PMC2761863 · Molecular cancer · 2009 · 8 claims · 7 setups
A 3-peak (m/z 9190, 6631, 8697 Da) SVM classification model discriminates PTC from non-cancer controls with high sensitivity and specificity
-
Full-text index only
BABELOMICS: a systems biology perspective in the functional annotation of genome-scale experiments.
PMID 16845052 · PMC1538844 · Nucleic acids research · 2006 · 8 claims · 8 setups
Babelomics is presented as an updated, complete suite of web tools for functional analysis of genome-scale experiments with new and improved modules
-
Has reproduction · 100
Gene signature discovery and systematic validation across diverse clinical cohorts for TB prognosis and response to treatment.
PMID 37471455 · PMC10393163 · PLoS computational biology · 2023 · 8 claims · 8 setups
A network-based meta-analysis across studies identifies a common 45-gene signature specific to active TB disease that accounts for cohort/population heterogeneity
-
Has reproduction · 42
CanCellCap: robust cancer cell capture across tissue types on single-cell RNA-seq data by multi-domain learning.
PMID 40739511 · PMC12312500 · BMC biology · 2025 · 8 claims · 8 setups
CanCellCap, a multi-domain learning framework integrating domain adversarial learning and Mixture of Experts, identifies cancer cells across all tissues, cancers, and sequencing platforms by extracting tissue-common and tissue-specific gene expression patterns.
-
Full-text index only
Association testing of novel type 2 diabetes risk alleles in the JAZF1, CDC123/CAMK1D, TSPAN8, THADA, ADAMTS9, and NOTCH2 loci with insulin release, insulin sensitivity, and obesity in a population-based sample of 4,516 glucose-tolerant middle-aged Danes.
PMID 18567820 · PMC2518507 · Diabetes · 2008 · 8 claims · 5 setups
CDC123/CAMK1D rs12779790 risk allele (homozygous) is associated with decreased insulinogenic index, corrected insulin response (CIR), and AUC-insulin/AUC-glucose ratio, indicating impaired insulin release
-
Full-text index only
Predicting failure rate of PCR in large genomes.
PMID 18492719 · PMC2441781 · Nucleic acids research · 2008 · 7 claims · 8 setups
The number of predicted primer-binding sites in genomic DNA is the most important factor determining PCR failure.
-
Has reproduction · 87
Forseti: a mechanistic and predictive model of the splicing status of scRNA-seq reads.
PMID 38940130 · PMC11256924 · Bioinformatics (Oxford, England) · 2024 · 7 claims · 5 setups
Forseti is the first probabilistic model for resolving the splicing status of exonic scRNA-seq reads by scoring putative fragments linking read alignments to proximate priming sites
-
Has reproduction · 67
Evaluating native-like structures of RNA-protein complexes through the deep learning method.
PMID 36828844 · PMC9958188 · Nature communications · 2023 · 8 claims · 7 setups
DRPScore identifies native-like RNA-protein structures with higher success rates than ITScore-PR, DARS-RNP, and 3dRPC across bound and unbound testing sets.
-
Full-text index only
Diagnostic proteomics: serum proteomic patterns for the detection of early stage cancers.
PMID 15258335 · PMC3851082 · Disease markers · 2003 · 8 claims · 8 setups
Proteomic pattern analysis of serum mass spectra, without identifying the underlying proteins, can distinguish cancer patients from healthy controls with high sensitivity and specificity.
-
Full-text index only
Numbers of mutations to different types of colorectal cancer.
PMID 16202134 · PMC1266026 · BMC cancer · 2005 · 8 claims · 5 setups
Different biologic subtypes of colorectal cancer require different numbers of oncogenic mutations before transformation
-
Full-text index only
Evaluation of models to predict BRCA germline mutations.
PMID 17016486 · PMC2360540 · British journal of cancer · 2006 · 7 claims · 7 setups
Four commonly used BRCA risk prediction models (BRCAPRO, Manchester, Penn, Myriad-Frank) have only modest ability to rule in or rule out BRCA1/2 germline mutation carrier status at a 10% probability threshold.
-
Has reproduction · 79
Interpretable prediction models for widespread m6A RNA modification across cell lines and tissues.
PMID 37995291 · PMC10697738 · Bioinformatics (Oxford, England) · 2023 · 7 claims · 6 setups
CLSM6A, a CNN-based model set, predicts single-nucleotide-resolution m6A RNA modification sites across eight cell lines and three tissues in H. sapiens
-
Has reproduction · 78
GenTB: A user-friendly genome-based predictor for tuberculosis resistance powered by machine learning.
PMID 34461978 · PMC8407037 · Genome medicine · 2021 · 8 claims · 6 setups
GenTB is a free, open, web-based application offering two ML predictors (Random Forest and WDNN) that predict resistance to 13 and 10 anti-TB drugs, respectively.
-
Full-text index only
Linking disease-associated genes to regulatory networks via promoter organization.
PMID 15701758 · PMC549397 · Nucleic acids research · 2005 · 8 claims · 7 setups
Pairs of TFBSs conserved both vertically (orthologous genes) and horizontally (co-regulated genes) can serve as seeds to build promoter models representing potential co-regulation networks
-
Full-text index only
The human phylome.
PMID 17567924 · PMC2394744 · Genome biology · 2007 · 6 claims · 5 setups
Reconstruction of the human phylome: evolutionary trees for all human proteins and their homologs among 39 fully sequenced eukaryotic genomes, using a pipeline combining alignment trimming, NJ, ML (PhyML) and Bayesian (MrBayes) methods.
-
Full-text index only
A space-efficient and accurate method for mapping and aligning cDNA sequences onto genomic sequence.
PMID 18344523 · PMC2377433 · Nucleic acids research · 2008 · 7 claims · 6 setups
Spaln maps and aligns large cDNA sequence sets onto whole mammalian genomes using substantially less memory than comparable existing tools
-
Full-text index only
The role of medical structural genomics in discovering new drugs for infectious diseases.
PMID 19855826 · PMC2756625 · PLoS computational biology · 2009 · 8 claims · 6 setups
Structure-based drug design using X-ray/NMR protein structures has contributed to the development of numerous approved and improved therapeutics.
-
Full-text index only
The FAS gene, brain volume, and disease progression in Alzheimer's disease.
PMID 19766542 · PMC3100774 · Alzheimer's & dementia : the journal of the Alzheimer's Association · 2010 · 8 claims · 4 setups
The minor (T) allele of rs1468063 in FAS is significantly associated with faster AD progression after permutation-based multiple-testing correction