Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
A comprehensive modular map of molecular interactions in RB/E2F pathway.
PMID 18319725 · PMC2290939 · Molecular systems biology · 2008 · 8 claims · 4 setups
A comprehensive, curated map of RB/E2F pathway molecular interactions was built using SBGN notation in CellDesigner and converted to BioPAX 2.0 format
-
Has reproduction · 98
maxATAC: Genome-scale transcription-factor binding prediction from ATAC-seq with deep neural networks.
PMID 36719906 · PMC9917285 · PLoS computational biology · 2023 · 8 claims · 6 setups
maxATAC is a suite of deep neural network models enabling state-of-the-art, genome-scale TFBS prediction from ATAC-seq, with models for 127 human transcription factors
-
Full-text index only
Sequence similarity network reveals common ancestry of multidomain proteins.
PMID 18475320 · PMC2377100 · PLoS computational biology · 2008 · 8 claims · 6 setups
Traditional homology definitions do not capture multidomain evolution; the authors extend the definition to include domain insertion via a common ancestral locus model.
-
Full-text index only
A parsimony approach to biological pathway reconstruction/inference for genomes and metagenomes.
PMID 19680427 · PMC2714467 · PLoS computational biology · 2009 · 8 claims · 6 setups
The naïve mapping approach (present if ≥1 associated function is found) leads to an inflated estimate of biological pathways and overestimates functional diversity of a sample.
-
Full-text index only
Integrating alternative splicing detection into gene prediction.
PMID 15705189 · PMC550657 · BMC bioinformatics · 2005 · 8 claims · 4 setups
An integrative intrinsic/extrinsic method was implemented in the gene finder EuGÈNE (as EuGÈNE-M) to detect AS evidence from aligned transcripts and generate alternative optimal gene predictions consistent with each detected AS event.
-
Full-text index only
The MAPPER database: a multi-genome catalog of putative transcription factor binding sites.
PMID 15608292 · PMC540057 · Nucleic acids research · 2005 · 8 claims · 6 setups
Built a library of 1134 HMM models (359 matrix-derived, 718 factor-derived, 57 JASPAR-derived), corresponding to 863 distinct TF names, from TRANSFAC and JASPAR binding site data
-
Full-text index only
Evola: Ortholog database of all human genes in H-InvDB with manual curation of phylogenetic trees.
PMID 17982176 · PMC2238928 · Nucleic acids research · 2008 · 6 claims · 7 setups
Evola combines genome synteny-based computational ortholog detection with manual curation of phylogenetic trees by experts to yield more reliable orthologs than automated pairwise methods
-
Full-text index only
Integrating proteomic, transcriptional, and interactome data reveals hidden components of signaling and regulatory networks.
PMID 19638617 · PMC2889494 · Science signaling · 2009 · 8 claims · 6 setups
Pathway reconstruction can be modeled as a prize-collecting Steiner tree problem, balancing penalties for excluding terminal nodes against costs for including edges, controlled by a parameter β.
-
Full-text index only
Atlas of nascent RNA transcripts reveals tissue-specific enhancer to gene linkages.
PMID 40281430 · PMC12032694 · BMC genomics · 2025 · 7 claims · 8 setups
A large repository of nascent run-on RNA-seq samples (DBNascent) was assembled and uniformly processed to identify sites of bidirectional transcription genome-wide.
-
Full-text index only
Much ado about nothing: modeling amino acid replacement with predicted protein structures.
PMID 42036821 · PMC13171170 · Bioinformatics (Oxford, England) · 2026 · 8 claims · 7 setups
AFSM was constructed from over 660,000 structural alignments across ~21,000 proteins (297 InterPro families), following the BLOSUM log-odds methodology.
-
Has reproduction · 79
Interpretable prediction models for widespread m6A RNA modification across cell lines and tissues.
PMID 37995291 · PMC10697738 · Bioinformatics (Oxford, England) · 2023 · 8 claims · 8 setups
CLSM6A is a set of CNN-based deep learning models that predict single-nucleotide-resolution m6A RNA modification sites across eight cell lines and three tissues in H. sapiens
-
Has reproduction
Comprehensive enhancer-target gene assignments improve gene set level interpretation of genome-wide regulatory data.
PMID 35473573 · PMC9044877 · Genome biology · 2022 · 8 claims · 8 setups
Combining multiple enhancer-definition and enhancer-gene link data sources yields 1860 genome-wide EnTDefs covering >500 cell types
-
Has reproduction · 77
Representing and querying disease networks using graph databases.
PMID 27462371 · PMC4960687 · BioData mining · 2016 · 7 claims · 8 setups
Graph databases are well suited for representing biological information because it is typically highly connected, semi-structured and unpredictable, unlike relational databases which require rigid schemas.
-
Has reproduction · 100
Intratumoral heterogeneity in microsatellite instability status at single-cell resolution.
PMID 41767255 · PMC12936829 · iScience · 2026 · 8 claims · 7 setups
A novel computational (Snakemake) pipeline quantifies intratumoral heterogeneity in MSI status at single-cell resolution
-
Has reproduction · 70
Spatial transcriptomics reveals the molecular signatures of prodromal and advanced α-synucleinopathy.
PMID 41736854 · PMC12927100 · iScience · 2026 · 7 claims · 6 setups
Early-stage (prodromal) aSyn pathology in M83+/+ mouse brainstem is associated with upregulation of ATP/energy metabolism pathways (glycolysis, oxidative phosphorylation, fatty acid metabolism)
-
Has reproduction · 45
Identifying and classifying trait linked polymorphisms in non-reference species by walking coloured de bruijn graphs.
PMID 23536903 · PMC3607606 · PloS one · 2013 · 8 claims · 9 setups
Bubbleparse detects sequence variants directly from NGS reads without a reference genome, using the coloured de Bruijn graph implementation of Cortex plus a new depth-first bubble-finding module.
-
Full-text index only
Systematic identification of pseudogenes through whole genome expression evidence profiling.
PMID 16945953 · PMC1636364 · Nucleic acids research · 2006 · 8 claims · 8 setups
Developed a novel bioinformatics method that identifies pseudogenes by profiling whole-genome transcript and protein expression evidence
-
Full-text index only
EGASP: Introduction.
PMID 16925831 · PMC1810546 · Genome biology · 2006 · 8 claims · 5 setups
Computational gene finding methods, when compared to the GENCODE golden standard annotation, show that the human genome annotation is nearly complete in terms of novel protein-coding loci.
-
Full-text index only
VIRGO: computational prediction of gene functions.
PMID 16845022 · PMC1538839 · Nucleic acids research · 2006 · 8 claims · 6 setups
VIRGO constructs a functional linkage network (FLN) from gene expression and molecular interaction data, labels genes with GO annotations, and propagates these labels to predict functions of unlabelled genes
-
Full-text index only
The DAVID Gene Functional Classification Tool: a novel biological module-centric algorithm to functionally analyze large gene lists.
PMID 17784955 · PMC2375021 · Genome biology · 2007 · 8 claims · 6 setups
Gene-gene functional similarity can be measured using kappa statistics applied to a binary gene-annotation-term matrix built from 14 annotation categories.