Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
Polymorphix: a sequence polymorphism database.
PMID 15608242 · PMC540030 · Nucleic acids research · 2005 · 8 claims · 5 setups
Polymorphix is an ACNUC-structured database that organizes EMBL/GenBank sequences into within-species homologous sequence families using similarity and bibliographic criteria, with alignments, outgroups and phylogenetic trees provided.
-
Has reproduction · 87
CoINcIDE: A framework for discovery of patient subtypes across multiple datasets.
PMID 26961683 · PMC4784276 · Genome medicine · 2016 · 8 claims · 6 setups
CoINcIDE is a methodological framework that discovers replicable patient subtypes (meta-clusters) across multiple datasets by finding consensus across dataset-specific clusterings, requiring no between-dataset transformations.
-
Has reproduction
Artificial Intelligence Meets Whole Slide Images: Deep Learning Model Shapes an Immune-Hot Tumor and Guides Precision Therapy in Bladder Cancer.
PMID 36245985 · PMC9553530 · Journal of oncology · 2022 · 6 claims · 8 setups
A deep learning WSI cluster (three-class mini batch K-means on Inception V3 features) is associated with overall survival (P<0.001) and is an independent prognostic predictor (P=0.031) in BLCA.
-
Has reproduction · 50
BiRNA-BERT allows efficient RNA language modeling with adaptive tokenization.
PMID 41266599 · PMC12635123 · Communications biology · 2025 · 8 claims · 8 setups
BiRNA-BERT uses adaptive dual-tokenization that dynamically selects nucleotide-level (NUC) or byte-pair encoding (BPE) tokens based on input sequence length
-
Full-text index only
The TIGR Gene Indices: clustering and assembling EST and known genes and integration with eukaryotic genomes.
PMID 15608288 · PMC540018 · Nucleic acids research · 2005 · 8 claims · 8 setups
The TIGR Gene Indices (TGI) are a collection of 77 species-specific databases that cluster and assemble EST and known gene sequences into tentative consensus (TC) sequences to identify and characterize expressed transcripts.
-
Full-text index only
ECgene: genome annotation for alternative splicing.
PMID 15608289 · PMC540072 · Nucleic acids research · 2005 · 8 claims · 5 setups
ECgene combines genome-based EST clustering with a graph-theoretic transcript assembly procedure to predict gene models including alternative splicing events.
-
Full-text index only
EPD in its twentieth year: towards complete promoter coverage of selected model organisms.
PMID 16381980 · PMC1347508 · Nucleic acids research · 2006 · 7 claims · 4 setups
EPD is an annotated, non-redundant collection of experimentally defined eukaryotic POL II promoters accessed via genome position pointers.
-
Full-text index only
DAVID Bioinformatics Resources: expanded annotation database and novel algorithms to better extract biology from large gene lists.
PMID 17576678 · PMC1933169 · Nucleic acids research · 2007 · 8 claims · 4 setups
The DAVID Gene Concept uses a single-linkage method to agglomerate tens of millions of gene/protein identifiers from NCBI, PIR, UniProt and other resources into unified DAVID genes.
-
Has reproduction · 90
Transcriptomic data meta-analysis reveals common and injury model specific gene expression changes in the regenerating zebrafish heart.
PMID 37012284 · PMC10070245 · Scientific reports · 2023 · 7 claims · 8 setups
Batch correction using sequencing platform as the correcting variable (via Combat-Seq) removes technical variability so that samples cluster by injury condition rather than dataset origin.
-
Full-text index only
Systems biology approach for mapping the response of human urothelial cells to infection by Enterococcus faecalis.
PMID 18047719 · PMC2099488 · BMC bioinformatics · 2007 · 8 claims · 5 setups
Deconvoluting gene expression variance into technical (Gaussian, ~6.5% relative SD) and biological components identifies hypervariable (HV) genes that reflect true biological response to infection without requiring replicates