Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
ChromBERT: A foundation model for learning interpretable representations for context-specific transcriptional regulatory networks.
PMID 41592570 · PMC13069865 · Cell genomics · 2026 · 8 claims · 7 setups
ChromBERT is pre-trained via masked reconstruction on the Cistrome-Human-6K dataset (6,391 cistromes, 991 transcription regulators) to learn genome-wide interaction syntax of transcription regulators
-
Full-text index only
CLAMP: predicting specific protein-mediated chromatin loops in diverse species with a chromatin accessibility language model.
PMID 41555433 · PMC12903630 · Genome biology · 2026 · 8 claims · 8 setups
CLAMP, a chromatin-accessibility language model, predicts protein-mediated chromatin loops across 10 species, 18 proteins, and 24 cell types with superior performance versus existing methods.
-
Full-text index only
CircleBase V2: an eccDNA annotation platform across cancers and species.
PMID 41273082 · PMC12807720 · Nucleic acids research · 2026 · 8 claims · 7 setups
CircleBase V2 provides a 12-fold increase in human eccDNA data, comprising over 3.8 million entries from >300 cell types/tissues
-
Full-text index only
A generic reference defined by consensus peaks for single-cell ATAC-seq data analysis.
PMID 41663439 · PMC12996591 · Nature communications · 2026 · 7 claims · 7 setups
Aggregating peaks from 624 high-quality bulk ATAC-seq datasets defines ~1.4 million observed consensus peaks (cPeaks) covering ~30% of the genome.
-
Full-text index only
Architectural and evolutionary features of TE-derived TSSs shape tissue-specific promoter activity in the human genome.
PMID 41620470 · PMC12963367 · Nature communications · 2026 · 8 claims · 8 setups
A three-step RAMPAGE-based pipeline can systematically identify TE-derived transcription start sites (TSSs) genome-wide, distinguishing them from autonomous TE transcription and background noise.
-
Full-text index only
Remodeling of XIST regulatory landscape during primate evolution.
PMID 41544163 · PMC12810636 · Science advances · 2026 · 8 claims · 8 setups
XIST regulation has diverged uniquely and rapidly between closely related primates: JPX is a major XIST regulator in human and marmoset ESCs but only a minor regulator in macaque ESCs
-
Has reproduction · 85
Ensembl 2013.
PMID 23203987 · PMC3531136 · Nucleic acids research · 2013 · 8 claims · 8 setups
Ensembl (http://www.ensembl.org) provides genome information for sequenced chordate genomes, currently supporting 70 species with a focus on human, mouse, zebrafish and rat.
-
Has reproduction · 98
maxATAC: Genome-scale transcription-factor binding prediction from ATAC-seq with deep neural networks.
PMID 36719906 · PMC9917285 · PLoS computational biology · 2023 · 8 claims · 6 setups
maxATAC is a suite of deep neural network models enabling state-of-the-art, genome-scale TFBS prediction from ATAC-seq, with models for 127 human transcription factors
-
Has reproduction · 79
Genome-wide prediction of DNase I hypersensitivity using gene expression.
PMID 29051481 · PMC5715040 · Nature communications · 2017 · 8 claims · 5 setups
Gene expression can, to a large extent, predict genome-wide DNase I hypersensitivity (chromatin accessibility)
-
Full-text index only
Identification of novel DNA sequence motifs that modulate transcription in T cells.
PMID 41514212 · PMC12879379 · BMC genomics · 2026 · 8 claims · 8 setups
Identified 2,036 novel DNA motifs enriched in regulatory regions of T-cell-specific genes
-
Has reproduction · 63
A methyl-sensitive element induces bidirectional transcription in TATA-less CpG island-associated promoters.
PMID 30332484 · PMC6192621 · PloS one · 2018 · 8 claims · 8 setups
The CGCG element (consensus TCTCGCGAGA) is a novel promoter motif enriched in TATA-less CpG island-associated promoters of ribosomal protein and housekeeping genes
-
Full-text index only
ENCODE whole-genome data in the UCSC Genome Browser.
PMID 19920125 · PMC2808953 · Nucleic acids research · 2010 · 7 claims · 8 setups
The UCSC ENCODE Data Coordination Center serves as the primary repository for ENCODE experimental results, providing access via Genome Browser, Table Browser, and FTP download.
-
Has reproduction
Comprehensive enhancer-target gene assignments improve gene set level interpretation of genome-wide regulatory data.
PMID 35473573 · PMC9044877 · Genome biology · 2022 · 8 claims · 8 setups
Combining multiple enhancer-definition and enhancer-gene link data sources yields 1860 genome-wide EnTDefs covering >500 cell types
-
Full-text index only
Exonic enhancers are a widespread class of dual-function regulatory elements.
PMID 41927541 · PMC13216554 · Nature communications · 2026 · 8 claims · 8 setups
Many protein-coding exons possess enhancer activity across species, forming a class of candidate Exonic Enhancers (cEEs)
-
Has reproduction · 89
Statistical framework for calling allelic imbalance in high-throughput sequencing data.
PMID 39966391 · PMC11836314 · Nature communications · 2025 · 8 claims · 6 setups
MIXALIME is a versatile computational framework for calling allele-specific variants (ASVs) from diverse high-throughput omics data
-
Full-text index only
Global atlas of enhancer-promoter interactome in cotton genome revealed by profiling RNA-RNA spatial interactions.
PMID 41484656 · PMC12857045 · Genome biology · 2026 · 8 claims · 8 setups
pRIC-seq, an adaptation of RIC-seq for plants, enables genome-wide mapping of RNA-RNA spatial interactions in diploid and tetraploid cotton
-
Full-text index only
EPInformer: scalable and integrative prediction of gene expression from promoter-enhancer sequences with multimodal epigenomic profiles.
PMID 41832145 · PMC13133354 · Nature communications · 2026 · 8 claims · 7 setups
EPInformer outperforms existing gene expression prediction models (Xpresso, CREaTor, Seq-GraphReg, Enformer, Borzoi) in rigorous 12-fold cross-chromosome validation for both RNA-seq and CAGE expression prediction
-
Full-text index only
Long-read assembly reveals vast transcriptional complexity in the placenta associated with metabolic and endocrine function.
PMID 41927596 · PMC13219539 · Nature communications · 2026 · 8 claims · 8 setups
Long-read RNA-seq of 72 term placentas yields a high-confidence reference of 37,661 isoforms across 12,302 genes, including thousands of previously unannotated isoforms and genes
-
Has reproduction · 74
ChIP-seq guidelines and practices of the ENCODE and modENCODE consortia.
PMID 22955991 · PMC3431496 · Genome research · 2012 · 8 claims · 8 setups
ENCODE/modENCODE define a set of working standards and guidelines for ChIP-seq covering antibody validation, experimental replication, sequencing depth, data/metadata reporting, and data quality assessment.
-
Full-text index only
Atlas of nascent RNA transcripts reveals tissue-specific enhancer to gene linkages.
PMID 40281430 · PMC12032694 · BMC genomics · 2025 · 7 claims · 8 setups
A large repository of nascent run-on RNA-seq samples (DBNascent) was assembled and uniformly processed to identify sites of bidirectional transcription genome-wide.