Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
Comparative phosphoproteomics reveals evolutionary and functional conservation of phosphorylation across eukaryotes.
PMID 18828897 · PMC2760871 · Genome biology · 2008 · 8 claims · 8 setups
The overlap between phosphoproteomes of six eukaryotes (human, mouse, fly, yeast, plant, zebrafish) is significantly greater than expected by chance.
-
Full-text index only
The association of Alu repeats with the generation of potential AU-rich elements (ARE) at 3' untranslated regions.
PMID 15610565 · PMC544599 · BMC genomics · 2004 · 6 claims · 4 setups
Alu repeats are a source of AREs at 3' UTRs of human mRNA, via poly-A regions of Alu generating complementary poly-T/poly-U regions that acquire regular adenine insertions to form ARE motifs.
-
Full-text index only
miRGen: a database for the study of animal microRNA genomic organization and function.
PMID 17108354 · PMC1669779 · Nucleic acids research · 2007 · 8 claims · 6 setups
miRGen is an integrated database combining Genomics, Targets, and Clusters interfaces to study miRNA genomic organization and function across 11 animal genomes
-
Full-text index only
Identifying cis-regulatory sequences by word profile similarity.
PMID 19730735 · PMC2731932 · PloS one · 2009 · 8 claims · 8 setups
WPH-finder identifies putative co-regulated CRMs by scanning the genome for sequences with word profiles similar to a known CRM, without explicitly defining binding sites
-
Has reproduction · 80
SLDMS: A Tool for Calculating the Overlapping Regions of Sequences.
PMID 35046988 · PMC8761809 · Frontiers in plant science · 2021 · 8 claims · 5 setups
SLDMS is a novel method for computing overlapping regions of sequencing reads using suffix array (SA), longest common prefix (LCP) array, document array (DA), and a monotonic stack.
-
Has reproduction · 98
Mutations in dnaA and a cryptic interaction site increase drug resistance in Mycobacterium tuberculosis.
PMID 33253310 · PMC7738170 · PLoS pathogens · 2020 · 7 claims · 8 setups
Non-synonymous mutations in dnaA are statistically associated with drug resistance (INH, RIF, SM) in clinical M. tuberculosis strains across two independent GWAS cohorts (China and Vietnam)
-
Full-text index only
Assessing the genomic evidence for conserved transcribed pseudogenes under selection.
PMID 19754956 · PMC2753554 · BMC genomics · 2009 · 8 claims · 8 setups
1750 transcribed pseudogene annotations (TPAs) were identified in the human genome, ~11.5% of all human pseudogene annotations.
-
Full-text index only
Genomic views of distant-acting enhancers.
PMID 19741700 · PMC2923221 · Nature · 2009 · 8 claims · 8 setups
Meta-analysis of ~1200 top GWAS SNPs found that in 40% of cases (472/1170) no known exons overlap the linked SNP or its haplotype block, implying noncoding variation causally contributes to many traits.
-
Full-text index only
Analyses of deep mammalian sequence alignments and constraint predictions for 1% of the human genome.
PMID 17567995 · PMC1891336 · Genome research · 2007 · 7 claims · 3 setups
Four different alignment methods show large-scale consistency but substantial differences in small-scale rearrangements, sensitivity, and specificity.
-
Full-text index only
POCUS: mining genomic sequence annotation to predict disease genes.
PMID 14611661 · PMC329128 · Genome biology · 2003 · 8 claims · 6 setups
Genes predisposing to the same disease tend to share functional annotation IDs (GO/InterPro) more than expected by chance
-
Full-text index only
Defining the proteome.
PMID 16356278 · PMC1414090 · Genome biology · 2005 · 8 claims · 8 setups
Integration of multi-omics data is an important step toward a systems-biology approach to the proteome
-
Full-text index only
L2L: a simple tool for discovering the hidden significance in microarray expression data.
PMID 16168088 · PMC1242216 · Genome biology · 2005 · 8 claims · 4 setups
L2L systematically compares a user's differentially expressed gene list against a database of published differentially expressed gene lists to find statistically significant overlaps and generate hypotheses about shared mechanisms
-
Full-text index only
Human and mouse introns are linked to the same processes and functions through each genome's most frequent non-conserved motifs.
PMID 18450818 · PMC2425492 · Nucleic acids research · 2008 · 8 claims · 5 setups
Pyknons (recurrent, genome-specific, ≥16nt motifs with ≥30 intact intergenic/intronic copies and ≥1 exonic copy) span a substantial fraction of previously uncharacterized intronic space (7.4% human, 4.4% mouse)
-
Full-text index only
A comprehensive modular map of molecular interactions in RB/E2F pathway.
PMID 18319725 · PMC2290939 · Molecular systems biology · 2008 · 8 claims · 4 setups
A comprehensive, curated map of RB/E2F pathway molecular interactions was built using SBGN notation in CellDesigner and converted to BioPAX 2.0 format
-
Has reproduction · 83
Public Omics Explorer (POE): Enabling integrative semantic search across GEO omics datasets based on PubMed publications.
PMID 41282419 · PMC12636342 · Computational and structural biotechnology journal · 2025 · 6 claims · 4 setups
POE is a web platform that semantically links GEO datasets and ENA records through their associated PubMed publications for literature-informed dataset retrieval
-
Has reproduction
Methylation patterns of the nasal epigenome of hospitalized SARS-CoV-2 positive patients reveal insights into molecular mechanisms of COVID-19.
PMID 40170038 · PMC11963311 · BMC medical genomics · 2025 · 8 claims · 5 setups
The nasal methylome shows differential DNA methylation in intergenic regions and low methylated regions (LMRs), highlighting distal regulatory/enhancer-like elements in COVID-19 gene regulation
-
Full-text index only
Mining expressed sequence tags identifies cancer markers of clinical interest.
PMID 17078886 · PMC1635568 · BMC bioinformatics · 2006 · 8 claims · 6 setups
An EST-mining approach (Fisher Exact Test on tumor vs. non-tumor library hit counts) identifies differentially expressed transcripts with an estimated false discovery rate below 22% when human and mouse screens are combined.
-
Has reproduction · 100
Computational modeling demonstrates that glioblastoma cells can survive spatial environmental challenges through exploratory adaptation.
PMID 31836713 · PMC6911112 · Nature communications · 2019 · 8 claims · 6 setups
Stochastic exploration of the gene-regulatory network structure confers enhanced adaptive capacity, enabling GBM cells to converge to new target phenotypes in novel environments.
-
Has reproduction
All of gene expression (AOE): An integrated index for public gene expression databases.
PMID 31978081 · PMC6980531 · PloS one · 2020 · 8 claims · 5 setups
AOE integrates publicly available gene expression data from GEO, ArrayExpress, and GEA into a single searchable index.
-
Has reproduction · 94
A Deluge of Complex Repeats: The Solanum Genome.
PMID 26241045 · PMC4524691 · PloS one · 2015 · 8 claims · 7 setups
~50–60% of the S. tuberosum and S. lycopersicum genomes are composed of repetitive elements