Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Has reproduction · 68
Rfam 15: RNA families database in 2025.
PMID 39526405 · PMC11701678 · Nucleic acids research · 2025 · 8 claims · 6 setups
Rfamseq was expanded to 26 106 genomes, a 76% increase, by incorporating the latest UniProt reference proteomes and additional viral genomes
-
Has reproduction · 83
Gene Expression Atlas update--a value-added database of microarray and sequencing-based functional genomics experiments.
PMID 22064864 · PMC3245177 · Nucleic acids research · 2012 · 8 claims · 5 setups
Gene Expression Atlas is an added-value database providing curated, re-annotated and statistically analysed gene expression data across cell types, organism parts, developmental stages, disease states and other biological/experimental conditions, derived from ArrayExpress Archive and the European Nucleotide Archive.
-
Has reproduction · 88
Human methylome variation across Infinium 450K data on the Gene Expression Omnibus.
PMID 33937763 · PMC8061458 · NAR genomics and bioinformatics · 2021 · 8 claims · 6 setups
Approximately two-thirds of compiled HM450K samples are from blood, one-quarter from brain, and roughly one-third from cancer patients.
-
Full-text index only
Assessing the gene space in draft genomes.
PMID 19042974 · PMC2615622 · Nucleic acids research · 2009 · 6 claims · 7 setups
The proportion of mapped CEGs in a draft genome assembly is a useful metric for describing gene space completeness, complementing N50 and x-fold coverage.
-
Full-text index only
Ensembl 2005.
PMID 15608235 · PMC540092 · Nucleic acids research · 2005 · 8 claims · 4 setups
Ensembl's automatic gene build system can flexibly and reliably annotate a wide variety of genomes with limited species-specific evidence.
-
Full-text index only
AceView: a comprehensive cDNA-supported gene and transcripts annotation.
PMID 16925834 · PMC1810549 · Genome biology · 2006 · 8 claims · 4 setups
At the mRNA level, AceView transcripts are the closest match to Gencode transcripts among all evaluated methods, including alternative splice variants
-
Has reproduction · 75
An informatics research platform to make public gene expression time-course datasets reusable for more scientific discoveries.
PMID 33247935 · PMC7698665 · Database : the journal of biological databases and curation · 2020 · 8 claims · 6 setups
GETc enables discovery and visualization of time-course gene expression data and analytical results from GEO
-
Full-text index only
Consolidating the set of known human protein-protein interactions in preparation for large-scale mapping of the human interactome.
PMID 15892868 · PMC1175952 · Genome biology · 2005 · 8 claims · 6 setups
Two quantitative benchmarks (functional-annotation-based and physical-interaction-based log likelihood ratio scores) can measure relative accuracy of human PPI datasets
-
Full-text index only
CanPredict: a computational tool for predicting cancer-associated missense mutations.
PMID 17537827 · PMC1933186 · Nucleic acids research · 2007 · 8 claims · 7 setups
CanPredict is a web application providing public access to a random forest classifier that combines SIFT, LogR.E-value, and GOSS scores to predict whether a missense mutation is cancer-associated
-
Has reproduction · 91
Chromosome-level genome assembly of agar-producing red seaweed Gracilaria vermiculophylla.
PMID 41629338 · PMC12966425 · Scientific data · 2026 · 8 claims · 8 setups
A chromosome-level genome assembly of G. vermiculophylla was generated by combining DNBSeq short reads, Nanopore long reads, and Hi-C data.
-
Full-text index only
Steps toward broad-spectrum therapeutics: discovering virulence-associated genes present in diverse human pathogens.
PMID 19874620 · PMC2774872 · BMC genomics · 2009 · 8 claims · 8 setups
Phylogenetic profiling of protein clusters across pathogen and non-pathogen genomes can identify candidate generic virulence factors