Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
AUGUSTUS at EGASP: using EST, protein and genomic alignments for improved gene prediction in the human genome.
PMID 16925833 · PMC1810548 · Genome biology · 2006 · 8 claims · 5 setups
AUGUSTUS predicted significantly more genes correctly than any other ab initio program in EGASP
-
Full-text index only
GeneAlign: a coding exon prediction tool based on phylogenetical comparisons.
PMID 16845010 · PMC1538901 · Nucleic acids research · 2006 · 8 claims · 5 setups
GeneAlign predicts coding exons by using signal detection (GeneSplicer/WMM) combined with CORAL, a heuristic linear-time alignment tool, to align candidate signal-flanked regions against annotated exons of a homologous organism's genes
-
Has reproduction · 100
Differentially expressed genes reflect disease-induced rather than disease-causing changes in the transcriptome.
PMID 34561431 · PMC8463674 · Nature communications · 2021 · 8 claims · 7 setups
revTWMR, a reverse transcriptome-wide Mendelian Randomization approach integrating GWAS and whole-blood trans-eQTL summary statistics, is proposed to estimate the causal effect of a phenotype on gene expression.
-
Has reproduction · 44
Dynamic Gene Attention Focus (DyGAF): Enhancing Biomarker Identification Through Dual-Model Attention Networks.
PMID 40160891 · PMC11951896 · Bioinformatics and biology insights · 2025 · 6 claims · 5 setups
DyGAF, a dual-model attention neural network (independent Model A + dependent Model B), identifies and ranks genes by significance for COVID-19 biomarker discovery more effectively than differential expression analysis (DEA) and random forest (RF) feature selection
-
Has reproduction · 75
geneshot: gene-level metagenomics identifies genome islands associated with immunotherapy response.
PMID 33952321 · PMC8097837 · Genome biology · 2021 · 8 claims · 4 setups
geneshot is a gene-level metagenomic bioinformatics tool that clusters de novo assembled protein-coding genes into co-abundant gene groups (CAGs) to reduce dimensionality and generate testable hypotheses from WGS microbiome data
-
Full-text index only
EGASP: the human ENCODE Genome Annotation Assessment Project.
PMID 16925836 · PMC1810551 · Genome biology · 2006 · 8 claims · 6 setups
Best-performing computational gene prediction methods correctly predict at least one transcript for close to 70% of annotated genes in the ENCODE regions.
-
Full-text index only
A modular analysis framework for blood genomics studies: application to systemic lupus erythematosus.
PMID 18631455 · PMC2727981 · Immunity · 2008 · 7 claims · 8 setups
Transcriptional modules constructed from coordinately expressed genes across 8 diseases form stable, biologically coherent functional units
-
Has reproduction · 51
Construction and Validation of an Immune Infiltration-Related Gene Signature for the Prediction of Prognosis and Therapeutic Response in Breast Cancer.
PMID 33986754 · PMC8110914 · Frontiers in immunology · 2021 · 7 claims · 8 setups
A 15-gene immune infiltration-related signature (IRS) predicts overall survival in breast cancer, with higher IRS indicating worse prognosis
-
Has reproduction · 92
Similarities and Differences in Gene Expression Networks Between the Breast Cancer Cell Line Michigan Cancer Foundation-7 and Invasive Human Breast Cancer Tissues.
PMID 34056582 · PMC8155268 · Frontiers in artificial intelligence · 2021 · 8 claims · 8 setups
MCF-7 cell lines and human breast cancer tissues share only minimal similarity in biological processes, though fundamental functions such as cell cycle are conserved
-
Has reproduction · 56
Identification of key genes in chickpea transcriptomics and the development of ChickpeaOmicsR as a comprehensive resource to advance breeding and genomic studies.
PMID 41909810 · PMC13022592 · Frontiers in bioinformatics · 2026 · 8 claims · 4 setups
ChickpeaOmicsR is the first comprehensive/specialized R package integrating transcriptomic, genomic, and proteomic (RNA-seq, GWAS, PPI) data within a unified, reproducible framework and standardizing fragmented chickpea gene nomenclature.
-
Full-text index only
Application of proteomics in the study of tumor metastasis.
PMID 15862116 · PMC5172469 · Genomics, proteomics & bioinformatics · 2004 · 8 claims · 8 setups
Cell function is directly regulated through proteins, not genes or mRNA, so metastasis-related gene findings need protein-level validation via proteomics.
-
Full-text index only
'Genome design' model and multicellular complexity: golden middle.
PMID 17062620 · PMC1635334 · Nucleic acids research · 2006 · 8 claims · 8 setups
Intermediately expressed human genes are the longest genes genome-wide, in both coding and intronic sequence, longer than housekeeping or tissue-specific genes.
-
Full-text index only
Reference based annotation with GeneMapper.
PMID 16600017 · PMC1557983 · Genome biology · 2006 · 7 claims · 6 setups
GeneMapper transfers reference gene annotations to target genomes with higher accuracy than GeneWise and Projector
-
Full-text index only
Cataloging coding sequence variations in human genome databases.
PMID 18974781 · PMC2570488 · PloS one · 2008 · 8 claims · 7 setups
A significant proportion of CVs overlap between HGMD and dbSNP (4.36% of HGMD CVs registered in dbSNP; 8.11% of dbSNP CVs registered in HGMD), warranting caution when interpreting phenotypic relevance of concurrent CVs.
-
Full-text index only
3' tag digital gene expression profiling of human brain and universal reference RNA using Illumina Genome Analyzer.
PMID 19917133 · PMC2781828 · BMC genomics · 2009 · 8 claims · 4 setups
3' tag DGE transcript profiles are highly reproducible between technical and biological replicates, across libraries made at different labs, and across two generations of Illumina Genome Analyzers (GA I and GA II).
-
Full-text index only
Identification of novel gene amplifications in breast cancer and coexistence of gene amplification with an activating mutation of PIK3CA.
PMID 19706770 · PMC2745517 · Cancer research · 2009 · 8 claims · 8 setups
Genome-wide DNA copy number analysis of 161 primary breast tumors identified six novel focally amplified genes: POLD3, IRAK4, IRX2, TBL1XR1, ASPH, and BRD4
-
Has reproduction · 97
Easy and efficient ensemble gene set testing with EGSEA.
PMID 29333246 · PMC5747338 · F1000Research · 2017 · 8 claims · 2 setups
EGSEA combines results from up to 12 prominent gene set testing algorithms to obtain a consensus ranking of biologically relevant gene sets.
-
Has reproduction · 86
Multi-INTACT: integrative analysis of the genome, transcriptome, and proteome identifies causal mechanisms of complex traits.
PMID 39901160 · PMC11789355 · Genome biology · 2025 · 8 claims · 2 setups
Multi-INTACT achieves higher power than existing single-gene-product methods while maintaining calibrated false discovery rates in simulations.
-
Has reproduction · 92
A network-guided protocol to discover susceptibility genes in genome-wide association studies using stability selection.
PMID 36609152 · PMC9850185 · STAR protocols · 2023 · 5 claims · 5 setups
The protocol identifies genes that are both statistically associated with a phenotype and functionally interconnected in a biological network
-
Has reproduction · 56
DESE: estimating driver tissues by selective expression of genes associated with complex diseases or traits.
PMID 31694669 · PMC6836538 · Genome biology · 2019 · 8 claims · 8 setups
DESE is a unified iterative framework that estimates driver tissues of complex diseases/traits from tissue-selective expression of GWAS-associated genes, and outputs prioritized susceptibility genes as a byproduct