Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
Uncovering information on expression of natural antisense transcripts in Affymetrix MOE430 datasets.
PMID 17598913 · PMC1929078 · BMC genomics · 2007 · 8 claims · 4 setups
Standard Affymetrix expression GeneChips (MOE430, HG-U133) contain probe sets that detect natural antisense transcripts (NATs)
-
Full-text index only
Reference based annotation with GeneMapper.
PMID 16600017 · PMC1557983 · Genome biology · 2006 · 7 claims · 6 setups
GeneMapper transfers reference gene annotations to target genomes with higher accuracy than GeneWise and Projector
-
Full-text index only
Speeding disease gene discovery by sequence based candidate prioritization.
PMID 15766383 · PMC1274252 · BMC bioinformatics · 2005 · 7 claims · 8 setups
Disease genes (OMIM) differ significantly from non-disease genes in sequence-based features including gene/cDNA/protein size, exon number, homolog conservation, secretion signal, 3' UTR length, CpG islands, and distance to nearest gene.
-
Full-text index only
Inferring combinatorial regulation of transcription in silico.
PMID 15647509 · PMC546154 · Nucleic acids research · 2005 · 8 claims · 5 setups
Combining Cluster-Buster (TFBS cluster prediction) with GOSSIP (rigorous GO enrichment statistics with multiple-testing/FDR correction) predicts biological functions controlled by combinatorial transcription factor action, without prior knowledge of factor targets
-
Full-text index only
The other side of comparative genomics: genes with no orthologs between the cow and other mammalian species.
PMID 20003425 · PMC2808326 · BMC genomics · 2009 · 7 claims · 4 setups
3,801 bovine genes have no orthologs in human, mouse and dog, and 1,010 human genes have no orthologs in cow despite having orthologs in mouse and dog
-
Full-text index only
Using structural bioinformatics to investigate the impact of non synonymous SNPs and disease mutations: scope and limitations.
PMID 19758473 · PMC2745591 · BMC bioinformatics · 2009 · 8 claims · 8 setups
None of 39 tested structural properties can be used as a sole classification criterion to separate neutral SNPs from disease mutations.
-
Full-text index only
Prioritization of candidate cancer genes--an aid to oncogenomic studies.
PMID 18710882 · PMC2566894 · Nucleic acids research · 2008 · 8 claims · 8 setups
Computational classifiers using combinations of protein conservation, gene structure, protein domains, protein interactions, and regulatory data can distinguish known cancer genes (CD/CR) from unlabelled human genes
-
Full-text index only
Ensembl 2007.
PMID 17148474 · PMC1761443 · Nucleic acids research · 2007 · 8 claims · 7 setups
Ensembl added 18 new chordate genomes this year, increasing total genomes available from 15 to 33, the largest yearly increase to date.
-
Full-text index only
Development of an integrated genome informatics, data management and workflow infrastructure: a toolbox for the study of complex disease genetics.
PMID 15601538 · PMC3525068 · Human genomics · 2004 · 8 claims · 8 setups
An integrated system combining Ensembl, ACeDB, Gbrowse and custom relational databases provides a scalable genome informatics and workflow infrastructure for complex disease gene discovery.
-
Full-text index only
GeneKeyDB: a lightweight, gene-centric, relational database to support data mining environments.
PMID 15790402 · PMC1274265 · BMC bioinformatics · 2005 · 8 claims · 6 setups
GeneKeyDB is a lightweight, gene-centric relational database that supports data mining and integration with computational analysis tools.
-
Full-text index only
Identification of the REST regulon reveals extensive transposable element-mediated binding site duplication.
PMID 16899447 · PMC1557810 · Nucleic acids research · 2006 · 8 claims · 8 setups
The RE1 PSSM identifies functional RE1 binding sites with greater sensitivity and selectivity than the previously used RE1 consensus sequence
-
Full-text index only
Functional nsSNPs from carcinogenesis-related genes expressed in breast tissue: potential breast cancer risk alleles and their distribution across human populations.
PMID 16595073 · PMC3500178 · Human genomics · 2006 · 7 claims · 5 setups
A bioinformatics strategy cross-referencing carcinogenesis-related gene lists with breast-tissue expression data can identify candidate breast cancer risk nsSNPs.
-
Full-text index only
Integration of text- and data-mining using ontologies successfully selects disease gene candidates.
PMID 15767279 · PMC1065256 · Nucleic acids research · 2005 · 7 claims · 6 setups
Integrating eVOC anatomical ontology-based text-mining of PubMed abstracts with data-mining of gene expression annotation successfully selects and prioritizes candidate disease genes
-
Full-text index only
Computational disease gene identification: a concert of methods prioritizes type 2 diabetes and obesity candidate genes.
PMID 16757574 · PMC1475747 · Nucleic acids research · 2006 · 6 claims · 8 setups
Applying seven independent computational disease-gene prioritization methods in concert to 9556 positional candidate genes identifies a prioritized set of likely T2D and obesity candidate genes
-
Full-text index only
JIGSAW, GeneZilla, and GlimmerHMM: puzzling out the features of human genes in the ENCODE regions.
PMID 16925843 · PMC1810558 · Genome biology · 2006 · 8 claims · 4 setups
Adding model states for specific biological features (signal peptides, CpG islands, etc.) to non-comparative GHMM gene finders did little or nothing to enhance predictive accuracy, sometimes reducing it.
-
Full-text index only
Benchmarking ortholog identification methods using functional genomics data.
PMID 16613613 · PMC1557999 · Genome biology · 2006 · 8 claims · 7 setups
InParanoid is the best overall ortholog identification method for identifying functionally equivalent proteins when sensitivity and selectivity are combined into an overall score.
-
Full-text index only
G2Cdb: the Genes to Cognition database.
PMID 18984621 · PMC2686544 · Nucleic acids research · 2009 · 7 claims · 7 setups
G2Cdb integrates experimentally validated synapse proteome datasets with mouse/human genomic annotation, phenotype, and human disease data in a gene-centric database.
-
Has reproduction · 88
Comprehensive benchmarking of large language models for RNA secondary structure prediction.
PMID 40205851 · PMC11982019 · Briefings in bioinformatics · 2025 · 7 claims · 4 setups
Existing RNA-LLMs had not previously been evaluated for secondary structure prediction in a unified, fair experimental setup with the same datasets and prediction model.