Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
Exhaustive prediction of disease susceptibility to coding base changes in the human genome.
PMID 18793467 · PMC2537574 · BMC bioinformatics · 2008 · 8 claims · 7 setups
Inter-species conservation is the strongest single predictor of disease-associated coding mutations among the factors tested.
-
Has reproduction · 87
Insights into the Evolution of the New World Diploid Cottons (Gossypium, Subgenus Houzingenia) Based on Genome Sequencing.
PMID 30476109 · PMC6320677 · Genome biology and evolution · 2019 · 8 claims · 8 setups
Subgenus Houzingenia likely originated via transoceanic dispersal from Africa about 6.6 Ma
-
Full-text index only
Conservation, variability and the modeling of active protein kinases.
PMID 17912359 · PMC1989141 · PloS one · 2007 · 7 claims · 5 setups
A novel sequence-order independent (fold-independent) structural alignment algorithm was developed that maximizes side-chain similarity to produce a consensus kinase structure.
-
Full-text index only
CpG_MI: a novel approach for identifying functional CpG islands in mammalian genomes.
PMID 19854943 · PMC2800233 · Nucleic acids research · 2010 · 8 claims · 6 setups
Functional ('bona fide') CGIs show distinct average/cumulative mutual information (AMI/CMI) distributions of neighboring CpG distances compared to non-functional CGIs and random genome segments
-
Has reproduction · 87
De Novo Transcriptome Meta-Assembly of the Mixotrophic Freshwater Microalga Euglena gracilis.
PMID 34072576 · PMC8227486 · Genes · 2021 · 7 claims · 8 setups
A new consensus transcriptome of E. gracilis was assembled by combining reads from five independent RNA-seq studies (23 samples)
-
Full-text index only
A global view of protein expression in human cells, tissues, and organs.
PMID 20029370 · PMC2824494 · Molecular systems biology · 2009 · 7 claims · 6 setups
A high fraction (>65%) of proteins is expressed in most human cells and tissues, while very few proteins (<2%) are detected in any single cell type.
-
Has reproduction · 94
Systematic assessment of pathway databases, based on a diverse collection of user-submitted experiments.
PMID 36088548 · PMC9487593 · Briefings in bioinformatics · 2022 · 8 claims · 6 setups
Well-established, hierarchically organized pathway annotation systems (e.g. GO, Reactome, KEGG) yield the best overall enrichment performance despite covering much of the human genome only in general terms.
-
Full-text index only
InParanoid 6: eukaryotic ortholog clusters with inparalogs.
PMID 18055500 · PMC2238924 · Nucleic acids research · 2008 · 8 claims · 3 setups
InParanoid 6 is an updated eukaryotic ortholog database covering 35 species (34 eukaryotes plus E. coli as outgroup), providing pairwise ortholog clusters with inparalogs for all species pairs.
-
Has reproduction · 80
Differential analysis of RNA structure probing experiments at nucleotide resolution: uncovering regulatory functions of RNA structure.
PMID 35869080 · PMC9307511 · Nature communications · 2022 · 7 claims · 4 setups
DiffScan is a computational framework combining a Normalization module and a Scan module to identify SVRs at nucleotide resolution from SP data.
-
Full-text index only
A re-annotation pipeline for Illumina BeadArrays: improving the interpretation of gene expression data.
PMID 19923232 · PMC2817484 · Nucleic acids research · 2010 · 8 claims · 7 setups
A Perl-based pipeline that BLASTs/BLATs Illumina probe sequences against genomes and transcript databases (RefSeq, UCSC Known Genes, UniGene/GenBank, Ensembl) can classify probes by quality grade (Perfect/Good/Bad/No match) and is applicable across 8 BeadArray platforms and other array types
-
Full-text index only
Discovering cancer genes by integrating network and functional properties.
PMID 19765316 · PMC2758898 · BMC medical genomics · 2009 · 8 claims · 6 setups
Cancer genes have distinct PPI network topology (higher connectivity, higher clustering coefficient, shorter path length to known cancer genes) compared to non-cancer genes
-
Has reproduction · 79
Symbiosis genes show a unique pattern of introgression and selection within a Rhizobium leguminosarum species complex.
PMID 32176601 · PMC7276703 · Microbial genomics · 2020 · 8 claims · 8 setups
196 R. leguminosarum sv. trifolii strains form a five-species complex (gsA-gsE) with generally little recent between-species gene transfer, aside from a few highly mobile genetic regions
-
Has reproduction · 90
Comparative Genomics Provides Insight into the Function of Broad-Host Range Sponge Symbionts.
PMID 34519538 · PMC8546597 · mBio · 2021 · 8 claims · 8 setups
Eleven new genomes were added to the Tethybacterales order and a novel family (Polydorabacteraceae) was identified
-
Full-text index only
A novel wavelet-based thresholding method for the pre-processing of mass spectrometry data that accounts for heterogeneous noise.
PMID 18615428 · PMC2855839 · Proteomics · 2008 · 6 claims · 4 setups
Noise in SELDI-TOF/MALDI-TOF mass spectrometry data is heteroscedastic across the m/z range, with larger variance at lower m/z values, contrary to the homogeneous noise assumption of existing wavelet denoising methods.
-
Full-text index only
An online database for brain disease research.
PMID 16594998 · PMC1489945 · BMC genomics · 2006 · 7 claims · 5 setups
SMRIDB is a comprehensive web-based database integrating gene expression data and clinical metadata to aid understanding of the genetic effects of brain disease (bipolar disorder, schizophrenia, depression)
-
Has reproduction · 67
Optimal scaling of digital transcriptomes.
PMID 24223126 · PMC3819321 · PloS one · 2013 · 8 claims · 8 setups
Fifteen existing and novel transcript-count normalization algorithms can be compared with two novel, mutually independent metrics: the number of "uniform" genes (sufficiently low coefficient of variation after normalization) and low average Spearman correlation between normalized expression profiles of gene pairs.
-
Full-text index only
Diversity of preferred nucleotide sequences around the translation initiation codon in eukaryote genomes.
PMID 18086709 · PMC2241899 · Nucleic acids research · 2008 · 8 claims · 5 setups
Preferred nucleotide sequences around the initiation codon are diverse among eukaryote species, but differences roughly reflect evolutionary relationships between species
-
Full-text index only
Information-based methods for predicting gene function from systematic gene knock-downs.
PMID 18959798 · PMC2596148 · BMC bioinformatics · 2008 · 8 claims · 4 setups
Information-based metrics, which incorporate a phenotype's genomic frequency, outperform non-information-based metrics for detecting gene-gene functional similarity from phenotypic knock-down profiles.
-
Full-text index only
The biological function of some human transcription factor binding motifs varies with position relative to the transcription start site.
PMID 18367472 · PMC2377430 · Nucleic acids research · 2008 · 8 claims · 5 setups
1226 eight-letter DNA words show statistically significant positional preferences relative to the TSS across 7914 human promoter regions
-
Has reproduction · 83
Public Omics Explorer (POE): Enabling integrative semantic search across GEO omics datasets based on PubMed publications.
PMID 41282419 · PMC12636342 · Computational and structural biotechnology journal · 2025 · 7 claims · 3 setups
POE performs literature-informed dataset retrieval by semantically linking GEO datasets and ENA records through associated PubMed publications