Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
Multiple whole genome alignments and novel biomedical applications at the VISTA portal.
PMID 17488840 · PMC1933192 · Nucleic acids research · 2007 · 8 claims · 4 setups
A novel multiple whole-genome alignment algorithm treats all genomes symmetrically, avoiding dependence on a single base/reference genome
-
Full-text index only
The DAVID Gene Functional Classification Tool: a novel biological module-centric algorithm to functionally analyze large gene lists.
PMID 17784955 · PMC2375021 · Genome biology · 2007 · 8 claims · 6 setups
Gene-gene functional similarity can be measured using kappa statistics applied to a binary gene-annotation-term matrix built from 14 annotation categories.
-
Full-text index only
Computer-aided identification of polymorphism sets diagnostic for groups of bacterial and viral genetic variants.
PMID 17672919 · PMC1973086 · BMC bioinformatics · 2007 · 6 claims · 8 setups
The Not-N algorithm, incorporated into the Minimum SNPs program, identifies small marker sets diagnostic for user-defined subgroups of genetic variants with 0% false negatives
-
Has reproduction · 79
TSUNAMI: Translational Bioinformatics Tool Suite for Network Analysis and Mining.
PMID 33705981 · PMC9403021 · Genomics, proteomics & bioinformatics · 2021 · 8 claims · 6 setups
TSUNAMI is a freely accessible web-based tool suite that mines gene co-expression network (GCN) modules from public (GEO, TCGA) or user-uploaded numerical omics data and performs downstream gene set enrichment analysis.
-
Has reproduction · 51
A platelet-related signature for predicting the prognosis and immunotherapy benefit in bladder cancer based on machine learning combinations.
PMID 39280688 · PMC11399026 · Translational andrology and urology · 2024 · 8 claims · 8 setups
An Enet (alpha=0.4) machine-learning algorithm built from 10 platelet-related genes yields the optimal platelet-related signature (PRS) for bladder cancer prognosis, with average C-index 0.73
-
Full-text index only
Genome-wide identification of specific oligonucleotides using artificial neural network and computational genomic analysis.
PMID 17518996 · PMC1892811 · BMC bioinformatics · 2007 · 7 claims · 4 setups
The IAB algorithm (integration of ANN and BLAST) identifies genome-wide specific oligos much faster than pure BLAST search while maintaining comparable success rate and cross homology
-
Full-text index only
Integrated analysis of genetic and proteomic data identifies biomarkers associated with adverse events following smallpox vaccination.
PMID 18923431 · PMC2692715 · Genes and immunity · 2009 · 7 claims · 6 setups
A two-stage strategy (Random Forest filtering followed by decision tree modeling) can integrate categorical genetic and continuous proteomic data to identify biomarkers of AE risk
-
Has reproduction · 85
Digital sorting of complex tissues for cell type-specific gene expression profiles.
PMID 23497278 · PMC3626856 · BMC bioinformatics · 2013 · 8 claims · 8 setups
The Digital Sorting Algorithm (DSA) deconvolves mixed tissue expression into cell type-specific profiles using only marker genes, without requiring prior knowledge of cell type frequencies or in vitro pure-cell profiles.
-
Full-text index only
A computational study of off-target effects of RNA interference.
PMID 15800213 · PMC1072799 · Nucleic acids research · 2005 · 8 claims · 5 setups
The chance of RNAi off-target effects is considerable, ranging from 5% to 80% depending on organism and parameters, when using exact sequence identity between siRNA and transcripts.
-
Full-text index only
Prioritization of candidate cancer genes--an aid to oncogenomic studies.
PMID 18710882 · PMC2566894 · Nucleic acids research · 2008 · 8 claims · 8 setups
Computational classifiers using combinations of protein conservation, gene structure, protein domains, protein interactions, and regulatory data can distinguish known cancer genes (CD/CR) from unlabelled human genes
-
Full-text index only
CGMIM: automated text-mining of Online Mendelian Inheritance in Man (OMIM) to identify genetically-associated cancers and candidate genes.
PMID 15796777 · PMC1274267 · BMC bioinformatics · 2005 · 8 claims · 2 setups
CGMIM is a Perl program that text-mines OMIM entries to identify cancer-gene associations and genetically-related cancer type pairs.
-
Full-text index only
Comprehensive search for intra- and inter-specific sequence polymorphisms among coding envelope genes of retroviral origin found in the human genome: genes and pseudogenes.
PMID 16150157 · PMC1236922 · BMC genomics · 2005 · 8 claims · 5 setups
HERV-W (envW) and HERV-FRD (envFRD) envelope genes, both specifically expressed in placenta, show strong sequence conservation with only two nonsynonymous SNPs identified across 91 individuals
-
Full-text index only
Variations in the transcriptome of Alzheimer's disease reveal molecular networks involved in cardiovascular diseases.
PMID 18842138 · PMC2760875 · Genome biology · 2008 · 8 claims · 6 setups
AD-related genes (APOE, A2M, PON2, MAP4) and CVD-associated genes (COMT, CBS, WNK1) congregate in a single co-expression module, linking AD and CVD at the transcriptional level
-
Full-text index only
Dissecting microregulation of a master regulatory network.
PMID 18294391 · PMC2289817 · BMC genomics · 2008 · 8 claims · 6 setups
143 human miRNAs (termed p53-miRs) each contain at least one putative p53 binding site within 10 kb flanking sequence and are predicted to target at least one known gene
-
Has reproduction · 70
GAUGE-Annotated Microbial Transcriptomic Data Facilitate Parallel Mining and High-Throughput Reanalysis To Form Data-Driven Hypotheses.
PMID 33758032 · PMC8547006 · mSystems · 2021 · 7 claims · 6 setups
GAUGE automatically annotates GEO microbial microarray and RNA-seq data sets, increasing the percentage amenable to analysis from 4% to 33%.
-
Has reproduction · 63
Clustering and machine learning-based integration identify cancer associated fibroblasts genes' signature in head and neck squamous cell carcinoma.
PMID 37065499 · PMC10098459 · Frontiers in genetics · 2023 · 8 claims · 8 setups
Clustering of 31 CAFs genes across 868 HNSCC samples identifies two distinct molecular patterns (C1, C2) with different survival outcomes
-
Full-text index only
A third approach to gene prediction suggests thousands of additional human transcribed regions.
PMID 16543943 · PMC1391917 · PLoS computational biology · 2006 · 8 claims · 7 setups
A third basic concept for gene prediction exists, based on detecting strand-specific 'transcription footprints' (mutational and selectional biases) rather than gene structure or sequence similarity.
-
Full-text index only
The Princeton Protein Orthology Database (P-POD): a comparative genomics analysis tool for biologists.
PMID 17712414 · PMC1942082 · PloS one · 2007 · 8 claims · 5 setups
P-POD is the first comparative genomics database to combine results from multiple computational ortholog/homolog prediction methods with manually curated literature-derived experimental evidence of functional conservation.
-
Full-text index only
CompMoby: comparative MobyDick for detection of cis-regulatory motifs.
PMID 18950538 · PMC2605473 · BMC bioinformatics · 2008 · 7 claims · 4 setups
CompMoby identifies cis-regulatory binding sites at both transcriptional and post-transcriptional levels in metazoans without prior knowledge of the trans-acting factor
-
Full-text index only
A modular analysis framework for blood genomics studies: application to systemic lupus erythematosus.
PMID 18631455 · PMC2727981 · Immunity · 2008 · 7 claims · 8 setups
Transcriptional modules constructed from coordinately expressed genes across 8 diseases form stable, biologically coherent functional units