Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
Transcription network construction for large-scale microarray datasets using a high-performance computing approach.
PMID 18366618 · PMC2386070 · BMC genomics · 2008 · 8 claims · 7 setups
RMT removes the random noise component of the gene expression correlation matrix by testing its eigenvalue statistics against a null hypothesis derived from a truly random correlation matrix
-
Full-text index only
CLEAN: CLustering Enrichment ANalysis.
PMID 19640299 · PMC2734555 · BMC bioinformatics · 2009 · 8 claims · 4 setups
The gene-specific CLEAN score improves reproducibility of cluster analysis conclusions across independent datasets compared to the traditional cluster-wide score (cwCLEAN).
-
Has reproduction · 63
Clustering and machine learning-based integration identify cancer associated fibroblasts genes' signature in head and neck squamous cell carcinoma.
PMID 37065499 · PMC10098459 · Frontiers in genetics · 2023 · 8 claims · 8 setups
Clustering of 31 CAFs genes across 868 HNSCC samples identifies two distinct molecular patterns (C1, C2) with different survival outcomes
-
Full-text index only
A taxonomy of epithelial human cancer and their metastases.
PMID 20017941 · PMC2806369 · BMC medical genomics · 2009 · 8 claims · 6 setups
Unsupervised hierarchical clustering of 1566 primary epithelial tumors yields large tissue-enriched clusters (breast, colon/GI, lung, ovary, kidney) plus smaller prostate, thyroid-kidney, and mixed clusters
-
Full-text index only
InParanoid 7: new algorithms and tools for eukaryotic orthology analysis.
PMID 19892828 · PMC2808972 · Nucleic acids research · 2010 · 8 claims · 7 setups
InParanoid 7 expands the database by an order of magnitude to 100 species, 1.3 million proteins, and 42.7 million pairwise ortholog groups.
-
Has reproduction
Artificial Intelligence Meets Whole Slide Images: Deep Learning Model Shapes an Immune-Hot Tumor and Guides Precision Therapy in Bladder Cancer.
PMID 36245985 · PMC9553530 · Journal of oncology · 2022 · 6 claims · 8 setups
A deep learning WSI cluster (three-class mini batch K-means on Inception V3 features) is associated with overall survival (P<0.001) and is an independent prognostic predictor (P=0.031) in BLCA.
-
Full-text index only
Clustering by neurocognition for fine mapping of the schizophrenia susceptibility loci on chromosome 6p.
PMID 19694819 · PMC4286260 · Genes, brain, and behavior · 2009 · 6 claims · 6 setups
A family-based clustering strategy using neurocognitive test scores (CPT, WCST) can identify more homogeneous subgroups of schizophrenia families for genetic association analysis
-
Full-text index only
Iterative class discovery and feature selection using Minimal Spanning Trees.
PMID 15355552 · PMC520744 · BMC bioinformatics · 2004 · 7 claims · 5 setups
Iterating between MST-based clustering and t-statistic feature selection removes noise genes step-wise while sharpening the sample clustering
-
Has reproduction · 95
A role for ColV plasmids in the evolution of pathogenic Escherichia coli ST58.
PMID 35115531 · PMC8813906 · Nature communications · 2022 · 8 claims · 8 setups
ST58 contains a major sub-lineage (BAP2, n=363) characterized by near-ubiquitous carriage of ColV plasmids
-
Full-text index only
ECgene: genome annotation for alternative splicing.
PMID 15608289 · PMC540072 · Nucleic acids research · 2005 · 8 claims · 5 setups
ECgene combines genome-based EST clustering with a graph-theoretic transcript assembly procedure to predict gene models including alternative splicing events.
-
Full-text index only
Comprehensive genome analysis of 203 genomes provides structural genomics with new insights into protein family space.
PMID 16481312 · PMC1373602 · Nucleic acids research · 2006 · 8 claims · 7 setups
The number of protein families continues to expand steadily as more genomes are sequenced, showing no sign of saturation.
-
Has reproduction · 56
Comparative Metagenomic Analysis of Biosynthetic Diversity across Sponge Microbiomes Highlights Metabolic Novelty, Conservation, and Diversification.
PMID 35862823 · PMC9426513 · mSystems · 2022 · 8 claims · 5 setups
The vast majority of recovered gene cluster families (GCFs) in sponge microbiomes show no similarity to any characterized BGC, revealing extreme biosynthetic novelty
-
Full-text index only
The Princeton Protein Orthology Database (P-POD): a comparative genomics analysis tool for biologists.
PMID 17712414 · PMC1942082 · PloS one · 2007 · 8 claims · 5 setups
P-POD is the first comparative genomics database to combine results from multiple computational ortholog/homolog prediction methods with manually curated literature-derived experimental evidence of functional conservation.
-
Full-text index only
MLST clustering of Campylobacter jejuni isolates from patients with gastroenteritis, reactive arthritis and Guillain-Barré syndrome.
PMID 19702866 · PMC3985121 · Journal of applied microbiology · 2010 · 8 claims · 4 setups
Danish C. jejuni isolates from gastroenteritis, RA and GBS patients are highly genetically diverse, comprising 51 STs within 18 clonal complexes among 122 isolates.
-
Has reproduction · 89
Identification of genes influencing the evolution of Escherichia coli ST372 in dogs and humans.
PMID 36752777 · PMC9997745 · Microbial genomics · 2023 · 8 claims · 8 setups
Dogs are the dominant host of E. coli ST372, and clusters within the ST372 population structure exhibit distinctive O:H types.
-
Full-text index only
The use of edge-betweenness clustering to investigate biological function in protein interaction networks.
PMID 15740614 · PMC555937 · BMC bioinformatics · 2005 · 8 claims · 7 setups
Edge-Betweenness clustering separates protein interaction graphs into subgraphs whose GO term distributions show significant correlations, revealing biologically meaningful functional modules.
-
Full-text index only
A protein interaction based model for schizophrenia study.
PMID 19091023 · PMC2638163 · BMC bioinformatics · 2008 · 8 claims · 4 setups
Products of 36 schizophrenia candidate genes cluster together into a single connected component within a PPI sub-network of 831 proteins
-
Has reproduction · 87
CoINcIDE: A framework for discovery of patient subtypes across multiple datasets.
PMID 26961683 · PMC4784276 · Genome medicine · 2016 · 8 claims · 6 setups
CoINcIDE is a methodological framework that discovers replicable patient subtypes (meta-clusters) across multiple datasets by finding consensus across dataset-specific clusterings, requiring no between-dataset transformations.
-
Full-text index only
Investigating hookworm genomes by comparative analysis of two Ancylostoma species.
PMID 15854223 · PMC1112591 · BMC genomics · 2005 · 8 claims · 8 setups
Nearly 20,000 ESTs from 7 cDNA libraries define nearly 7,000 hookworm genes across A. caninum and A. ceylanicum
-
Full-text index only
Clustering of phosphorylation site recognition motifs can be exploited to predict the targets of cyclin-dependent kinase.
PMID 17316440 · PMC1852407 · Genome biology · 2007 · 8 claims · 6 setups
CDK consensus motifs are frequently clustered (closely spaced) in known CDK substrate proteins rather than uniformly distributed