Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
Coiled-coil protein composition of 22 proteomes--differences and common themes in subcellular infrastructure and traffic control.
PMID 16288662 · PMC1322226 · BMC evolutionary biology · 2005 · 7 claims · 5 setups
Proteins with extended coiled-coil domains (>250 amino acids) are largely absent from bacterial genomes but present in archaea and eukaryotes.
-
Has reproduction · 49
oPOSSUM-3: advanced analysis of regulatory motif over-representation across genes or ChIP-Seq datasets.
PMID 22973536 · PMC3429929 · G3 (Bethesda, Md.) · 2012 · 8 claims · 6 setups
oPOSSUM-3 is a web-accessible system that identifies over-represented TFBS and TFBS families in DNA sequences of co-expressed genes or in sequences from high-throughput methods such as ChIP-Seq.
-
Full-text index only
Identifying cis-regulatory sequences by word profile similarity.
PMID 19730735 · PMC2731932 · PloS one · 2009 · 8 claims · 8 setups
WPH-finder identifies putative co-regulated CRMs by scanning the genome for sequences with word profiles similar to a known CRM, without explicitly defining binding sites
-
Full-text index only
Mutation analysis of 24 known cancer genes in the NCI-60 cell line set.
PMID 17088437 · PMC2705832 · Molecular cancer therapeutics · 2006 · 8 claims · 3 setups
137 oncogenic mutations were identified across 14 of 24 screened cancer genes (APC, BRAF, CDKN2A, CTNNB1, HRAS, KRAS, NRAS, SMAD4, PIK3CA, PTEN, RB1, STK11, TP53, VHL) in the NCI-60 panel
-
Full-text index only
Rapid detection and curation of conserved DNA via enhanced-BLAT and EvoPrinterHD analysis.
PMID 18307801 · PMC2268679 · BMC genomics · 2008 · 8 claims · 8 setups
eBLAT detects up to 75% more conserved bases than original BLAT alignments, with the largest gains between evolutionarily distant orthologs
-
Full-text index only
Empirical codon substitution matrix.
PMID 15927081 · PMC1173088 · BMC bioinformatics · 2005 · 8 claims · 5 setups
The authors present the first empirical codon substitution matrix built entirely from alignments of vertebrate coding DNA sequences.
-
Full-text index only
EPGD: a comprehensive web resource for integrating and displaying eukaryotic paralog/paralogon information.
PMID 17984073 · PMC2238967 · Nucleic acids research · 2008 · 8 claims · 8 setups
EPGD is a gene-centered, internet-accessible database integrating paralog family and paralogon information for 26 eukaryotic genomes.
-
Full-text index only
A procedure for the detection of linkage with high density SNP arrays in a large pedigree with colorectal cancer.
PMID 17222328 · PMC1784097 · BMC cancer · 2007 · 7 claims · 8 setups
A workflow combining Alohomora, Mega2, MENDEL, SNPLINK and SimWalk2 enables linkage analysis with high-density SNP arrays in large pedigrees (>35-40 bits) that exceed the capacity of single existing programs
-
Full-text index only
CLEAN: CLustering Enrichment ANalysis.
PMID 19640299 · PMC2734555 · BMC bioinformatics · 2009 · 8 claims · 4 setups
The gene-specific CLEAN score improves reproducibility of cluster analysis conclusions across independent datasets compared to the traditional cluster-wide score (cwCLEAN).
-
Has reproduction · 71
RNAmountAlign: Efficient software for local, global, semiglobal pairwise and multiple RNA sequence/structure alignment.
PMID 31978147 · PMC6980424 · PloS one · 2020 · 7 claims · 6 setups
RNAmountAlign performs pairwise local, global, and semiglobal (query search) alignment and progressive multiple alignment (global and local) using incremental ensemble mountain height, running in O(n^3) time and O(n^2) space for two sequences of length n
-
Full-text index only
The whole alignment and nothing but the alignment: the problem of spurious alignment flanks.
PMID 18796526 · PMC2566872 · Nucleic acids research · 2008 · 8 claims · 4 setups
Some common scoring schemes tend to overextend alignments, generating spurious alignment flanks up to hundreds of bp/amino acids in length
-
Full-text index only
MODBASE, a database of annotated comparative protein structure models and associated resources.
PMID 18948282 · PMC2686492 · Nucleic acids research · 2009 · 8 claims · 8 setups
MODBASE contains 5,152,695 reliable comparative protein structure models for 1,593,209 unique protein sequences.
-
Full-text index only
CTCF binding site classes exhibit distinct evolutionary, genomic, epigenomic and transcriptomic features.
PMID 19922652 · PMC3091324 · Genome biology · 2009 · 8 claims · 8 setups
CTCF binding sites can be classified into three occupancy-based classes (LowOc, MedOc, HighOc) based on similarity to the CTCF PWM motif
-
Has reproduction · 79
TSUNAMI: Translational Bioinformatics Tool Suite for Network Analysis and Mining.
PMID 33705981 · PMC9403021 · Genomics, proteomics & bioinformatics · 2021 · 8 claims · 6 setups
TSUNAMI is a freely accessible web-based tool suite that mines gene co-expression network (GCN) modules from public (GEO, TCGA) or user-uploaded numerical omics data and performs downstream gene set enrichment analysis.
-
Has reproduction · 89
TrEMOLO: accurate transposable element allele frequency estimation using long-read sequencing data combining assembly and mapping-based approaches.
PMID 37013657 · PMC10069131 · Genome biology · 2023 · 6 claims · 6 setups
TrEMOLO combines an assembly-based INSIDER module and a mapping-based OUTSIDER module to detect TE insertions/deletions from long-read sequencing data and estimate their allele frequency
-
Full-text index only
Automatic discovery of cross-family sequence features associated with protein function.
PMID 16409628 · PMC1395344 · BMC bioinformatics · 2006 · 8 claims · 6 setups
A self-supervised data mining approach can find relationships between sequence features and functional annotations without preconceived functional categories.
-
Has reproduction · 87
Enhanced Generalizability of RNA Secondary Structure Prediction via Convolutional Block Attention Network and Ensemble Learning.
PMID 40871599 · PMC12388828 · Molecules (Basel, Switzerland) · 2025 · 8 claims · 8 setups
TrioFold integrates base-pairing clues from thermodynamic- and DL-based methods via ensemble learning and a convolutional block attention mechanism to enhance RSS prediction generalizability.
-
Full-text index only
Finding disease candidate genes by liquid association.
PMID 17915034 · PMC2246280 · Genome biology · 2007 · 7 claims · 6 setups
LA can detect functionally associated genes that are not directly co-expressed by identifying a mediating gene Z whose expression level changes the correlation between X and Y.
-
Has reproduction · 78
Emergent dynamics of underlying regulatory network links EMT and androgen receptor-dependent resistance in prostate cancer.
PMID 36851919 · PMC9957767 · Computational and structural biotechnology journal · 2023 · 8 claims · 7 setups
Simulations of the EMT-AR crosstalk network reveal four possible phenotypes: epithelial-sensitive (ES), epithelial-resistant (ER), mesenchymal-resistant (MR), and mesenchymal-sensitive (MS), with MS occurring rarely
-
Has reproduction · 79
Symbiosis genes show a unique pattern of introgression and selection within a Rhizobium leguminosarum species complex.
PMID 32176601 · PMC7276703 · Microbial genomics · 2020 · 8 claims · 8 setups
196 R. leguminosarum sv. trifolii strains form a five-species complex (gsA-gsE) with generally little recent between-species gene transfer, aside from a few highly mobile genetic regions