Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
CompMoby: comparative MobyDick for detection of cis-regulatory motifs.
PMID 18950538 · PMC2605473 · BMC bioinformatics · 2008 · 7 claims · 4 setups
CompMoby identifies cis-regulatory binding sites at both transcriptional and post-transcriptional levels in metazoans without prior knowledge of the trans-acting factor
-
Full-text index only
Optimal step length EM algorithm (OSLEM) for the estimation of haplotype frequency and its application in lipoprotein lipase genotyping.
PMID 12529185 · PMC149347 · BMC bioinformatics · 2003 · 5 claims · 4 setups
OSLEM (Optimal Step Length EM), which approximates an optimal step length via a fixed-point search (D_N = D_{N-1} + λ(D_preN - D_{N-1})), runs about twice as fast as standard EM while producing the same haplotype frequency estimates.
-
Full-text index only
Assignment of Streptococcus agalactiae isolates to clonal complexes using a small set of single nucleotide polymorphisms.
PMID 18710585 · PMC2533671 · BMC microbiology · 2008 · 7 claims · 6 setups
A four-SNP set (glnA36, glnA429, glcK180, adhP111) identified via the Not-N algorithm plus empirical testing divides GBS into 10 groups concordant with eBURST-defined population structure.
-
Has reproduction · 67
binny: an automated binning algorithm to recover high-quality genomes from complex metagenomic datasets.
PMID 36239393 · PMC9677464 · Briefings in bioinformatics · 2022 · 8 claims · 8 setups
binny outperforms or is highly competitive with commonly used and state-of-the-art binning methods (MetaBAT2, MaxBin2, CONCOCT, VAMB, SemiBin, MetaDecoder)
-
Full-text index only
InParanoid 7: new algorithms and tools for eukaryotic orthology analysis.
PMID 19892828 · PMC2808972 · Nucleic acids research · 2010 · 8 claims · 7 setups
InParanoid 7 expands the database by an order of magnitude to 100 species, 1.3 million proteins, and 42.7 million pairwise ortholog groups.
-
Full-text index only
Local combinational variables: an approach used in DNA-binding helix-turn-helix motif prediction with sequence information.
PMID 19651875 · PMC2761287 · Nucleic acids research · 2009 · 8 claims · 7 setups
The LCV approach predicts HTH motifs with 93.29% accuracy, 93.93% sensitivity and 92.66% specificity using only primary sequence information
-
Full-text index only
Swarm intelligence based wavelet coefficient feature selection for mass spectral classification: an application to proteomics data.
PMID 19733729 · PMC2748225 · Analytica chimica acta · 2009 · 8 claims · 4 setups
ACA-based wavelet coefficient feature selection can achieve up to 100% classification accuracy on training, validating, and independent testing sets using only 5 selected features.
-
Has reproduction · 79
TSUNAMI: Translational Bioinformatics Tool Suite for Network Analysis and Mining.
PMID 33705981 · PMC9403021 · Genomics, proteomics & bioinformatics · 2021 · 8 claims · 6 setups
TSUNAMI is a freely accessible web-based tool suite that mines gene co-expression network (GCN) modules from public (GEO, TCGA) or user-uploaded numerical omics data and performs downstream gene set enrichment analysis.
-
Full-text index only
MiPred: classification of real and pseudo microRNA precursors using random forest prediction model with combined features.
PMID 17553836 · PMC1933124 · Nucleic acids research · 2007 · 8 claims · 8 setups
A hybrid feature combining local contiguous triplet structure-sequence composition, MFE of the secondary structure, and P-value of a randomization test improves classification of real vs pseudo pre-miRNAs
-
Has reproduction · 51
A platelet-related signature for predicting the prognosis and immunotherapy benefit in bladder cancer based on machine learning combinations.
PMID 39280688 · PMC11399026 · Translational andrology and urology · 2024 · 8 claims · 8 setups
An Enet (alpha=0.4) machine-learning algorithm built from 10 platelet-related genes yields the optimal platelet-related signature (PRS) for bladder cancer prognosis, with average C-index 0.73
-
Has reproduction · 71
Assessment tool based on fatty acid metabolic signatures for predicting the prognosis and treatment response in bladder cancer.
PMID 38076064 · PMC10703629 · Heliyon · 2023 · 8 claims · 8 setups
Consensus clustering of prognosis-related fatty acid metabolism genes (FAMGs) identifies three molecular subtypes of BLCA (FAMC1, FAMC2, FAMC3) with distinct prognoses and tumor microenvironments
-
Full-text index only
Evaluating the performance of Affymetrix SNP Array 6.0 platform with 400 Japanese individuals.
PMID 18803882 · PMC2566316 · BMC genomics · 2008 · 8 claims · 5 setups
About 20% of the 909,622 SNPs on the SNP Array 6.0 are monomorphic in the Japanese population
-
Full-text index only
CpG_MI: a novel approach for identifying functional CpG islands in mammalian genomes.
PMID 19854943 · PMC2800233 · Nucleic acids research · 2010 · 8 claims · 6 setups
Functional ('bona fide') CGIs show distinct average/cumulative mutual information (AMI/CMI) distributions of neighboring CpG distances compared to non-functional CGIs and random genome segments
-
Has reproduction · 83
Multiomic machine learning on lactylation for molecular typing and prognosis of lung adenocarcinoma.
PMID 39856156 · PMC11760357 · Scientific reports · 2025 · 8 claims · 8 setups
Ten multiomics clustering algorithms identify two distinct lactylation cancer subtypes (CS1 and CS2) in LUAD
-
Full-text index only
Integration of text- and data-mining using ontologies successfully selects disease gene candidates.
PMID 15767279 · PMC1065256 · Nucleic acids research · 2005 · 7 claims · 6 setups
Integrating eVOC anatomical ontology-based text-mining of PubMed abstracts with data-mining of gene expression annotation successfully selects and prioritizes candidate disease genes
-
Has reproduction · 71
Protein structure quality assessment based on the distance profiles of consecutive backbone Cα atoms.
PMID 24555103 · PMC3892923 · F1000Research · 2013 · 8 claims · 8 setups
The distance between consecutive backbone Cα atoms in high-quality structures is normally distributed with mean 3.8 Å and standard deviation 0.04 Å, justifying a reference state in which all consecutive Cα atoms are 3.8 Å apart.
-
Has reproduction · 96
Deep learning based protocol to construct an immune-related gene network of host-pathogen interactions in plants.
PMID 36525344 · PMC9791427 · STAR protocols · 2023 · 6 claims · 6 setups
A deep-learning protocol (DLNet) ranks genes by their contribution to classifying treatment versus control expression data, identifying genes involved in host defense against pathogens.
-
Has reproduction · 81
Macrophages on the run: Exercise balances macrophage polarization for improved health.
PMID 39476967 · PMC11585839 · Molecular metabolism · 2024 · 8 claims · 7 setups
Immediate/acute exercise triggers an M1 (pro-inflammatory) macrophage polarization surge.
-
Has reproduction · 66
A global database for modeling tumor-immune cell communication.
PMID 37438390 · PMC10338499 · Scientific data · 2023 · 7 claims · 6 setups
TICCom integrates 739 experimentally-validated or manually-curated TIC interactions collected from more than 3,000 literatures
-
Has reproduction · 77
spotter: a single-nucleotide resolution stochastic simulation model of supercoiling-mediated transcription and translation in prokaryotes.
PMID 37602419 · PMC10516669 · Nucleic acids research · 2023 · 8 claims · 4 setups
spotter is the first simulation model to integrate transcription, DNA supercoiling, and translation simultaneously in a single stochastic framework for prokaryotes.