Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
Optimal step length EM algorithm (OSLEM) for the estimation of haplotype frequency and its application in lipoprotein lipase genotyping.
PMID 12529185 · PMC149347 · BMC bioinformatics · 2003 · 5 claims · 4 setups
OSLEM (Optimal Step Length EM), which approximates an optimal step length via a fixed-point search (D_N = D_{N-1} + λ(D_preN - D_{N-1})), runs about twice as fast as standard EM while producing the same haplotype frequency estimates.
-
Full-text index only
SNP-RFLPing: restriction enzyme mining for SNPs in genomes.
PMID 16503968 · PMC1386656 · BMC genomics · 2006 · 8 claims · 2 setups
SNP-RFLPing accepts three flexible input types (dbSNP rs#/ss# IDs, HUGO gene name/Entrez gene ID, or free-form SNP-in-sequence including IUPAC or [dNTP1/dNTP2] formats) for human, rat, and mouse genomes
-
Full-text index only
htSNPer1.0: software for haplotype block partition and htSNPs selection.
PMID 15740612 · PMC1274247 · BMC bioinformatics · 2005 · 6 claims · 1 setups
The GBB algorithm finds the globally optimal minimal htSNP set with far less computing time than exhaustive/enumeration search.
-
Full-text index only
Evolutionary algorithms for the selection of single nucleotide polymorphisms.
PMID 12875658 · PMC183839 · BMC bioinformatics · 2003 · 8 claims · 3 setups
Evolutionary algorithms are well suited to multiobjective optimization problems with large, intractable search spaces such as SNP selection, unlike exact methods (exhaustive enumeration) or single-objective search techniques (tabu search, simulated annealing).
-
Full-text index only
PPC: an algorithm for accurate estimation of SNP allele frequencies in small equimolar pools of DNA using data from high density microarrays.
PMID 16199750 · PMC1240117 · Nucleic acids research · 2005 · 7 claims · 6 setups
The PPC algorithm, which applies a probe-pair-specific second-degree polynomial correction, increases the accuracy of allele frequency estimates from pooled DNA compared with previously described algorithms
-
Full-text index only
InParanoid 7: new algorithms and tools for eukaryotic orthology analysis.
PMID 19892828 · PMC2808972 · Nucleic acids research · 2010 · 8 claims · 7 setups
InParanoid 7 expands the database by an order of magnitude to 100 species, 1.3 million proteins, and 42.7 million pairwise ortholog groups.
-
Full-text index only
ASPIC: a web resource for alternative splicing prediction and transcript isoforms characterization.
PMID 16845044 · PMC1538898 · Nucleic acids research · 2006 · 8 claims · 2 setups
The ASPIC algorithm, using an optimization procedure that minimizes splice site predictions and transcript isoforms from multiple EST-genome alignments, outperforms other similar AS-prediction tools in sensitivity and selectivity
-
Has reproduction · 85
Digital sorting of complex tissues for cell type-specific gene expression profiles.
PMID 23497278 · PMC3626856 · BMC bioinformatics · 2013 · 8 claims · 8 setups
The Digital Sorting Algorithm (DSA) deconvolves mixed tissue expression into cell type-specific profiles using only marker genes, without requiring prior knowledge of cell type frequencies or in vitro pure-cell profiles.
-
Full-text index only
Performance assessment of promoter predictions on ENCODE regions in the EGASP experiment.
PMID 16925837 · PMC1810552 · Genome biology · 2006 · 6 claims · 3 setups
Promoter predictors that combine promoter prediction with gene prediction (N-SCAN, Fprom) achieve better performance than pure ab initio promoter predictors, mainly by reducing the promoter search space and false positives
-
Has reproduction · 80
SLDMS: A Tool for Calculating the Overlapping Regions of Sequences.
PMID 35046988 · PMC8761809 · Frontiers in plant science · 2021 · 8 claims · 5 setups
SLDMS is a novel method for computing overlapping regions of sequencing reads using suffix array (SA), longest common prefix (LCP) array, document array (DA), and a monotonic stack.
-
Full-text index only
ADaCGH: A parallelized web-based application and R package for the analysis of aCGH data.
PMID 17710137 · PMC1940324 · PloS one · 2007 · 8 claims · 4 setups
ADaCGH implements eight CNA detection methods, including the best-performing ones from recent reviews (CBS, GLAD, CGHseg, HMM)
-
Has reproduction · 76
Tracing human genetic histories and natural selection with precise local ancestry inference.
PMID 40379651 · PMC12084304 · Nature communications · 2025 · 7 claims · 7 setups
Orchestra, a two-stage LAI method combining a recombination-distance base layer with a deep learning (convolutional + attention) smoothing module, outperforms RFmix, FLARE and Gnomix in precision and recall across simulated admixture generations.
-
Full-text index only
MAZIE: a mass and charge inference engine to enhance database searching of tandem mass spectra.
PMID 19850495 · PMC2818324 · Journal of the American Society for Mass Spectrometry · 2010 · 7 claims · 4 setups
MAZIE is a post-acquisition Perl algorithm that determines precursor ion monoisotopic mass and charge (+1 to +4) from MS1 zoom scan isotopic distributions on a Thermo LTQ-XL
-
Full-text index only
Computational tradeoffs in multiplex PCR assay design for SNP genotyping.
PMID 16042802 · PMC1190169 · BMC genomics · 2005 · 7 claims · 6 setups
Achieving high-multiplexing/high-coverage multiplex PCR designs is subject to a computational phase transition as the SNP-pair compatibility probability crosses a critical threshold
-
Full-text index only
SeqBuster, a bioinformatic tool for the processing and analysis of small RNAs datasets, reveals ubiquitous miRNA modifications in human embryonic cells.
PMID 20008100 · PMC2836562 · Nucleic acids research · 2010 · 8 claims · 6 setups
SeqBuster is a versatile web-based and stand-alone bioinformatic toolkit for processing and analyzing large-scale small RNA deep sequencing datasets.
-
Has reproduction · 63
Community assessment of methods to deconvolve cellular composition from bulk gene expression.
PMID 39191725 · PMC11350143 · Nature communications · 2024 · 8 claims · 4 setups
Most deconvolution methods accurately predict coarse-grained immune/stromal cell populations from bulk expression.
-
Has reproduction · 71
Protein structure quality assessment based on the distance profiles of consecutive backbone Cα atoms.
PMID 24555103 · PMC3892923 · F1000Research · 2013 · 8 claims · 8 setups
The distance between consecutive backbone Cα atoms in high-quality structures is normally distributed with mean 3.8 Å and standard deviation 0.04 Å, justifying a reference state in which all consecutive Cα atoms are 3.8 Å apart.
-
Full-text index only
Evolutionary trace annotation of protein function in the structural proteome.
PMID 20036248 · PMC2831211 · Journal of molecular biology · 2010 · 8 claims · 7 setups
ET-ranked residue clusters can be used to build 3D templates that predict GO function in enzymes and non-enzymes alike, without prior knowledge of functional mechanism.
-
Full-text index only
CpG_MI: a novel approach for identifying functional CpG islands in mammalian genomes.
PMID 19854943 · PMC2800233 · Nucleic acids research · 2010 · 8 claims · 6 setups
Functional ('bona fide') CGIs show distinct average/cumulative mutual information (AMI/CMI) distributions of neighboring CpG distances compared to non-functional CGIs and random genome segments
-
Has reproduction · 78
QuasiFlow: a Nextflow pipeline for analysis of NGS-based HIV-1 drug resistance data.
PMID 36699347 · PMC9722223 · Bioinformatics advances · 2022 · 6 claims · 8 setups
QuasiFlow is a Nextflow pipeline that runs entirely locally via command-line tools and a local HIVdb database copy to analyze NGS-based HIV-1 drug resistance testing data.