Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
Ensembl 2006.
PMID 16381931 · PMC1347495 · Nucleic acids research · 2006 · 8 claims · 5 setups
Ensembl now provides annotation for 19 genomes, up from 4 the previous year, including new mammalian (Rhesus macaque, Opossum), chordate (Ciona intestinalis), and yeast genomes.
-
Full-text index only
PPC: an algorithm for accurate estimation of SNP allele frequencies in small equimolar pools of DNA using data from high density microarrays.
PMID 16199750 · PMC1240117 · Nucleic acids research · 2005 · 7 claims · 6 setups
The PPC algorithm, which applies a probe-pair-specific second-degree polynomial correction, increases the accuracy of allele frequency estimates from pooled DNA compared with previously described algorithms
-
Has reproduction · 92
Prognostic biomarker discovery in pancreatic cancer through hybrid ensemble feature selection and multi-omics data.
PMID 41957754 · PMC13188360 · BioData mining · 2026 · 7 claims · 3 setups
The hEFS framework integrates data subsampling with multiple prognostic models (embedded and wrapper-based), aggregates feature rankings via a voting-theory-inspired approach, and selects the optimal feature subset via Pareto front optimization, eliminating user-defined thresholds.
-
Full-text index only
InParanoid 7: new algorithms and tools for eukaryotic orthology analysis.
PMID 19892828 · PMC2808972 · Nucleic acids research · 2010 · 8 claims · 7 setups
InParanoid 7 expands the database by an order of magnitude to 100 species, 1.3 million proteins, and 42.7 million pairwise ortholog groups.
-
Has reproduction · 53
spliceJAC: transition genes and state-specific gene regulation from single-cell transcriptome data.
PMID 36321549 · PMC9627675 · Molecular systems biology · 2022 · 8 claims · 6 setups
spliceJAC quantifies multivariate mRNA splicing from unspliced/spliced count matrices to construct cell state-specific gene-gene (Jacobian) interaction matrices.
-
Has reproduction · 53
Combining evidence of preferential gene-tissue relationships from multiple sources.
PMID 23950964 · PMC3741196 · PloS one · 2013 · 8 claims · 8 setups
A high-level integration approach combining three methods across four human microarray datasets, merged by consensus voting and a rule-based inner/total score, predicts preferentially expressed genes while reducing method- and study-specific bias.
-
Full-text index only
MultiPhyl: a high-throughput phylogenomics webserver using distributed computing.
PMID 17553837 · PMC1933173 · Nucleic acids research · 2007 · 8 claims · 8 setups
MultiPhyl is the first high-throughput distributed phylogenetics platform capable of using idle computational resources of many heterogeneous non-dedicated machines to form a phylogenetics supercomputer
-
Full-text index only
Grammar-based distance in progressive multiple sequence alignment.
PMID 18616828 · PMC2478692 · BMC bioinformatics · 2008 · 7 claims · 3 setups
A grammar-based (LZ complexity) distance metric can be used to determine the order in which sequences are progressively pairwise aligned
-
Has reproduction · 76
Tracing human genetic histories and natural selection with precise local ancestry inference.
PMID 40379651 · PMC12084304 · Nature communications · 2025 · 7 claims · 7 setups
Orchestra, a two-stage LAI method combining a recombination-distance base layer with a deep learning (convolutional + attention) smoothing module, outperforms RFmix, FLARE and Gnomix in precision and recall across simulated admixture generations.
-
Full-text index only
Benchmarking tools for the alignment of functional noncoding DNA.
PMID 14736341 · PMC344529 · BMC bioinformatics · 2004 · 8 claims · 4 setups
Global alignment tools (Avid, ClustalW, Lagan, Needle, DiAlign-G) typically have higher sensitivity over entire noncoding sequences and within constrained blocks than local tools
-
Full-text index only
ADaCGH: A parallelized web-based application and R package for the analysis of aCGH data.
PMID 17710137 · PMC1940324 · PloS one · 2007 · 8 claims · 4 setups
ADaCGH implements eight CNA detection methods, including the best-performing ones from recent reviews (CBS, GLAD, CGHseg, HMM)
-
Has reproduction · 89
Spatial information matters: are traditional imputation methods effective for spatial transcriptomics data?
PMID 41627342 · PMC12862982 · Briefings in bioinformatics · 2026 · 7 claims · 3 setups
No single existing SOTA imputation method consistently performs well across newer SRT platforms/datasets
-
Full-text index only
SNP-RFLPing: restriction enzyme mining for SNPs in genomes.
PMID 16503968 · PMC1386656 · BMC genomics · 2006 · 8 claims · 2 setups
SNP-RFLPing accepts three flexible input types (dbSNP rs#/ss# IDs, HUGO gene name/Entrez gene ID, or free-form SNP-in-sequence including IUPAC or [dNTP1/dNTP2] formats) for human, rat, and mouse genomes
-
Has reproduction · 86
LMAS: evaluating metagenomic short de novo assembly methods through defined communities.
PMID 36576131 · PMC9795473 · GigaScience · 2022 · 8 claims · 5 setups
LMAS (Last Metagenomic Assembler Standing) is a flexible, Nextflow-based, Docker-containerized automated workflow for benchmarking de novo metagenomic assemblers against defined mock communities, producing an interactive HTML report.
-
Full-text index only
Optimal step length EM algorithm (OSLEM) for the estimation of haplotype frequency and its application in lipoprotein lipase genotyping.
PMID 12529185 · PMC149347 · BMC bioinformatics · 2003 · 5 claims · 4 setups
OSLEM (Optimal Step Length EM), which approximates an optimal step length via a fixed-point search (D_N = D_{N-1} + λ(D_preN - D_{N-1})), runs about twice as fast as standard EM while producing the same haplotype frequency estimates.
-
Full-text index only
GeneKeyDB: a lightweight, gene-centric, relational database to support data mining environments.
PMID 15790402 · PMC1274265 · BMC bioinformatics · 2005 · 8 claims · 6 setups
GeneKeyDB is a lightweight, gene-centric relational database that supports data mining and integration with computational analysis tools.
-
Full-text index only
Exogean: a framework for annotating protein-coding genes in eukaryotic genomic DNA.
PMID 16925841 · PMC1810556 · Genome biology · 2006 · 8 claims · 5 setups
Exogean is a framework using directed acyclic coloured multigraphs (DACMs) to represent biological objects (mRNA, ESTs, protein alignments, exons) and iteratively combine them into complex protein-coding transcript models.
-
Full-text index only
SNPmasker: automatic masking of SNPs and repeats across eukaryotic genomes.
PMID 16845091 · PMC1538889 · Nucleic acids research · 2006 · 8 claims · 4 setups
SNPmasker is a web service combining SNP masking and repeat masking, supporting both coordinate-defined and homology-search-defined input regions, a combination not offered by prior tools
-
Full-text index only
Performance assessment of promoter predictions on ENCODE regions in the EGASP experiment.
PMID 16925837 · PMC1810552 · Genome biology · 2006 · 6 claims · 3 setups
Promoter predictors that combine promoter prediction with gene prediction (N-SCAN, Fprom) achieve better performance than pure ab initio promoter predictors, mainly by reducing the promoter search space and false positives
-
Full-text index only
Genomics--from Neanderthals to high-throughput sequencing.
PMID 16934106 · PMC1779599 · Genome biology · 2006 · 8 claims · 8 setups
Next-generation sequencing platforms (GS20/454 and Solexa) can deliver the throughput and cost reductions needed for population-scale and medical resequencing.