Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
Trimmomatic: a decade of feature-rich, high-performance NGS read preprocessing.
PMID 42178219 · PMC13242794 · Bioinformatics (Oxford, England) · 2026 · 8 claims · 4 setups
A high-performance multithreading architecture allows batches of read pairs to be processed independently by a pool of worker threads, scaling efficiently with available hardware.
-
Full-text index only
A map of human protein interactions derived from co-expression of human mRNAs and their orthologs.
PMID 18414481 · PMC2387231 · Molecular systems biology · 2008 · 8 claims · 6 setups
Comparing human mRNA co-expression with co-expression of orthologous gene pairs in five other organisms identifies proteins that physically associate
-
Has reproduction · 74
SpaGene: A Deep Adversarial Framework for Spatial Gene Imputation.
PMID 42146899 · PMC13176606 · Computational and structural biotechnology journal · 2026 · 8 claims · 6 setups
SpaGene improves average PCC and SSIM and reduces RMSE compared to 6 baseline methods (SpaGE, gimVI, Tangram, VISTA, spRefine, stDiff) across 8 diverse ST-SC dataset pairs under gene-holdout evaluation.
-
Full-text index only
SNAP: predict effect of non-synonymous polymorphisms on function.
PMID 17526529 · PMC1920242 · Nucleic acids research · 2007 · 7 claims · 8 setups
SNAP, a neural network-based method using sequence-derived information, predicts whether a non-synonymous SNP is neutral or non-neutral for protein function
-
Has reproduction · 88
Comprehensive benchmarking of large language models for RNA secondary structure prediction.
PMID 40205851 · PMC11982019 · Briefings in bioinformatics · 2025 · 7 claims · 4 setups
Existing RNA-LLMs had not previously been evaluated for secondary structure prediction in a unified, fair experimental setup with the same datasets and prediction model.
-
Full-text index only
Accurate splice site prediction using support vector machines.
PMID 18269701 · PMC2230508 · BMC bioinformatics · 2007 · 8 claims · 5 setups
Weighted degree (WD) kernel SVMs outperform Markov Chains, GeneSplicer and SpliceMachine for genome-wide splice site recognition
-
Full-text index only
Improved mutation tagging with gene identifiers applied to membrane protein stability prediction.
PMID 19758467 · PMC2745585 · BMC bioinformatics · 2009 · 8 claims · 4 setups
MutationTagger achieves 87% F-measure for the mutation retrieval task on a benchmark dataset
-
Full-text index only
Robust and efficient annotation of cell states through gene signature scoring.
PMID 41708334 · PMC12951948 · Genome research · 2026 · 8 claims · 8 setups
Established scoring methods (Seurat, SCANPY, UCell, JASMINE) fail to provide robust and comparable score distributions across diverse signatures and experimental conditions, precluding accurate unsupervised cell-state annotation.
-
Full-text index only
MatchMiner: a tool for batch navigation among gene and gene product identifiers.
PMID 12702208 · PMC154578 · Genome biology · 2003 · 8 claims · 3 setups
MatchMiner's LookUp function automates batch translation of an input list of gene identifiers into a matching list of a different identifier type.
-
Full-text index only
Predicting the phenotypic effects of non-synonymous single nucleotide polymorphisms based on support vector machines.
PMID 18005451 · PMC2216041 · BMC bioinformatics · 2007 · 8 claims · 5 setups
Parepro, an SVM-based method integrating three attribute sets (RD, MI, IE) derived from evolutionary and residue-property information, predicts whether an nsSNP is deleterious or neutral.
-
Full-text index only
A space-efficient and accurate method for mapping and aligning cDNA sequences onto genomic sequence.
PMID 18344523 · PMC2377433 · Nucleic acids research · 2008 · 7 claims · 6 setups
Spaln maps and aligns large cDNA sequence sets onto whole mammalian genomes using substantially less memory than comparable existing tools
-
Full-text index only
EpiXFormer: a cross-attention neural network for predicting cell type-specific transcription factor binding sites.
PMID 41527854 · PMC12796812 · Briefings in bioinformatics · 2026 · 8 claims · 8 setups
EpiXFormer achieves high accuracy (mean AUROC ~0.99) predicting binding sites of both TFs and non-sequence-specific DBPs across 199 DBP-cell type pairs
-
Has reproduction · 49
EDGE COVID-19: a web platform to generate submission-ready genomes from SARS-CoV-2 sequencing efforts.
PMID 35561186 · PMC9113274 · Bioinformatics (Oxford, England) · 2022 · 7 claims · 5 setups
EDGE COVID-19 (EC-19) is a web-based platform that automates QC, reference-based variant/consensus calling, lineage determination, and submission of SARS-CoV-2 genomes and metadata to GenBank, GISAID and INSDC for both Illumina and ONT data.
-
Full-text index only
CLAMP: predicting specific protein-mediated chromatin loops in diverse species with a chromatin accessibility language model.
PMID 41555433 · PMC12903630 · Genome biology · 2026 · 8 claims · 8 setups
CLAMP, a chromatin-accessibility language model, predicts protein-mediated chromatin loops across 10 species, 18 proteins, and 24 cell types with superior performance versus existing methods.
-
Has reproduction · 61
TEMP: a computational method for analyzing transposable element polymorphism in populations.
PMID 24753423 · PMC4066757 · Nucleic acids research · 2014 · 8 claims · 8 setups
TEMP combines pair-end (discordant) read and split (soft-clipped) read information to identify both presence and absence of TE insertions in genomic DNA from heterogeneous/pooled samples.
-
Full-text index only
CellPredX, a computational framework for cross-data type, cross-sample, and cross-protocol cell type annotation through domain adaptation and deep metric learning.
PMID 41481570 · PMC12758788 · PLoS computational biology · 2026 · 8 claims · 7 setups
CellPredX is a unified semi-supervised framework integrating domain adaptation and deep metric learning to align heterogeneous embeddings for cross-modality cell type annotation.
-
Has reproduction · 70
Predicting enhancers in mammalian genomes using supervised hidden Markov models.
PMID 30917778 · PMC6437899 · BMC bioinformatics · 2019 · 8 claims · 8 setups
eHMM predicts enhancers with high precision and recall comparable to state-of-the-art methods and consistently outperforms them in accuracy and resolution
-
Full-text index only
miRGen: a database for the study of animal microRNA genomic organization and function.
PMID 17108354 · PMC1669779 · Nucleic acids research · 2007 · 8 claims · 6 setups
miRGen is an integrated database combining Genomics, Targets, and Clusters interfaces to study miRNA genomic organization and function across 11 animal genomes
-
Full-text index only
Analysis of protein sequence and interaction data for candidate disease gene prediction.
PMID 17020920 · PMC1636487 · Nucleic acids research · 2006 · 8 claims · 7 setups
Combining CPS and CMP using known disease genes as input achieves sensitivity 0.52 and specificity 0.97, reducing candidate lists 13-fold
-
Full-text index only
StrainMake: reproducible hybrid metagenomics with MAG recovery and strain-level resolution.
PMID 42097292 · PMC13188985 · Bioinformatics (Oxford, England) · 2026 · 8 claims · 5 setups
StrainMake is a Snakemake-based, Conda-managed workflow for de novo metagenomic analysis from short, long, or hybrid sequencing data.