Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
A map of human protein interactions derived from co-expression of human mRNAs and their orthologs.
PMID 18414481 · PMC2387231 · Molecular systems biology · 2008 · 8 claims · 6 setups
Comparing human mRNA co-expression with co-expression of orthologous gene pairs in five other organisms identifies proteins that physically associate
-
Has reproduction · 50
MEDUSA: A Pipeline for Sensitive Taxonomic Classification and Flexible Functional Annotation of Metagenomic Shotgun Sequences.
PMID 35330728 · PMC8940201 · Frontiers in genetics · 2022 · 7 claims · 6 setups
MEDUSA correctly identifies more species than MEGAN 6 CE, especially less abundant species.
-
Full-text index only
Improved mutation tagging with gene identifiers applied to membrane protein stability prediction.
PMID 19758467 · PMC2745585 · BMC bioinformatics · 2009 · 8 claims · 4 setups
MutationTagger achieves 87% F-measure for the mutation retrieval task on a benchmark dataset
-
Full-text index only
rMAP 2.0: a modular, reproducible, and scalable WDL-Cromwell-Docker workflow for genomic analysis of ESKAPEE pathogens.
PMID 41782684 · PMC12955837 · Bioinformatics advances · 2026 · 8 claims · 8 setups
rMAP 2.0 standardizes end-to-end bacterial WGS analysis (QC, trimming, assembly, annotation, AMR/virulence/mobile-element profiling, sequence typing, pangenome inference, phylogenetics) via containerized WDL/Cromwell execution
-
Full-text index only
CLAMP: predicting specific protein-mediated chromatin loops in diverse species with a chromatin accessibility language model.
PMID 41555433 · PMC12903630 · Genome biology · 2026 · 8 claims · 8 setups
CLAMP, a chromatin-accessibility language model, predicts protein-mediated chromatin loops across 10 species, 18 proteins, and 24 cell types with superior performance versus existing methods.
-
Full-text index only
Using multiple alignments to improve seeded local alignment algorithms.
PMID 16100379 · PMC1185574 · Nucleic acids research · 2005 · 8 claims · 2 setups
Using information implicit in a multiple alignment to dynamically build a spaced-seed index weighted toward promising regions increases sensitivity of local alignment search compared to indexing a sequence alone
-
Full-text index only
A space-efficient and accurate method for mapping and aligning cDNA sequences onto genomic sequence.
PMID 18344523 · PMC2377433 · Nucleic acids research · 2008 · 7 claims · 6 setups
Spaln maps and aligns large cDNA sequence sets onto whole mammalian genomes using substantially less memory than comparable existing tools
-
Full-text index only
miRGen: a database for the study of animal microRNA genomic organization and function.
PMID 17108354 · PMC1669779 · Nucleic acids research · 2007 · 8 claims · 6 setups
miRGen is an integrated database combining Genomics, Targets, and Clusters interfaces to study miRNA genomic organization and function across 11 animal genomes
-
Full-text index only
Large-scale estimation of bacterial and archaeal DNA prevalence in metagenomes reveals biome-specific patterns.
PMID 41854267 · PMC13098197 · mSystems · 2026 · 8 claims · 6 setups
SPF scalably and robustly estimates the fraction of bacterial and archaeal reads in a metagenome using detection of prokaryotic single-copy marker genes, without requiring eukaryotic or viral reference genomes
-
Full-text index only
StrainMake: reproducible hybrid metagenomics with MAG recovery and strain-level resolution.
PMID 42097292 · PMC13188985 · Bioinformatics (Oxford, England) · 2026 · 8 claims · 5 setups
StrainMake is a Snakemake-based, Conda-managed workflow for de novo metagenomic analysis from short, long, or hybrid sequencing data.
-
Has reproduction · 70
Predicting enhancers in mammalian genomes using supervised hidden Markov models.
PMID 30917778 · PMC6437899 · BMC bioinformatics · 2019 · 8 claims · 8 setups
eHMM predicts enhancers with high precision and recall comparable to state-of-the-art methods and consistently outperforms them in accuracy and resolution
-
Has reproduction · 42
CanCellCap: robust cancer cell capture across tissue types on single-cell RNA-seq data by multi-domain learning.
PMID 40739511 · PMC12312500 · BMC biology · 2025 · 8 claims · 7 setups
CanCellCap identifies cancer cells in scRNA-seq data across 13 tissue types, 23 cancer types, and 7 sequencing platforms with 0.977 average accuracy
-
Full-text index only
CONTRAST: a discriminative, phylogeny-free approach to multiple informant de novo gene prediction.
PMID 18096039 · PMC2246271 · Genome biology · 2007 · 8 claims · 5 setups
CONTRAST predicts exact coding region structures for 65% more human genes than the previous state-of-the-art de novo predictor (N-SCAN)
-
Full-text index only
Umi-pipeline-nf: a modular and scalable workflow for UMI-tagged nanopore amplicon analysis with real-time sequencing integration and GPU-acceleration.
PMID 41923360 · PMC13070649 · Bioinformatics (Oxford, England) · 2026 · 8 claims · 6 setups
umi-pipeline-nf is a portable, fully containerized, modular Nextflow DSL2 workflow that generates single-molecule consensus sequences from UMI-tagged nanopore amplicon data and scales linearly from single samples to large cohorts.
-
Full-text index only
Benchmarking LLM-based agents for single-cell omics analysis.
PMID 41742311 · PMC13064268 · Genome biology · 2026 · 8 claims · 8 setups
Introduces a comprehensive benchmarking evaluation system comprising an open-source agent platform, 18 evaluation metrics across four dimensions, and 50 real-world single-cell omics tasks