Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
Uncovering Cas9 PAM diversity through metagenomic mining and machine learning.
PMID 41656299 · PMC12996302 · Nature communications · 2026 · 8 claims · 6 setups
CRISPR-PAMdb is a publicly accessible database compiling Cas9 protein sequences from 3.8 million bacterial/archaeal genomes and PAM profiles from 7.4 million phage/plasmid sequences
-
Has reproduction · 83
Accurate prediction of metagenome-assembled genome completeness by MAGISTA, a random forest model built on alignment-free intra-bin statistics.
PMID 35248155 · PMC8898458 · Environmental microbiome · 2022 · 7 claims · 7 setups
MAGISTA, a random forest model built on alignment-free intra-bin distance-distribution statistics, can estimate MAG completeness and purity without relying on reference marker genes.
-
Has reproduction · 86
LMAS: evaluating metagenomic short de novo assembly methods through defined communities.
PMID 36576131 · PMC9795473 · GigaScience · 2022 · 8 claims · 5 setups
LMAS (Last Metagenomic Assembler Standing) is a flexible, Nextflow-based, Docker-containerized automated workflow for benchmarking de novo metagenomic assemblers against defined mock communities, producing an interactive HTML report.
-
Has reproduction · 83
MetaGT: A pipeline for de novo assembly of metatranscriptomes with the aid of metagenomic data.
PMID 36386613 · PMC9651917 · Frontiers in microbiology · 2022 · 7 claims · 4 setups
MetaGT is a pipeline that combines metatranscriptomic and metagenomic data from the same sample to assemble complete transcript sequences
-
Full-text index only
BaGPipe: an automated, reproducible, and flexible pipeline for bacterial genome-wide association studies.
PMID 41896736 · PMC13147680 · BMC microbiology · 2026 · 7 claims · 8 setups
BaGPipe is an automated, reproducible Nextflow pipeline that integrates pre-processing, Pyseer-based association analysis, and downstream visualisation into a unified bacterial GWAS workflow
-
Full-text index only
NCBI Reference Sequences: current status, policy and new initiatives.
PMID 18927115 · PMC2686572 · Nucleic acids research · 2009 · 7 claims · 5 setups
RefSeq is a curated, non-redundant, explicitly linked database of nucleotide and protein sequences spanning genomes, transcripts and proteins across prokaryotes, eukaryotes and viruses
-
Has reproduction · 99
getSequenceInfo: a suite of tools allowing to get genome sequence information from public repositories.
PMID 35804320 · PMC9264741 · BMC bioinformatics · 2022 · 8 claims · 8 setups
getSequenceInfo (gSeqI) allows programmatic (CLI) or GUI-based retrieval of sequence data and metadata from GenBank, RefSeq, and ENA across Linux, MacOS, and Windows.
-
Has reproduction · 88
Wochenende - modular and flexible alignment-based shotgun metagenome analysis.
PMID 36368923 · PMC9650795 · BMC genomics · 2022 · 8 claims · 6 setups
Wochenende is a modular, transparent alignment-based pipeline for shotgun metagenome analysis supporting short and long reads across all kingdoms of life
-
Full-text index only
The post-genomic era for a select few.
PMID 14759254 · PMC395745 · Genome biology · 2004 · 8 claims · 8 setups
The Exofish comparative-genomics tool identifies protein-coding DNA segments by comparing two genome sequences and was used to compare pufferfish (Takifugu, Tetraodon) genomes with mammalian genomes, improving annotation of the human and mouse genomes.
-
Has reproduction · 95
MetaMap: an atlas of metatranscriptomic reads in human disease-related RNA-seq data.
PMID 29901703 · PMC6025204 · GigaScience · 2018 · 8 claims · 7 setups
The MetaMap pipeline recapitulates known infection agents in bona fide dual RNA-seq validation studies (Salmonella, HPV, HSV, rhinovirus)
-
Full-text index only
MODBASE, a database of annotated comparative protein structure models and associated resources.
PMID 18948282 · PMC2686492 · Nucleic acids research · 2009 · 8 claims · 8 setups
MODBASE contains 5,152,695 reliable comparative protein structure models for 1,593,209 unique protein sequences.
-
Has reproduction · 78
Determining the quality and complexity of next-generation sequencing data without a reference genome.
PMID 25514851 · PMC4298064 · Genome biology · 2014 · 8 claims · 8 setups
kPAL, an open-source alignment-free package, assesses sequencing data quality and complexity using k-mer frequency profiles and pairwise distances between them, without a reference sequence.
-
Full-text index only
Diversity of tRNA genes in eukaryotes.
PMID 17088292 · PMC1693877 · Nucleic acids research · 2006 · 8 claims · 6 setups
The number of tRNA genes having the same anticodon but different sequences elsewhere (isodecoder genes) varies significantly (10–246) across 11 eukaryotes despite isoacceptor numbers being similar (41–55)
-
Full-text index only
Genomewide pattern of synonymous nucleotide substitution in two complete genomes of Mycobacterium tuberculosis.
PMID 12453367 · PMC2738538 · Emerging infectious diseases · 2002 · 8 claims · 6 setups
Genomewide comparison of two complete M. tuberculosis genomes reveals substantially more nucleotide diversity than prior studies based on few loci suggested
-
Full-text index only
Metagenomic study of the oral microbiota by Illumina high-throughput sequencing.
PMID 19796657 · PMC3568755 · Journal of microbiological methods · 2009 · 8 claims · 6 setups
The 16S rRNA V5 hypervariable region, amplified as a short ~82-base segment, provides reliable taxonomic identification of oral bacteria against public databases like HOMD.
-
Has reproduction · 100
Analysis of the Taxonomy, Synteny, and Virulence Factors for Soft Rot Pathogen Pectobacterium aroidearum in Amorphophallus konjac Using Comparative Genomics.
PMID 35910650 · PMC9326479 · Frontiers in microbiology · 2022 · 8 claims · 8 setups
The causal agent of konjac soft rot in China is Pectobacterium aroidearum, confirmed via in vitro/in vivo pathogenicity tests, ANI, dDDH, and phylogenomic analysis.
-
Has reproduction · 62
Metatranscriptomics of the human oral microbiome during health and disease.
PMID 24692635 · PMC3977359 · mBio · 2014 · 8 claims · 8 setups
Disease-associated periodontal communities display conserved community-level metabolic gene expression profiles between patients, whereas the metabolic gene expression of individual species is highly variable between patients.
-
Full-text index only
The sequence and de novo assembly of the giant panda genome.
PMID 20010809 · PMC3951497 · Nature · 2010 · 8 claims · 8 setups
A draft giant panda genome was successfully generated and assembled de novo using only Illumina Genome Analyser short-read sequencing
-
Full-text index only
Dynamics of gut bacteriophage in diversity outbred mice studied over lifespan and during extreme caloric restriction.
PMID 41772715 · PMC12983593 · Microbiome · 2026 · 8 claims · 8 setups
Quiescent prophages dominate gut viral metagenomes, consistent with 'piggyback-the-winner' dynamics
-
Full-text index only
Pseudoalteromonas is a symbiont of marine invertebrates that exhibits broad patterns of phylosymbiosis.
PMID 42007585 · PMC13245728 · The ISME journal · 2026 · 8 claims · 8 setups
Pseudoalteromonas is a symbiont with substantial evidence of phylosymbiosis across at least three marine invertebrate phyla (Nematoda, Mollusca, and Cnidaria)