Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Has reproduction · 78
A case study for large-scale human microbiome analysis using JCVI's metagenomics reports (METAREP).
PMID 22719821 · PMC3374610 · PloS one · 2012 · 8 claims · 7 setups
METAREP version 1.3.1 is an open-source, scalable tool for querying, browsing and comparing extremely large volumes of metagenomic annotations, with an extended data model, dynamic weighting, distributed searches and advanced clustering.
-
Has reproduction · 73
Vespucci: a system for building annotated databases of nascent transcripts.
PMID 24304890 · PMC3936758 · Nucleic acids research · 2014 · 8 claims · 7 setups
Existing ChIP-seq and RNA-seq analysis platforms (e.g. Cufflinks, peak callers) are unsuited to GRO-seq because they assume spliced/exonic reads, uniform density and paired-end data, and cannot identify transcriptional units de novo across the whole genome.
-
Full-text index only
The diploid genome sequence of an Asian individual.
PMID 18987735 · PMC2716080 · Nature · 2008 · 8 claims · 8 setups
First diploid genome sequence of an Asian (Han Chinese) individual generated using massively parallel Illumina sequencing
-
Full-text index only
High throughput sequencing and proteomics to identify immunogenic proteins of a new pathogen: the dirty genome approach.
PMID 20037647 · PMC2793016 · PloS one · 2009 · 7 claims · 7 setups
A dirty genome approach using unfinished, unclosed genome sequences combined with proteomics can rapidly identify immunogenic proteins useful for diagnostic tool development
-
Full-text index only
Cancer genome standards for long-read sequencing using cancer cell line mixtures.
PMID 41934171 · PMC13137868 · GigaScience · 2026 · 8 claims · 6 setups
Long-read variant calling tools achieve recall rates comparable to short-read gold standards
-
Has reproduction · 83
Macrel: antimicrobial peptide screening in genomes and metagenomes.
PMID 33384902 · PMC7751412 · PeerJ · 2020 · 8 claims · 8 setups
Macrel introduces a novel set of 22 peptide features (6 local, 16 global), including a new Free Energy Transition (FET) feature group, for AMP and hemolytic activity classification
-
Has reproduction · 73
Genomics of Environmental Salmonella: Engaging Students in the Microbiology and Bioinformatics of Foodborne Pathogens.
PMID 33967968 · PMC8100199 · Frontiers in microbiology · 2021 · 8 claims · 8 setups
An undergraduate CURE combining field sampling, wet-lab microbiology, and genomic bioinformatics can be used to isolate and characterize environmental S. enterica strains.
-
Has reproduction · 89
A near complete genome for goat genetic and genomic research.
PMID 34507524 · PMC8434745 · Genetics, selection, evolution : GSE · 2021 · 8 claims · 8 setups
Saanen_v1 is a near-complete de novo goat genome assembly generated from 117x PacBio and 118x Hi-C data, including the first goat Y chromosome scaffold
-
Has reproduction · 71
A crowdsourced set of curated structural variants for the human genome.
PMID 32559231 · PMC7329145 · PLoS computational biology · 2020 · 8 claims · 8 setups
1235 manually curated SVs were produced that can be used to evaluate SV callers or train machine learning models
-
Has reproduction · 76
Organelle Genomes and Transcriptomes of Nymphaea Reveal the Interplay between Intron Splicing and RNA Editing.
PMID 34576004 · PMC8466565 · International journal of molecular sciences · 2021 · 8 claims · 8 setups
Both cis- and trans-splicing group II introns in Nymphaea organelle genomes are spliced in random order, generating diverse co-existing intermediates rather than following a fixed splicing sequence.
-
Has reproduction · 44
Population differentiation and epidemic tracking of Bursaphelenchus xylophilus in China based on chromosome-level assembly and whole-genome sequencing data.
PMID 34839581 · PMC9300093 · Pest management science · 2022 · 6 claims · 8 setups
Generated the first chromosome-level genome assembly (AH1) of B. xylophilus using PacBio, Illumina, BioNano, and Hi-C data
-
Has reproduction · 78
annotate_my_genomes: an easy-to-use pipeline to improve genome annotation and uncover neglected genes by hybrid RNA sequencing.
PMID 36472574 · PMC9724561 · GigaScience · 2022 · 7 claims · 8 setups
annotate_my_genomes is an easy-to-use genome-guided pipeline that uses hybrid (PacBio+Illumina) assembled transcripts to distinguish coding genes from long non-coding RNAs and reconcile them with prior annotations.
-
Has reproduction · 99
Evaluation of taxonomic classification and profiling methods for long-read shotgun metagenomic sequencing datasets.
PMID 36513983 · PMC9749362 · BMC bioinformatics · 2022 · 8 claims · 7 setups
Long-read classifiers generally performed best among the 11 methods tested
-
Has reproduction · 100
Integrative transcriptome sequencing identifies trans-splicing events with important roles in human embryonic stem cell pluripotency.
PMID 24131564 · PMC3875859 · Genome research · 2014 · 8 claims · 8 setups
TSscan, a computational pipeline integrating long- and short-read transcriptome sequencing from multiple hESC lines, can detect trans-splicing while minimizing false positives from experimental artifacts and genetic rearrangements.
-
Full-text index only
The Chromosome-Scale Genome Assembly of the Redlip Blenny, Ophioblennius macclurei (Blenniidae).
PMID 41378738 · PMC12758960 · Genome biology and evolution · 2026 · 8 claims · 12 setups
A chromosome-scale genome assembly of O. macclurei was generated (529.6 Mb, scaffold N50 23.7 Mb, GC 43.49%) using ONT long reads, Illumina short reads, and Hi-C scaffolding.
-
Full-text index only
WeavePop: a bioinformatics workflow to explore and analyze genomic variants of eukaryotic populations.
PMID 41685638 · PMC13042275 · G3 (Bethesda, Md.) · 2026 · 8 claims · 7 setups
WeavePop is a novel Snakemake-based, reproducible, scalable workflow that performs reference-based read mapping, assembly, annotation, small variant calling/effect prediction, and CNV detection for eukaryotic haploid organisms
-
Full-text index only
The dynamic distribution of genetic tandem amplifications in a heteroresistant Escherichia coli population revealed by ultra-deep long read sequencing.
PMID 41760616 · PMC12953903 · Nature communications · 2026 · 8 claims · 6 setups
Ultra-deep Nanopore sequencing of I-SceI-linearized plasmid DNA can detect and quantify full tandem amplification arrays at single-molecule resolution down to frequencies of 10^-5
-
Full-text index only
Benchmarking methods for genome annotation using nanopore direct RNA in a non-model crop plant.
PMID 41800382 · PMC12967217 · Bioinformatics advances · 2026 · 6 claims · 8 setups
Annotation tools show substantial variation in isoform detection, structural completeness, splicing classification, and handling of 5' read truncation when applied to plant dRNA-seq data.
-
Full-text index only
Manual validation finds ultra-long-read sequencing best enables faithful, population-level structural variant calling in Drosophila melanogaster euchromatin with nanopore.
PMID 41806374 · PMC13148403 · G3 (Bethesda, Md.) · 2026 · 8 claims · 5 setups
Only ultra-long long-reads (N50 > 50 kb) are capable of accurately calling structural variants of any size in D. melanogaster euchromatin
-
Full-text index only
Rapid identification of microbial pathogens and antimicrobial resistance from bloodstream infections using long-read sequencing.
PMID 42274466 · PMC13256323 · Microbial genomics · 2026 · 8 claims · 8 setups
A novel ONT long-read sequencing laboratory and bioinformatic workflow rapidly identifies bacterial and fungal organisms and AMR determinants from positive blood cultures