Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Has reproduction · 100
poreCov-An Easy to Use, Fast, and Robust Workflow for SARS-CoV-2 Genome Reconstruction via Nanopore Sequencing.
PMID 34394197 · PMC8355734 · Frontiers in genetics · 2021 · 8 claims · 8 setups
poreCov is an easy-to-use, fast, and robust Nextflow-based workflow for reference-based SARS-CoV-2 genome reconstruction and lineage determination from nanopore sequencing data
-
Has reproduction · 55
Genome-Wide Survey and Development of the First Microsatellite Markers Database (AnCorDB) in Anemone coronaria L.
PMID 35328546 · PMC8949970 · International journal of molecular sciences · 2022 · 8 claims · 8 setups
Generated the first draft genome assembly of A. coronaria by Illumina sequencing a haploid androgenetic plant
-
Has reproduction · 50
BiRNA-BERT allows efficient RNA language modeling with adaptive tokenization.
PMID 41266599 · PMC12635123 · Communications biology · 2025 · 8 claims · 8 setups
BiRNA-BERT uses adaptive dual-tokenization that dynamically selects nucleotide-level (NUC) or byte-pair encoding (BPE) tokens based on input sequence length
-
Full-text index only
Computational comparison of two mouse draft genomes and the human golden path.
PMID 12537546 · PMC151282 · Genome biology · 2003 · 8 claims · 7 setups
The Celera and public mouse genome assemblies differ in about 10% of the mouse genome, with complementary strengths (Celera higher base-pair accuracy and overall coverage; public assembly higher quality in some finished BAC regions and freely accessible)
-
Full-text index only
Pegasys: software for executing and integrating analyses of biological sequences.
PMID 15096276 · PMC406494 · BMC bioinformatics · 2004 · 8 claims · 7 setups
Pegasys is a flexible, modular, customizable software system for executing and integrating heterogeneous biological sequence analysis tools
-
Full-text index only
HapMap-based study of the 17q21 ERBB2 amplicon in susceptibility to breast cancer.
PMID 17117180 · PMC2360759 · British journal of cancer · 2006 · 6 claims · 5 setups
Common genetic variation (tSNPs and haplotypes) across the 400-kb 17q21 ERBB2 amplicon is not associated with breast cancer risk in British women.
-
Full-text index only
PigGIS: Pig Genomic Informatics System.
PMID 17090590 · PMC1669765 · Nucleic acids research · 2007 · 7 claims · 7 setups
PigGIS identified 15,700 pig consensus sequences covering 18.5 Mb of homologous human exons
-
Full-text index only
SNPmasker: automatic masking of SNPs and repeats across eukaryotic genomes.
PMID 16845091 · PMC1538889 · Nucleic acids research · 2006 · 8 claims · 4 setups
SNPmasker is a web service combining SNP masking and repeat masking, supporting both coordinate-defined and homology-search-defined input regions, a combination not offered by prior tools
-
Full-text index only
The Chromosome-Scale Genome Assembly of the Redlip Blenny, Ophioblennius macclurei (Blenniidae).
PMID 41378738 · PMC12758960 · Genome biology and evolution · 2026 · 8 claims · 12 setups
A chromosome-scale genome assembly of O. macclurei was generated (529.6 Mb, scaffold N50 23.7 Mb, GC 43.49%) using ONT long reads, Illumina short reads, and Hi-C scaffolding.
-
Full-text index only
CholeraSeq: a comprehensive genomic pipeline for cholera surveillance and near real-time outbreak investigation.
PMID 41400832 · PMC12790814 · Bioinformatics (Oxford, England) · 2026 · 7 claims · 6 setups
CholeraSeq is an automated, V. cholerae-specific Nextflow pipeline that processes WGS outbreak data from raw reads/assemblies to high-quality SNPs and phylogenies in near-real-time.
-
Full-text index only
First genome assemblies of Neotropical Thoracobombus bumblebees Bombus pauloensis and Bombus pullatus.
PMID 41436027 · PMC12958814 · G3 (Bethesda, Md.) · 2026 · 7 claims · 8 setups
This study produced the first genome assemblies of Neotropical Bombus (Thoracobombus) species, B. pauloensis and B. pullatus
-
Full-text index only
Integration of a neuronal RNAseq dataset with the draft Gryllus bimaculatus transcriptome refines gene predictions and highlights potential systematic response to injury.
PMID 42054377 · PMC13127959 · PloS one · 2026 · 8 claims · 7 setups
Integrating prothoracic ganglion RNAseq data with the draft genome refines gene predictions, adding 3,868 novel genes and 9,172 new transcript isoforms (including non-coding transcripts)
-
Full-text index only
Diversity and evolution of a phase-variable multi-locus antigen in Neisseria gonorrhoeae.
PMID 42113870 · PMC13183285 · PLoS pathogens · 2026 · 8 claims · 8 setups
Each N. gonorrhoeae genome has on average 7 distinct opa alleles at 9-12 opa loci
-
Full-text index only
Evaluating the performance of ancient DNA genetic relatedness estimation methods using high-fidelity pedigree simulations.
PMID 41796349 · PMC13081257 · Genome biology · 2026 · 8 claims · 5 setups
BADGER, an automated snakemake pipeline, was developed to simulate high-fidelity pedigrees and raw ancient DNA sequence data for benchmarking genetic relatedness methods
-
Has reproduction · 52
epiGBS2: Improvements and evaluation of highly multiplexed, epiGBS-based reduced representation bisulfite sequencing.
PMID 35178872 · PMC9311447 · Molecular ecology resources · 2022 · 8 claims · 8 setups
epiGBS2 provides a laboratory protocol and revised bioinformatics pipeline for de novo cytosine methylation and SNP calling in species with or without a reference genome
-
Has reproduction · 92
Chromosome-scale genome sequencing, assembly and annotation of six genomes from subfamily Leishmaniinae.
PMID 34489462 · PMC8421402 · Scientific data · 2021 · 8 claims · 8 setups
Chromosome-scale genomes of six Leishmaniinae species (five L. (Mundinia) species and one Porcisia species) were sequenced, assembled and annotated, providing genome, proteome, transcriptome and GFF outputs for taxa previously lacking public reference genomes
-
Full-text index only
The Vertebrate Genome Annotation (Vega) database.
PMID 15608237 · PMC540089 · Nucleic acids research · 2005 · 8 claims · 8 setups
Vega is a community database for browsing manual annotation of finished vertebrate genome sequences, based on an extended Ensembl-style schema.
-
Full-text index only
Highly contiguous chromosome-level assembly of the rock goby (Gobius paganellus) genome.
PMID 41611730 · PMC12957465 · Scientific data · 2026 · 8 claims · 8 setups
Chromosome-level genome assembly of Gobius paganellus spans 813 Mb with >99.9% of sequence anchored to 23 pseudochromosomes
-
Has reproduction · 85
ScLRTC: imputation for single-cell RNA-seq data via low-rank tensor completion.
PMID 34844559 · PMC8628418 · BMC genomics · 2021 · 8 claims · 8 setups
scLRTC imputes dropout entries closest to the original expression values on simulated datasets, outperforming other state-of-the-art methods by SSE and PCC.
-
Full-text index only
A clustering property of highly-degenerate transcription factor binding sites in the mammalian genome.
PMID 16670430 · PMC1456330 · Nucleic acids research · 2006 · 8 claims · 7 setups
Highly-degenerate RE1 sites are significantly enriched in promoters of validated and putative REST target genes compared to control promoters