Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Has reproduction · 90
CONSULT: accurate contamination removal using locality-sensitive hashing.
PMID 34377979 · PMC8340999 · NAR genomics and bioinformatics · 2021 · 8 claims · 4 setups
CONSULT is a k-mer read-matching tool that uses locality-sensitive hashing (LSH) to allow inexact k-mer matches (within a user-defined Hamming distance) between query reads and a reference dataset.
-
Full-text index only
Rapid identification of microbial pathogens and antimicrobial resistance from bloodstream infections using long-read sequencing.
PMID 42274466 · PMC13256323 · Microbial genomics · 2026 · 8 claims · 8 setups
A novel ONT long-read sequencing laboratory and bioinformatic workflow rapidly identifies bacterial and fungal organisms and AMR determinants from positive blood cultures
-
Full-text index only
Prediction of candidate primary immunodeficiency disease genes using a support vector machine learning approach.
PMID 19801557 · PMC2780952 · DNA research : an international journal for rapid publication of reports on genes and genomes · 2009 · 6 claims · 3 setups
An SVM trained on 69 binary features of known PID genes can accurately classify PID vs non-PID genes and predict novel candidate PID genes
-
Full-text index only
Four genomic islands that mark post-1995 pandemic Vibrio parahaemolyticus isolates.
PMID 16672049 · PMC1464126 · BMC genomics · 2006 · 8 claims · 7 setups
Seven genomic islands (VPaI-1 to VPaI-7, 10-81 kb) were identified in V. parahaemolyticus RIMD2210633 by aberrant GC content, presence of integrases/transposases, flanking direct repeats, and absence from related Vibrionaceae genomes.
-
Full-text index only
The UCSC Genome Browser Database: update 2009.
PMID 18996895 · PMC2686463 · Nucleic acids research · 2009 · 8 claims · 6 setups
The UCSC Genome Browser Database (GBD) is a publicly available, integrated collection of genome assembly sequences and annotations across many organisms, including extensive comparative-genomic resources.
-
Has reproduction
Methylation patterns of the nasal epigenome of hospitalized SARS-CoV-2 positive patients reveal insights into molecular mechanisms of COVID-19.
PMID 40170038 · PMC11963311 · BMC medical genomics · 2025 · 7 claims · 7 setups
Differential DNA methylation occurs predominantly in intergenic regions and low methylated regions (LMRs), highlighting the role of distal regulatory elements in COVID-19 severity.
-
Full-text index only
Geographically Distinct Circulation of Genotype II and III St. Louis Encephalitis Virus, Texas, USA, 2009-2024.
PMID 41986946 · PMC13094854 · Emerging infectious diseases · 2026 · 7 claims · 7 setups
Genotype II and genotype III SLEV circulated concurrently in Texas during 2009–2024 but were geographically segregated, with no county having both.
-
Has reproduction
Systematic analysis of CNGCs in cotton and the positive role of GhCNGC32 and GhCNGC35 in salt tolerance.
PMID 35931984 · PMC9356423 · BMC genomics · 2022 · 8 claims · 8 setups
114 CNGC genes were identified across the genomes of four cotton species (G. arboreum, G. raimondii, G. barbadense, G. hirsutum)
-
Full-text index only
Assessment of algorithms for high throughput detection of genomic copy number variation in oligonucleotide microarray data.
PMID 17910767 · PMC2148068 · BMC bioinformatics · 2007 · 8 claims · 4 setups
Different CNV analysis software packages produce highly variable numbers and types of candidate CNVs from the same data
-
Full-text index only
SNP@Evolution: a hierarchical database of positive selection on the human genome.
PMID 19732458 · PMC2755008 · BMC evolutionary biology · 2009 · 7 claims · 6 setups
SNP@Evolution is a hierarchical database integrating HET, FST, and iHS from HapMap Phase II and III to identify genome-wide positive selection signals
-
Has reproduction · 89
TrEMOLO: accurate transposable element allele frequency estimation using long-read sequencing data combining assembly and mapping-based approaches.
PMID 37013657 · PMC10069131 · Genome biology · 2023 · 6 claims · 6 setups
TrEMOLO combines an assembly-based INSIDER module and a mapping-based OUTSIDER module to detect TE insertions/deletions from long-read sequencing data and estimate their allele frequency
-
Full-text index only
Commonality of functional annotation: a method for prioritization of candidate genes from genome-wide linkage studies.
PMID 18263617 · PMC2275105 · Nucleic acids research · 2008 · 8 claims · 7 setups
Genes correlated with a common complex trait are more likely to share GO functional annotations than genes not correlated with that trait
-
Full-text index only
Characterizing natural variation using next-generation sequencing technologies.
PMID 19801172 · PMC3994700 · Trends in genetics : TIG · 2009 · 8 claims · 8 setups
Next-generation sequencing enables complete, genome-wide surveys of genetic variation at unprecedented resolution, overcoming limitations of genotyping panels and microarrays.
-
Full-text index only
Single-molecule sequencing of an individual human genome.
PMID 19668243 · PMC4117198 · Nature biotechnology · 2009 · 8 claims · 7 setups
Single-molecule sequencing without cloning, amplification or ligation can sequence an individual human genome on one instrument by a single operator in four runs
-
Full-text index only
Ensembl 2005.
PMID 15608235 · PMC540092 · Nucleic acids research · 2005 · 8 claims · 4 setups
Ensembl's automatic gene build system can flexibly and reliably annotate a wide variety of genomes with limited species-specific evidence.
-
Has reproduction · 79
RetroSnake: A modular pipeline to detect human endogenous retroviruses in genome sequencing data.
PMID 36339261 · PMC9626663 · iScience · 2022 · 8 claims · 4 setups
RetroSnake is an end-to-end, modular, computationally efficient Snakemake pipeline for detecting HERV-K insertions in short-read NGS data, from raw alignment files to an annotated interactive HTML report
-
Full-text index only
Phylogenomic approaches to common problems encountered in the analysis of low copy repeats: the sulfotransferase 1A gene family example.
PMID 15752422 · PMC555591 · BMC evolutionary biology · 2005 · 8 claims · 8 setups
A previously unidentified fourth human SULT1A gene (SULT1A4) exists on chromosome 16 and is transcriptionally active
-
Full-text index only
Comprehensive genome analysis of 203 genomes provides structural genomics with new insights into protein family space.
PMID 16481312 · PMC1373602 · Nucleic acids research · 2006 · 8 claims · 7 setups
The number of protein families continues to expand steadily as more genomes are sequenced, showing no sign of saturation.
-
Full-text index only
Ab initio identification of human microRNAs based on structure motifs.
PMID 18088431 · PMC2238772 · BMC bioinformatics · 2007 · 8 claims · 7 setups
MiRPred predicts miRNA precursors ab initio using only predicted secondary structure motifs, ignoring nucleotide sequence
-
Full-text index only
Design and analysis issues in genome-wide somatic mutation studies of cancer.
PMID 18692126 · PMC2820387 · Genomics · 2009 · 6 claims · 4 setups
Two-stage (discovery + validation) sequencing designs efficiently allocate resources and can produce highly informative candidate driver gene lists even with relatively small sample sizes.