Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
SNPdetector: a software tool for sensitive and accurate SNP detection.
PMID 16261194 · PMC1274293 · PLoS computational biology · 2005 · 7 claims · 7 setups
SNPdetector, which models human visual inspection of sequencing traces, achieves low false positive and false negative rates in automated SNP and mutation detection
-
Full-text index only
Completing the map of human genetic variation.
PMID 17495918 · PMC2685471 · Nature · 2007 · 8 claims · 5 setups
A community resource initiative will sequence fosmid and BAC clone libraries from 62 HapMap individuals to systematically discover and resolve structural genetic variants at nucleotide resolution
-
Full-text index only
QuantiSNP: an Objective Bayes Hidden-Markov Model to detect and accurately map copy number variation using SNP genotyping data.
PMID 17341461 · PMC1874617 · Nucleic acids research · 2007 · 8 claims · 7 setups
QuantiSNP (OB-HMM) provides probabilistic quantification of copy number states and significantly improves accuracy of segmental aneuploidy identification and breakpoint mapping relative to existing tools (BeadStudio/Illumina)
-
Full-text index only
VarDetect: a nucleotide sequence variation exploratory tool.
PMID 19091032 · PMC2638149 · BMC bioinformatics · 2008 · 8 claims · 2 setups
VarDetect is a stand-alone software tool that automatically detects nucleotide variation (SNPs) from fluorescence-based chromatogram traces using pre-calculated peak content ratios and artifact-handling rules.
-
Full-text index only
A high throughput method for genome-wide analysis of retroviral integration.
PMID 17028098 · PMC1636494 · Nucleic acids research · 2006 · 8 claims · 8 setups
VITA uses MmeI to cleave DNA at a fixed distance from its recognition site, generating 21-22 bp genomic tags that serve as signatures of lentiviral integration sites.
-
Full-text index only
Capturing genomic signatures of DNA sequence variation using a standard anonymous microarray platform.
PMID 17000641 · PMC1636412 · Nucleic acids research · 2006 · 8 claims · 6 setups
An anonymous SHyP oligonucleotide microarray can capture genomic signatures of DNA sequence variation from any organism, including a previously unsequenced species
-
Full-text index only
Microdroplet-based PCR enrichment for large-scale targeted sequencing.
PMID 19881494 · PMC2779736 · Nature biotechnology · 2009 · 7 claims · 5 setups
Microdroplet PCR enables massively parallel singleplex amplification (up to ~1.5 million reactions, up to 4,000 targets) for targeted sequencing enrichment
-
Full-text index only
Speeding disease gene discovery by sequence based candidate prioritization.
PMID 15766383 · PMC1274252 · BMC bioinformatics · 2005 · 7 claims · 8 setups
Disease genes (OMIM) differ significantly from non-disease genes in sequence-based features including gene/cDNA/protein size, exon number, homolog conservation, secretion signal, 3' UTR length, CpG islands, and distance to nearest gene.
-
Full-text index only
A clustering property of highly-degenerate transcription factor binding sites in the mammalian genome.
PMID 16670430 · PMC1456330 · Nucleic acids research · 2006 · 8 claims · 7 setups
Highly-degenerate RE1 sites are significantly enriched in promoters of validated and putative REST target genes compared to control promoters
-
Has reproduction · 90
CONSULT: accurate contamination removal using locality-sensitive hashing.
PMID 34377979 · PMC8340999 · NAR genomics and bioinformatics · 2021 · 8 claims · 4 setups
CONSULT is a k-mer read-matching tool that uses locality-sensitive hashing (LSH) to allow inexact k-mer matches (within a user-defined Hamming distance) between query reads and a reference dataset.
-
Full-text index only
CARAT: a novel method for allelic detection of DNA copy number changes using high density oligonucleotide arrays.
PMID 16504045 · PMC1402331 · BMC bioinformatics · 2006 · 8 claims · 5 setups
CARAT is a novel algorithm that uses SNP probe intensity and genotype-based allelic dosage response in a regression framework to estimate allele-specific copy number genome-wide.
-
Full-text index only
Human and mouse introns are linked to the same processes and functions through each genome's most frequent non-conserved motifs.
PMID 18450818 · PMC2425492 · Nucleic acids research · 2008 · 8 claims · 5 setups
Pyknons (recurrent, genome-specific, ≥16nt motifs with ≥30 intact intergenic/intronic copies and ≥1 exonic copy) span a substantial fraction of previously uncharacterized intronic space (7.4% human, 4.4% mouse)
-
Full-text index only
Single-molecule sequencing of an individual human genome.
PMID 19668243 · PMC4117198 · Nature biotechnology · 2009 · 8 claims · 7 setups
Single-molecule sequencing without cloning, amplification or ligation can sequence an individual human genome on one instrument by a single operator in four runs
-
Has reproduction · 62
Application of alternative de novo motif recognition models for analysis of structural heterogeneity of transcription factor binding sites: a case study of FOXA2 binding sites.
PMID 34547062 · PMC8408018 · Vavilovskii zhurnal genetiki i selektsii · 2021 · 8 claims · 4 setups
MultiDeNA pipeline combines PWM, diPWM, BaMM and InMoDe models to train, evaluate, threshold, and classify ChIP-seq peaks for TFBS structural heterogeneity
-
Has reproduction · 89
TrEMOLO: accurate transposable element allele frequency estimation using long-read sequencing data combining assembly and mapping-based approaches.
PMID 37013657 · PMC10069131 · Genome biology · 2023 · 6 claims · 6 setups
TrEMOLO combines an assembly-based INSIDER module and a mapping-based OUTSIDER module to detect TE insertions/deletions from long-read sequencing data and estimate their allele frequency
-
Full-text index only
Characterizing natural variation using next-generation sequencing technologies.
PMID 19801172 · PMC3994700 · Trends in genetics : TIG · 2009 · 8 claims · 8 setups
Next-generation sequencing enables complete, genome-wide surveys of genetic variation at unprecedented resolution, overcoming limitations of genotyping panels and microarrays.
-
Full-text index only
Benchmarking of methods to analyse data derived from GBS-MeDIP.
PMID 41555215 · PMC12829230 · BMC bioinformatics · 2026 · 7 claims · 4 setups
featureCounts is the most reliable tool for count matrix generation from GBS-MeDIP data, outperforming MEDIPS
-
Has reproduction · 78
Single duplex DNA sequencing with CODEC detects mutations with high sensitivity.
PMID 37106072 · PMC10181940 · Nature genetics · 2023 · 8 claims · 8 setups
CODEC concatenates both strands of an original DNA duplex into a single NGS read pair via an adapter quadruplex and strand-displacing extension, enabling single-duplex resolution
-
Full-text index only
Combined subtractive cDNA cloning and array CGH: an efficient approach for identification of overexpressed genes in DNA amplicons.
PMID 15018647 · PMC365025 · BMC genomics · 2004 · 8 claims · 8 setups
Combined SSH subtractive cloning and array CGH is an efficient strategy to identify overexpressed genes located within DNA amplicons.
-
Full-text index only
Comparative genomics search for losses of long-established genes on the human lineage.
PMID 18085818 · PMC2134963 · PLoS computational biology · 2007 · 8 claims · 6 setups
A novel comparative genomics method (TransMap-based syntenic mapping of gene structures between human, mouse, and dog) can detect losses of well-established single-copy genes without relying on sequence homology to a parental gene, distinguishing them from typical duplication- or retrotransposition-derived pseudogenes.