Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Has reproduction · 98
Massively parallel genomic perturbations with multi-target CRISPR interrogates Cas9 activity and DNA repair at endogenous sites.
PMID 36064968 · PMC9481459 · Nature cell biology · 2022 · 8 claims · 6 setups
Multi-target gRNAs (mgRNAs) can direct Cas9 to over a hundred well-mapped endogenous genomic sites simultaneously, enabling massively parallel, high-throughput interrogation of Cas9 activity via short-read sequencing
-
Full-text index only
Origin and diversification of the basic helix-loop-helix gene family in metazoans: insights from comparative genomics.
PMID 17335570 · PMC1828162 · BMC evolutionary biology · 2007 · 8 claims · 4 setups
An initial diversification of bHLHs occurred in the pre-Cambrian, prior to metazoan cladogenesis
-
Full-text index only
Cryptic loxP sites in mammalian genomes: genome-wide distribution and relevance for the efficiency of BAC/PAC recombineering techniques.
PMID 17284462 · PMC1865043 · Nucleic acids research · 2007 · 6 claims · 6 setups
Cryptic lox P sites occur frequently and are homogeneously distributed across the mouse genome (1.2 primary sites per megabase).
-
Full-text index only
A genomics-based approach to biodefence preparedness.
PMID 14708013 · PMC7097618 · Nature reviews. Genetics · 2004 · 7 claims · 8 setups
Genome sequence data are now available for essentially all 25-30 principal human bacterial pathogens, including most CDC category A-C bioterror agents
-
Full-text index only
Kangaroo--a pattern-matching program for biological sequences.
PMID 12150718 · PMC119856 · BMC bioinformatics · 2002 · 7 claims · 2 setups
Kangaroo is a web-based regular expression pattern-matching program that searches DNA, protein, or coding-region sequences across ten organisms with no restriction on query length or complexity.
-
Full-text index only
In vitro and in silico analysis reveals an efficient algorithm to predict the splicing consequences of mutations at the 5' splice sites.
PMID 17726045 · PMC2094079 · Nucleic acids research · 2007 · 8 claims · 6 setups
Two exonic mutations, PINK1 E417G and PARK7 E64D, disrupt binding to U1 snRNA and cause skipping of the mutation-harboring exon
-
Full-text index only
Transduplication resulted in the incorporation of two protein-coding sequences into the turmoil-1 transposable element of C. elegans.
PMID 18842128 · PMC2572040 · Biology direct · 2008 · 8 claims · 6 setups
The Turmoil-1 transposable element in C. elegans incorporated two unrelated protein-coding sequences into its inverted terminal repeats (ITRs)
-
Full-text index only
In silico whole-genome screening for cancer-related single-nucleotide polymorphisms located in human mRNA untranslated regions.
PMID 17201911 · PMC1774567 · BMC genomics · 2007 · 8 claims · 5 setups
A computational EST-based pipeline can identify UTR-SNPs that are statistically over-represented in cancerous versus normal tissue libraries
-
Full-text index only
Spontaneous symmetry breaking in genome evolution.
PMID 18367477 · PMC2377439 · Nucleic acids research · 2008 · 6 claims · 3 setups
Exon size distributions in sequenced genomes follow a lognormal pattern typical of a random Kolmogoroff fractioning process
-
Full-text index only
Design and analysis issues in genome-wide somatic mutation studies of cancer.
PMID 18692126 · PMC2820387 · Genomics · 2009 · 6 claims · 4 setups
Two-stage (discovery + validation) sequencing designs efficiently allocate resources and can produce highly informative candidate driver gene lists even with relatively small sample sizes.
-
Has reproduction · 93
Population genomics of the Wolbachia endosymbiont in Drosophila melanogaster.
PMID 23284297 · PMC3527207 · PLoS genetics · 2012 · 8 claims · 8 setups
Wolbachia infection status can be accurately predicted in silico from whole-genome shotgun sequence of individual host strains, showing 99% concordance with diagnostic PCR.
-
Has reproduction · 98
maxATAC: Genome-scale transcription-factor binding prediction from ATAC-seq with deep neural networks.
PMID 36719906 · PMC9917285 · PLoS computational biology · 2023 · 8 claims · 6 setups
maxATAC is a suite of deep neural network models enabling state-of-the-art, genome-scale TFBS prediction from ATAC-seq, with models for 127 human transcription factors
-
Has reproduction · 96
A bioinformatic pipeline for simulating viral integration data.
PMID 35496474 · PMC9046613 · Data in brief · 2022 · 7 claims · 3 setups
A snakemake-based pipeline was developed to simulate integration of a viral or vector genome into a host genome, including sub-genomic fragment integration, structural variation, and host-site deletions.
-
Full-text index only
MtSNPscore: a combined evidence approach for assessing cumulative impact of mitochondrial variations in disease.
PMID 19758471 · PMC2745589 · BMC bioinformatics · 2009 · 8 claims · 5 setups
MtSNPscore, a weighted scoring pipeline combining literature evidence, in silico predictions, and case/control frequency, can prioritize likely pathogenic mtDNA variations
-
Has reproduction · 86
Prediction of Antibiotic Susceptibility Profiles of Vibrio cholerae Isolates From Whole Genome Illumina and Nanopore Sequencing Data: CholerAegon.
PMID 35814690 · PMC9257098 · Frontiers in microbiology · 2022 · 6 claims · 6 setups
CholerAegon, a Nextflow-based pipeline, predicts AMR profiles of V. cholerae from assembled genomes using CARD ontology
-
Full-text index only
EPD in its twentieth year: towards complete promoter coverage of selected model organisms.
PMID 16381980 · PMC1347508 · Nucleic acids research · 2006 · 7 claims · 4 setups
EPD is an annotated, non-redundant collection of experimentally defined eukaryotic POL II promoters accessed via genome position pointers.
-
Full-text index only
Mycoplasma genitalium: an efficient strategy to generate genetic variation from a minimal genome.
PMID 17784912 · PMC2169797 · Molecular microbiology · 2007 · 8 claims · 6 setups
MG192 is highly variable among and within M. genitalium strains both in vitro and in vivo
-
Full-text index only
Enrichment of sequencing targets from the human genome by solution hybridization.
PMID 19835619 · PMC2784331 · Genome biology · 2009 · 8 claims · 5 setups
Solution hybridization with 120-mer capture probes efficiently enriches targeted genomic sequences for next-generation sequencing
-
Has reproduction · 76
Transcriptional landscape of repetitive elements in normal and cancer human cells.
PMID 25012247 · PMC4122776 · BMC genomics · 2014 · 8 claims · 8 setups
RepEnrich, a computational method that uses all mapping reads (uniquely mapping plus multi-mapping reads assigned to repetitive element subfamily assemblies/pseudogenomes), quantifies genome-wide repetitive element enrichment
-
Full-text index only
Comparative analysis of eccDNA and circRNA tools shows increased accuracy of tool combination.
PMID 41738836 · PMC13154841 · GigaScience · 2026 · 8 claims · 6 setups
Detection accuracy of eccDNA/circRNA tools is highly influenced by sequencing depth, alignment algorithm, and experimental enrichment protocol