Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Has reproduction · 85
Optimizing open data to support one health: best practices to ensure interoperability of genomic data from bacterial pathogens.
PMID 33103064 · PMC7568946 · One health outlook · 2020 · 8 claims · 3 setups
An open-access pathogen surveillance database (NCBI Pathogen Detection) plus contributor Best Practices enables FAIR, interoperable genomic data across human, animal, food, and environmental sources for One Health surveillance.
-
Has reproduction · 87
Insights into the Evolution of the New World Diploid Cottons (Gossypium, Subgenus Houzingenia) Based on Genome Sequencing.
PMID 30476109 · PMC6320677 · Genome biology and evolution · 2019 · 8 claims · 8 setups
Subgenus Houzingenia originated via transoceanic dispersal from Africa ~6.6 Ma, with most biodiversity arising from rapid mid-Pleistocene (0.5–2.0 Ma) diversification plus multiple long-distance dispersals.
-
Full-text index only
Evaluating the performance of commercial whole-genome marker sets for capturing common genetic variation.
PMID 17562002 · PMC1914356 · BMC genomics · 2007 · 8 claims · 5 setups
Commercial SNP panels provide levels of coverage in a non-reference Caucasian (Estonian) population similar to those seen in the HapMap CEPH (CEU) population sample
-
Has reproduction · 23
Analysis of whole-genome re-sequencing data of ducks reveals a diverse demographic history and extensive gene flow between Southeast/South Asian and Chinese populations.
PMID 33849442 · PMC8042899 · Genetics, selection, evolution : GSE · 2021 · 8 claims · 8 setups
Whole-genome resequencing reveals three geographically distinct genetic groups: local Chinese, wild, and local Southeast/South Asian duck populations
-
Has reproduction · 73
Genetic polyploid phasing from low-depth progeny samples.
PMID 35692633 · PMC9184567 · iScience · 2022 · 8 claims · 7 setups
WH-PPG phases polyploid parental samples by scoring informative variant pairs with a Bayesian log-likelihood model of progeny allele depths, clustering alleles by co-occurrence likelihood, and assigning clusters to haplotypes via interval scheduling
-
Full-text index only
Calibrating the performance of SNP arrays for whole-genome association studies.
PMID 18584036 · PMC2432039 · PLoS genetics · 2008 · 8 claims · 7 setups
Previous SNP array genetic coverage estimates are inflated due to SNP overfitting and sample overfitting, since they were evaluated on the same HapMap SNPs/individuals used to design the arrays.
-
Has reproduction · 92
Acquisition and loss of CTX-M plasmids in Shigella species associated with MSM transmission in the UK.
PMID 34427554 · PMC8549364 · Microbial genomics · 2021 · 8 claims · 8 setups
bla_CTX-M-27 is located on IncFII pKSR100-like plasmids, flanked by IS26 and IS903B
-
Full-text index only
Slider--maximum use of probability information for alignment of short sequence reads and SNP detection.
PMID 18974170 · PMC2638935 · Bioinformatics (Oxford, England) · 2009 · 7 claims · 3 setups
Slider aligns reads using all bases above a probability threshold (baseMinPrb) from prb files, generating all possible read sequences above a read probability threshold (read_0_MinPrb), rather than only the most probable sequence
-
Full-text index only
Telomere-to-telomere assembly of a complete human X chromosome.
PMID 32663838 · PMC7484160 · Nature · 2020 · 8 claims · 8 setups
Produced the first gapless, telomere-to-telomere assembly of a human chromosome (the X chromosome) using the CHM13 cell line
-
Full-text index only
Rise of the machines.
PMID 18670625 · PMC2467494 · PLoS genetics · 2008 · 8 claims · 4 setups
New short-read sequencing platforms (Illumina Genome Analyzer, 454 FLX, ABI SOLiD) enable rapid, scalable whole-genome resequencing that was previously restricted to dedicated sequencing centers using Sanger methods.
-
Full-text index only
An evaluation of the performance of HapMap SNP data in a Shanghai Chinese population: analyses of allele frequency, linkage disequilibrium pattern and tagging SNPs transferability on chromosome 1q21-q25.
PMID 18302794 · PMC2292209 · BMC genetics · 2008 · 7 claims · 5 setups
Among the four HapMap populations, CHB shows the best correlation with the Shanghai population on allele frequencies, LD, and haplotype frequencies
-
Full-text index only
Searching for SNPs with cloud computing.
PMID 19930550 · PMC3091327 · Genome biology · 2009 · 8 claims · 4 setups
Crossbow combines the Bowtie short-read aligner and SOAPsnp SNP caller into a seamless, automatic Hadoop/MapReduce pipeline for whole-genome resequencing analysis
-
Full-text index only
A robust approach to identifying tissue-specific gene expression regulatory variants using personalized human induced pluripotent stem cells.
PMID 19911041 · PMC2766639 · PLoS genetics · 2009 · 8 claims · 7 setups
Padlock probes combined with high-throughput sequencing enable accurate, quantitative, low-bias digital RNA allelotyping of allele-specific expression
-
Has reproduction · 83
Accurate prediction of metagenome-assembled genome completeness by MAGISTA, a random forest model built on alignment-free intra-bin statistics.
PMID 35248155 · PMC8898458 · Environmental microbiome · 2022 · 7 claims · 7 setups
MAGISTA, a random forest model built on alignment-free intra-bin distance-distribution statistics, can estimate MAG completeness and purity without relying on reference marker genes.
-
Full-text index only
Targeted next-generation sequencing of a cancer transcriptome enhances detection of sequence variants and novel fusion transcripts.
PMID 19835606 · PMC2784330 · Genome biology · 2009 · 7 claims · 2 setups
Hybrid selection of cDNA dramatically increases the specificity of sequencing reads mapping to targeted cancer-related transcripts.
-
Has reproduction · 49
EDGE COVID-19: a web platform to generate submission-ready genomes from SARS-CoV-2 sequencing efforts.
PMID 35561186 · PMC9113274 · Bioinformatics (Oxford, England) · 2022 · 7 claims · 5 setups
EDGE COVID-19 (EC-19) is a web-based platform that automates QC, reference-based variant/consensus calling, lineage determination, and submission of SARS-CoV-2 genomes and metadata to GenBank, GISAID and INSDC for both Illumina and ONT data.
-
Has reproduction · 82
Ordinal-level phylogenomics of the arthropod class Diplopoda (millipedes) based on an analysis of 221 nuclear protein-coding loci generated using next-generation sequence analyses.
PMID 24236165 · PMC3827447 · PloS one · 2013 · 8 claims · 8 setups
An ordinal-level phylogeny of Diplopoda reconstructed from 221 nuclear protein-coding loci (61,641 aligned amino acid columns) differs from existing classifications in fundamental ways.
-
Has reproduction · 69
Genomic insights into the diversity, virulence, and antimicrobial resistance of group B Streptococcus clinical isolates from Saudi Arabia.
PMID 38711928 · PMC11070470 · Frontiers in cellular and infection microbiology · 2024 · 8 claims · 8 setups
Sequenced GBS isolates from Saudi Arabia show high genetic diversity, with 28 sequence types and nine distinct serotypes including uncommon serotypes VII and VIII
-
Has reproduction · 78
annotate_my_genomes: an easy-to-use pipeline to improve genome annotation and uncover neglected genes by hybrid RNA sequencing.
PMID 36472574 · PMC9724561 · GigaScience · 2022 · 7 claims · 8 setups
annotate_my_genomes is an easy-to-use genome-guided pipeline that uses hybrid (PacBio+Illumina) assembled transcripts to distinguish coding genes from long non-coding RNAs and reconcile them with prior annotations.
-
Has reproduction · 67
Optimal scaling of digital transcriptomes.
PMID 24223126 · PMC3819321 · PloS one · 2013 · 8 claims · 8 setups
Fifteen existing and novel transcript-count normalization algorithms can be compared with two novel, mutually independent metrics: the number of "uniform" genes (sufficiently low coefficient of variation after normalization) and low average Spearman correlation between normalized expression profiles of gene pairs.