Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
NCBI Reference Sequences: current status, policy and new initiatives.
PMID 18927115 · PMC2686572 · Nucleic acids research · 2009 · 7 claims · 5 setups
RefSeq is a curated, non-redundant, explicitly linked database of nucleotide and protein sequences spanning genomes, transcripts and proteins across prokaryotes, eukaryotes and viruses
-
Has reproduction · 71
A crowdsourced set of curated structural variants for the human genome.
PMID 32559231 · PMC7329145 · PLoS computational biology · 2020 · 7 claims · 6 setups
A crowdsourcing web application (SVCurator) enables curators to manually review and label large indels and SVs by displaying short, long, and linked read sequencing evidence from GIAB HG002.
-
Full-text index only
A bioinformatics pipeline for a tick pathogen surveillance multiplex amplicon sequencing assay.
PMID 37247570 · PMC10878300 · Ticks and tick-borne diseases · 2023 · 7 claims · 3 setups
The MPAS pipeline is a portable, reproducible Nextflow-based bioinformatics pipeline that identifies and summarizes amplicon sequences produced by the MPAS assay.
-
Full-text index only
Analysis of recent segmental duplications in the bovine genome.
PMID 19951423 · PMC2796684 · BMC genomics · 2009 · 8 claims · 6 setups
Recently duplicated sequence (≥1 kb, ≥90% identity) comprises 3.11% (94.4 Mb) of the bovine genome assembly (Btau_4.0)
-
Has reproduction · 85
An extensive evaluation of read trimming effects on Illumina NGS data analysis.
PMID 24376861 · PMC3871669 · PloS one · 2013 · 8 claims · 8 setups
Read trimming increases the quality and reliability of downstream NGS analyses (RNA-Seq mapping, SNP identification, genome assembly) while reducing execution time and computational resources.
-
Full-text index only
Slider--maximum use of probability information for alignment of short sequence reads and SNP detection.
PMID 18974170 · PMC2638935 · Bioinformatics (Oxford, England) · 2009 · 7 claims · 3 setups
Slider aligns reads using all bases above a probability threshold (baseMinPrb) from prb files, generating all possible read sequences above a read probability threshold (read_0_MinPrb), rather than only the most probable sequence
-
Has reproduction · 96
A bioinformatic pipeline for simulating viral integration data.
PMID 35496474 · PMC9046613 · Data in brief · 2022 · 7 claims · 3 setups
A snakemake-based pipeline was developed to simulate integration of a viral or vector genome into a host genome, including sub-genomic fragment integration, structural variation, and host-site deletions.
-
Has reproduction · 58
Genomic Correlates of Virulence Attenuation in the Deadly Amphibian Chytrid Fungus, Batrachochytrium dendrobatidis.
PMID 26333840 · PMC4632049 · G3 (Bethesda, Md.) · 2015 · 8 claims · 8 setups
Virulence attenuation in the longer-passaged Bd isolate (JEL427-P39) is associated with loss of chromosome copy number relative to the shorter-passaged, more virulent isolate (JEL427-P9)
-
Full-text index only
First report of an HIV-1 triple recombinant of subtypes B, C and F in Buenos Aires, Argentina.
PMID 16959032 · PMC1570496 · Retrovirology · 2006 · 8 claims · 6 setups
Nearly full-length sequencing of 10 HIV-1 seroincident MSM samples revealed 6 subtype B, 3 unique BF recombinants, and 1 novel B/C/F triple recombinant
-
Full-text index only
SeqBuster, a bioinformatic tool for the processing and analysis of small RNAs datasets, reveals ubiquitous miRNA modifications in human embryonic cells.
PMID 20008100 · PMC2836562 · Nucleic acids research · 2010 · 8 claims · 6 setups
SeqBuster is a versatile web-based and stand-alone bioinformatic toolkit for processing and analyzing large-scale small RNA deep sequencing datasets.
-
Has reproduction · 86
Assessing Bos taurus introgression in the UOA Bos indicus assembly.
PMID 34922445 · PMC8684283 · Genetics, selection, evolution : GSE · 2021 · 7 claims · 6 setups
Aligning divergent (cross-subspecies) sequence data detects substantially more SNVs than aligning to a same-subspecies reference, indicating reference/assembly bias in variant calling.
-
Has reproduction · 51
Evaluation of the Available Variant Calling Tools for Oxford Nanopore Sequencing in Breast Cancer.
PMID 36140751 · PMC9498802 · Genes · 2022 · 7 claims · 6 setups
Clair3 and Human-SNP-wf (which incorporates Clair3) achieved the highest performance among the six variant callers tested.
-
Has reproduction · 45
Identifying and classifying trait linked polymorphisms in non-reference species by walking coloured de bruijn graphs.
PMID 23536903 · PMC3607606 · PloS one · 2013 · 8 claims · 9 setups
Bubbleparse detects sequence variants directly from NGS reads without a reference genome, using the coloured de Bruijn graph implementation of Cortex plus a new depth-first bubble-finding module.
-
Has reproduction · 58
A comparative study of techniques for differential expression analysis on RNA-Seq data.
PMID 25119138 · PMC4132098 · PloS one · 2014 · 8 claims · 8 setups
edgeR performs slightly better than DESeq and Cuffdiff2 in terms of the ability to uncover true positives.
-
Has reproduction · 86
Plasmid transmission dynamics and evolution of partner quality in a natural population of Rhizobium leguminosarum.
PMID 41212030 · PMC12691615 · mBio · 2025 · 8 claims · 8 setups
Of the four most frequent plasmid types, types II and III have more stable size, larger core genomes, and track the chromosomal phylogeny (more vertical transmission), while types I and IV (pSym) vary in size and gene content with phylogenies consistent with frequent horizontal transmission.
-
Full-text index only
Codon usage comparison of novel genes in clinical isolates of Haemophilus influenzae.
PMID 15983137 · PMC1160521 · Nucleic acids research · 2005 · 8 claims · 4 setups
A codon usage similarity statistic (ε, based on squared/absolute differences of codon frequencies with an optimized amino acid usage factor) was developed to compare ORFs against a set of 80 reference genomes.
-
Full-text index only
Designating eukaryotic orthology via processed transcription units.
PMID 18445630 · PMC2425467 · Nucleic acids research · 2008 · 8 claims · 5 setups
Existing ortholog databases discard/ignore alternative splicing via all-against-all protein comparisons, causing ambiguous ortholog calls and misclassification of AS isoforms as in-paralogs
-
Full-text index only
Evidence of recombination in Hepatitis C Virus populations infecting a hemophiliac patient.
PMID 19922637 · PMC2784780 · Virology journal · 2009 · 7 claims · 6 setups
A new intragenotypic recombinant HCV strain (1b/1a), named H23, was detected in 1 of 10 hemophiliac patients studied
-
Full-text index only
The UCSC Genome Browser database: update 2010.
PMID 19906737 · PMC2808870 · Nucleic acids research · 2010 · 8 claims · 5 setups
The UCSC Genome Browser provides a large database of publicly available sequence and annotation data with an integrated tool set for examining, comparing, aligning, and displaying genomes
-
Has reproduction · 45
Accurate sequence variant genotyping in cattle using variation-aware genome graphs.
PMID 31092189 · PMC6521551 · Genetics, selection, evolution : GSE · 2019 · 8 claims · 7 setups
Graphtyper outperformed GATK and SAMtools in genotype concordance, non-reference sensitivity, and non-reference discrepancy compared to microarray genotypes