Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
Shotgun haplotyping: a novel method for surveying allelic sequence variation.
PMID 16221968 · PMC1253838 · Nucleic acids research · 2005 · 8 claims · 7 setups
A novel shotgun haplotyping method generates haplotypic sequences from long PCR products by shotgun sequencing both alleles concurrently and using read-pair information to separate alleles during assembly
-
Has reproduction · 83
Hobbes: optimized gram-based methods for efficient read alignment.
PMID 22199254 · PMC3315303 · Nucleic acids research · 2012 · 8 claims · 4 setups
Hobbes, a gram-based short-read mapper supporting Hamming and edit distance, is faster than all other read-mapping programs tested while maintaining high mapping quality.
-
Full-text index only
Nucleotide-resolution analysis of structural variants using BreakSeq and a breakpoint library.
PMID 20037582 · PMC2951730 · Nature biotechnology · 2010 · 8 claims · 7 setups
A standardized, non-redundant library of 1,889 breakpoint-resolved SVs was assembled from eight published surveys
-
Has reproduction · 71
polishCLR: A Nextflow Workflow for Polishing PacBio CLR Genome Assemblies.
PMID 36792366 · PMC9985148 · Genome biology and evolution · 2023 · 8 claims · 8 setups
polishCLR is a reproducible, containerized Nextflow workflow that implements best practices for polishing PacBio CLR genome assemblies.
-
Has reproduction · 78
annotate_my_genomes: an easy-to-use pipeline to improve genome annotation and uncover neglected genes by hybrid RNA sequencing.
PMID 36472574 · PMC9724561 · GigaScience · 2022 · 7 claims · 8 setups
annotate_my_genomes is an easy-to-use genome-guided pipeline that uses hybrid (PacBio+Illumina) assembled transcripts to distinguish coding genes from long non-coding RNAs and reconcile them with prior annotations.
-
Has reproduction · 86
Assessing Bos taurus introgression in the UOA Bos indicus assembly.
PMID 34922445 · PMC8684283 · Genetics, selection, evolution : GSE · 2021 · 7 claims · 6 setups
Aligning divergent (cross-subspecies) sequence data detects substantially more SNVs than aligning to a same-subspecies reference, indicating reference/assembly bias in variant calling.
-
Has reproduction · 97
CRISPRbuilder-TB: "CRISPR-builder for tuberculosis". Exhaustive reconstruction of the CRISPR locus in mycobacterium tuberculosis complex using SRA.
PMID 33667225 · PMC7968741 · PLoS computational biology · 2021 · 8 claims · 7 setups
CRISPRbuilder-TB is a new pipeline that reconstructs MTC CRISPR-Cas loci directly from short SRA reads without requiring genome assembly
-
Has reproduction · 58
A comparative study of techniques for differential expression analysis on RNA-Seq data.
PMID 25119138 · PMC4132098 · PloS one · 2014 · 8 claims · 8 setups
edgeR performs slightly better than DESeq and Cuffdiff2 in terms of the ability to uncover true positives.
-
Full-text index only
The sequence and de novo assembly of the giant panda genome.
PMID 20010809 · PMC3951497 · Nature · 2010 · 8 claims · 8 setups
A draft giant panda genome was successfully generated and assembled de novo using only Illumina Genome Analyser short-read sequencing
-
Has reproduction · 83
Gene Expression Atlas update--a value-added database of microarray and sequencing-based functional genomics experiments.
PMID 22064864 · PMC3245177 · Nucleic acids research · 2012 · 8 claims · 5 setups
Gene Expression Atlas is an added-value database providing curated, re-annotated and statistically analysed gene expression data across cell types, organism parts, developmental stages, disease states and other biological/experimental conditions, derived from ArrayExpress Archive and the European Nucleotide Archive.
-
Has reproduction · 85
Ensembl 2013.
PMID 23203987 · PMC3531136 · Nucleic acids research · 2013 · 8 claims · 8 setups
Ensembl (http://www.ensembl.org) provides genome information for sequenced chordate genomes, currently supporting 70 species with a focus on human, mouse, zebrafish and rat.
-
Has reproduction · 82
Landscape of allele-specific transcription factor binding in the human genome.
PMID 33980847 · PMC8115691 · Nature communications · 2021 · 8 claims · 6 setups
A novel statistical framework (ADASTRA) calls allele-specific TF binding from existing ChIP-Seq alignments by jointly correcting for background allelic dosage (BAD, from aneuploidy/CNVs) and reference mapping bias.
-
Has reproduction · 67
binny: an automated binning algorithm to recover high-quality genomes from complex metagenomic datasets.
PMID 36239393 · PMC9677464 · Briefings in bioinformatics · 2022 · 8 claims · 8 setups
binny outperforms or is highly competitive with commonly used and state-of-the-art binning methods (MetaBAT2, MaxBin2, CONCOCT, VAMB, SemiBin, MetaDecoder)
-
Has reproduction · 67
A consensus approach to vertebrate de novo transcriptome assembly from RNA-seq data: assembly of the duck (Anas platyrhynchos) transcriptome.
PMID 25009556 · PMC4070175 · Frontiers in genetics · 2014 · 8 claims · 8 setups
Multiple k-mer (MK) assemblies are more complete than single k-mer (SK) assemblies, showing higher reads-mapped-back-to-transcripts (RMBT) and higher CEGMA complete-gene percentages for all three tools.
-
Full-text index only
Ensembl 2008.
PMID 18000006 · PMC2238821 · Nucleic acids research · 2008 · 8 claims · 6 setups
The Ensembl regulatory build integrates multiple genome-wide functional genomics datasets to automatically annotate regulatory regions and assign putative functions across the genome.
-
Has reproduction · 83
Macrel: antimicrobial peptide screening in genomes and metagenomes.
PMID 33384902 · PMC7751412 · PeerJ · 2020 · 8 claims · 8 setups
Macrel is an end-to-end pipeline that predicts high-quality AMP candidates from peptides, contigs, or reads of (meta)genomes
-
Has reproduction · 27
Transcriptome profiling of radish (Raphanus sativus L.) root and identification of genes involved in response to Lead (Pb) stress with next generation sequencing.
PMID 23840502 · PMC3688795 · PloS one · 2013 · 8 claims · 5 setups
A de novo radish root transcriptome of 68,940 assembled transcripts including 33,337 unigenes was generated, providing the first comprehensive molecular characterization of the radish root response to Pb stress.
-
Has reproduction · 68
LaSSO, a strategy for genome-wide mapping of intronic lariats and branch points using RNA-seq.
PMID 24709818 · PMC4079972 · Genome research · 2014 · 8 claims · 8 setups
LaSSO (Lariat Sequence Site Origin) identifies intronic lariat reads and pinpoints branch points genome-wide from RNA-seq data by considering every intronic base as a potential branch point and including all possible exon-skipping lariats.
-
Has reproduction · 74
ChIP-seq guidelines and practices of the ENCODE and modENCODE consortia.
PMID 22955991 · PMC3431496 · Genome research · 2012 · 8 claims · 8 setups
ENCODE/modENCODE define a set of working standards and guidelines for ChIP-seq covering antibody validation, experimental replication, sequencing depth, data/metadata reporting, and data quality assessment.
-
Has reproduction · 91
Insights into the evolution of cotton diploids and polyploids from whole-genome re-sequencing.
PMID 23979935 · PMC3789805 · G3 (Bethesda, Md.) · 2013 · 8 claims · 8 setups
An index of 23,859,893 (~24 million) homoeo-SNPs distinguishing A-genome from D-genome cotton was constructed at a density of one SNP per 32.3 bases of the D5 reference.