Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Has reproduction · 58
A comparative study of techniques for differential expression analysis on RNA-Seq data.
PMID 25119138 · PMC4132098 · PloS one · 2014 · 8 claims · 8 setups
edgeR performs slightly better than DESeq and Cuffdiff2 in terms of the ability to uncover true positives.
-
Full-text index only
BFAST: an alignment tool for large scale genome resequencing.
PMID 19907642 · PMC2770639 · PloS one · 2009 · 7 claims · 4 setups
BFAST is a new algorithm and freely available software tool for aligning large-scale short-read sequencing data to large reference genomes with user-customizable speed and accuracy
-
Has reproduction · 69
TC-hunter: identification of the insertion site of a transgenic gene within the host genome.
PMID 35184734 · PMC8859905 · BMC genomics · 2022 · 7 claims · 4 setups
TC-hunter is an open-source Nextflow pipeline that identifies transgene insertion sites using chimeric reads and discordant read pairs from NGS data.
-
Has reproduction · 100
Intra-Host Co-Existing Strains of SARS-CoV-2 Reference Genome Uncovered by Exhaustive Computational Search.
PMID 37243151 · PMC10224212 · Viruses · 2023 · 8 claims · 7 setups
An exhaustive-search workflow can recover intra-host co-existing SARS-CoV-2 strains from the reference-genome read set (SRR11092062) that de Bruijn-graph assemblers discard.
-
Full-text index only
BOAT: Basic Oligonucleotide Alignment Tool.
PMID 19958483 · PMC2788372 · BMC genomics · 2009 · 7 claims · 3 setups
BOAT can accurately and efficiently map sequencing reads to a reference genome while handling several substitutions and indels simultaneously
-
Full-text index only
Rise of the machines.
PMID 18670625 · PMC2467494 · PLoS genetics · 2008 · 8 claims · 4 setups
New short-read sequencing platforms (Illumina Genome Analyzer, 454 FLX, ABI SOLiD) enable rapid, scalable whole-genome resequencing that was previously restricted to dedicated sequencing centers using Sanger methods.
-
Has reproduction · 95
transXpress: a Snakemake pipeline for streamlined de novo transcriptome assembly and annotation.
PMID 37016291 · PMC10074830 · BMC bioinformatics · 2023 · 6 claims · 7 setups
transXpress is a Snakemake pipeline that streamlines de novo transcriptome assembly, quantification, and annotation for non-model organisms
-
Has reproduction · 85
An extensive evaluation of read trimming effects on Illumina NGS data analysis.
PMID 24376861 · PMC3871669 · PloS one · 2013 · 8 claims · 8 setups
Read trimming increases the quality and reliability of downstream NGS analyses (RNA-Seq mapping, SNP identification, genome assembly) while reducing execution time and computational resources.
-
Has reproduction · 51
Evaluation of the Available Variant Calling Tools for Oxford Nanopore Sequencing in Breast Cancer.
PMID 36140751 · PMC9498802 · Genes · 2022 · 7 claims · 6 setups
Clair3 and Human-SNP-wf (which incorporates Clair3) achieved the highest performance among the six variant callers tested.
-
Has reproduction · 73
Gapless provides combined scaffolding, gap filling, and assembly correction with long reads.
PMID 37142439 · PMC10166144 · Life science alliance · 2023 · 8 claims · 5 setups
gapless is a new tool that combines assembly correction, scaffolding, and gap filling in one pipeline using PacBio or Oxford Nanopore long reads.
-
Has reproduction · 95
nf-rnaSeqCount: A Nextflow pipeline for obtaining raw read counts from RNA-seq data.
PMID 35574063 · PMC9097006 · South African computer journal = Suid-Afrikaanse rekenaartydskrif · 2021 · 7 claims · 5 setups
nf-rnaSeqCount is a portable, reproducible Nextflow pipeline that maps RNA-seq reads to a reference genome and quantifies gene abundance for differential expression analysis
-
Has reproduction · 63
Comparative transcriptome analysis of tomato (Solanum lycopersicum) in response to exogenous abscisic acid.
PMID 24289302 · PMC4046761 · BMC genomics · 2013 · 8 claims · 7 setups
Exogenous ABA alters the expression of a majority (54.73%) of expressed tomato leaf transcripts, with 2,787 significantly differentially expressed genes, predominantly up-regulated.
-
Full-text index only
Searching for SNPs with cloud computing.
PMID 19930550 · PMC3091327 · Genome biology · 2009 · 8 claims · 4 setups
Crossbow combines the Bowtie short-read aligner and SOAPsnp SNP caller into a seamless, automatic Hadoop/MapReduce pipeline for whole-genome resequencing analysis
-
Has reproduction · 67
Optimal scaling of digital transcriptomes.
PMID 24223126 · PMC3819321 · PloS one · 2013 · 8 claims · 8 setups
Fifteen existing and novel transcript-count normalization algorithms can be compared with two novel, mutually independent metrics: the number of "uniform" genes (sufficiently low coefficient of variation after normalization) and low average Spearman correlation between normalized expression profiles of gene pairs.
-
Has reproduction · 45
De novo transcriptomic analysis of leaf and fruit tissue of Cornus officinalis using Illumina platform.
PMID 29451882 · PMC5815590 · PloS one · 2018 · 7 claims · 7 setups
This is the first de novo transcriptomic analysis of Cornus officinalis, providing fundamental gene and biosynthetic pathway information.
-
Has reproduction · 63
hgtseq: A Standard Pipeline to Study Horizontal Gene Transfer.
PMID 36498841 · PMC9738810 · International journal of molecular sciences · 2022 · 8 claims · 8 setups
hgtseq is a fully automated, portable, and scalable Nextflow/nf-core pipeline for detecting horizontal gene transfer signatures from unmapped sequencing reads.
-
Has reproduction · 73
Vespucci: a system for building annotated databases of nascent transcripts.
PMID 24304890 · PMC3936758 · Nucleic acids research · 2014 · 8 claims · 7 setups
Existing ChIP-seq and RNA-seq analysis platforms (e.g. Cufflinks, peak callers) are unsuited to GRO-seq because they assume spliced/exonic reads, uniform density and paired-end data, and cannot identify transcriptional units de novo across the whole genome.
-
Has reproduction · 61
Comprehensive transcriptome study to develop molecular resources of the copepod Calanus sinicus for their potential ecological applications.
PMID 24982883 · PMC4055022 · BioMed research international · 2014 · 8 claims · 8 setups
Illumina RNA-Seq with Trinity de novo assembly produced a C. sinicus transcriptome of 69,751 contigs (average 928.8 bp, N50 1,127 bp) from 58.9 million reads.
-
Has reproduction · 69
Manual curation for improved genome annotation of the functionally extinct northern white rhinoceros (Ceratotherium simum cottoni).
PMID 41490125 · PMC12768360 · PloS one · 2026 · 7 claims · 7 setups
Manual curation of RNA-seq-derived de novo transcripts increased the number of functional genes in the NWR annotation by 81% (from 8,701 to 15,738).
-
Has reproduction · 71
A crowdsourced set of curated structural variants for the human genome.
PMID 32559231 · PMC7329145 · PLoS computational biology · 2020 · 7 claims · 6 setups
A crowdsourcing web application (SVCurator) enables curators to manually review and label large indels and SVs by displaying short, long, and linked read sequencing evidence from GIAB HG002.