Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Has reproduction · 95
transXpress: a Snakemake pipeline for streamlined de novo transcriptome assembly and annotation.
PMID 37016291 · PMC10074830 · BMC bioinformatics · 2023 · 6 claims · 7 setups
transXpress is a Snakemake pipeline that streamlines de novo transcriptome assembly, quantification, and annotation for non-model organisms
-
Has reproduction · 100
Intra-Host Co-Existing Strains of SARS-CoV-2 Reference Genome Uncovered by Exhaustive Computational Search.
PMID 37243151 · PMC10224212 · Viruses · 2023 · 8 claims · 7 setups
An exhaustive-search workflow can recover intra-host co-existing SARS-CoV-2 strains from the reference-genome read set (SRR11092062) that de Bruijn-graph assemblers discard.
-
Has reproduction · 85
An extensive evaluation of read trimming effects on Illumina NGS data analysis.
PMID 24376861 · PMC3871669 · PloS one · 2013 · 8 claims · 8 setups
Read trimming increases the quality and reliability of downstream NGS analyses (RNA-Seq mapping, SNP identification, genome assembly) while reducing execution time and computational resources.
-
Has reproduction · 50
Viewing RNA-seq data on the entire human genome.
PMID 28979763 · PMC5605993 · F1000Research · 2017 · 8 claims · 3 setups
RNA-Seq Viewer is a web application that visualizes genome-wide RNA-seq expression data pulled from NCBI's SRA and GEO databases using Ideogram.js
-
Has reproduction · 95
nf-rnaSeqCount: A Nextflow pipeline for obtaining raw read counts from RNA-seq data.
PMID 35574063 · PMC9097006 · South African computer journal = Suid-Afrikaanse rekenaartydskrif · 2021 · 7 claims · 5 setups
nf-rnaSeqCount is a portable, reproducible Nextflow pipeline that maps RNA-seq reads to a reference genome and quantifies gene abundance for differential expression analysis
-
Has reproduction · 50
RNA-Seq alignment to individualized genomes improves transcript abundance estimates in multiparent populations.
PMID 25236449 · PMC4174954 · Genetics · 2014 · 8 claims · 7 setups
Genetic variants distinguishing an individual genome from the reference cause read misalignment and biased transcript abundance estimates, and fine-tuning of alignment algorithms does not correct this problem.
-
Has reproduction · 57
Diapause vs. reproductive programs: transcriptional phenotypes in a keystone copepod.
PMID 33782539 · PMC8007741 · Communications biology · 2021 · 8 claims · 7 setups
t-SNE clustering of all-gene expression data groups field-collected (diapause program) samples into one cluster while early and late culture (reproductive program) samples separate into two distinct phenotypes
-
Has reproduction · 30
IsoSCM: improved and alternative 3' UTR annotation using multiple change-point inference.
PMID 25406361 · PMC4274634 · RNA (New York, N.Y.) · 2015 · 8 claims · 6 setups
Existing ab initio assemblers (Cufflinks, Scripture) annotate at most one 3' boundary per terminal exon and therefore cannot assemble coexpressed tandem 3' UTR isoforms.
-
Has reproduction · 58
The Li2 mutation results in reduced subgenome expression bias in elongating fibers of allotetraploid cotton (Gossypium hirsutum L.).
PMID 24598808 · PMC3944810 · PloS one · 2014 · 8 claims · 7 setups
The Li2 mutation significantly reduces subgenome (homeolog) expression bias in the elongating fiber transcriptome.
-
Has reproduction · 76
Transcriptional landscape of repetitive elements in normal and cancer human cells.
PMID 25012247 · PMC4122776 · BMC genomics · 2014 · 8 claims · 8 setups
RepEnrich, a computational method that uses all mapping reads (uniquely mapping plus multi-mapping reads assigned to repetitive element subfamily assemblies/pseudogenomes), quantifies genome-wide repetitive element enrichment
-
Has reproduction · 82
Landscape of allele-specific transcription factor binding in the human genome.
PMID 33980847 · PMC8115691 · Nature communications · 2021 · 8 claims · 6 setups
A novel statistical framework (ADASTRA) calls allele-specific TF binding from existing ChIP-Seq alignments by jointly correcting for background allelic dosage (BAD, from aneuploidy/CNVs) and reference mapping bias.
-
Has reproduction · 73
Proteogenomic analysis prioritises functional single nucleotide variants in cancer samples.
PMID 29221171 · PMC5707065 · Oncotarget · 2017 · 8 claims · 6 setups
A customised SAAV peptide database built from RNA-seq/WGS variant calls can be used to search proteomics data and detect single amino acid variant (SAAV)-containing peptides at the protein level
-
Has reproduction · 86
RNASEQR--a streamlined and accurate RNA-seq sequence analysis program.
PMID 22199257 · PMC3315322 · Nucleic acids research · 2012 · 8 claims · 7 setups
RNASEQR is a new RNA-seq mapper/aligner that combines a BWT-based (Bowtie) transcriptomic/genomic alignment with hash-based BLAT local alignment in three sequential steps: transcriptome mapping, novel exon detection, and anchor-and-align novel splice junction identification.
-
Full-text index only
Eukan: a fully automated nuclear genome annotation pipeline for less studied and divergent eukaryotes.
PMID 41567515 · PMC12817076 · NAR genomics and bioinformatics · 2026 · 8 claims · 7 setups
Eukan automatically leverages RNA-Seq coverage to inform generalized Hidden Markov Model gene prediction and intron lengths to inform protein sequence alignments
-
Full-text index only
A generic reference defined by consensus peaks for single-cell ATAC-seq data analysis.
PMID 41663439 · PMC12996591 · Nature communications · 2026 · 7 claims · 7 setups
Aggregating peaks from 624 high-quality bulk ATAC-seq datasets defines ~1.4 million observed consensus peaks (cPeaks) covering ~30% of the genome.
-
Full-text index only
A haplotype-complete chromosome-level assembly of octoploid Urochloa humidicola cv. Tully reveals multiple genomic compositions and evolutionary histories in the species.
PMID 41678351 · PMC13042314 · G3 (Bethesda, Md.) · 2026 · 8 claims · 8 setups
A haplotype-complete, chromosome-level genome assembly of octoploid U. humidicola cv. Tully was generated from PacBio HiFi reads alone
-
Has reproduction · 90
PrimerSeq: Design and visualization of RT-PCR primers for alternative splicing using RNA-seq data.
PMID 24747190 · PMC4411361 · Genomics, proteomics & bioinformatics · 2014 · 8 claims · 3 setups
PrimerSeq is a user-friendly stand-alone software with a GUI for systematic design and visualization of RT-PCR primers for alternative splicing analysis using user-provided RNA-seq data.
-
Has reproduction · 53
PulmonDB: a curated lung disease gene expression database.
PMID 31949184 · PMC6965635 · Scientific reports · 2020 · 7 claims · 6 setups
PulmonDB is a curated, web-based relational database (plus R package) integrating microarray and RNA-seq gene expression data with controlled-vocabulary annotation for COPD and IPF.
-
Full-text index only
Early feature extraction drives model performance in high-resolution chromatin accessibility prediction.
PMID 41526189 · PMC12951969 · Genome research · 2026 · 8 claims · 6 setups
Early feature extraction (via ConvNeXt V2 blocks), rather than downstream architecture type, is the primary determinant of prediction accuracy in high-resolution chromatin accessibility prediction.
-
Has reproduction · 74
Transcriptome profiling of Giardia intestinalis using strand-specific RNA-seq.
PMID 23555231 · PMC3610916 · PLoS computational biology · 2013 · 8 claims · 8 setups
Most of the G. intestinalis genome is transcribed in in vitro-grown trophozoites, but at vastly different expression levels.