Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Full-text index only
Target SNP selection in complex disease association studies.
PMID 15248903 · PMC487897 · BMC bioinformatics · 2004 · 7 claims · 3 setups
A computational pipeline can retrieve gene sequence, collect SNP variation data, and annotate SNPs falling in functional motifs (promoter, exon-intron structure, AU-rich elements, TF binding sites, splice sites) with expression in target tissue
-
Has reproduction · 85
scSAMAC: saliency-adjusted masking induced attention contrastive learning for single-cell clustering.
PMID 40131310 · PMC11934584 · Briefings in bioinformatics · 2025 · 8 claims · 1 setups
scSAMAC integrates contrastive learning and negative binomial (NB) losses into a VAE, extracting features via contrastive unit similarity while preserving intrinsic data characteristics to enhance robustness and generalization in clustering.
-
Full-text index only
Nucleotide-resolution analysis of structural variants using BreakSeq and a breakpoint library.
PMID 20037582 · PMC2951730 · Nature biotechnology · 2010 · 8 claims · 7 setups
A standardized, non-redundant library of 1,889 breakpoint-resolved SVs was assembled from eight published surveys
-
Full-text index only
NCBI Reference Sequence (RefSeq): a curated non-redundant sequence database of genomes, transcripts and proteins.
PMID 15608248 · PMC539979 · Nucleic acids research · 2005 · 7 claims · 5 setups
RefSeq provides a curated, non-redundant, explicitly linked collection of genomic, transcript and protein sequences spanning prokaryotes, eukaryotes and viruses.
-
Full-text index only
Genome informatics: taming the avalanche of genomic data.
PMID 15642109 · PMC549058 · Genome biology · 2005 · 8 claims · 7 setups
Ultraconserved regions (>100 bp, 100% conserved among mammals) exist in the genome and their function remains unknown
-
Has reproduction · 95
MetaMap: an atlas of metatranscriptomic reads in human disease-related RNA-seq data.
PMID 29901703 · PMC6025204 · GigaScience · 2018 · 6 claims · 7 setups
A two-step 'omni' RNA-seq pipeline (MetaMap) combining STAR human alignment with CLARK-S metagenomic classification can quantify archaeal, bacterial, and viral reads from the non-human read fraction of human RNA-seq data
-
Has reproduction · 50
RNA modifications detection by comparative Nanopore direct RNA sequencing.
PMID 34893601 · PMC8664944 · Nature communications · 2021 · 7 claims · 5 setups
Nanocompore is a model-free comparative method that uses a 2-component Gaussian mixture model (GMM) and univariate statistical tests on signal intensity/dwell time to detect RNA modifications in Nanopore direct RNA sequencing data without needing a training set
-
Has reproduction · 79
RetroSnake: A modular pipeline to detect human endogenous retroviruses in genome sequencing data.
PMID 36339261 · PMC9626663 · iScience · 2022 · 8 claims · 4 setups
RetroSnake is an end-to-end, modular, computationally efficient Snakemake pipeline for detecting HERV-K insertions in short-read NGS data, from raw alignment files to an annotated interactive HTML report
-
Has reproduction · 80
VGEA: an RNA viral assembly toolkit.
PMID 34567846 · PMC8428259 · PeerJ · 2021 · 8 claims · 5 setups
VGEA is a Snakemake workflow that chains existing tools (fastp, BWA, SAMtools, IVA, shiver, SeqKit, QUAST, MultiQC) into an all-in-one RNA viral genome assembly pipeline
-
Full-text index only
Sushi gets serious: the draft genome sequence of the pufferfish Fugu rubripes.
PMID 12225591 · PMC139409 · Genome biology · 2002 · 8 claims · 7 setups
The Fugu rubripes draft genome sequence was generated by whole-genome shotgun sequencing assembled to ~5.6x coverage using the JAZZ pipeline.
-
Full-text index only
TassDB: a database of alternative tandem splice sites.
PMID 17142241 · PMC1669710 · Nucleic acids research · 2007 · 7 claims · 3 setups
TassDB is a relational database storing GYNGYN donor and NAGNAG acceptor tandem splice sites across eight species
-
Full-text index only
LOCATE: a mammalian protein subcellular localization database.
PMID 17986452 · PMC2238969 · Nucleic acids research · 2008 · 8 claims · 6 setups
LOCATE is a curated, web-accessible database housing membrane organization and subcellular localization data for mouse and human proteins.
-
Has reproduction · 85
An extensive evaluation of read trimming effects on Illumina NGS data analysis.
PMID 24376861 · PMC3871669 · PloS one · 2013 · 8 claims · 8 setups
Read trimming increases the quality and reliability of downstream NGS analyses (RNA-Seq mapping, SNP identification, genome assembly) while reducing execution time and computational resources.
-
Has reproduction · 68
Bayesian transcriptome assembly.
PMID 25367074 · PMC4397945 · Genome biology · 2014 · 8 claims · 8 setups
Bayesembler, a probabilistic transcriptome assembler built on a Bayesian model of the RNA sequencing process with Gibbs sampling over expressed candidates, abundances and read assignments, is introduced.