Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Has reproduction · 95
nf-rnaSeqCount: A Nextflow pipeline for obtaining raw read counts from RNA-seq data.
PMID 35574063 · PMC9097006 · South African computer journal = Suid-Afrikaanse rekenaartydskrif · 2021 · 7 claims · 5 setups
nf-rnaSeqCount is a portable, reproducible Nextflow pipeline that maps RNA-seq reads to a reference genome and quantifies gene abundance for differential expression analysis
-
Full-text index only
Correlation of microsynteny conservation and disease gene distribution in mammalian genomes.
PMID 19909546 · PMC2779822 · BMC genomics · 2009 · 7 claims · 8 setups
Density of mouse orthologs of human disease genes correlates with regions of conserved microsynteny in the mouse genome
-
Full-text index only
Ensembl 2007.
PMID 17148474 · PMC1761443 · Nucleic acids research · 2007 · 8 claims · 7 setups
Ensembl added 18 new chordate genomes this year, increasing total genomes available from 15 to 33, the largest yearly increase to date.
-
Has reproduction · 68
Enhancing cell subpopulation discovery in cancer by integrating single-cell transcriptome and expressed variants.
PMID 41647537 · PMC12869734 · Fundamental research · 2026 · 6 claims · 3 setups
scCluster, an end-to-end deep clustering model integrating gene expression and expressed variant (eSNP) features, stratifies cell subpopulations in cancer scRNA-seq data.
-
Has reproduction · 88
nf-core/isoseq: simple gene and isoform annotation with PacBio Iso-Seq long-read sequencing.
PMID 36961337 · PMC10199315 · Bioinformatics (Oxford, England) · 2023 · 7 claims · 4 setups
nf-core/isoseq is a new automated Nextflow-based pipeline that processes raw Iso-Seq subreads through to genome annotation (BED format) without requiring transcriptome assembly.
-
Has reproduction · 94
Manually curated transcriptomics data collection for toxicogenomic assessment of engineered nanomaterials.
PMID 33558569 · PMC7870661 · Scientific data · 2021 · 7 claims · 7 setups
A unified collection of 101 manually curated and homogenized transcriptomics datasets covering human, mouse, and rat ENM exposures in vitro and in vivo was compiled.
-
Has reproduction · 64
Blood Transcriptome Analysis of Septic Patients Reveals a Long Non-Coding Alu-RNA in the Complement C5a Receptor 1 Gene.
PMID 35447887 · PMC9027897 · Non-coding RNA · 2022 · 6 claims · 7 setups
A computational pipeline intersecting immune gene coordinates with Alu element coordinates can identify candidate Alu-lncRNAs
-
Has reproduction · 51
Polyploidy and the petal transcriptome of Gossypium.
PMID 24393201 · PMC3890615 · BMC plant biology · 2014 · 8 claims · 8 setups
Most homoeologous gene pairs in polyploid cotton petals are expressed at equal levels, indicating a surprising level of expression homeostasis; only ~20% of expressed genes show significant genome bias.
-
Full-text index only
A statistical framework for consolidating "sibling" probe sets for Affymetrix GeneChip data.
PMID 18435860 · PMC2397416 · BMC genomics · 2008 · 7 claims · 4 setups
A two-way ANOVA model with a treatment x probe-set interaction term can automatically determine whether sibling probe sets for a gene behave similarly (non-significant interaction, consolidate) or differently (significant interaction, treat as independent)
-
Has reproduction · 30
First step toward gene expression data integration: transcriptomic data acquisition with COMMAND>_.
PMID 30691411 · PMC6348648 · BMC bioinformatics · 2019 · 7 claims · 3 setups
COMMAND>_ is a flexible multi-user web application that searches, downloads, parses, re-annotates, and imports gene expression experiments into a coherent data model.
-
Has reproduction · 59
WASP: a versatile, web-accessible single cell RNA-Seq processing platform.
PMID 33736596 · PMC7977290 · BMC genomics · 2021 · 7 claims · 7 setups
WASP is a software platform for processing Drop-Seq-based scRNA-seq data generated with ddSEQ or 10x protocols, combining a Snakemake pre-processing pipeline with an R Shiny post-processing application.
-
Has reproduction · 90
Optimal Dual RNA-Seq Mapping for Accurate Pathogen Detection in Complex Eukaryotic Hosts.
PMID 39959292 · PMC11825298 · Bio-protocol · 2025 · 7 claims · 6 setups
Mapping adapter-trimmed reads first to the pathogen genome recovers more pathogen reads than the traditional host-first mapping approach.
-
Has reproduction · 95
Determining virus-host interactions and glycerol metabolism profiles in geographically diverse solar salterns with metagenomics.
PMID 28097058 · PMC5228507 · PeerJ · 2017 · 8 claims · 8 setups
Similar virus-host interactions and glycerol metabolism gene associations (notably dihydroxyacetone kinase with Haloquadratum/Halorubrum) exist across geographically diverse solar salterns
-
Has reproduction · 76
Correcting scale distortion in RNA sequencing data.
PMID 39875825 · PMC11776150 · BMC bioinformatics · 2025 · 8 claims · 8 setups
Local averaging reveals expression-level-dependent biases that differ from sample to sample across all RNA-seq datasets studied, and are not corrected by conventional normalization (TPM/FPKM)
-
Has reproduction · 75
Identification of Key Differentially Expressed Genes in Arabidopsis thaliana Under Short- and Long-Term High Light Stress.
PMID 40869111 · PMC12386182 · International journal of molecular sciences · 2025 · 7 claims · 5 setups
Short- and long-term HL responses in Arabidopsis leaves are driven by distinct transcriptional programs, with duration of HL treatment as the primary factor separating transcriptomic clusters.
-
Has reproduction · 84
Expression Atlas update--a database of gene and transcript expression from microarray- and sequencing-based functional genomics experiments.
PMID 24304889 · PMC3964963 · Nucleic acids research · 2014 · 8 claims · 6 setups
Expression Atlas is a value-added database providing gene, protein and splice variant expression across cell types, organism parts, developmental stages, diseases and other biological/experimental conditions, built from manually curated high-quality microarray and RNA-sequencing experiments from ArrayExpress.
-
Has reproduction · 75
Graph-Based Approaches Significantly Improve the Recovery of Antibiotic Resistance Genes From Complex Metagenomic Datasets.
PMID 34690959 · PMC8528159 · Frontiers in microbiology · 2021 · 8 claims · 6 setups
GraphAMR, a Nextflow pipeline that aligns AMR profile HMMs (or AA sequences) to metagenomic assembly graphs via PathRacer, then dereplicates and annotates hits, recovers more and more complete AMR genes than contig-based or read-based methods.
-
Has reproduction · 85
ScLRTC: imputation for single-cell RNA-seq data via low-rank tensor completion.
PMID 34844559 · PMC8628418 · BMC genomics · 2021 · 8 claims · 8 setups
scLRTC imputes dropout entries closest to the original expression values on simulated datasets, outperforming other state-of-the-art methods by SSE and PCC.
-
Has reproduction · 75
A step forward for Shiga toxin-producing Escherichia coli identification and characterization in raw milk using long-read metagenomics.
PMID 36748417 · PMC9836091 · Microbial genomics · 2022 · 8 claims · 6 setups
Long-read metagenomics enables isolation-independent identification and characterization of eae-positive STEC directly from raw milk.
-
Has reproduction · 89
DFAST and DAGA: web-based integrated genome annotation tools and resources.
PMID 27867804 · PMC5107635 · Bioscience of microbiota, food and health · 2016 · 8 claims · 7 setups
DFAST is a web-based bacterial genome annotation and DDBJ submission pipeline with integrated CheckM quality assessment and ANI taxonomic assessment.