Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Has reproduction · 74
Evaluation of classification and forecasting methods on time series gene expression data.
PMID 33156855 · PMC7647064 · PloS one · 2020 · 6 claims · 3 setups
Deep learning based methods generally outperform traditional approaches for time series gene expression classification
-
Has reproduction · 80
SLDMS: A Tool for Calculating the Overlapping Regions of Sequences.
PMID 35046988 · PMC8761809 · Frontiers in plant science · 2021 · 8 claims · 5 setups
SLDMS is a novel method for computing overlapping regions of sequencing reads using suffix array (SA), longest common prefix (LCP) array, document array (DA), and a monotonic stack.
-
Has reproduction · 75
An informatics research platform to make public gene expression time-course datasets reusable for more scientific discoveries.
PMID 33247935 · PMC7698665 · Database : the journal of biological databases and curation · 2020 · 8 claims · 6 setups
GETc enables discovery and visualization of time-course gene expression data and analytical results from GEO
-
Full-text index only
Visualization of three-way comparisons of omics data.
PMID 17335588 · PMC1831488 · BMC bioinformatics · 2007 · 7 claims · 3 setups
A novel HSB (hue, saturation, brightness) color-coding scheme can represent three-way comparisons of corresponding datapoints from three datasets.
-
Full-text index only
Processing and population genetic analysis of multigenic datasets with ProSeq3 software.
PMID 19797407 · PMC2778335 · Bioinformatics (Oxford, England) · 2009 · 8 claims · 7 setups
ProSeq3 is a program with a graphic user interface that simplifies preparation and basic population genetic analysis of multigenic DNA polymorphism datasets
-
Has reproduction · 100
Shiny-Calorie: a context-aware application for indirect calorimetry data analysis and visualization using R.
PMID 41640623 · PMC12867577 · Bioinformatics advances · 2026 · 8 claims · 8 setups
Shiny-Calorie is an open-source interactive application for transparent data and metadata integration, statistical analysis, and visualization of indirect calorimetry datasets.
-
Full-text index only
BTW: a web server for Boltzmann time warping of gene expression time series.
PMID 16845055 · PMC1538860 · Nucleic acids research · 2006 · 5 claims · 4 setups
Symmetric time warping distance is more flexible than Euclidean distance or correlation coefficient for identifying genes with similar temporal expression profiles, especially across sequences of different length.
-
Full-text index only
Quadratic regression analysis for gene discovery and pattern recognition for non-cyclic short time-course microarray experiments.
PMID 15850479 · PMC1127068 · BMC bioinformatics · 2005 · 8 claims · 8 setups
A step-down quadratic regression method (fitting quadratic, then linear, then null models per gene) identifies differentially expressed genes and classifies them into 9 temporal expression patterns using continuous time information.
-
Has reproduction · 80
VGEA: an RNA viral assembly toolkit.
PMID 34567846 · PMC8428259 · PeerJ · 2021 · 8 claims · 5 setups
VGEA is a Snakemake workflow that chains existing tools (fastp, BWA, SAMtools, IVA, shiver, SeqKit, QUAST, MultiQC) into an all-in-one RNA viral genome assembly pipeline
-
Has reproduction · 43
Compression of structured high-throughput sequencing data.
PMID 24260313 · PMC3832420 · PloS one · 2013 · 8 claims · 7 setups
Leveraging an explicit data schema (separate field encoding, field modeling, template compression, domain modeling) enables stronger compression of HTS alignment data than general-purpose compression of serialized bytes.
-
Full-text index only
Inconsistencies in Neanderthal genomic DNA sequences.
PMID 17937503 · PMC2014787 · PLoS genetics · 2007 · 8 claims · 6 setups
The Noonan et al. and Green et al. Neanderthal nuclear DNA datasets yield mutually inconsistent estimates of population split time and Neanderthal admixture proportion when analyzed with the same method
-
Full-text index only
Analyses and comparison of accuracy of different genotype imputation methods.
PMID 18958166 · PMC2569208 · PloS one · 2008 · 8 claims · 3 setups
Stronger LD produces higher imputation accuracy rates for all five methods
-
Has reproduction · 100
miR-4478 Accelerates Nucleus Pulposus Cells Apoptosis Induced by Oxidative Stress by Targeting MTH1.
PMID 36130054 · PMC9897280 · Spine · 2023 · 7 claims · 8 setups
miR-4478 is upregulated in nucleus pulposus tissues from IVDD patients and is degeneration-degree-associated
-
Has reproduction · 64
metaGEM: reconstruction of genome scale metabolic models directly from metagenomes.
PMID 34614189 · PMC8643649 · Nucleic acids research · 2021 · 8 claims · 8 setups
metaGEM enables end-to-end reconstruction of FBA-ready GEMs directly from metagenomes without relying on reference genomes
-
Has reproduction · 44
Detecting DNA modifications from SMRT sequencing data by modeling sequence context dependence of polymerase kinetic.
PMID 23516341 · PMC3597545 · PLoS computational biology · 2013 · 8 claims · 7 setups
Local sequence context strongly determines position-specific polymerase kinetic rate: roughly 80% of IPD variation is explained by a 10 bp context (7 bases upstream, 2 bases downstream of the incorporation site), saturating at 7 bases upstream.
-
Full-text index only
Independent component analysis reveals new and biologically significant structures in micro array data.
PMID 16762055 · PMC1557674 · BMC bioinformatics · 2006 · 7 claims · 8 setups
ICA applied to three microarray datasets reveals many biologically significant components, including low-ranking ones not obvious by rank alone
-
Full-text index only
Grammar-based distance in progressive multiple sequence alignment.
PMID 18616828 · PMC2478692 · BMC bioinformatics · 2008 · 7 claims · 3 setups
A grammar-based (LZ complexity) distance metric can be used to determine the order in which sequences are progressively pairwise aligned
-
Full-text index only
Size matters: just how big is BIG?: Quantifying realistic sample size requirements for human genome epidemiology.
PMID 18676414 · PMC2639365 · International journal of epidemiology · 2009 · 7 claims · 2 setups
Conventional power calculations for case-control studies disregard analytic complexity (e.g. clinical assessment errors, unmeasured aetiological determinants) and can seriously underestimate true sample size requirements
-
Has reproduction · 81
Deubiquitination enzyme USP35 negatively regulates MAVS signaling to inhibit anti-tumor immunity.
PMID 40016186 · PMC11868397 · Cell death & disease · 2025 · 8 claims · 8 setups
USP35 interacts with MAVS and removes its K63-linked polyubiquitin chains, inhibiting viral-induced MAVS-TBK1-IRF3 activation and downstream inflammatory gene expression
-
Has reproduction · 83
Public Omics Explorer (POE): Enabling integrative semantic search across GEO omics datasets based on PubMed publications.
PMID 41282419 · PMC12636342 · Computational and structural biotechnology journal · 2025 · 6 claims · 4 setups
POE is a web platform that semantically links GEO datasets and ENA records through their associated PubMed publications for literature-informed dataset retrieval