Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Has reproduction · 83
Gene Expression Analysis Platform (GEAP): A highly customizable, fast, versatile and ready-to-use microarray analysis platform.
PMID 34927664 · PMC8754388 · Genetics and molecular biology · 2021 · 8 claims · 2 setups
GEAP is a GUI-based microarray analysis platform combining a C# front-end with an R (RTerm) back-end via the rgeap package, enabling analysis independent of manufacturer/platform.
-
Full-text index only
RAId_DbS: mass-spectrometry based peptide identification web server with knowledge integration.
PMID 18954448 · PMC2605478 · BMC genomics · 2008 · 7 claims · 4 setups
Constructed enhanced protein databases integrating annotated SAPs, PTMs, and disease associations for 17 organisms.
-
Full-text index only
Compressing DNA sequence databases with coil.
PMID 18489794 · PMC2426707 · BMC bioinformatics · 2008 · 8 claims · 1 setups
coil achieves higher compression ratio than state-of-the-art general-purpose compression tools on a large GenBank EST database file
-
Full-text index only
GOLD.db: genomics of lipid-associated disorders database.
PMID 15588328 · PMC544894 · BMC genomics · 2004 · 8 claims · 4 setups
GOLD.db integrates annotated pathways, gene/protein reference information, and curated gene expression datasets for lipid-associated disorders research
-
Has reproduction · 95
The archives are half-empty: an assessment of the availability of microbial community sequencing data.
PMID 32859925 · PMC7455719 · Communications biology · 2020 · 8 claims · 5 setups
More than half of surveyed amplicon sequencing studies were affected by lack of data deposition, improper file formatting, or inconsistent labeling that impede reuse.
-
Has reproduction · 100
A workflow reproducibility scale for automatic validation of biological interpretation results.
PMID 37150537 · PMC10164546 · GigaScience · 2022 · 8 claims · 4 setups
Comparing output files by checksum alone is insufficient to verify reproducibility, since checksums can differ even when the underlying biological interpretation is unchanged
-
Full-text index only
Asterias: integrated analysis of expression and aCGH data using an open-source, web-based, parallelized software suite.
PMID 17488846 · PMC1933128 · Nucleic acids research · 2007 · 8 claims · 2 setups
Asterias is an integrated, open-source, web-based software suite for analysis of gene expression and aCGH data
-
Has reproduction · 95
Pathway-targeting gene matrix for Drosophila gene set enrichment analysis.
PMID 34710184 · PMC8553153 · PloS one · 2021 · 8 claims · 4 setups
Gene matrix files for GSEA are largely unavailable for Drosophila, limiting pathway-level enrichment analysis in this model organism
-
Has reproduction · 50
Workflow sharing with automated metadata validation and test execution to improve the reusability of published workflows.
PMID 36810800 · PMC9944229 · GigaScience · 2022 · 8 claims · 5 setups
Yevis is a system that builds a workflow registry which automatically validates and tests workflows prior to publication, ensuring they are 'reusable with confidence'.
-
Has reproduction · 71
Hyb: a bioinformatics pipeline for the analysis of CLASH (crosslinking, ligation and sequencing of hybrids) data.
PMID 24211736 · PMC3969109 · Methods (San Diego, Calif.) · 2014 · 8 claims · 6 setups
The 'hyb' pipeline detects, calls, folds and annotates chimeric reads from CLASH high-throughput sequencing data.
-
Has reproduction · 43
Compression of structured high-throughput sequencing data.
PMID 24260313 · PMC3832420 · PloS one · 2013 · 8 claims · 7 setups
Leveraging an explicit data schema (separate field encoding, field modeling, template compression, domain modeling) enables stronger compression of HTS alignment data than general-purpose compression of serialized bytes.
-
Full-text index only
A SNP-centric database for the investigation of the human genome.
PMID 15046636 · PMC395999 · BMC bioinformatics · 2004 · 8 claims · 3 setups
SNPper is a web-based, integrated SNP database combining dbSNP, the Human Genome sequence (Goldenpath), LocusLink, GeneOntology, and SWISS-PROT data with querying, visualization, and export tools.
-
Has reproduction · 58
MZPAQ: a FASTQ data compression tool.
PMID 31171931 · PMC6547476 · Source code for biology and medicine · 2019 · 8 claims · 4 setups
MZPAQ, a hybrid of MFCompress and ZPAQ, achieves the highest compression ratio compared to all evaluated state-of-the-art and general-purpose tools on all benchmark datasets.
-
Has reproduction · 69
Manual curation for improved genome annotation of the functionally extinct northern white rhinoceros (Ceratotherium simum cottoni).
PMID 41490125 · PMC12768360 · PloS one · 2026 · 7 claims · 7 setups
Manual curation of RNA-seq-derived de novo transcripts increased the number of functional genes in the NWR annotation by 81% (from 8,701 to 15,738).
-
Has reproduction · 78
QuasiFlow: a Nextflow pipeline for analysis of NGS-based HIV-1 drug resistance data.
PMID 36699347 · PMC9722223 · Bioinformatics advances · 2022 · 6 claims · 8 setups
QuasiFlow is a Nextflow pipeline that runs entirely locally via command-line tools and a local HIVdb database copy to analyze NGS-based HIV-1 drug resistance testing data.
-
Full-text index only
Applications for protein sequence-function evolution data: mRNA/protein expression analysis and coding SNP scoring tools.
PMID 16912992 · PMC1538848 · Nucleic acids research · 2006 · 7 claims · 8 setups
PANTHER HMMs built from family/subfamily multiple sequence alignments can classify novel protein sequences into functional groups based on statistically significant HMM match scores
-
Has reproduction · 58
Revised annotations, sex-biased expression, and lineage-specific genes in the Drosophila melanogaster group.
PMID 25273863 · PMC4267930 · G3 (Bethesda, Md.) · 2014 · 8 claims · 6 setups
Revised RNA-seq-based gene models for D. ananassae, D. yakuba, and D. simulans include UTRs, empirically verified intron-exon boundaries, and previously unannotated novel exons, improving on r1.3 comparative-genomics annotations that lack UTRs.
-
Has reproduction · 96
A bioinformatic pipeline for simulating viral integration data.
PMID 35496474 · PMC9046613 · Data in brief · 2022 · 7 claims · 3 setups
A snakemake-based pipeline was developed to simulate integration of a viral or vector genome into a host genome, including sub-genomic fragment integration, structural variation, and host-site deletions.
-
Has reproduction · 77
Accurate chromatin marks peak calling with Omnipeak.
PMID 41521664 · PMC12784980 · Nucleic acids research · 2026 · 8 claims · 6 setups
Omnipeak is a universal unsupervised peak-calling algorithm based on a constrained three-state hidden Markov model (zero, noise, signal states)
-
Has reproduction · 99
getSequenceInfo: a suite of tools allowing to get genome sequence information from public repositories.
PMID 35804320 · PMC9264741 · BMC bioinformatics · 2022 · 8 claims · 8 setups
getSequenceInfo (gSeqI) allows programmatic (CLI) or GUI-based retrieval of sequence data and metadata from GenBank, RefSeq, and ENA across Linux, MacOS, and Windows.