Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Has reproduction · 50
MEDUSA: A Pipeline for Sensitive Taxonomic Classification and Flexible Functional Annotation of Metagenomic Shotgun Sequences.
PMID 35330728 · PMC8940201 · Frontiers in genetics · 2022 · 6 claims · 6 setups
MEDUSA is an automated, Conda-installable and Snakemake-managed pipeline performing preprocessing, assembly, alignment, taxonomic classification, and functional annotation on shotgun data.
-
Has reproduction · 76
nf-core/circrna: a portable workflow for the quantification, miRNA target prediction and differential expression analysis of circular RNAs.
PMID 36694127 · PMC9875403 · BMC bioinformatics · 2023 · 8 claims · 4 setups
Existing circRNA workflows are limited: none delineate circRNA-miRNA interactions and only one performs differential expression analysis, requiring users to supplement missing analysis types with in-house expertise
-
Has reproduction · 75
PHA4GE quality control contextual data tags: standardized annotations for sharing public health sequence datasets with known quality issues to facilitate testing and training.
PMID 38860884 · PMC11261899 · Microbial genomics · 2024 · 7 claims · 3 setups
PHA4GE developed a set of standardized contextual data tags (five fields plus controlled-vocabulary terms) for annotating pathogen sequence datasets with known quality issues.
-
Full-text index only
A SNP-centric database for the investigation of the human genome.
PMID 15046636 · PMC395999 · BMC bioinformatics · 2004 · 8 claims · 3 setups
SNPper is a web-based, integrated SNP database combining dbSNP, the Human Genome sequence (Goldenpath), LocusLink, GeneOntology, and SWISS-PROT data with querying, visualization, and export tools.
-
Has reproduction · 89
Evaluating sequence data quality from the Swift Accel-Amplicon CFTR Panel.
PMID 31913291 · PMC6949293 · Scientific data · 2020 · 6 claims · 7 setups
The Accel-Amplicon CFTR panel generates sequencing data with high coverage depth and near 100% on-target reads.
-
Has reproduction · 51
Cell type-specific eQTL analysis of COVID-19 based on single-cell transcriptomic data.
PMID 41064594 · PMC12501775 · NAR genomics and bioinformatics · 2025 · 8 claims · 8 setups
Single-cell eQTL analysis across eight immune cell types identified 2593 genes whose expression is significantly associated with common genetic polymorphisms, with most genes showing cell type-specific effects
-
Has reproduction · 66
RiboTaxa: combined approaches for rRNA genes taxonomic resolution down to the species level from metagenomics data revealing novelties.
PMID 36159175 · PMC9492272 · NAR genomics and bioinformatics · 2022 · 8 claims · 6 setups
RiboTaxa, combining BBTools, FastQC, SortMeRNA, MetaRib, EMIRGE, VSEARCH, BBMap and QIIME 2's Sklearn classifier, was built as a pipeline for SSU rRNA-based taxonomic profiling of metagenomics data.
-
Has reproduction · 95
transXpress: a Snakemake pipeline for streamlined de novo transcriptome assembly and annotation.
PMID 37016291 · PMC10074830 · BMC bioinformatics · 2023 · 6 claims · 7 setups
transXpress is a Snakemake pipeline that streamlines de novo transcriptome assembly, quantification, and annotation for non-model organisms
-
Has reproduction · 50
rMAP: the Rapid Microbial Analysis Pipeline for ESKAPE bacterial group whole-genome sequence data.
PMID 34110280 · PMC8461470 · Microbial genomics · 2021 · 8 claims · 8 setups
rMAP is a pipeline capable of profiling the resistomes of ESKAPE pathogens using Illumina WGS data
-
Has reproduction · 85
An extensive evaluation of read trimming effects on Illumina NGS data analysis.
PMID 24376861 · PMC3871669 · PloS one · 2013 · 8 claims · 8 setups
Read trimming increases the quality and reliability of downstream NGS analyses (RNA-Seq mapping, SNP identification, genome assembly) while reducing execution time and computational resources.
-
Full-text index only
Pigs in sequence space: a 0.66X coverage pig genome survey based on shotgun sequencing.
PMID 15885146 · PMC1142312 · BMC genomics · 2005 · 8 claims · 7 setups
Pig sequence is closer to human than mouse is, across exons, UTRs, introns, intergenic regions, ultra-conserved elements, and miRNAs
-
Has reproduction · 71
Protein structure quality assessment based on the distance profiles of consecutive backbone Cα atoms.
PMID 24555103 · PMC3892923 · F1000Research · 2013 · 8 claims · 8 setups
The distance between consecutive backbone Cα atoms in high-quality structures is normally distributed with mean 3.8 Å and standard deviation 0.04 Å, justifying a reference state in which all consecutive Cα atoms are 3.8 Å apart.
-
Full-text index only
A clinical genetic method to identify mechanisms by which pain causes depression and anxiety.
PMID 16623937 · PMC1488826 · Molecular pain · 2006 · 8 claims · 4 setups
Pain-triggered depression/anxiety are mediated by distinct spino-parabrachial-hypothalamic-amygdalar neurochemical pathways compared to mood disorders independent of pain
-
Has reproduction · 85
Optimizing open data to support one health: best practices to ensure interoperability of genomic data from bacterial pathogens.
PMID 33103064 · PMC7568946 · One health outlook · 2020 · 8 claims · 3 setups
An open-access pathogen surveillance database (NCBI Pathogen Detection) plus contributor Best Practices enables FAIR, interoperable genomic data across human, animal, food, and environmental sources for One Health surveillance.
-
Full-text index only
DAVID Knowledgebase: a gene-centered database integrating heterogeneous gene annotation resources to facilitate high-throughput gene functional analysis.
PMID 17980028 · PMC2186358 · BMC bioinformatics · 2007 · 7 claims · 3 setups
The DAVID Gene Concept, a single-linkage algorithm, merges gene clusters from Entrez Gene, UniRef100, and PIR-NREF100 that share protein IDs and species into unified DAVID gene clusters, improving cross-referencing between NCBI and UniProt systems
-
Full-text index only
COMUS: Clinician-Oriented locus-specific MUtation detection and deposition System.
PMID 19958500 · PMC2788389 · BMC genomics · 2009 · 8 claims · 6 setups
COMUS is a bioinformatics system for detecting and depositing new mutations from patient DNA with a clinician-friendly interface
-
Full-text index only
The personal genome project.
PMID 16729065 · PMC1681452 · Molecular systems biology · 2005 · 8 claims · 1 setups
A Personal Genome Project (PGP) should be established as the natural successor to the Human Genome Project, providing integrated genome and phenome data for genetically diverse subjects.
-
Full-text index only
NCBI Reference Sequence (RefSeq): a curated non-redundant sequence database of genomes, transcripts and proteins.
PMID 15608248 · PMC539979 · Nucleic acids research · 2005 · 7 claims · 5 setups
RefSeq provides a curated, non-redundant, explicitly linked collection of genomic, transcript and protein sequences spanning prokaryotes, eukaryotes and viruses.
-
Full-text index only
Ensembl 2006.
PMID 16381931 · PMC1347495 · Nucleic acids research · 2006 · 8 claims · 5 setups
Ensembl now provides annotation for 19 genomes, up from 4 the previous year, including new mammalian (Rhesus macaque, Opossum), chordate (Ciona intestinalis), and yeast genomes.
-
Full-text index only
In silico meets in vivo.
PMID 18304380 · PMC2374716 · Genome biology · 2008 · 8 claims · 8 setups
About 10% of positions in multiple sequence alignments of the human genome with other vertebrate genomes are likely incorrect.