Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Has reproduction · 84
Deep transcriptomics reveals cell-specific isoforms of pan-neuronal genes.
PMID 40379625 · PMC12084633 · Nature communications · 2025 · 8 claims · 5 setups
Pan-neuronal genes (expressed in many/all neurons) harbor highly cell-specific splice variants/isoforms restricted to single or few neuron types.
-
Full-text index only
BFAST: an alignment tool for large scale genome resequencing.
PMID 19907642 · PMC2770639 · PloS one · 2009 · 7 claims · 4 setups
BFAST is a new algorithm and freely available software tool for aligning large-scale short-read sequencing data to large reference genomes with user-customizable speed and accuracy
-
Has reproduction · 100
nf-core/mag: a best-practice pipeline for metagenome hybrid assembly and binning.
PMID 35118380 · PMC8808542 · NAR genomics and bioinformatics · 2022 · 8 claims · 7 setups
nf-core/mag is a Nextflow/nf-core pipeline for hybrid metagenome assembly, binning and taxonomic classification of MAGs.
-
Has reproduction · 71
A crowdsourced set of curated structural variants for the human genome.
PMID 32559231 · PMC7329145 · PLoS computational biology · 2020 · 8 claims · 8 setups
1235 manually curated SVs were produced that can be used to evaluate SV callers or train machine learning models
-
Has reproduction · 86
LMAS: evaluating metagenomic short de novo assembly methods through defined communities.
PMID 36576131 · PMC9795473 · GigaScience · 2022 · 8 claims · 5 setups
LMAS (Last Metagenomic Assembler Standing) is a flexible, Nextflow-based, Docker-containerized automated workflow for benchmarking de novo metagenomic assemblers against defined mock communities, producing an interactive HTML report.
-
Has reproduction
Genome-wide signatures of convergent evolution in echolocating mammals.
PMID 24005325 · PMC3836225 · Nature · 2013 · 8 claims · 8 setups
Genome-wide convergent sequence evolution between echolocating lineages is not rare but widespread and continuously distributed, with signatures consistent with convergence in nearly 200 loci out of 2,326 examined.
-
Has reproduction · 78
A case study for large-scale human microbiome analysis using JCVI's metagenomics reports (METAREP).
PMID 22719821 · PMC3374610 · PloS one · 2012 · 8 claims · 7 setups
METAREP version 1.3.1 is an open-source, scalable tool for querying, browsing and comparing extremely large volumes of metagenomic annotations, with an extended data model, dynamic weighting, distributed searches and advanced clustering.
-
Has reproduction · 93
Population genomics of the Wolbachia endosymbiont in Drosophila melanogaster.
PMID 23284297 · PMC3527207 · PLoS genetics · 2012 · 8 claims · 8 setups
Wolbachia infection status can be accurately predicted in silico from whole-genome shotgun sequence of individual host strains, showing 99% concordance with diagnostic PCR.
-
Has reproduction · 94
BaRTv2: a highly resolved barley reference transcriptome for accurate transcript-specific RNA-seq quantification.
PMID 35704392 · PMC9546494 · The Plant journal : for cell and molecular biology · 2022 · 8 claims · 6 setups
BaRTv2.18 is the most comprehensive and resolved reference transcriptome in barley to date, containing 39,434 genes and 148,260 transcripts
-
Has reproduction · 78
annotate_my_genomes: an easy-to-use pipeline to improve genome annotation and uncover neglected genes by hybrid RNA sequencing.
PMID 36472574 · PMC9724561 · GigaScience · 2022 · 7 claims · 8 setups
annotate_my_genomes is an easy-to-use genome-guided pipeline that uses hybrid (PacBio+Illumina) assembled transcripts to distinguish coding genes from long non-coding RNAs and reconcile them with prior annotations.
-
Full-text index only
High throughput sequencing and proteomics to identify immunogenic proteins of a new pathogen: the dirty genome approach.
PMID 20037647 · PMC2793016 · PloS one · 2009 · 7 claims · 7 setups
A dirty genome approach using unfinished, unclosed genome sequences combined with proteomics can rapidly identify immunogenic proteins useful for diagnostic tool development
-
Has reproduction · 99
A platinum standard pan-genome resource that represents the population structure of Asian rice.
PMID 32265447 · PMC7138821 · Scientific data · 2020 · 6 claims · 6 setups
The 3,000 Rice Genomes (3K-RG) dataset can be subdivided into 15 subpopulations (K=15), refining the previous K=9 population structure.
-
Has reproduction · 83
Macrel: antimicrobial peptide screening in genomes and metagenomes.
PMID 33384902 · PMC7751412 · PeerJ · 2020 · 8 claims · 8 setups
Macrel introduces a novel set of 22 peptide features (6 local, 16 global), including a new Free Energy Transition (FET) feature group, for AMP and hemolytic activity classification
-
Has reproduction · 89
Evolution of Highly Repetitive Silk Genes in the Luna Moth, Actias luna.
PMID 41738778 · PMC12962854 · Genome biology and evolution · 2026 · 8 claims · 6 setups
Eight sericin genes were identified in the A. luna genome, including two clusters of closely related paralogs (serB-D and serE-G)
-
Has reproduction · 83
Gene Expression Atlas update--a value-added database of microarray and sequencing-based functional genomics experiments.
PMID 22064864 · PMC3245177 · Nucleic acids research · 2012 · 8 claims · 5 setups
Gene Expression Atlas is an added-value database providing curated, re-annotated and statistically analysed gene expression data across cell types, organism parts, developmental stages, disease states and other biological/experimental conditions, derived from ArrayExpress Archive and the European Nucleotide Archive.
-
Has reproduction · 43
TransFlow: a Snakemake workflow for transmission analysis of Mycobacterium tuberculosis whole-genome sequencing data.
PMID 36469333 · PMC9825751 · Bioinformatics (Oxford, England) · 2023 · 8 claims · 8 setups
TransFlow is a Snakemake- and Conda-based workflow that combines state-of-the-art tools into a single, fast, scalable pipeline for MTBC WGS-based transmission analysis.