Experiments
Searchable full-text extractions: founding hypothesis, core claims, experimental setups, key results and statistics — pulled out of each paper as structure. Search a cell line, an assay or an entity (e.g. HUH7) and find every paper that worked with it. This corpus stands on its own: most entries carry no reproduction assessment (yet).
-
Has reproduction · 100
Recurrent RNA edits in human preimplantation potentially enhance maternal mRNA clearance.
PMID 36543858 · PMC9772385 · Communications biology · 2022 · 8 claims · 7 setups
Compiled the largest human embryonic A-to-I editome to date from 2071 RNA-seq transcriptomes and identified thousands of per-stage Recurrent Embryonic Edits (REEs, present in ≥50% of samples per stage)
-
Has reproduction · 95
nf-rnaSeqCount: A Nextflow pipeline for obtaining raw read counts from RNA-seq data.
PMID 35574063 · PMC9097006 · South African computer journal = Suid-Afrikaanse rekenaartydskrif · 2021 · 7 claims · 5 setups
nf-rnaSeqCount is a portable, reproducible Nextflow pipeline that maps RNA-seq reads to a reference genome and quantifies gene abundance for differential expression analysis
-
Has reproduction · 67
GAVISUNK: genome assembly validation via inter-SUNK distances in Oxford Nanopore reads.
PMID 36321867 · PMC9805576 · Bioinformatics (Oxford, England) · 2023 · 7 claims · 4 setups
GAVISUNK is an open-source pipeline that validates phased diploid HiFi assemblies by assessing concordance of inter-SUNK distances against orthogonal Oxford Nanopore (ONT) reads.
-
Full-text index only
NCBI Reference Sequence (RefSeq): a curated non-redundant sequence database of genomes, transcripts and proteins.
PMID 15608248 · PMC539979 · Nucleic acids research · 2005 · 7 claims · 5 setups
RefSeq provides a curated, non-redundant, explicitly linked collection of genomic, transcript and protein sequences spanning prokaryotes, eukaryotes and viruses.
-
Has reproduction · 98
Large-scale quality assessment of prokaryotic genomes with metashot/prok-quality.
PMID 35136576 · PMC8804904 · F1000Research · 2021 · 8 claims · 6 setups
metashot/prok-quality is a container-enabled Nextflow pipeline for quality assessment and dereplication of draft prokaryotic genomes
-
Full-text index only
Nucleotide-resolution analysis of structural variants using BreakSeq and a breakpoint library.
PMID 20037582 · PMC2951730 · Nature biotechnology · 2010 · 8 claims · 7 setups
A standardized, non-redundant library of 1,889 breakpoint-resolved SVs was assembled from eight published surveys
-
Has reproduction · 79
RetroSnake: A modular pipeline to detect human endogenous retroviruses in genome sequencing data.
PMID 36339261 · PMC9626663 · iScience · 2022 · 8 claims · 4 setups
RetroSnake is an end-to-end, modular, computationally efficient Snakemake pipeline for detecting HERV-K insertions in short-read NGS data, from raw alignment files to an annotated interactive HTML report
-
Full-text index only
MrHAMER yields highly accurate single molecule viral sequences enabling analysis of intra-host evolution.
PMID 33849057 · PMC8266615 · Nucleic acids research · 2021 · 8 claims · 7 setups
MrHAMER yields >1000s of viral genomes per sample at 99.9% accuracy
-
Has reproduction · 92
A network-guided protocol to discover susceptibility genes in genome-wide association studies using stability selection.
PMID 36609152 · PMC9850185 · STAR protocols · 2023 · 5 claims · 5 setups
The protocol identifies genes that are both statistically associated with a phenotype and functionally interconnected in a biological network
-
Full-text index only
Eduomics: a Nextflow pipeline to simulate -omics data for education.
PMID 41816779 · PMC12972896 · NAR genomics and bioinformatics · 2026 · 8 claims · 4 setups
Eduomics is a Nextflow DSL2 pipeline that automates generation of validated variant-calling and RNA-seq datasets for education while abstracting away technical requirements
-
Has reproduction · 58
iCOMIC: a graphical interface-driven bioinformatics pipeline for analyzing cancer omics data.
PMID 35899080 · PMC9310080 · NAR genomics and bioinformatics · 2022 · 8 claims · 4 setups
iCOMIC provides a GUI-driven, Snakemake-based pipeline integrating multiple tools for DNA-Seq and RNA-Seq analysis with minimal command-line interaction.
-
Full-text index only
MobiCT: a UMI-based circulating tumor DNA analysis pipeline.
PMID 41503160 · PMC12770973 · NAR genomics and bioinformatics · 2026 · 7 claims · 7 setups
MobiCT is a Nextflow/nf-core UMI-based ctDNA pipeline (deduplication, alignment, variant calling with VarDict, annotation with VEP) achieving sensitivity, precision, and F1-score around 90% after comprehensive filtering.
-
Full-text index only
PeakPrime: a peak-guided primer design pipeline for target enrichment in 3'-end RNA-seq.
PMID 41919010 · PMC13034549 · Bioinformatics advances · 2026 · 8 claims · 7 setups
PeakPrime is a reproducible Nextflow pipeline that calls 3′ RNA-seq coverage peaks (MACS2), selects exonic windows, designs strand-appropriate primers (Primer3), and screens specificity (Bowtie2)
-
Full-text index only
Identification of Key Genes in Atherosclerosis by Combined DNA Methylation and miRNA Expression Analyses.
PMID 35949126 · PMC9682560 · Anatolian journal of cardiology · 2022 · 7 claims · 7 setups
10 key genes (TCF7L2, CACNA1C, NRP1, GABBR2, FANCC, DCK, CCDC88C, TCF12, ABLIM1, PBX1) are regulated by both aberrant DNA methylation and miRNA activity in atherosclerosis
-
Full-text index only
Metapipeline-DNA: A comprehensive germline and somatic genomics Nextflow pipeline.
PMID 41850291 · PMC13030954 · Cell reports methods · 2026 · 8 claims · 7 setups
Metapipeline-DNA automates germline and somatic DNA sequencing analysis end-to-end, from raw reads through preprocessing, feature detection, QC, and visualization.
-
Full-text index only
Variation resources at UC Santa Cruz.
PMID 17151077 · PMC1781230 · Nucleic acids research · 2007 · 8 claims · 8 setups
The UCSC Genome Browser variation resources integrate polymorphism data from public collections (dbSNP, HapMap, Affymetrix, Perlegen, SeattleSNPs) into a common format with additional annotations and genomic context.
-
Has reproduction
Fast, accurate, and racially unbiased pan-cancer tumor-only variant calling with tabular machine learning.
PMID 36611079 · PMC9825621 · NPJ precision oncology · 2023 · 8 claims · 8 setups
Tree-based (XGBoost, LightGBM) and deep-learning (TabNet) tabular ML classifiers achieve state-of-the-art somatic vs germline classification in tumor-only WES samples, outperforming PureCN.
-
Full-text index only
Integration with the human genome of peptide sequences obtained by high-throughput mass spectrometry.
PMID 15642101 · PMC549070 · Genome biology · 2005 · 8 claims · 4 setups
PeptideAtlas, a public database integrating MS/MS-derived peptide identifications with the human genome, was built as an expandable resource for proteomic data.
-
Full-text index only
Sushi gets serious: the draft genome sequence of the pufferfish Fugu rubripes.
PMID 12225591 · PMC139409 · Genome biology · 2002 · 8 claims · 7 setups
The Fugu rubripes draft genome sequence was generated by whole-genome shotgun sequencing assembled to ~5.6x coverage using the JAZZ pipeline.
-
Full-text index only
GENCODE: producing a reference annotation for ENCODE.
PMID 16925838 · PMC1810553 · Genome biology · 2006 · 8 claims · 8 setups
GENCODE annotation combines initial manual annotation by HAVANA, experimental validation, and refinement based on results to identify protein-coding genes in ENCODE regions