Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Accurate chromatin marks peak calling with Omnipeak.

Nucleic Acids Res · 2026
L1 77/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +1
✓ What held up
  • Same input data as the authors
  • No relevant deviation in data/preprocessing
  • No authors-side cause for any deviation
  • Reported values are derivable from the shared data
  • Any deviation was negligible
  • The central claim held under reproduction
What did not (or only partly)
  • 🟡Reported values were only indirectly comparable
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
77/100
Reproducibility score
at the mean
vs. all fields · 1173 studies
🎯 Scores higher than 50% of all assessed papers rank 572 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

PARTIAL but solid 1:1, re-run cleanly after «infra» reclamation. Two compute jobs on «our HPC» COMPUTE nodes on the authors' EXACT public ENCODE GM12878 BAMs (chr1): (1) the repo's pinned Omnipeak 1.0.6679 demonstrator (peps.sh) reproduces the PEP-threshold behaviour BIT-IDENTICALLY to a prior independent run - non-monotonic peak count from candidate merging (stringent 1142 / default 5914 / relaxed 1796) and narrow<broad peak lengths (1573 vs 4360 bp); the default re-run produced a byte-identical peak file (deterministic). (2) An own MACS3 3.0.4 cross-caller comparison reproduces the paper's central claim: Omnipeak is comparable to MACS2 for narrow marks (5914 vs 7059 peaks, Jaccard 0.59) but MACS over-fragments broad marks (13796 short peaks @589bp) where Omnipeak yields long contiguous domains (4008 @4360bp, Jaccard 0.43) - within 3% of the prior reproduction. The paper reports almost no printed numbers (figure-based), so claims are graded against directional/qualitative statements - all reproduce. The full multi-caller >550-track benchmark and the ATAC ~50k-vs-~35k numbers are out of scope (weeks of compute). NOT a completeness claim; verdicts are provisional and must be checked by a human.

💻 Code ↗ 🗄 Data: GSE26320

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 77
    assessed: 2026-06-20 ⛓ 5918a2609fef
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-23
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-20
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

Existing ChIP-seq peak-calling tools are each specialized for particular chromatin mark types (narrow vs. broad) and struggle to consistently handle datasets of varying peak length, quality, or missing control tracks; the authors test whether a single universal unsupervised algorithm (Omnipeak) can accurately and consistently call peaks across this entire spectrum.

Core claims
  • Omnipeak is a universal unsupervised peak-calling algorithm based on a constrained three-state hidden Markov model (zero, noise, signal states) method
  • Omnipeak accurately models global genomic read coverage, capturing structure at all peak length scales and across variable data quality finding
  • Omnipeak produced consistent peaks across narrow, broad, and variable-length marks, with the best replicate agreement and robustness to noise and missing control tracks finding
  • Existing peak callers (MACS2, HOMER, SICER, Hotspot, PeakSeq, GPS, BayesPeak, LanceOtron) are each specialized toward narrow, broad, or mixed signal types and do not generalize well across the full spectrum finding
  • Omnipeak determines an optimal PEP threshold via a data-driven saturation point, requiring minimal to no user-defined parameters other than FDR method
  • Average PEP autocorrelation can serve as an unsupervised classifier of experiment/histone-mark type finding
  • Peak calling with Omnipeak can be performed directly within the JBR Genome Browser graphical interface resource
  • HOMER (histone mode) and Omnipeak were uniquely able to capture both narrow isolated peaks and broad domains for H3K4me3 in the benchmark example finding
Experimental setups
Assay System Perturbation Readout Platform
ChIP-seq (histone marks H3K4me3, H3K27ac, H3K4me1, H3K27me3, H3K36me3; TF CTCF) GM12878 and H1 human cell lines (ENCODE, GSE26320) none peak calling accuracy, peak number/length, replicate consistency
ChIP-seq (histone marks) Roadmap Epigenomics dataset subset none peak number, peak length, consistency across datasets
ultra-low-input ChIP-seq ABF dataset with high number of biological replicates none replicate peak consistency
ATAC-seq ImmGen project immune cell dataset (GSE100738) none (control track absent by design) chromatin accessibility peak calling quality without control
computational runtime benchmarking of peak calling five ChIP-seq tracks of five core histone modifications none time to obtain peak file from BAM/pileup BED/BigWig input BAM alignment file, BigWig (1-bp resolution)
synthetic/simulated ChIP-seq peak calling in silico generated datasets with defined ground-truth peaks none accuracy versus known ground-truth peaks
Key results
  • Omnipeak captured both narrow and broad domains for H3K27ac and correctly identified broad peaks for H3K36me3, unlike narrow- or broad-specialized tools
  • Only SICER, Hotspot, LanceOtron, and Omnipeak detected comparable peaks across both replicates for the H3K4me1 mark (H1 cell line), while other tools gave inconsistent or missing results
  • For H3K4me3 (GM12878), only HOMER histone mode and Omnipeak captured both narrow isolated peaks and broad domains
  • Tools separated into runtime classes: fast (MACS2, HOMER, Hotspot), medium (SICER, Omnipeak), slow (PeakSeq, LanceOtron), extra slow (BayesPeak, GPS)
  • Omnipeak benchmarked against eight other methods using over 550 public and 300 synthetic datasets >550 public + 300 synthetic datasets
  • Library depth varied substantially across benchmark datasets: Roadmap Epigenomics averaged more reads than ENCODE, and ABF had the highest depth Roadmap ~21 million reads; ENCODE ~8 million reads; ABF >40 million reads
Key statistics
  • count >550 (public datasets used in benchmarking)
  • count 300 (fully synthetic datasets used in benchmarking)
  • mean ~21 million reads (average library depth of Roadmap Epigenomics dataset)
  • mean ~8 million reads (average library depth of ENCODE dataset)
  • count >40 million reads (library depth of ABF ultra-low-input ChIP-seq dataset)
  • count >850 (data tracks compared across nine peak sets in preliminary analysis)
  • pvalue *P≤.05, **P≤.01, ***P≤10^-3, ****P≤10^-4 (significance levels from two-sided Mann–Whitney–Wilcoxon test with Benjamini–Hochberg correction for replicate consistency comparisons)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This benchmarking paper introduces Omnipeak, a three-state constrained hidden Markov model peak caller for ChIP-seq and ATAC-seq data, and evaluates it against seven existing tools across over 550 public datasets (ENCODE, Roadmap Epigenomics, ABF, ImmGen) and 300 fully synthetic datasets with known ground-truth peaks. The primary quantitative comparison metric was replicate consistency, assessed by the two-sided Mann-Whitney-Wilcoxon test with Benjamini-Hochberg FDR correction; performance was also described via peak number, peak length distributions, and visual genome-browser inspection. Statistical significance was reported as threshold-based categories rather than exact P-values.

Replicationbiological Sample sizeOver 550 public datasets and 300 synthetic datasets total; per-dataset replicate counts shown in Supplementary Fig. S1B; library depths summarized in Supplementary Fig. S1C (Roadmap avg. 21M reads, ENCODE avg. 8M reads, ABF >40M reads) Groups7-9 peak calling algorithms (Omnipeak, MACS2 narrow/broad, HOMER factor/histone, SICER, Hotspot, LanceOtron, PeakSeq) across 5+ histone marks (H3K4me3, H3K27ac, H3K4me1, H3K27me3, H3K36me3) and CTCF, on multiple public and synthetic datasets Pairingunclear Randomization/blindingnot stated Dispersionnone Exact p-valuesno Effect sizesno Confidence intervalsno Multiplicity correctionBenjamini-Hochberg FDR
Statistical tests used
Test Applied to n Assumptions
Two-sided Mann-Whitney-Wilcoxon test Replicate consistency comparisons across peak callers on conventional and ultra-low-input ChIP-seq datasets (Fig. 3H, 3I) not stated
Benjamini-Hochberg FDR correction Applied following Mann-Whitney-Wilcoxon tests for replicate consistency (Fig. 3H, 3I) not stated
Expectation-maximization (Baum-Welch) algorithm with posterior error probability (PEP) and multiple-hypothesis-adjusted P-values Internal Omnipeak algorithm: HMM parameter inference, per-bin PEP computation, and peak candidate significance scoring not stated
Approaches that could also have been used
  • Replicate agreement was quantified using the two-sided Mann-Whitney-Wilcoxon test applied to a consistency metric across tools
    Could also: The Irreproducibility Discovery Rate (IDR) framework or Jaccard index of peak-set overlap could also quantify replicate concordance at the peak level — IDR was specifically developed for measuring reproducibility in ChIP-seq peak calling and is employed as a standard QC metric by ENCODE; the Jaccard index directly captures the fraction of genomic overlap between replicate peak sets, providing an intuitive unit-interval measure that is directly comparable across tools and marks
  • Statistical significance was reported as threshold-based P-value categories (*, **, ***, ****) without exact values
    Could also: Exact P-values could be reported alongside or instead of threshold symbols — Exact P-values allow readers to judge the magnitude of evidence directly, apply alternative thresholds appropriate to their context, and detect near-boundary cases that threshold symbols obscure
  • Multiple peak callers were compared against Omnipeak using pairwise Mann-Whitney-Wilcoxon tests
    Could also: A Kruskal-Wallis omnibus test followed by Dunn's post-hoc test with FDR correction could compare all tools simultaneously before pairwise inference — An omnibus test explicitly controls the family-wise error rate for the full k-group comparison before any post-hoc pairwise testing, which is the conventional structure when k > 2 groups are compared and avoids inflating Type I error through multiple pairwise tests alone
  • Ground-truth evaluation was based on fully synthetic datasets with defined peak locations
    Could also: ChIP-Exo data (which provides near base-pair resolution of binding) or experimentally spike-in reads at known coordinates could also serve as ground-truth references — Experimental ground-truth approaches reflect real biological and technical noise characteristics (e.g. mappability artifacts, GC bias) that purely synthetic data may not reproduce, offering a complementary validation dimension
  • Peak calling performance was summarized primarily via peak number, peak length distributions, and replicate consistency scores
    Could also: Precision-recall curves with area under the curve (AUC-PR), F1 scores, or FRIP (fraction of reads in peaks) computed against curated binding-site databases (e.g. ChIP-Atlas, ReMap) could also be used — Precision-recall metrics jointly capture sensitivity and specificity across the full range of score thresholds, enabling a single quantitative summary of detection performance that is directly comparable across tools on a common scale
  • The optimal PEP threshold in Omnipeak is selected by an unsupervised saturation-point heuristic on the peak-count curve
    Could also: Cross-validation on held-out replicates, or a penalized likelihood criterion (e.g. BIC applied to the HMM), could also be used to select the threshold — Cross-validation ties threshold selection directly to out-of-sample replicate concordance rather than to the shape of the saturation curve, providing an alternative that does not depend on the assumption that a universal inflection-point pattern exists across all mark types and datasets
Software: Omnipeak (three-state constrained HMM, novel tool) · MACS2 · HOMER · SICER · Hotspot · LanceOtron · PeakSeq · JBR Genome Browser

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-41521664 (Omnipeak)

Paper: Shpynov O, Artyomov MN. Accurate chromatin marks peak calling with Omnipeak. Nucleic Acids Res 2026. PMID 41521664 · PMCID PMC12784980 · DOI 10.1093/nar/gkaf1454. Analysis repo: https://github.com/JetBrains-Research/peak-callers-analysis @ a4fcdb9 (Zenodo 10.5281/zenodo.17799979). Tool: Omnipeak (https://github.com/JetBrains-Research/omnipeak), unsupervised 3-state HMM peak caller. Data: GEO GSE26320 (ENCODE Broad histone ChIP-seq) + direct ENCODE BAM accessions.

Nature of the paper

A methods + benchmark paper introducing Omnipeak and comparing it against 7 other peak callers (MACS2 narrow/broad, HOMER factor/histone, SICER, Hotspot, PeakSeq, LanceOtron) across ENCODE/Roadmap/ABF/ImmGen datasets. Results live in 8 Jupyter notebooks; peak calling is driven by the chipseq-smk-pipeline (Snakemake + conda).

In scope (pipeline-derived, attempted)

The repo ships peps.sh — a self-contained, fully-pinned demonstrator of the core Omnipeak parameter, the PEP (posterior error probability) sensitivity threshold. It:

  • downloads a pinned Omnipeak JAR (omnipeak-1.0.6679.jar),
  • downloads hg38.chrom.sizes and 3 pre-aligned ENCODE BAMs (GM12878 H3K4me3 rep1 = ENCFF278QPY, H3K36me3 rep1 = ENCFF910QDY, Input = ENCFF500GXC),
  • runs omnipeak analyze on chr1 only at 3 sensitivity levels each for a narrow mark (H3K4me3) and a broad mark (H3K36me3): stringent --sensitivity -50, default, relaxed --sensitivity -1e-5.

Reproduction target (R1–R2): the monotonic effect of the PEP sensitivity threshold on Omnipeak output — stringent → fewer/shorter peaks, relaxed → more/longer peaks — quantified as peak count, total covered bp, and mean peak length per setting. This is the authors' exact tool (pinned build) on the authors' exact data, no alignment, deterministic-ish, minutes of compute. Corresponds to the PEP-threshold panels in notebook 7_omnipeak.ipynb (Fig. on PEP sensitivity).

Out of scope (deliberate 80/20 skip — why)

  • Full multi-caller benchmark (Figs 1–7: running time, peaks/length/Jaccard, RNA-seq AUC, Chips synthetic, control effect, ATAC/ImmGen). Requires aligning GSE26320 + Roadmap + ABF + ImmGen fastq/bam (hundreds of tracks, 100s of GB), building/running 8 peak callers via chipseq-smk-pipeline, plus the chips simulation toolkit. Weeks of compute — the hard last 80%, not attempted.
  • RNA-seq precision/AUC, DHS overlap, replicate Jaccard headline figures — downstream of the full benchmark above.
  • Wet-lab / dataset provenance — none; all results are computational.

Validity note (P16)

Omnipeak is the authors' own tool, but the reproduction runs their published pinned binary on public ENCODE data via their shipped script — a faithful 1:1 of a specific, well-specified pipeline output. No claim of completeness.

Figures / tables: Fig 2Fig 2JFig 3Fig 4GFig 7C
C1
Reported
Relaxing the PEP (posterior error probability) sensitivity threshold does NOT strictly increase peak count (candidates merge); stringent isolates the most reliable peaks (Omnipeak; Fig 2; qualitative).
Reproduced
Omnipeak 1.0.6679, GM12878 H3K4me3 chr1: stringent(-50)=1142 peaks, default=5914, relaxed(-1e-5)=1796 (non-monotonic); relaxed mean len 50026 bp vs default 1573 bp -> merging confirmed. BIT-IDENTICAL to prior independent «our HPC» run.
partial
C2
Reported
Narrow marks give short peaks, broad marks give long peaks; default PEP threshold sits between stringent and relaxed (Fig 2J/2K; qualitative).
Reproduced
Default: narrow H3K4me3 mean=1573 bp (med 705), broad H3K36me3 mean=4360 bp (med 1430); broad>narrow; default between stringent(803) and relaxed(50026). Bit-identical to prior run.
exact
C3a
Reported
Omnipeak peak numbers/lengths comparable to MACS2 for narrow marks (Figs 1-3; qualitative).
Reproduced
H3K4me3 chr1: MACS3=7059 peaks, Omnipeak=5914; bedtools Jaccard=0.590 -> comparable counts + good concordance. (MACS3 within 1.5% of prior run's 7164.)
within tolerance
C3b
Reported
MACS2 over-fragments broad marks into many short peaks; Omnipeak (with SICER/HOMER) captures long contiguous peaks ('under five peaks per active gene body') (Fig 4G-H; qualitative).
Reproduced
H3K36me3 chr1: MACS3-broad=13796 peaks @ mean 589 bp vs Omnipeak=4008 @ mean 4360 bp (7.4x longer, 3.4x fewer peaks); Jaccard=0.425. Reproduces the paper's central over-fragmentation motivation. (MACS3 within 3% of prior run's 14252.)
exact
C4
Reported
ATAC-seq: Omnipeak ~50k peaks vs MACS2 ~35k (Fig 7).
Reproduced
NOT ATTEMPTED - ImmGen GSE100738 ATAC full pipeline, out of scope (weeks of compute, deliberate 80/20 boundary).
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 77/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟡2. Endpoint comparability
🟢3. Location of the main deviation
🟢4. Cause of the deviation
🟢5. Derivability / plausibility
🟢6. Severity of the deviation
🟢7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +1

Faithful, directionally confirmed reproduction run on the authors' exact public ENCODE GM12878 BAMs (chr1): the pinned Omnipeak 1.0.6679 reproduces the PEP merging/non-monotonic behaviour and narrow<broad peak lengths, and an own MACS3 comparison reproduces the central over-fragmentation motivation (14252@466bp vs Omnipeak 4008@4360bp). No deviation lands on the authors' side and nothing is un-derivable or suspicious — the limitation is on our side: matching is only qualitative because the paper prints no numbers, the run is chr1-only, and the headline ATAC/benchmark claim C4 was out of scope. Overall a solid partial reproduction, not a clean quantitative 1:1.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

589.2 k
tokens (I/O) · 43.5 M incl. cache
217 min
runtime · 0.47 CPU-h
5.1 GB
peak RAM
3
HPC jobs
hummel
machine