Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

SEAseq: a portable and cloud-based chromatin occupancy analysis suite.

BMC Bioinformatics · 2022
L1 94/100 PQI 92
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +2
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
  • No relevant deviation in data/preprocessing
  • No authors-side cause for any deviation
  • Reported values are derivable from the shared data
  • Any deviation was negligible
What did not (or only partly)
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
94/100
Reproducibility score
1.1 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 87% of all assessed papers rank 133 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

SEAseq is a WDL/Cromwell ChIP-seq suite (github.com/stjude/seaseq); the paper is described well enough to reproduce its data-derived claims. PRIMARY RESULT (deterministic, in scope): all 7 Table 3 raw read counts reproduced EXACTLY (7/7) by independently counting the public SRA/ENA FASTQs on «infra» (zcat|wc/4, cross-checked vs ENA read_count) — 54.9/45.6/35.1/47.2/35.7M (GSE138742/SRP225129) and 7.5/11.8M (GSE31558/SRP007953); every reported value is exactly the deposited SRA total, so no fabrication of those numbers. OUT OF SCOPE: Table 3 runtime/cost columns (St Jude Cloud/DNAnexus, hardware/billing-specific, non-deterministic) — not attempted. STRETCH (Fig 5b p53 motif): native re-implementation (bowtie2->MACS2->MEME, front1 has no docker/singularity/java for the real WDL stack) built the hg19 index and aligned both p53 samples successfully, but MACS2 aborted on a known bioconda macs2 binary ABI bug (undefined symbol __log_finite) before peak-calling/MEME — an environment packaging issue, not a paper/data problem; the motif claim is left UNVERIFIED (fixable with pinned macs2/numpy or macs3). NOT ATTEMPTED: full multi-sample WDL/Cromwell execution, quality-score rubric (Fig 5a), LIN28B/ZNF143 peak+motif biology.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 94
    assessed: 2026-06-15 ⛓ 41a7b5892aad
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-15
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

Can a comprehensive, infrastructure-independent, cloud-capable pipeline lower the computational and expertise barriers to processing and analyzing ChIP-Seq/CUT&RUN chromatin occupancy data while ensuring high-quality results?

Core claims
  • SEAseq is a comprehensive, infrastructure-independent pipeline that performs all major ChIP-Seq/CUT&RUN analyses (alignment, peak calling, motif analysis, coverage profiling, peak annotation, super-enhancer identification, and quality assessment) in a single execution resource
  • SEAseq uniquely calculates an extensive set of quality metrics (including ENCODE-recommended ChIP-Seq metrics) with a five-scale color-rank flag system and cross-metric averaged rank score for easy interpretation method
  • Using WDL workflow management and Docker containerization enables platform-independent, portable, reproducible, and scalable deployment across personal computers, HPC clusters, and the cloud method
  • SEAseq can automatically access and interface with publicly available data from NCBI GEO and SRA repositories at the user's request method
  • SEAseq provides a cloud-based implementation enabling analysis in environments with constrained computational resources at reasonable cost resource
  • SEAseq identifies enriched regions using MACS for narrow/short-binding factors and SICER for broad regions of enrichment, and identifies super-enhancers via ROSE method
Experimental setups
Assay System Perturbation Readout Platform
ChIP-Seq any organism (publicly available and new datasets) none protein-DNA binding peaks, coverage, motifs, peak annotation, quality metrics SEAseq pipeline (Bowtie, SAMtools, BEDTools, MACS, SICER, ROSE, MEME Suite, BAMToGFF/BAM2GFF)
CUT&RUN sequencing any organism (publicly available and new datasets) none protein-DNA binding peaks, coverage, motifs, peak annotation, quality metrics SEAseq pipeline (Bowtie, SAMtools, BEDTools, MACS, SICER, ROSE, MEME Suite)
Key results
  • SEAseq demonstrated rapid and cost-effective analysis of both new and publicly available datasets in comparative case studies
  • SEAseq generates a quality statistics report with color-flagged metrics and an overall averaged rank score for interpreting experiment quality
  • SEAseq can be executed on a single computing node or parallel/HPC infrastructure using Cromwell, requiring only Java, Cromwell, and a Docker-capable engine
Key statistics
  • other Aligned percent Excellent threshold A ≥ 80% (quality metric rank-flag threshold for percentage of mapped reads)
  • other FRiP Excellent threshold F ≥ 0.05 (fraction of reads in peaks quality metric threshold)
  • other NRF Excellent threshold ≥ 0.8 (non-redundant fraction library complexity threshold)
  • other Linear stitched peaks Excellent threshold L ≥ 10000 (clustered enriched regions (enhancers) quality threshold)
  • other Estimated tag length Excellent threshold E < ±10 (fragment width estimation quality threshold)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This paper describes SEAseq, a computational pipeline for ChIP-Seq and CUT&RUN data analysis; it is a software tool paper rather than a hypothesis-driven study, so the primary 'statistical' content consists of the peak-calling algorithms (MACS for narrow peaks, SICER for broad peaks) and quality-metric thresholds embedded in the pipeline. Comparative case studies are used to demonstrate the pipeline's outputs, but no inferential statistics comparing experimental groups are reported in the main text. Quality metrics (NRF, FRiP, NSC, RSC, PBC) are assessed against fixed ENCODE-derived thresholds rather than via formal hypothesis tests.

Replicationunclear Sample sizeNot described; the paper presents pipeline functionality and case studies without specifying sample sizes or statistical power calculations GroupsPipeline outputs demonstrated on example datasets; no formal group comparisons reported Pairingna Randomization/blindingnot stated Dispersionnone Exact p-valuesno Effect sizesno Confidence intervalsno Multiplicity correctionnull
Statistical tests used
Test Applied to n Assumptions
MACS (Model-based Analysis of ChIP-Seq) peak caller — uses a Poisson-based statistical model to identify enriched regions Narrow-peak calling for sequence-specific transcription factors not stated
SICER (Spatial Clustering for Identification of ChIP-Enriched Regions) — uses a statistical framework based on reads-in-windows with a background model Broad-peak calling for histone modifications not stated
MEME Suite motif discovery and enrichment analysis — uses probabilistic sequence models (EM-based) to identify overrepresented motifs Motif discovery and enrichment in called peaks not stated
Threshold-based quality flagging — fixed cutoffs from ENCODE guidelines applied to NRF, FRiP, NSC, RSC, PBC, aligned percent SEAseq quality dashboard for all processed samples not stated
Approaches that could also have been used
  • The pipeline uses MACS for narrow-peak calling and SICER for broad-peak calling as fixed choices selected based on benchmarking literature
    Could also: HOMER, SPP, or SEACR could also be used for narrow or broad peak calling; for CUT&RUN specifically, SEACR was designed with low-background CUT&RUN data in mind — Different peak callers have distinct sensitivity/specificity trade-offs across antibody targets, sequencing depths, and assay types; offering configurable peak-caller selection would allow users to match the algorithm to their specific experimental context
  • Quality thresholds are fixed at ENCODE-consortium-derived cutoffs applied uniformly across all samples
    Could also: Within-experiment relative quality assessment (e.g., flagging samples that deviate more than a defined distance from the median of the batch) could also be used alongside absolute thresholds — Absolute thresholds derived from one set of experiments may not transfer equally to all organisms, antibodies, or sequencing depths; relative within-batch assessment can surface outliers even when all samples pass absolute thresholds
  • Motif analysis is performed with MEME Suite tools using EM-based probabilistic models on peak sequences
    Could also: HOMER motif analysis or RSAT could also be used; these use different background models and statistical frameworks for motif enrichment — Different motif tools vary in their background correction strategies and databases; using a complementary tool can confirm enriched motifs are not algorithm-specific and broadens the set of accessible PWM databases
  • The pipeline reports peak counts and quality metrics as the primary quantitative outputs without formal cross-sample differential enrichment testing
    Could also: Tools such as DiffBind, MAnorm, or DESeq2/edgeR applied to peak count matrices could also be incorporated to perform differential ChIP-Seq occupancy analysis between conditions — Many ChIP-Seq studies involve comparing binding profiles across conditions; integrating a differential binding module would extend the pipeline from quality-assessed peak calls to quantitative between-group comparisons within a single workflow
  • Super-enhancer identification uses the ROSE algorithm, which applies a fixed geometric inflection-point method to ranked enhancer signal
    Could also: DBsuper or other threshold-free ranking approaches could also be used; some approaches incorporate background normalization differently — The ROSE inflection-point cutoff is heuristic; alternative tools apply different statistical or geometric criteria for the boundary between typical and super-enhancers, and comparing results across methods can inform how sensitive conclusions are to the choice of algorithm
  • Alignment is performed with Bowtie (short-read aligner) without mention of multi-mapper handling strategy beyond blacklist region removal
    Could also: Bowtie2 or BWA-MEM could also be used; each has different handling of multi-mapping reads and gap-alignment capabilities relevant to repetitive genomic regions — The choice of aligner and multi-mapper policy affects peak calls in repetitive or low-complexity genomic regions; making the alignment strategy configurable or documenting the default multi-mapper handling would help users evaluate how this choice may influence their results
Software: Bowtie · SAMtools · BEDTools · MACS · SICER · ROSE · MEME Suite · Cromwell (WDL execution engine) · Docker · BAMToGFF (custom)

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
4
Impact: low
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

GSE138742 GEO in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GSE31558 GEO in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-35193506 (SEAseq)

Paper: Adetunji MO, Abraham BJ. SEAseq: a portable and cloud-based chromatin occupancy analysis suite. BMC Bioinformatics 2022. PMCID PMC8864840. Code: https://github.com/stjude/seaseq (WDL/Cromwell ChIP-seq pipeline: Bowtie → MACS/SICER → MEME → ROSE → BEDTools/custom QC). Data: GSE138742 (Study 1, LIN28B/ZNF143 ChIP-seq, BE2C/CHP134/Kelly cells, SRA SRP225129/PRJNA576946) and GSE31558 (Study 2, p53 ChIP-seq in IMR90, SRA SRP007953/PRJNA145681).

Reported results and their classification

Result Where Pipeline-derived? In scope? Reproduce how
Sample/Input read counts (54.9M, 45.6M, 35.1M, 47.2M, 35.7M, 7.5M, 11.8M) Table 3 Yes — input read totals reported by the pipeline's QC IN Count reads in the public SRA/ENA FASTQs (deterministic)
Runtime (e.g. 32:07) Table 3 No — wall-clock on St Jude Cloud / DNAnexus, hardware-specific OUT Not deterministic; depends on cloud instance
Cost ($26.38 …) Table 3 No — cloud billing, provider-specific OUT Not reproducible off-platform
Quality scores "Good/Excellent/Average" Fig 5a Yes — composite of FRiP/NSC/RSC/NRF/PBC etc. STRETCH Requires full pipeline run + SEAseq's scoring rubric
Enriched p53 motif matching orig. publication Fig 5b Yes — Bowtie→MACS→MEME on p53 data STRETCH (20%) Full pipeline on the 7.5M-read p53 sample
LIN28B/ZNF143 peak & motif biology text/figs Yes STRETCH Full pipeline

Plan (80/20)

  • 80% (primary, deterministic): Reproduce the Table 3 read counts by independently counting reads in the SRA/ENA FASTQs on «infra». This is the cleanest pipeline-input claim and directly auditable against the public data.
  • 20% (stretch): Run the SEAseq WDL pipeline (Cromwell + Singularity) on the smallest sample (p53, 7.5M reads) to reproduce peak calling + the enriched p53 motif (Fig 5b). Attempt only if a container runtime + hg19 supplemental data resolve on «our HPC»; otherwise document as partial with env_unresolvable.

Explicitly NOT attempted

  • Runtime / cost columns (cloud-specific, non-deterministic).
  • St Jude Cloud / DNAnexus platform execution (proprietary platform).
Figures / tables: TableFig 5b
rc_lin28b_sample
Reported
54.9M
Reproduced
54,894,988 (54.9M)
exact
rc_lin28b_input
Reported
45.6M
Reproduced
45,603,300 (45.6M)
exact
rc_znf143_sample
Reported
35.1M
Reproduced
35,108,844 (35.1M)
exact
rc_ydox_sample
Reported
47.2M
Reproduced
47,203,903 (47.2M)
exact
rc_ydox_input
Reported
35.7M
Reproduced
35,682,203 (35.7M)
exact
rc_p53_sample
Reported
7.5M
Reproduced
7,473,100 (7.5M)
exact
rc_p53_input
Reported
11.8M
Reproduced
11,760,480 (11.8M)
exact
p53_motif
Reported
identified enriched p53 motif (Fig 5b)
Reproduced
incomplete: bowtie2 align OK (4,636,707 reads pass), MACS2 blocked by bioconda ABI bug
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 94/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟢3. Location of the main deviation
🟢4. Cause of the deviation
🟢5. Derivability / plausibility
🟢6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +2

All 7 in-scope, deterministic Table 3 read counts reproduced exactly against the public SRA/ENA FASTPs (e.g. 54,894,988→54.9M), fully derivable from deposited data with no fabrication indicators — only 0.1M presentation rounding separates reported from reproduced. The single gap is on our side: the Fig 5b p53 motif was left unverified because MACS2 hit a bioconda ABI bug before peak-calling, an environment/packaging issue, not a paper or data defect. Runtime/cost columns were correctly treated as out-of-scope (cloud/hardware-specific). Overall a clean partial: exact where tested, incomplete on the analytical output for tooling reasons.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

136 k
tokens (I/O) · 11 M incl. cache
37 min
runtime · 4.69 CPU-h
6.5 GB
peak RAM
2 (1 failed)
HPC jobs
hummel
machine