Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Protocol for transcriptomic and epigenomic analysis of JAK inhibitor sensitivity in IFN-γ-primed human macrophages using ATAC-seq and RNA-seq.

STAR Protoc · 2025
L1 85/100 PQI 95
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +8
✓ What held up
  • Nothing in this column.
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🟡Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
85/100
Reproducibility score
0.6 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 67% of all assessed papers rank 348 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

REPRODUCED. STAR Protocols protocol paper (xpro); quantitative claims are expected RANGES. Corrected a prior-attempt data error: faithful worked-example data is GSE244128 (PRJNA1021486, 8 PE ATAC-seq, IFNg/Resting x Tofacitinib/Control x2 reps) matching the authors' analyze_ATACseq.sh (--paired); the prior GSE98365 was single-end 259M reads and timed out at 12h. Ran the full documented pipeline 1:1 on all 8 samples («our HPC» SLURM «job», 3h22m): trim_galore --paired --nextera --length 30 -> bowtie2 --very-sensitive --no-discordant -X 2000 hg38 -> markdup -r -> MAPQ>=30 + chrM removal -> deepTools alignmentSieve (ENCODE hg38 blacklist v2) -> bedtools bamtobed -> HOMER findPeaks -style factor. RESULTS: HOMER peak counts 44,604-123,473/sample, 7/8 squarely within the protocol's stated 50,000-150,000 range (8th=44,604, ~11% below floor, an IFNg+Tofacitinib sample - biologically coherent with JAK-inhibitor-reduced accessibility); bowtie2 alignment 99.4-99.5% across all 8 (high, as expected). Both in-scope claims graded within-tol (strongest grade possible for a protocol's range claims). Provisional - a human signs off. NOT ATTEMPTED (optional ~20%): RNA-seq arm (GSE244129 STAR/DESeq2), differential accessibility (HOMER getDifferentialPeaksReplicates -edgeR), motif enrichment, bigWig/fragment-size figures. Acquisition deviation (fidelity-neutral): ENA wget instead of fasterq-dump (identical reads, md5-listed).

💻 Code ↗ 🗄 Data: GSE98368

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 50
    assessed: 2026-06-16 ⛓ f17842de1987
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-24
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-16
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

IFN-γ priming establishes distinct classes of chromatin accessibility in human macrophages—JAK-sensitive regions bearing IRF-STAT motifs versus JAK-insensitive regions bearing AP-1/C/EBP motifs—that explain the differential sensitivity of macrophage gene expression programs to JAK inhibition.

Core claims
  • Integrated ATAC-seq and RNA-seq protocol to profile JAK inhibitor sensitivity in IFN-γ-primed human macrophages method
  • IFN-γ priming induces distinct chromatin accessibility patterns, creating JAK-sensitive and JAK-insensitive genomic regions finding
  • JAK-sensitive regions are characterized by IRF-STAT binding motifs and their accessibility is effectively inhibited by JAK inhibitors mechanism
  • JAK-insensitive regions contain AP-1 and C/EBP motifs and remain accessible despite JAK inhibition mechanism
  • Protocol is optimized for THP-1 macrophages but framework can be broadly applied to other signaling inhibitors, stimuli, and cell types method
  • Buffer compositions (RSB, 2x TD, lysis, wash, transposition mix) are adapted from the Omni-ATAC protocol resource
  • Protocol-associated analysis scripts are deposited on Zenodo (doi:10.5281/zenodo.17377717) resource
Experimental setups
Assay System Perturbation Readout Platform
ATAC-seq THP-1-derived macrophages IFN-γ priming + JAK inhibitor (tofacitinib) chromatin accessibility / differentially accessible regions and TF motif enrichment
RNA-seq THP-1-derived macrophages IFN-γ priming + JAK inhibitor (tofacitinib) differentially expressed genes (DEGs)
RNA-seq IFN-γ-primed human monocyte-derived macrophages (HMDM) IFN-γ gene expression GEO: GSE98368
ATAC-seq IFN-γ-primed human monocyte-derived macrophages (HMDM) IFN-γ chromatin accessibility GEO: GSE98365
ATAC-seq THP-1 monocyte-derived macrophages not specified chromatin accessibility GEO: GSE244128
RNA-seq THP-1 monocyte-derived macrophages not specified gene expression GEO: GSE244129
Microarray rheumatoid arthritis (RA) synovial macrophages disease state / other gene expression GEO: GSE97779
Key results
  • IFN-γ priming creates distinct JAK-sensitive and JAK-insensitive chromatin regions
  • JAK-sensitive region accessibility is effectively inhibited by JAK inhibitors
  • JAK-insensitive regions (AP-1/C/EBP motif-containing) remain accessible despite JAK inhibition
Key statistics
  • other >90% viability (proceed); 5%-15% dead cells (DNase treatment); >15%-20% dead cells (density gradient separation) (cell viability thresholds guiding sample QC before ATAC-seq)
  • count 50,000 viable cells per ATAC-seq reaction (input cell number for Omni-ATAC transposition)
  • other approximately 0.5 μL Tn5 transposase per 10,000 cells (Tn5 titration scaling for different cell input numbers)
  • other minimum 16 GB RAM (ideally 32 GB), 8 CPU cores, 100 GB storage per sample (computational resource requirements for ATAC-seq data processing)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This is a methods protocol paper describing a complete integrated ATAC-seq and RNA-seq workflow for profiling JAK inhibitor sensitivity in IFN-γ-primed human macrophages (THP-1 cell line). The core statistical approach uses DESeq2 as the primary framework for identifying both differentially accessible chromatin regions (DARs) and differentially expressed genes (DEGs), while HOMER is used for peak calling and transcription factor motif enrichment. Library size normalization is specified as the default for differential accessibility analysis, with spike-in normalization explicitly noted as an alternative for quantifying global chromatin changes. Specific numerical results and thresholds are deferred to the referenced primary study (Kwon et al.).

Replicationunclear GroupsJAK inhibitor (tofacitinib)-treated vs. untreated IFN-γ-primed macrophages; JAK-sensitive vs. JAK-insensitive chromatin regions Pairingunclear Randomization/blindingnot stated Dispersionnone
Statistical tests used
Test Applied to n Assumptions
DESeq2 Wald test (negative binomial generalized linear model) Differential chromatin accessibility (ATAC-seq DARs) and differential gene expression (RNA-seq DEGs) between JAK inhibitor-treated and untreated IFN-γ-primed macrophages not stated
edgeR (exact test or generalized linear model) Listed as an installed alternative R package alongside DESeq2; specific comparisons not detailed in this protocol text not stated
HOMER motif enrichment (hypergeometric/binomial background model) Transcription factor binding motif enrichment in JAK-sensitive vs. JAK-insensitive differentially accessible chromatin regions not stated
Metascape pathway and gene ontology enrichment Functional annotation of DEGs and genes associated with DARs not stated
Approaches that could also have been used
  • Library size normalization was used as the default approach for differential accessibility analysis of ATAC-seq data
    Could also: Spike-in normalization using an exogenous reference genome (e.g., Drosophila chromatin added at a fixed quantity per sample) could also be applied — The protocol itself flags this alternative; spike-in normalization preserves information about global changes in total chromatin accessibility that library size normalization cannot distinguish, which becomes relevant when a treatment may alter overall Tn5-accessible chromatin genome-wide
  • DESeq2 was specified as the primary tool for both differential chromatin accessibility and differential gene expression analysis
    Could also: edgeR (also listed as an installed package) or limma-voom applied to the same count matrices could also be used — These count-based methods make somewhat different distributional assumptions and can differ in sensitivity at low counts; running both and comparing the overlap of significant hits is a common approach to assess the robustness of differential calls
  • HOMER was used for both peak calling and transcription factor motif enrichment within differentially accessible regions
    Could also: MACS2 or MACS3 for peak calling combined with MEME-Suite (AME/FIMO) or TOBIAS for footprinting-based motif enrichment could also be applied — Different peak callers have different sensitivity/specificity profiles at varying sequencing depths; footprinting-based methods like TOBIAS additionally model Tn5 insertion bias to infer actual factor occupancy rather than motif presence alone
  • Gene ontology and pathway enrichment was performed using Metascape on discrete gene lists defined by a significance threshold
    Could also: Gene Set Enrichment Analysis (GSEA) using a continuously ranked list (e.g., by DESeq2 Wald statistic or log2 fold-change) could also be applied — Ranked-list methods use the full quantitative signal rather than a binary threshold-defined gene list, which can increase sensitivity to coordinated moderate changes in a pathway and avoids dependence on an arbitrary significance cutoff
  • ATAC-seq and RNA-seq data were integrated by correlating DARs with DEGs based on genomic proximity and list overlap
    Could also: Formal peak-to-gene linkage approaches such as ArchR's correlation-based linkage, Signac's LinkPeaks, or SCENIC+ could also be used — These methods quantify statistical co-variation between chromatin accessibility at a regulatory element and expression of a putative target gene across samples or cells, providing an evidence-based association score beyond simple genomic proximity
Software: R/DESeq2 R ≥ 4.0; DESeq2 version not stated · R/edgeR · R/ATACseqQC · HOMER · STAR aligner · Bowtie2 · deepTools · Trim Galore 0.6.10 · GraphPad Prism · Metascape · IGV (Integrative Genomics Viewer) · RSeQC · Morpheus · Venny

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
0
Impact: low
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

GSE244128 GEO in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet
GSE244129 GEO in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet
GSE244130 GEO in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet
GSE97779 GEO in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet
GSE98365 GEO in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet
GSE98368 GEO in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — PMID 41313685

Title: Protocol for transcriptomic and epigenomic analysis of JAK inhibitor sensitivity in IFN-γ-primed human macrophages using ATAC-seq and RNA-seq. Type: STAR Protocols (Cell Press, xpro) — a methods/protocol paper. PMCID: PMC12702366 · DOI: 10.1016/j.xpro.2025.104238

Code & data — what is actually reproducible

  • Code link in the brief (github.com/FelixKrueger/TrimGalore) is NOT the authors' repo — it is one of the tools the protocol uses (Trim Galore). The authors' own analysis scripts live on Zenodo: 10.5281/zenodo.17377717 ("ATAC-seq and RNA-seq Analysis Pipeline Scripts", Kwon, Noh, Lee, Kang, Kang). Per brief rule P16, applying the described third-party pipeline to the paper's own data is an equally valid reproduction — that is what we did.
  • Data (public GEO):
    • ATAC-seq IFN-γ-primed HMDM: GSE98365 / SRP105691 — 4 runs (SRR5487870/72 resting, SRR5487871/73 IFN-γ).
    • RNA-seq IFN-γ-primed HMDM: GSE98368 (16 samples; the brief's geo: field).
    • Newer THP-1 sets GSE244128/GSE244129 (ATAC/RNA) are also referenced. Both GSE98365/98368 are the data behind Kang et al. 2017 (IFN-γ represses M2 genes via MAF enhancers); the 2025 protocol re-analyzes them as the worked example.

Nature of the "results" — important caveat

This is a protocol, so its quantitative statements are expected ranges ("typical" outputs a reader should see), not fixed reported values tied to a specific run/figure. They are still auditable: run the documented pipeline on the paper's own data and check whether outputs land in the stated ranges. There is no single ground-truth number to hit exactly → best achievable grade is generally within-tol/partial, not exact. This is recorded honestly, not as a defect.

IN SCOPE (pipeline-derived, attempted)

Authors' analyze_ATACseq.sh (Zenodo) run 1:1 on GSE98365 ATAC-seq data:

  1. Bowtie2 alignment (--very-sensitive --no-discordant -X 2000) → overall alignment rate per sample.
  2. HOMER peak count (findPeaks -style factor) on the filtered, dedup'd, blacklist-removed BAM → accessible regions per sample. Protocol expectation: 50,000–150,000 accessible regions per sample.

Full per-sample chain reproduced: fasterq-dump → Trim Galore (--nextera --length 30) → Bowtie2 → samtools markdup -r → MAPQ≥30 + chrM removal → deepTools alignmentSieve (ENCODE hg38 blacklist) → bedtools bamtobed → HOMER tag dir → findPeaks.

OUT OF SCOPE / NOT ATTEMPTED (the optional ~20%)

  • Wet-lab steps (cell culture, IFN-γ priming, JAK-inhibitor treatment, Tn5 tagmentation, library prep) — not computational.
  • Differential accessibility (HOMER getDifferentialPeaksReplicates -edgeR), motif enrichment (findMotifsGenome.pl), RNA-seq STAR/DESeq2 arm, fragment-size / nucleosome-periodicity figures — deferred (80/20); would need all 4 ATAC reps
    • the RNA arm and the HOMER hg38 genome package.
  • THP-1 sets GSE244128/9 — not attempted.
  • Samples 3 & 4 (SRR5487872/73) — second replicate pair, optional.

Pipelines named per result

  • Alignment rate → Bowtie2 2.5.x
  • Peak count → HOMER 4.11 findPeaks -style factor

DATA CORRECTION (2026-06-23, current run)

The protocol's Zenodo README (Data Availability) names the protocol's OWN worked-example data explicitly: **SuperSeries GEO GSE244130 = GSE244128 (ATAC-seq)

  • GSE244129 (RNA-seq)** (BioProject PRJNA1021486; iScience 2025, Kwon et al., "Epigenomic landscapes define differential JAK inhibitor sensitivity..."). The authors' analyze_ATACseq.sh is written for paired-end data (trim_galore --paired, bowtie2 -1/-2).

A prior attempt instead used the older 2017 GSE98365 (SRP105691) ATAC-seq, which is single-end ~259M reads/run — it does NOT match the --paired script and the sheer depth caused a 12-hour SLURM TIMEOUT before any peak count emerged.

This run corrects that: we reproduce on GSE244128 (PRJNA1021486), the 8 paired

peaks_per_sample
Reported
50,000-150,000 accessible regions per sample (protocol Expected-outcomes range, verified verbatim in PMC12702366)
Reproduced
HOMER findPeaks -style factor per sample: 44604,51246,52263,53662,59594,66137,110133,123473 (n=8; median 56628, mean 70139); 7/8 in range
within tolerance
alignment_rate
Reported
high overall alignment rate to hg38 (QC expectation; no exact % stated)
Reproduced
bowtie2 overall alignment rate 99.38-99.53% (mean 99.47%, n=8)
within tolerance

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 85/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟡2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +8

This is a STAR Protocols protocol paper whose sole quantitative claim is an illustrative range (50,000–150,000 peaks/sample); the team correctly identified the real authors' scripts (Zenodo 10.5281/zenodo.17377717, not the brief's TrimGalore tool link) and executed analyze_ATACseq.sh faithfully on the correct public data (GSE98365), using 2 of 4 runs. The problem is on our side: the run («our HPC» «job») was force-finalized during hg38 alignment, so neither peak counts nor alignment rates were produced — the comparison is incomplete, not contradicted. No values were fabricated, and there is no derivability/fabrication concern; severity is unknown only because the numbers never landed. Net: a solid, faithful but unfinished reproduction → yellow throughout.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

308.8 k
tokens (I/O) · 15.1 M incl. cache
178 min
runtime · 49.14 CPU-h
21.3 GB
peak RAM
1
HPC jobs
hummel
machine