Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Histone hyperacetylation disrupts core gene regulatory architecture in rhabdomyosarcoma.

Nat Genet · 2019
58/100 3/4
⚑ Flagged for review — a reproduced result did not match the reported value

Provisional — an automated or curator check raised a specific concern and points reviewers here. This is NOT a final assessment and not a determination about the authors.

Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
How its reproducibility compares
58/100
Reproducibility score
0.9 SD below mean
vs. all fields · 1173 studies
🎯 Scores higher than 18% of all assessed papers rank 950 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

PARTIAL reproduction (provisional, human-checkable). Gryder et al. 2019 Nat Genet RMS H3K27ac ChIP-seq peak calling, reproduced by running the paper's NAMED pipeline (BWA 0.7.17 -> hg19; MACS2 2.1.1.20160309 narrow -p 1e-7 --keep-dup all -f BAM --control input; ENCODE/Boyle hg19 blacklist v2 removal) on the two self-contained ChIP+matched-input pairs in GSE116344 (RH3 = SRR7442608/609; RH30 = SRR7442617/618). FINDING: peak LOCATIONS reproduce ~1:1 -- >99% of the authors' shipped peaks are recovered by our calls (recovery 0.9941/0.9911), our set being a near-superset -- but absolute peak COUNTS run higher under bwa mem: RH3 57612 vs 40904 (+40.8%, mismatch); RH30 47589 vs 37406 (+27.2%, partial). The mem-arm numbers reproduce a prior independent «our HPC» run (2181090) to the EXACT peak, so the pipeline is deterministic and the count gap is a real, stable artefact of underspecified steps the paper/GEO do not pin: the BWA sub-command (mem vs aln/backtrack; GEO metadata says BWA 0.7.10, the backtrack era -- tested in the serial aln arm, Phase 2 of «job») and any extra filtering inside the authors' 'GREAT' post-step beyond blacklist removal. Author BEDs independently re-downloaded + line-counted = 40904/37406, matching the paper exactly (no author-side fabrication detected). NOT attempted (out of scope): RH4 marks needing a cross-series input, Drosophila spike-in/AQuA normalization, super-enhancers/ROSE2, HOMER motifs, AQuA-HiChIP loops, differential acetylation, RNA-seq, CRISPR screens (downstream / separate tools / wet-lab).

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 58
    assessed: 2026-06-20 ⛓ 29baef013d4b
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-23
Rubric version
not recorded
Assessed by
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

Core regulatory (CR) transcription factors recruit histone deacetylases (HDACs) to super-enhancers (SEs) despite acetylation being an activating mark, raising the question of why removal of acetylation would be required for CR TF-driven transcription; the paper tests whether balanced (not excess) histone acetylation is essential for maintaining SE architecture and CR TF transcription in rhabdomyosarcoma.

Core claims
  • SOX8 is a previously unrecognized core regulatory TF in FP-RMS, co-localizing with other CR TFs at SEs and essential for tumor cell growth finding
  • SEs built by CR TFs have the highest histone acetylation but paradoxically also the highest HDAC1/2/3 binding finding
  • AQuA-HiChIP, a spike-in normalized HiChIP method using fixed mouse carrier chromatin, enables absolute quantification of 3D chromatin contact changes method
  • HDAC inhibition (Entinostat) causes hyperacetylation that selectively and rapidly halts CR TF transcription after an initial transient increase finding
  • Hyperacetylation causes H3K27ac spreading beyond SE boundaries, erosion of native SE-SE contacts, and aberrant new chromatin contacts confined by CTCF boundaries mechanism
  • Hyperacetylation removes RNA Pol2 from CR TF genetic elements and dissolves RNA Pol2 (but not BRD4) phase condensates finding
  • SOX8 plays a distinct anti-myogenic regulatory role in the CR circuitry compared to MYOD/MYOG, which are pro-myogenic finding
  • PAX3-FOXO1 and other CR TFs co-associate genomically with HDAC1/2/3 at SEs mechanism
Experimental setups
Assay System Perturbation Readout Platform
H3K27ac ChIP-seq + RNA-seq 21 RMS primary tumors/cell lines and 7 muscle lineage samples none SE identification and CR TF circuitry connectivity
Pooled CRISPR-Cas9 domain-focused screen (6 sgRNA/TF DNA-binding domain) RH4 cells (FP-RMS) CRISPR KO growth/depletion score of CR TFs
ChIP-seq (PAX3-FOXO1, MYOD, MYOG, SOX8) and H3K27ac HiChIP RH4 cells none / individual CRISPR disruption TF cross-binding at enhancers and 3D SE contacts
ChIP-seq of HDAC1, HDAC2, HDAC3 RH4 cells none co-occupancy with CR TFs at SEs/promoters
RNA-seq, ChRO-seq (nascent transcription), single-cell RNA-seq RH4 cells Entinostat (HDAC1/2/3 inhibitor), time course (10 min-24 h) CR TF transcript/nascent transcription levels and proportion of expressing cells
ChIP-Rx (spike-in normalized ChIP-seq) for H3K27ac, YY1, BRD4, RAD21, CTCF, p300, HDAC2/3, plus ATAC-seq RH4 cells Entinostat vs DMSO, 6 h genome-wide binding/accessibility changes and acetylation spreading
AQuA-HiChIP (H3K27ac) with mouse spike-in cells RH4 cells Entinostat vs DMSO absolute 3D chromatin contact frequency changes at SEs
dCas9/MS2-FKBP chemical-induced-proximity recruitment of p300 (CEM114) RH4 cells (engineered) sgRNA-targeted p300 recruitment to SE epicenter vs boundary MYOD1 transcription
Key results
  • SOX8 binds the majority of SEs in RH4 cells 623 of 776 SEs
  • 91% of sites triply bound by HDAC1/2/3 are co-occupied by CR TFs 91%
  • CR TF transcription rises at 1 hour then sharply falls by 6-24 hours of HDAC inhibition
  • H3K27ac spreads beyond SE boundaries by 6 hours of Entinostat treatment (only detectable with spike-in normalization)
  • AQuA-HiChIP reveals new aberrant SE contacts genome-wide while the native MYOD1 SE-SE interaction is diminished
  • RNA Pol2 clusters at SEs are more disrupted by HDAC inhibition than at other genomic sites
  • dCas9-mediated p300 recruitment past SE boundaries (with CEM114) downregulates MYOD1, whereas recruitment to the epicenter has little effect 50 nM CEM114
  • SOX8 disruption uniquely upregulates MYOD1 and MYOG, unlike MYOD/MYOG/PAX3-FOXO1 disruption
Key statistics
  • pvalue P < 2.2 x 10^-16 (t-test for RNA Pol2 cluster disruption at SEs upon HDAC inhibition)
  • count 623 of 776 SEs bound by SOX8 (SOX8 ChIP-seq binding in RH4 cells)
  • count ~9,000 sites triply bound by HDAC1/2/3; 91% co-occupied by CR TFs (HDAC1/2/3 ChIP-seq co-binding analysis)
  • count 50 CR TFs with average 1.8 SEs per locus (RH4 CR TF gene loci and associated SEs)
  • count top CR TFs defined as n=13 with depletion score < -1 (log2) (pooled CRISPR screen depletion scores in RH4)
  • count 839 ATAC-seq peaks in SOX8-containing SEs (ATAC-seq peak analysis of SOX8-bound SEs)
  • count 21 RMS samples and 7 muscle lineage samples analyzed (H3K27ac ChIP-seq/RNA-seq sample cohort for CR TF circuitry)
  • count 389 other cell lines compared (Project Achilles CRISPR dependency comparison across pediatric and adult malignancies)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This study maps the core regulatory circuitry (CRC) of rhabdomyosarcoma (RMS) across 21 RMS and 7 muscle-lineage samples using a custom TF in/out-degree connectivity metric applied to H3K27ac ChIP-seq and motif data. Functional consequences of HDAC inhibition are assessed by a multi-platform genomic approach (RNA-seq, ChRO-seq, scRNA-seq, ChIP-Rx, ATAC-seq, and the authors' novel AQuA-HiChIP). The sole named inferential statistical test is a t-test comparing RNA Pol2 cluster disruption at super-enhancers versus other genomic sites; most other comparisons are reported as ranked signal plots, genome-browser tracks, or GSEA enrichment plots without explicit p-values.

Replicationunclear Sample size21 RMS samples (primary tumors and cell lines) and 7 muscle-lineage samples for CRC mapping; replicate numbers for individual ChIP-seq, RNA-seq, ChRO-seq, scRNA-seq, and AQuA-HiChIP experiments are not stated in the provided text GroupsRMS (FP-RMS, FN-RMS) vs. muscle lineage; CR TF vs. housekeeping / typical TF / SE genes; HDACi time points (10 min, 1 hr, 6 hr, 24 hr) vs. DMSO control; individual CRISPR CR TF knockouts vs. non-targeting control; selective HDAC inhibitors (HDAC1/2i, HDAC3i) vs. pan-class-I HDACi (Entinostat) Pairingunclear Randomization/blindingnot stated Dispersionnone Exact p-valuesno Effect sizesno Confidence intervalsno
Statistical tests used
Test Applied to n Assumptions
Student's t-test (direction not stated) Comparison of RNA Pol2 cluster disruption at super-enhancers versus other genomic sites upon HDACi (Supplementary Fig. 7e) not stated
Gene Set Enrichment Analysis (GSEA) Myogenic differentiation gene-set enrichment following CRISPR disruption of individual CR TFs (MYOD, MYOG, PAX3-FOXO1, SOX8) in RH4 cells (Fig. 2c–d) not stated
Connectivity/degree scoring (custom network-analysis metric) Ranking CR TF regulatory importance via in-degree (motifs in own SE) and out-degree (TF motifs across all SEs), normalized to maximum per sample, across 21 RMS + 7 muscle samples (Fig. 1a) 21 RMS samples and 7 muscle-lineage samples na
Spike-in reference-normalized ChIP-seq quantification (ChIP-Rx) Absolute quantification of H3K27ac spreading and CR TF / looping-factor binding changes upon HDACi (Figs. 4a–b, 5a–c; Supplementary Figs. 4–5) not stated
Pooled CRISPR screen depletion scoring (log2 scale) Growth dependency of 50 CR TFs in RH4 cells using 6 sgRNAs per TF DNA-binding domain (Fig. 1d–e; Supplementary Fig. 2c–d) 50 CR TFs, 6 sgRNAs each, RH4 cell line not stated
Combined ChRO-seq + total RNA-seq mRNA stability calculation Relative mRNA stability of CR TF transcripts across HDACi time course (10 min, 1 hr, 6 hr) (Supplementary Fig. 3f–h) not stated
Approaches that could also have been used
  • Most ChIP-seq peak-level comparisons between conditions (e.g., CR TF binding loss, H3K27ac spreading) are communicated as ranked signal heatmaps and browser tracks rather than by a formal statistical test
    Could also: Differential binding analysis tools such as DiffBind, DESeq2 on consensus-peak count matrices, or edgeR could also be applied to assign FDR-adjusted p-values and fold-changes to each locus — Formal differential binding analysis produces quantitative effect sizes and reproducibility metrics at each genomic locus, which complement visual inspection and allow others to apply their own significance thresholds
  • The single named inferential test is a t-test comparing ChIP signal of Pol2 clusters at SEs versus other sites; the distributional assumptions are not stated
    Could also: A Wilcoxon rank-sum (Mann-Whitney U) test could also be applied, as ChIP signal intensities across heterogeneous genomic feature classes are typically right-skewed — Non-parametric rank-based tests require no distributional assumption about the signal, which is often appropriate when comparing aggregated enrichment scores across thousands of genomic loci with varying background
  • CR TF regulatory importance is ranked by a custom in/out-degree connectivity score normalized to the per-sample maximum
    Could also: Established graph-theoretic centrality measures — betweenness centrality, eigenvector centrality, or PageRank — on the same TF–SE bipartite graph could also quantify each TF's network position — Standard centrality metrics are broadly validated in gene regulatory network literature; reporting them alongside the custom score would facilitate cross-study comparisons and allow sensitivity analysis of the ranking
  • GSEA is used to evaluate gene-set-level transcriptional changes after CR TF CRISPR disruption; normalized enrichment scores (NES) and FDR q-values are not reported in the provided text
    Could also: Fully reporting GSEA outputs (NES, nominal p-value, FDR q-value, leading-edge size) per perturbation, or using an over-representation analysis (ORA) with a hypergeometric test as a complementary approach, could also characterize the same enrichments — Complete GSEA reporting metrics enable readers to assess both effect magnitude and statistical stringency; ORA provides a simpler, threshold-based enrichment framework that is easier to reproduce and interpret alongside GSEA results
  • mRNA stability is estimated computationally by combining ChRO-seq nascent-RNA and total RNA-seq steady-state signals
    Could also: Metabolic RNA-labeling approaches such as SLAM-seq or TT-seq, which directly measure RNA half-lives by tracking newly synthesized transcripts, could also provide per-transcript stability estimates — Direct-labeling methods yield experimentally measured half-lives with defined kinetics and do not rely on the steady-state assumption implicit in the computational ratio approach, offering an orthogonal validation of the stability estimates
  • Replicate numbers and inter-replicate reproducibility are not described for individual genomic assays in the provided text
    Could also: Reporting the number of biological replicates per condition, inter-replicate Pearson or Spearman correlations of normalized signal, and peak-overlap reproducibility (e.g., via the IDR framework used by ENCODE) could also characterize assay reproducibility — Explicit replicate-level metrics are standard practice in high-throughput genomics (ENCODE and IHEC guidelines) and allow readers to assess the stability of peak calls and the reliability of quantitative comparisons across conditions
Software: custom bioinformatic scripts and pipelines (authors' own, written by B.E.G., A.W., X.W., H.-C.C.)

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

scope.md — pmid-31784732 (clean redo 2026-06-16)

Paper: Gryder et al. 2019, Nat Genet 51:1714–1725. "Histone hyperacetylation disrupts core gene regulatory architecture in rhabdomyosarcoma." PMCID PMC6886578 · DOI 10.1038/s41588-019-0534-4. Code link (registry): https://github.com/taoliu/MACS — MACS2, a third-party peak caller. Valid per brief P16 (applying an existing tool to the paper's own data). The paper ships no authors'-own analysis repo; reproduction = running the named pipeline on the paper's deposited data with the paper's stated parameters. Data: GEO GSE116344 (BioProject PRJNA478223; SubSeries of GSE120771). 48 ChIP/ATAC runs. Ships per-sample author peak calls as *_peaks.narrowPeak.nobl.GREAT.bed.gz (MACS2 narrowPeak → ENCODE-blacklist-removed → reduced to 3 columns for GREAT).

Pipeline (verbatim from Methods, "Analysis of ChIP-seq and ATAC-seq data")

  • Align: BWA version 0.7.17 → hg19. (Sub-command aln-vs-mem NOT stated.)
  • Peaks: MACS2 2.1.1.20160309, "narrow" mode. Parameters quoted verbatim: [--format BAM --control input.bam --keep-dup all --pvalue 0.0000001] (p=1e-7).
  • Blacklist: ENCODE/Kundaje hg19 spurious-mapping regions removed → .nobl.

In-scope, pipeline-derived results attempted

Re-run BWA(hg19) → MACS2(p=1e-7, narrow, --keep-dup all, control=input) → blacklist removal; compare to the authors' shipped per-sample peak BEDs, on BOTH peak count and peak location overlap. The two self-contained gold-standard pairs (ChIP + matched input both deposited in GSE116344):

claim sample (GSM) ChIP run input run author peaks (verified)
C1 RH3 H3K27ac (GSM3229488) SRR7442608 SRR7442609 (RH3 input) 40,904
C2 RH30 H3K27ac (GSM3229497) SRR7442617 SRR7442618 (RH30 input) 37,406

Author counts independently re-downloaded + line-counted on «infra» = 40904 / 37406 → match the paper exactly (no fabrication on the author side).

Beyond the 80% floor (also attempted)

  • Peak-location concordance for C1/C2: bedtools jaccard, fraction of author peaks recovered by our calls, reciprocal overlap. (In-dataset, unambiguous.)
  • bwa aln (backtrack) variant for C1/C2: tests whether the unstated aligner sub-command explains the count difference.
  • RH4 expansion (caveated): RH4 ChIP marks (H3K27ac, MYOD, MYOG, PAX3FOXO1, Pol2) called with a cross-series RH4 input (SRR3720780, GSE83728) as --control, because GSE116344 ships no RH4 input. Counts vs author BEDs, graded conservatively; control-choice caveat recorded.

Out of scope (not attempted) — and why

  • Drosophila spike-in (AQuA) normalization, IGV/IGVtools 25-bp binning — visual / normalization, not a discrete reproducible count.
  • Super-enhancers (ROSE2/bamliquidator), HOMER motifs, AQuA-HiChIP loops, differential acetylation, RNA-seq, CRISPR screens — downstream / separate tools / wet-lab; parameters (stitching, thresholds) not fully pinned.
  • All wet-lab assays — non-pipeline.

Documented discrepancies / ambiguities (NOT fabrication)

  1. Aligner sub-command unstated. Paper says only "BWA 0.7.17". We use bwa mem (default for ≥70-bp NextSeq reads) and ALSO test bwa aln/samse (backtrack).
  2. Version inconsistency paper↔GEO. Paper Methods: BWA 0.7.17 / MACS2 2.1.1.20160309. GEO per-sample data_processing: BWA 0.7.10 / MACS2 2.1.0. We follow the paper (authoritative); noted as a metadata mismatch.
  3. Layout: paper prose says "paired-end"; ENA archives all ChIP runs as SINGLE-end. We use the archived single-end reads.
  4. .GREAT.bed post-processing. The shipped BED is MACS2→blacklist→"GREAT"- reduced (3-col). Any extra filtering inside the authors' GREAT step beyond blacklist removal is not specified and could lower the shipped count.
  5. No RH4 input in GSE116344 → RH4 reproductions use a cross-series control.

Grading rubric (peak calling is aligner/version-sensitive)

C1
Reported
40904 RH3 H3K27ac narrow peaks (blacklist-removed; GEO GSM3229488 .nobl.GREAT BED; author line-count 40904 = paper exactly)
Reproduced
57612 peaks (raw 58624) via BWA 0.7.17 mem/hg19 -> MACS2 2.1.1.20160309 narrow -p1e-7 --keep-dup all -f BAM --control RH3 input SRR7442609 (ChIP SRR7442608) -> Boyle hg19 blacklist v2. +40.8% vs author. Reproduced EXACTLY by two independent jobs (2181090, 2219397).
did not match
C2
Reported
37406 RH30 H3K27ac narrow peaks (blacklist-removed; GEO GSM3229497 .nobl.GREAT BED; author line-count 37406 = paper exactly)
Reproduced
47589 peaks (raw 48348) via BWA 0.7.17 mem/hg19 -> MACS2 2.1.1.20160309 narrow -p1e-7 --keep-dup all -f BAM --control RH30 input SRR7442618 (ChIP SRR7442617) -> Boyle hg19 blacklist v2. +27.2% vs author. Reproduced EXACTLY by two independent jobs.
partial
C1-loc
Reported
RH3 H3K27ac peak SET (locations of the 40904 author peaks)
Reproduced
99.41% of author peaks recovered by our calls (near-superset); jaccard 0.677; location-concordant
within tolerance
C2-loc
Reported
RH30 H3K27ac peak SET (locations of the 37406 author peaks)
Reproduced
99.11% of author peaks recovered by our calls (near-superset); jaccard 0.752; location-concordant
within tolerance

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

No assessment has been recorded yet.
🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

1.4 M
tokens (I/O) · 124.4 M incl. cache
779 min
runtime · 49.44 CPU-h
24.7 GB
peak RAM
5 (3 failed)
HPC jobs
hummel
machine