Transcription-coupled and epigenome-encoded mechanisms direct H3K4 methylation.
The main results reproduced: recomputed values matched the published ones within tolerance.
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
REPRODUCED. Oya et al. 2022 Nat Commun (H3K4 methylation in Arabidopsis) is described well enough to reproduce: the authors' repo Satoyo08/Arabidopsis_H3K4me1@11015fe is a self-contained figure-reproduction package shipping all derived intermediates + reference tables. Layer-A re-ran the authors' own base-R computations on those intermediates on a «our HPC» compute node (n093) and regenerated all 11 in-scope headline numbers (C1-C11), deterministically (identical SHA256 across re-runs). Strong/exact matches: Fig1f hypergeometric under-representation (expected overlap 328.37 == paper's n=328; all pairs p<1e-4), Fig5 H3K4me1-transcription Spearman directions + Welch p-values (atx1/2 stronger p=0.003, atxr7 weaker p=0.012), Fig3 RF top-importance features (H3K36me3 #1 for ATX1/ATX2; RNAP2/TES-channel #1 for ATXR7). High-AUC claims confirmed quantitatively (RF 0.88-0.98; SVM 0.82-0.85), ATX1/ATX2 SVM weight correlation r=0.70, TATA-negative/GAGA-positive weight signs, and sppRNA over-representation (p<1e-12) all reproduced. Two items graded partial: C1 (ATX1/2 summit-TSS 'most peaks' window holds as modal/peak localization, but strict in-window fraction is ~30% due to a downstream tail) and C11 (mouse Fig7 localization pattern reproduced; RF AUCs high but shipped mouse objects use opposite class-label polarity, so reported MW values are ~1-AUC). NOT attempted: Layer-B from-raw realignment of PRJNA732996 (162 runs / 3.36e9 reads) and wet-lab/figure-rendering items (out of scope). No fabrication indicators: every paper-printed value checked is derivable from the deposited data via the authors' code. Dataset PRJNA732996 profiled: 162 open Illumina A. thaliana runs (145 ChIP / 15 RNA / 1 ATAC / 1 WGBS) whose 4 library strategies map cleanly onto the analyses (grade A, delivers_promised=yes).
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-26
- Rubric version
- not recorded
- Assessed by
- —
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetIt is unresolved whether H3K4 methylation (H3K4me1/2/3) is a cause of transcription or merely a downstream consequence of it, and in Arabidopsis no methyltransferase responsible for H3K4me1 had been identified, so the paper asks which enzymes direct H3K4me1 and by what mechanism (transcription-coupled vs. epigenome/genome-encoded) they are targeted to chromatin.
- ★ ATX1, ATX2, and ATXR7 redundantly mediate H3K4 monomethylation genome-wide finding
- ★ ATXR3 mediates H3K4me3 and ATX3/ATX4/ATX5 mediate H3K4me2, largely independent of H3K4me1 control finding
- ★ ATXR7 localizes near transcription termination sites and colocalizes with RNA Polymerase II (including Ser2/Ser5-phosphorylated forms), indicating transcription-coupled recruitment finding
- ★ ATX1 and ATX2 localize near transcription start sites and their binding is predicted by chromatin modifications (H3K36me3, H2Bub, H4K16ac) rather than by RNAP2 abundance finding
- ★ DNA sequence (6-mer) composition of TSS regions predicts ATX1/ATX2 binding, but TTS sequence does not predict ATXR7 binding finding
- ★ Target genes of the ATX1/2/R7, ATX3/4/5, and ATXR3 groups are mutually exclusive finding
- ★ Two distinct modes of H3K4 methyltransferase chromatin targeting (transcription-coupled vs. epigenome/genome-encoded) generalize to other Arabidopsis and mouse H3K4 methyltransferases mechanism
- Random forest and linear SVM machine learning models were used to identify chromatin and genomic determinants of methyltransferase localization method
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| ChIP-seq (H3K4me1/me2/me3) | Arabidopsis thaliana, single and multiple atx1-5/atxr7 mutants | KO (T-DNA mutants) | genome-wide H3K4me1/2/3 levels per gene | — |
| Western blot | Arabidopsis, bulk histone extract from atx1/2/r7 and other mutants | KO | H3K4me1/me2/me3 abundance | — |
| ChIP-seq (H3K4me0) | Arabidopsis, atx1/2/r7 vs WT | KO | unmethylated H3K4 signal at ATX1/2/R7-marked genes | — |
| ChIP-seq of FLAG-tagged proteins | Arabidopsis transgenic lines expressing FLAG-ATX1, ATX2, ATXR7 | tagged transgene expression | genome-wide protein localization (TSS/TTS positioning) | MACS2 peak calling |
| Random forest machine learning | Arabidopsis genomic/chromatin feature dataset | none (computational classification) | feature importance (mean decrease in Gini) predicting ATX1/ATX2/ATXR7-bound vs unbound genes; AUC | — |
| Linear support vector machine on 6-mer DNA sequence vectors | Arabidopsis TSS/TTS DNA sequences | none (computational classification) | prediction accuracy (ROC/AUC) of ATX1/ATX2-bound TSS and ATXR7-bound TTS sequences; 6-mer weights | — |
| ChIP-seq (RNAP2 total and Ser2/Ser5-phosphorylated CTD) | Arabidopsis | none | RNAP2 occupancy at ATXR7-bound vs unbound genes | — |
| ChIP-seq (H3K36me3, H2Bub, H4K16ac) | Arabidopsis, hub1/hub2, ashh2/ashh3, fld mutants | KO (hub1, hub2, ashh2, ashh3, fld) | chromatin mark levels and ATX1/ATX2 localization changes | — |
- ▼ atx1/2/r7 triple mutant shows substantial genome-wide loss of H3K4me1 while H3K4me2/3 are largely unaffected
- ▼ atxr7 single mutant shows H3K4me1 loss concentrated in the 3' half of gene bodies, while atx2 mutant shows a broader decrease
- ▲ ATXR7-bound genes show higher levels of total RNAP2 and Ser2/Ser5-phosphorylated RNAP2 compared to unbound genes
- – H3K36me3 is the top random-forest predictive feature for ATX1 and ATX2 localization, whereas RNAP2 features are the top predictor for ATXR7
- ▼ Pairwise overlaps between ATX1/2/R7-, ATX3/4/5-, and ATXR3-marked genes are significantly fewer than expected by chance p<1e-4
- – Linear SVM models trained on 6-mer DNA sequence achieve good prediction accuracy for ATX1- and ATX2-bound TSS regions but not for ATXR7-bound TTS regions
- – atx1/2/r7 mutation causes concomitant loss of H3K36me3 along with H3K4me1, while the ashh2 mutant (which strongly reduces H3K36me3) does not reduce H3K4me1
- ▼ hub2 mutation (H2Bub loss) decreases ATX2 localization at genes predicted to lose ATX2 by in silico H2Bub-loss simulation, and genes losing ATX1/2 binding tend to lose H3K4me1
- pvalue p < 1e-4 (hypergeometric test for under-representation of overlap among ATX1/2/R7-, ATX3/4/5-, and ATXR3-marked genes)
- count n = 3000 (top genes defined as 'marked genes' for each methyltransferase group (ATX1/2/R7, ATX3/4/5, ATXR3) based on H3K4me loss/signal)
- other TSS region defined as -150 to +300 bp from TSS; TTS region defined as -200 to +200 bp from TTS (genomic windows used to define ATX1/ATX2 and ATXR7 ChIP-seq peak-enriched regions)
- count 5 repeats (random forest and linear SVM models trained with 5 independent repeats/cross-fold validations, each with independently sampled negative gene sets)
- other ROC/AUC calculated on held-out Chromosome 5 test data (evaluation scheme for random forest and linear SVM localization prediction models)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
The study is a genomics/epigenomics investigation combining ChIP-seq profiling of histone methylation and transcription factors in Arabidopsis wild-type versus mutant lines, with computational classification (random forest and linear SVM) used to identify chromatin and sequence features predictive of methyltransferase binding sites. Statistical treatment centers on hypergeometric tests for gene-set overlap, machine-learning model performance reported via ROC/AUC with repeated-training error bars (SD), and a Pearson correlation between model weights; results are otherwise presented largely as descriptive genome-wide plots (metaplots, heat maps, violin plots) rather than classical hypothesis tests. Based on the visible text, a detailed statistics subsection of the Methods was not included, so some procedural details (e.g., software packages, exact model configuration) are not fully specified here.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Hypergeometric test | Testing under-representation (mutual exclusivity) of overlaps among ATX1/2/R7-, ATX3/4/5-, and ATXR3-marked gene sets (Fig. 1f) | n = 328 | not stated |
| Random forest classification (feature importance via mean decrease in Gini; performance via ROC/AUC) | Distinguishing ATX1-, ATX2-, and ATXR7-bound genes from unbound genes using chromatin/genomic features (Fig. 3a–f) | Top n = 3000 bound genes per protein; 5 repeats of training with independently chosen negative samples; Chr. 5 held out as test data | not stated |
| Linear support vector machine (lSVM) classification (performance via ROC/AUC) | Distinguishing ATX1- and ATX2-bound TSS regions and ATXR7-bound TTS regions from unbound regions using 6-mer sequence frequency vectors (Fig. 4a–c) | 5-fold cross-validation; Chr. 5 held out as test data | not stated |
| Pearson correlation coefficient | Correlation of averaged SVM feature weights between ATX1 and ATX2 models (Fig. 4d) | — | not stated |
-
Three pairwise hypergeometric tests for gene-set overlap were reported together with a shared p < 1e-4 threshold, without a stated multiple-comparison correction.↳ Could also: A multiple-testing correction such as Bonferroni or Benjamini-Hochberg FDR applied across the family of pairwise overlap tests — When several related tests are reported together, an explicit correction step conveys how the family-wise or false-discovery error rate was controlled, which can complement a fixed significance threshold.
-
Hypergeometric test significance was reported as a threshold (p < 1e-4) rather than as exact p-values.↳ Could also: Reporting the exact computed p-value for each pairwise comparison — Exact p-values allow readers to gauge the precise strength of evidence and to perform their own multiple-comparison adjustments if desired.
-
Random forest and SVM model performance (ROC/AUC, feature importance) was summarized as an average with SD across 5 repeats of training or 5-fold cross-validation.↳ Could also: Reporting a 95% confidence interval (e.g., via bootstrap resampling) for AUC or feature-importance estimates, in addition to or instead of SD across repeats — A confidence interval directly communicates the precision of the estimated model performance and can be more interpretable than SD when the number of repeats is small.
-
Model generalization was assessed using a single held-out chromosome (Chr. 5) as test data.↳ Could also: Genome-wide k-fold cross-validation (e.g., leave-one-chromosome-out across all chromosomes, or standard k-fold across the full gene set) — Testing on multiple held-out folds spanning the genome can provide a distribution of performance estimates and reduce the influence of any single chromosome's characteristics on the reported AUC.
-
Pearson's correlation coefficient was used to relate SVM weights between the ATX1 and ATX2 models (Fig. 4d).↳ Could also: A rank-based measure such as Spearman's correlation, alongside a reported p-value or confidence interval for the correlation — Spearman's correlation is robust to non-linear monotonic relationships and outliers in weight distributions, and accompanying inferential statistics would quantify the certainty of the association.
-
Differences in chromatin feature abundance (e.g., RNAP2, H3K36me3, H2Bub, H4K16ac levels) between bound and unbound gene groups were shown descriptively via violin plots (Fig. 3g–i) without an accompanying formal statistical test.↳ Could also: A nonparametric two-group test such as the Mann-Whitney U test, or a permutation test, to formally compare the distributions shown in the violin plots — Adding a formal test alongside the descriptive violin plots would provide a quantitative statistic and p-value to accompany the visual comparison of distributions.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
scope.md — pmid-35953471
Paper: Oya, Takahashi, Takashima, Kakutani, Inagaki. Transcription-coupled and epigenome-encoded mechanisms direct H3K4 methylation. Nat Commun 13, 4521 (2022). PMID 35953471 · PMC9372134 · DOI 10.1038/s41467-022-32165-8.
Repo: https://github.com/Satoyo08/Arabidopsis_H3K4me1 (MIT, default branch main,
analysis commit 11015fe3904cefb4cbfe98a1b1f28774d724184b, 2026-05-29). Authors' own
code = P16 "own repo".
Data: NCBI BioProject PRJNA732996 (Illumina ChIP-seq / mRNA-seq, A. thaliana; plus reuse of mouse mm9 ChIP-seq for Fig.7 and Chen 2017 Arabidopsis for Fig.5).
Key insight for reproducibility
The repo is a self-contained figure-reproduction package: alongside the per-figure
R scripts it ships ALL derived intermediate objects AND reference files
(data/refs/araport11_all_sorted.bed, cDNA lengths, etc.). Therefore the paper's
headline pipeline-derived numbers can be regenerated from shipped data using the
authors' own R code, without re-aligning raw reads. The from-raw alignment
(PRJNA732996 → bowtie → MACS2 → RPM/RPKM) is a deeper, optional validation layer.
Two reproduction layers:
- Layer A (shipped-data, fast floor + well beyond): run the authors' R scripts on the shipped derived tables to regenerate the reported statistics. No raw reads.
- Layer B (from-raw, heavy, «our HPC»): download PRJNA732996 FASTQs on front1, align
with
bowtie -v 2 -m 1to TAIR10, MACS2-g 1.3e8 [-q 0.3 --nomodel], regenerate RPM/RPKM per gene, and check they match the shippedChIP1_RPM.txt/ChIP1_RPKM.txt.
IN SCOPE (pipeline-derived results)
| id | result (paper location) | reported | pipeline | layer | shipped inputs |
|---|---|---|---|---|---|
| C1 | ATX1, ATX2 peak summits localize −150 to +300 bp from TSS (Fig.2d) | window −150..+300 from TSS; mode near TSS | MACS2 --nomodel -q 0.3 + bedtools closest, summit→TSS distance histogram |
A | data/Figure2/covaris_re_ATX{1,2}_..._closest_TSS_hist |
| C2 | ATXR7 peak summits localize −200 to +200 bp from TTS/TES (Fig.2d) | window −200..+200 from TES | same, summit→TES | A | data/Figure2/covaris_re_ATXR7_..._closest_TES_hist |
| C3 | ATX1/2/R7-, ATX3/4/5-, ATXR3-marked gene groups overlap significantly less than expected, p<1e-4 (Fig.1f, hypergeometric) | p < 1e-4; ~constant expectation | top-3000 differential gene sets + phyper |
A | ChIP1_RPKM.txt, ChIP1_RPM.txt, araport11_all_sorted.bed |
| C4 | H3K4me1–transcription Spearman correlation strengthens in atx1/2 (P<0.05), weakens in atxr7 (P<0.05) vs WT (Fig.5) | direction + significance | per-gene FPKM vs RPM, cor.test(method=spearman), t.test across reps |
A | PerGene_RNA8_araport.txt, atx12r7_me3H3_RPM.txt, cDNA_length_araport.txt |
| C5 | Random-Forest models predict ATX1/ATX2/ATXR7 target genes with high AUC (Fig.3d–f) | AUC values printed in panels (computed round(mean(auc),3)) |
trained RF objects → ROCR AUC | A | data/Figure3/ATX{1,2}_with_K36me3_ChIP8_rep, ATXR7_..._FLD |
| C6 | RNAP2 (Pol II) is the top importance feature for ATXR7; H3K36me3/H2Bub/H4K16ac top for ATX1/2 (Fig.3a–c) | feature ranking (mean decrease Gini) | RF $Importance colMeans ranking |
A | same RF objects |
| C7 | Linear-SVM on 6-mer TSS sequence predicts ATX1/ATX2 targets (AUC, Fig.4a–c) | auc_score stored in ROC file headers | lSVM ROC (precomputed FPR/TPR + AUC) | A | data/Figure4/ROC/*_FPRTPR |
| C8 | ATX1 and ATX2 SVM 6-mer weight vectors are highly correlated (Fig.4d) | Pearson r (cor.test) | cor.test(weights1$V7, weights2$V7) |
A | data/Figure4/weights/ChIP8_ATX{1,2}_weights |
| C9 | TATA-box 6-mers negatively, GA/GAGA-repeat 6-mers positively weighted in ATX1/2 SVM (Fig.4e–f) | sign of weights for TATA vs GAGA k-mers | inspect signed weights | A | data/Figure4/weights/* |
| C10 | ATX1/ATX2-bound genes overlap **sppRNA (Hen2) genes more tha |
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.