Effects of replication domains on genome-wide UV-induced DNA damage and repair.
The main results reproduced, with only marginal, non-material deviations.
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Described well enough to reproduce. The named code artifact boquila (P16, bioconda 0.6.1, repo commit 948787d) reproduces cleanly on the authors' own E. coli example: deterministic under --seed (EXACT) and simulated reads match the input nucleotide distribution (Pearson r=0.984). Fig 1B XR-seq excised-oligomer length reproduces within 1 nt (median 25 vs reported 26; 92.3% of reads in the reported 22-30 nt window) on HeLa CPD XR-seq (SRR11147228). Both datasets profiled (PRJNA608124 open/complete/A; PRJEB25180 reused OK-seq open/partial/B). NOT attempted: C3 Damage-seq dipyrimidine motif and C4 ERD/LRD repair log2FC (both genome-scale, need hg19 + full xr-ds-seq + replicationRepair + EdU windows); Fig 5 ICGC melanoma is controlled-access (out of scope). No fabrication indicators: every checked value is derivable from the shipped tool + open data. Verdict provisional, human-checkable.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-29
- Rubric version
- not recorded
- Assessed by
- —
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetThe paper tests how DNA replication timing and ongoing replication fork progression directly affect the genome-wide distribution and efficiency of nucleotide excision repair (NER) of UV-induced DNA lesions, and whether this crosstalk contributes to strand-asymmetric mutation patterns seen in skin cancers.
- ★ Ongoing replication stimulates local nucleotide excision repair in both early and late replication domains as those regions become replicated finding
- ★ Early replication domains (ERDs) are repaired faster than late replication domains (LRDs) overall, consistent with open chromatin accessibility finding
- ★ Lesions on lagging strand templates are repaired more slowly than leading strand templates in late replication domains, likely due to imbalanced sequence/AT context finding
- ★ The observed asymmetric relative repair around replication initiation zones parallels the strand bias of melanoma mutations finding
- ★ Genome-wide replication (EdU-seq), damage (Damage-seq), and repair (XR-seq) maps were generated in cell-cycle synchronized, UV-irradiated HeLa cells method
- At 12 minutes after UV, there is no detectable transcription-coupled repair contribution (no template vs non-template strand difference), isolating replication-associated effects from transcription effects finding
- ★ Ongoing replication preferentially promotes CPD repair in inactive (e.g., Polycomb repressed, heterochromatin) and transcription-associated chromatin states, with less effect in already-active states finding
- ★ Repair rate shows strand asymmetry favoring the leading strand template independent of active replication around initiation zones in LRDs finding
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Damage-seq | HeLa cells, synchronized (early/late S phase) | UV irradiation (20 J/m2 UVC) | genome-wide distribution of (6-4)PP and CPD DNA lesions at nucleotide resolution | — |
| XR-seq | HeLa cells, synchronized (early/late S phase) | UV irradiation (20 J/m2 UVC) | genome-wide nucleotide excision repair events (excised oligomers) at 12 min and 2 h after UV | — |
| EdU-seq (EdU incorporation and sequencing) | HeLa cells synchronized via double thymidine block | release into S phase for defined times | genome-wide early and late replication domains (ERDs/LRDs) | — |
| Flow cytometry | HeLa cells, synchronized | double thymidine block/release | cell cycle phase distribution to confirm early/late S phase enrichment | — |
| OK-seq (Okazaki fragment sequencing, retrieved dataset) | HeLa cells | none | leading/lagging strand assignment and replication initiation zones | — |
| ChromHMM chromatin state segmentation (retrieved from UCSC) | HeLa cells | none | chromatin state annotation integrated with replication domains and repair rates | — |
| Repli-seq (retrieved dataset, comparison) | HeLa cells, asynchronized and synchronized | none | validation of EdU-seq-defined replication domains | — |
- – Repair rate increases in a domain once it begins active replication (early S phase for ERDs, late S phase for LRDs) log2FC (6-4)PPs: ERDs 0.05, LRDs -0.13; CPDs: ERDs 0.2, LRDs -0.3
- – Normalized CPD repair rates peak at the center of ERDs and show the opposite (dip) pattern at LRDs
- – XR-seq excised oligomers range 22-30 nucleotides with a median of 26 nucleotides, matching known NER product size 22-30 nt, median 26 nt
- – No difference in repair rate between transcription template and non-template strands at 12 minutes post-UV
- ▲ Preferential repair of template strand (transcription-coupled repair) becomes apparent by 2 hours post-UV
- ▲ Replication elevates CPD repair rate more strongly in inactive/transcription-associated chromatin states than in already active states
- – Repair rate around initiation zones in LRDs is asymmetric, favoring the leading strand template
- – Initiation zones in LRDs show high AT content with more T-tracts on lagging strands, correlating repair asymmetry with sequence context
- other 20 J/m2 UVC (UV dose used to irradiate synchronized HeLa cells)
- other 22-30 nucleotides, median 26 nt (XR-seq excised oligomer length distribution)
- fold_change log2FC (6-4)PPs in ERDs = 0.05 (repair rate change between early and late S phase in ERDs for (6-4)PPs)
- fold_change log2FC (6-4)PPs in LRDs = -0.13 (repair rate change between early and late S phase in LRDs for (6-4)PPs)
- fold_change log2FC CPDs in ERDs = 0.2 (repair rate change between early and late S phase in ERDs for CPDs)
- fold_change log2FC CPDs in LRDs = -0.3 (repair rate change between early and late S phase in LRDs for CPDs)
- count n=118 ERDs, n=237 LRDs (number of replication domains analyzed in Fig 2A repair rate profiles)
- count 2130 initiation zones in ERDs, 1450 in LRDs (number of replication initiation zones analyzed around ERDs and LRDs)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
The study used genome-wide sequencing assays (Damage-seq, XR-seq, EdU-seq, OK-seq) in synchronized, UV-irradiated HeLa cells to compare DNA repair rates across replication domains, chromatin states, and DNA strands. Comparisons between conditions (e.g., early vs. late S phase, plus vs. minus strand) were assessed with the Wilcoxon test (paired or unpaired depending on the comparison), and results were summarized primarily with boxplots and log2 fold-change values. Two replicates (A and B) were generated and either shown separately or combined across figures.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Wilcoxon test (unpaired, rank-sum) | Comparison of normalized repair rates between early and late S phase for ERDs and LRDs (Fig 2B) | ERDs n=266, LRDs n=541 (genomic windows/regions), replicates A and B combined | not stated |
| Wilcoxon test | Significance of the relative difference in repair rate [log2(RR Early/RR Late)] versus 0 for each chromatin state (Fig 3B) | not explicitly stated per chromatin state | not stated |
| Paired Wilcoxon test | Comparison of repair rates between plus and minus strands around initiation zones in ERDs and LRDs (Fig 4C-D) | ERDs n=2130, LRDs n=1450 initiation zones | not stated |
-
Differences in repair rate between early and late S phase (and between strands) were assessed with the nonparametric Wilcoxon test.↳ Could also: A linear mixed-effects model treating genomic window/region as a random effect, or a paired/unpaired t-test if the distribution of repair rates approximates normality — Genomic windows compared across conditions are often spatially correlated (nearby windows share dependency), and a mixed-effects framework can explicitly model this non-independence while a t-test could add parametric efficiency if normality assumptions are reasonable.
-
Wilcoxon tests were applied separately to many chromatin states and genomic comparisons (Fig 3, S6-S8 Figs) without a stated multiple-testing correction.↳ Could also: A Benjamini-Hochberg FDR or Bonferroni correction applied across the family of chromatin-state comparisons — When many tests are performed across chromatin categories, an FDR or family-wise correction would also help control the overall false-positive rate across that comparison family.
-
Effect sizes were reported as log2 fold-change point estimates without accompanying confidence intervals.↳ Could also: Bootstrap resampling or analytic confidence intervals around the log2 fold-change estimates — Adding interval estimates would convey the precision/uncertainty of the fold-change values alongside the point estimate.
-
Genomic windows and initiation zones were treated as the unit of comparison in Wilcoxon tests.↳ Could also: A permutation test or block-bootstrap approach that accounts for spatial autocorrelation among neighboring genomic windows — Permutation- or block-based resampling is a standard alternative in genomics for data with positional dependence, and could complement the rank-based test already used.
-
Repair rate distributions were summarized visually via boxplots (median/IQR).↳ Could also: Reporting mean ± SD or mean with a 95% CI alongside the boxplots — Providing both median-based and mean-based summaries can make results more directly comparable to studies that use parametric summary statistics.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-36155646
"Effects of replication domains on genome-wide UV-induced DNA damage and repair." Huang et al., PLoS Genet 2022. PMID 36155646 / PMC9536635 / DOI 10.1371/journal.pgen.1010426.
Datasets the paper relies on
- PRJNA608124 (NCBI SRA) — the paper's OWN newly generated data: EdU-seq (early/late S, ±UV), Damage-seq (CPD, 6-4PP), XR-seq (CPD, 6-4PP; 12min & 2h), input gDNA libraries. 56 runs.
- PRJEB25180 (ENA) — REUSED OK-seq (Okazaki fragment) data (Chen lab; HeLa/IMR90/GM06990 etc.), used for leading/lagging strand assignment. 92 runs. (This is the accession named in the brief.)
- Repli-seq (UCSC), ChromHMM HeLa states (UCSC) — reused reference tracks.
- ICGC MELA-AU melanoma mutations (183 tumors, release 28) — controlled/registered access.
Code artifacts
- github.com/compGenomeLab/boquila (named in brief) — Rust NGS read simulator (v0.6.1, bioconda). Generates synthetic reads with same nucleotide distribution as model reads. Used to build the "simulation" background for normalization: RR = (XR_real/XR_sim)/(Dmg_real/Dmg_sim), log2.
- github.com/CompGenomeLab/xr-ds-seq-snakemake — pre-analysis (cutadapt, bowtie2 -X1000, samtools -q20, picard dedup, bedtools). [pipeline, in scope for a slice]
- github.com/CompGenomeLab/replicationRepair — downstream R/Python figure generation. Ships SCRIPTS ONLY (no processed data); needs full upstream output + hg19 + ICGC melanoma.
In scope (pipeline-derived, attempted)
| id | result | pipeline | tractability |
|---|---|---|---|
| C1 | boquila simulated reads match input nucleotide/k-mer distribution; --seed determinism | boquila 0.6.1 | HIGH — deterministic, self-contained (P16 named tool) |
| C2 | Fig 1B: XR-seq excised oligomer length 22–30 nt, median 26 nt | cutadapt adapter trim + length histogram | HIGH — one FASTQ |
| C3 | Damage-seq reads enriched for dipyrimidines (TT/TC/CT/CC) at damage site | bowtie2 + bedtools flank/slop + motif | MEDIUM |
| C4 | Repair-rate normalization in ERDs vs LRDs (Fig 2 log2 fold changes ~+0.2/-0.3) | full xr-ds-seq + replicationRepair | LOW — needs all samples + hg19 + EdU windows |
Out of scope (not attempted)
- ICGC melanoma mutation asymmetry (Fig 5) — controlled-access data (registered, MELA-AU).
- Wet-lab steps (EdU pulse, IP, antibody), microscopy.
- Full multi-panel figure regeneration requiring all 56 samples + melanoma + ChromHMM joins.
Strategy
Reach the quick minimum with C1 (boquila, the named code) + C2 (a clean quantitative claim), profile both accessions in the same pass, then push toward C3/C4 as compute allows. All heavy compute on «our HPC»; downloads on front1 into «infra»; «host» holds small results only.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.