Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Effects of replication domains on genome-wide UV-induced DNA damage and repair.

PLoS Genet · 2022
74/100 3/4
Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
How its reproducibility compares
74/100
Reproducibility score
at the mean
vs. all fields · 1173 studies
🎯 Scores higher than 43% of all assessed papers rank 644 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough to reproduce. The named code artifact boquila (P16, bioconda 0.6.1, repo commit 948787d) reproduces cleanly on the authors' own E. coli example: deterministic under --seed (EXACT) and simulated reads match the input nucleotide distribution (Pearson r=0.984). Fig 1B XR-seq excised-oligomer length reproduces within 1 nt (median 25 vs reported 26; 92.3% of reads in the reported 22-30 nt window) on HeLa CPD XR-seq (SRR11147228). Both datasets profiled (PRJNA608124 open/complete/A; PRJEB25180 reused OK-seq open/partial/B). NOT attempted: C3 Damage-seq dipyrimidine motif and C4 ERD/LRD repair log2FC (both genome-scale, need hg19 + full xr-ds-seq + replicationRepair + EdU windows); Fig 5 ICGC melanoma is controlled-access (out of scope). No fabrication indicators: every checked value is derivable from the shipped tool + open data. Verdict provisional, human-checkable.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-29
Rubric version
not recorded
Assessed by
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The paper tests how DNA replication timing and ongoing replication fork progression directly affect the genome-wide distribution and efficiency of nucleotide excision repair (NER) of UV-induced DNA lesions, and whether this crosstalk contributes to strand-asymmetric mutation patterns seen in skin cancers.

Core claims
  • Ongoing replication stimulates local nucleotide excision repair in both early and late replication domains as those regions become replicated finding
  • Early replication domains (ERDs) are repaired faster than late replication domains (LRDs) overall, consistent with open chromatin accessibility finding
  • Lesions on lagging strand templates are repaired more slowly than leading strand templates in late replication domains, likely due to imbalanced sequence/AT context finding
  • The observed asymmetric relative repair around replication initiation zones parallels the strand bias of melanoma mutations finding
  • Genome-wide replication (EdU-seq), damage (Damage-seq), and repair (XR-seq) maps were generated in cell-cycle synchronized, UV-irradiated HeLa cells method
  • At 12 minutes after UV, there is no detectable transcription-coupled repair contribution (no template vs non-template strand difference), isolating replication-associated effects from transcription effects finding
  • Ongoing replication preferentially promotes CPD repair in inactive (e.g., Polycomb repressed, heterochromatin) and transcription-associated chromatin states, with less effect in already-active states finding
  • Repair rate shows strand asymmetry favoring the leading strand template independent of active replication around initiation zones in LRDs finding
Experimental setups
Assay System Perturbation Readout Platform
Damage-seq HeLa cells, synchronized (early/late S phase) UV irradiation (20 J/m2 UVC) genome-wide distribution of (6-4)PP and CPD DNA lesions at nucleotide resolution
XR-seq HeLa cells, synchronized (early/late S phase) UV irradiation (20 J/m2 UVC) genome-wide nucleotide excision repair events (excised oligomers) at 12 min and 2 h after UV
EdU-seq (EdU incorporation and sequencing) HeLa cells synchronized via double thymidine block release into S phase for defined times genome-wide early and late replication domains (ERDs/LRDs)
Flow cytometry HeLa cells, synchronized double thymidine block/release cell cycle phase distribution to confirm early/late S phase enrichment
OK-seq (Okazaki fragment sequencing, retrieved dataset) HeLa cells none leading/lagging strand assignment and replication initiation zones
ChromHMM chromatin state segmentation (retrieved from UCSC) HeLa cells none chromatin state annotation integrated with replication domains and repair rates
Repli-seq (retrieved dataset, comparison) HeLa cells, asynchronized and synchronized none validation of EdU-seq-defined replication domains
Key results
  • Repair rate increases in a domain once it begins active replication (early S phase for ERDs, late S phase for LRDs) log2FC (6-4)PPs: ERDs 0.05, LRDs -0.13; CPDs: ERDs 0.2, LRDs -0.3
  • Normalized CPD repair rates peak at the center of ERDs and show the opposite (dip) pattern at LRDs
  • XR-seq excised oligomers range 22-30 nucleotides with a median of 26 nucleotides, matching known NER product size 22-30 nt, median 26 nt
  • No difference in repair rate between transcription template and non-template strands at 12 minutes post-UV
  • Preferential repair of template strand (transcription-coupled repair) becomes apparent by 2 hours post-UV
  • Replication elevates CPD repair rate more strongly in inactive/transcription-associated chromatin states than in already active states
  • Repair rate around initiation zones in LRDs is asymmetric, favoring the leading strand template
  • Initiation zones in LRDs show high AT content with more T-tracts on lagging strands, correlating repair asymmetry with sequence context
Key statistics
  • other 20 J/m2 UVC (UV dose used to irradiate synchronized HeLa cells)
  • other 22-30 nucleotides, median 26 nt (XR-seq excised oligomer length distribution)
  • fold_change log2FC (6-4)PPs in ERDs = 0.05 (repair rate change between early and late S phase in ERDs for (6-4)PPs)
  • fold_change log2FC (6-4)PPs in LRDs = -0.13 (repair rate change between early and late S phase in LRDs for (6-4)PPs)
  • fold_change log2FC CPDs in ERDs = 0.2 (repair rate change between early and late S phase in ERDs for CPDs)
  • fold_change log2FC CPDs in LRDs = -0.3 (repair rate change between early and late S phase in LRDs for CPDs)
  • count n=118 ERDs, n=237 LRDs (number of replication domains analyzed in Fig 2A repair rate profiles)
  • count 2130 initiation zones in ERDs, 1450 in LRDs (number of replication initiation zones analyzed around ERDs and LRDs)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The study used genome-wide sequencing assays (Damage-seq, XR-seq, EdU-seq, OK-seq) in synchronized, UV-irradiated HeLa cells to compare DNA repair rates across replication domains, chromatin states, and DNA strands. Comparisons between conditions (e.g., early vs. late S phase, plus vs. minus strand) were assessed with the Wilcoxon test (paired or unpaired depending on the comparison), and results were summarized primarily with boxplots and log2 fold-change values. Two replicates (A and B) were generated and either shown separately or combined across figures.

Replicationunclear Sample sizeSample sizes given as counts of genomic windows/regions or initiation zones (e.g., n=118, 237, 266, 541, 2130, 1450); no formal power analysis or biological replicate count described in the excerpted text Groupsearly vs. late S phase repair rates; chromatin states (ChromHMM segments); leading vs. lagging (plus vs. minus) strand repair around replication initiation zones Pairingmixed Randomization/blindingnot stated DispersionIQR Effect sizesyes
Statistical tests used
Test Applied to n Assumptions
Wilcoxon test (unpaired, rank-sum) Comparison of normalized repair rates between early and late S phase for ERDs and LRDs (Fig 2B) ERDs n=266, LRDs n=541 (genomic windows/regions), replicates A and B combined not stated
Wilcoxon test Significance of the relative difference in repair rate [log2(RR Early/RR Late)] versus 0 for each chromatin state (Fig 3B) not explicitly stated per chromatin state not stated
Paired Wilcoxon test Comparison of repair rates between plus and minus strands around initiation zones in ERDs and LRDs (Fig 4C-D) ERDs n=2130, LRDs n=1450 initiation zones not stated
Approaches that could also have been used
  • Differences in repair rate between early and late S phase (and between strands) were assessed with the nonparametric Wilcoxon test.
    Could also: A linear mixed-effects model treating genomic window/region as a random effect, or a paired/unpaired t-test if the distribution of repair rates approximates normality — Genomic windows compared across conditions are often spatially correlated (nearby windows share dependency), and a mixed-effects framework can explicitly model this non-independence while a t-test could add parametric efficiency if normality assumptions are reasonable.
  • Wilcoxon tests were applied separately to many chromatin states and genomic comparisons (Fig 3, S6-S8 Figs) without a stated multiple-testing correction.
    Could also: A Benjamini-Hochberg FDR or Bonferroni correction applied across the family of chromatin-state comparisons — When many tests are performed across chromatin categories, an FDR or family-wise correction would also help control the overall false-positive rate across that comparison family.
  • Effect sizes were reported as log2 fold-change point estimates without accompanying confidence intervals.
    Could also: Bootstrap resampling or analytic confidence intervals around the log2 fold-change estimates — Adding interval estimates would convey the precision/uncertainty of the fold-change values alongside the point estimate.
  • Genomic windows and initiation zones were treated as the unit of comparison in Wilcoxon tests.
    Could also: A permutation test or block-bootstrap approach that accounts for spatial autocorrelation among neighboring genomic windows — Permutation- or block-based resampling is a standard alternative in genomics for data with positional dependence, and could complement the rank-based test already used.
  • Repair rate distributions were summarized visually via boxplots (median/IQR).
    Could also: Reporting mean ± SD or mean with a 95% CI alongside the boxplots — Providing both median-based and mean-based summaries can make results more directly comparable to studies that use parametric summary statistics.

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-36155646

"Effects of replication domains on genome-wide UV-induced DNA damage and repair." Huang et al., PLoS Genet 2022. PMID 36155646 / PMC9536635 / DOI 10.1371/journal.pgen.1010426.

Datasets the paper relies on

  • PRJNA608124 (NCBI SRA) — the paper's OWN newly generated data: EdU-seq (early/late S, ±UV), Damage-seq (CPD, 6-4PP), XR-seq (CPD, 6-4PP; 12min & 2h), input gDNA libraries. 56 runs.
  • PRJEB25180 (ENA) — REUSED OK-seq (Okazaki fragment) data (Chen lab; HeLa/IMR90/GM06990 etc.), used for leading/lagging strand assignment. 92 runs. (This is the accession named in the brief.)
  • Repli-seq (UCSC), ChromHMM HeLa states (UCSC) — reused reference tracks.
  • ICGC MELA-AU melanoma mutations (183 tumors, release 28) — controlled/registered access.

Code artifacts

  • github.com/compGenomeLab/boquila (named in brief) — Rust NGS read simulator (v0.6.1, bioconda). Generates synthetic reads with same nucleotide distribution as model reads. Used to build the "simulation" background for normalization: RR = (XR_real/XR_sim)/(Dmg_real/Dmg_sim), log2.
  • github.com/CompGenomeLab/xr-ds-seq-snakemake — pre-analysis (cutadapt, bowtie2 -X1000, samtools -q20, picard dedup, bedtools). [pipeline, in scope for a slice]
  • github.com/CompGenomeLab/replicationRepair — downstream R/Python figure generation. Ships SCRIPTS ONLY (no processed data); needs full upstream output + hg19 + ICGC melanoma.

In scope (pipeline-derived, attempted)

id result pipeline tractability
C1 boquila simulated reads match input nucleotide/k-mer distribution; --seed determinism boquila 0.6.1 HIGH — deterministic, self-contained (P16 named tool)
C2 Fig 1B: XR-seq excised oligomer length 22–30 nt, median 26 nt cutadapt adapter trim + length histogram HIGH — one FASTQ
C3 Damage-seq reads enriched for dipyrimidines (TT/TC/CT/CC) at damage site bowtie2 + bedtools flank/slop + motif MEDIUM
C4 Repair-rate normalization in ERDs vs LRDs (Fig 2 log2 fold changes ~+0.2/-0.3) full xr-ds-seq + replicationRepair LOW — needs all samples + hg19 + EdU windows

Out of scope (not attempted)

  • ICGC melanoma mutation asymmetry (Fig 5) — controlled-access data (registered, MELA-AU).
  • Wet-lab steps (EdU pulse, IP, antibody), microscopy.
  • Full multi-panel figure regeneration requiring all 56 samples + melanoma + ChromHMM joins.

Strategy

Reach the quick minimum with C1 (boquila, the named code) + C2 (a clean quantitative claim), profile both accessions in the same pass, then push toward C3/C4 as compute allows. All heavy compute on «our HPC»; downloads on front1 into «infra»; «host» holds small results only.

Figures / tables: Fig normalizationFig 1BFig 2
C1a
Reported
boquila simulated reads have same nucleotide distribution as model reads
Reproduced
per-position nucleotide profile Pearson r=0.984 (max|Δ|=0.060)
within tolerance
C1b
Reported
boquila --seed gives deterministic/reproducible simulation
Reproduced
seed7 two runs byte-identical; seed99 differs
exact
C2
Reported
XR-seq excised oligomers 22-30 nt, median 26 nt (Fig 1B)
Reproduced
median 25 nt, 92.3% within 22-30 nt (SRR11147228 CPD XR-seq, 27M reads)
within tolerance
C3
Reported
Damage-seq TT/TC dipyrimidine enrichment at damage sites
Reproduced
not attempted (needs hg19 alignment + motif)
partial
C4
Reported
ERD/LRD repair log2FC CPDs +0.2/-0.3 (Fig 2)
Reproduced
not attempted (full genome pipeline + EdU windows)
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

No assessment has been recorded yet.
🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.