Bisulfite sequencing of chromatin immunoprecipitated DNA (BisChIP-seq) directly informs methylation status of histone-modified DNA.
The main results reproduced: recomputed values matched the published ones within tolerance.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Same input data as the authors
- ✓Reported values are derivable from the shared data
- ✓Any deviation was negligible
- 🟡Reported values were only indirectly comparable
- 🟡A deviation arose in the data or preprocessing
- 🟡A deviation was attributed to the published material
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Independently re-ran the astatham/Bisulfite-seq-pipeline (bowtie + hg18 in-silico bisulfite conversion + pileup-based CpG calling) on all 7 raw SRA lanes (SRR309230-236, LNCaP/PrEC BisChIP-seq, GSE30558) end-to-end on «our HPC»/SLURM, then compared the reproduced CpG methylation calls against the author-deposited processed files (GSM758360/758361). Plus-strand calls (the dominant, genome-wide output) matched with r=0.997 and ~0.7pp mean absolute difference in methylation fraction at ~7.3M/5.1M commonly-covered CpG sites -- a strong, independent reproduction of the pipeline's core output. Minus-strand calls agreed just as well at overlapping sites (r>0.999) but covered ~40% fewer sites than the original deposit (partial). Total unique aligned read counts matched within ~1.7%. The paper's headline ChromaBlocks enriched-region counts matched the paper text exactly, but that check only re-read the author's own deposited BED file -- ChromaBlocks/Repitools itself was not independently re-run (out of scope this pass). Two ancillary self-devised checks were inconclusive rather than failed: bisulfite conversion efficiency (formula ambiguity -- could not isolate the paper's non-CpG-specific denominator from available intermediate files) and a plus/minus strand-concordance recomputation (the assumed pos+1 join convention found almost no overlapping sites, so the paper's reported r=0.91/0.94 strand concordance remains untested by this analysis, not contradicted). Also flagging a room-brief data-accession error: the brief states 'geo:GSE19726', which is actually an unrelated earlier companion paper's MeDIP-chip/expression-array dataset, not this paper's BisChIP-seq data (correct: GSE34403 -> GSE30558 + GSE34340). Not attempted: GSE34340 Infinium 450K array validation (out of scope, array-based not pipeline-based) and any downstream biological/wet-lab claims.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-07-29
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-07-31no human curator yet
- Last updated
- 2026-07-31
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusThe paper asks whether DNA methylation and Polycomb H3K27me3-marked histones are mutually exclusive on the same DNA molecule, and develops a direct method (BisChIP-seq) to test this rather than relying on correlative overlays of independent maps. It tests whether the relationship is genomic-region dependent and whether it is remodeled in cancer.
- ★ BisChIP-seq — bisulfite sequencing of chromatin immunoprecipitated DNA — enables direct genome-wide, base-resolution interrogation of DNA methylation on histone-modified DNA molecules method
- ★ DNA methylation and H3K27me3 are not always mutually exclusive but can co-occur in a genomic region-dependent manner finding
- ★ In prostate cancer cells the co-dependency of the dual repressive marks is redistributed: increased dual marks at CpG islands and TSSs, with loss of DNA methylation in intergenic, intronic, and exonic H3K27me3-marked regions finding
- ★ Both methylated and unmethylated alleles can simultaneously be associated with H3K27me3 histones, so DNA methylation status in these regions is not dependent on Polycomb chromatin status finding
- ★ In normal PrEC cells, H3K27me3-enriched TSSs and CpG islands show a bimodal methylation distribution (predominantly either fully unmethylated or fully methylated), whereas exons, introns, and intergenic H3K27me3 regions are primarily methylated finding
- ★ H3K27me3-marked genes are transcriptionally repressed regardless of the level of underlying DNA methylation, suggesting H3K27me3-mediated silencing is mechanistically distinct from DNA methylation-associated silencing mechanism
- Bisulfite treatment can be performed successfully on <100 ng of sonicated formaldehyde-fixed ChIP DNA with sufficient yield for Illumina library generation, using the Clark et al. (2006) 4-h protocol with Microcon YM50 desalting and 14 library PCR cycles method
- A custom analysis pipeline for paired-end bisulfite ChIP reads plus ChromaBlocks region calling identifies marked regions and computes their methylation status resource
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| BisChIP-seq (bisulfite sequencing of H3K27me3 ChIP DNA, paired-end) | LNCaP prostate cancer cell line and normal prostate epithelial cells (PrEC) | none (H3K27me3 antibody immunoprecipitation; formaldehyde cross-linking) | CpG methylation level per site within H3K27me3-enriched regions; H3K27me3 enrichment regions; allele-specific methylation at heterozygous SNPs | Illumina GAIIx, 75-bp paired-end (3 lanes PrEC, 4 lanes LNCaP); methylated Illumina paired-end adaptors; PfuTurbo Cx hotstart DNA polymerase; Agilent 2100 Bioanalyser for QC |
| Chromatin immunoprecipitation (ChIP) | LNCaP and PrEC cells (~1 × 10^6 cells per 10-cm dish) | none; 1% formaldehyde fixation 10 min at 37°C, sonication to 200–500 bp; no-antibody control included | H3K27me3-enriched DNA | Millipore ChIP assay kit; anti-tri-methyl-histone H3(Lys27) Millipore #07-449 lot #DAM 1514011, 10 μL per IP |
| ChIP-qPCR validation | LNCaP and PrEC cells | H3K27me3 ChIP vs control | H3K27me3 enrichment at known candidate genes | — |
| Bisulfite conversion method comparison with methylation-specific headloop suppression PCR (MSH-PCR) | 100 ng sonicated ChIP input DNA from formaldehyde-fixed cells | QIAGEN EpiTect Bisulfite Kit (5-h) vs modified Clark et al. (2006) method (4-h) with Microcon YM50 desalting | Yield of bisulfite-converted DNA at positive control loci GSTP1 and EN1 | QIAGEN EpiTect Bisulfite Kit; Microcon YM50 |
| Library preparation optimization PCR | Adaptor-modified bisulfite-treated DNA (60 μL total) | Input volume titration (5, 10, 22.5 μL) and PCR cycle titration (10, 12, 14, 18 cycles) | Library yield/size distribution | Illumina paired-end primers PE 1.0 and PE 2.0; PfuTurbo Cx hotstart DNA polymerase; 50-μL reaction; Agilent 2100 Bioanalyser |
| Gene expression microarray | LNCaP and PrEC cells | none | Expression values of H3K27me3-marked vs -unmarked genes | Affymetrix Gene 1.0 ST |
| Array-based DNA methylation profiling (validation) | Native (non-cross-linked) LNCaP and PrEC genomic DNA | none | CpG methylation beta values for comparison with BisChIP-seq methylation | Illumina Infinium 450K |
| Clonal bisulfite (PCR) genomic sequencing | Native LNCaP and PrEC DNA | none | Single-molecule CpG methylation at selected loci, to test whether cross-linking affects bisulfite results | — |
- ▲ H3K27me3-enriched regions called by ChromaBlocks: 53,749 regions covering 148.6 Mb in PrEC and 52,677 regions covering 139.9 Mb in LNCaP (FDR < 0.001), preferentially enriched at TSSs and CpG islands but not exons, introns, or intergenic regions 53,749 / 148.6 Mb (PrEC); 52,677 / 139.9 Mb (LNCaP)
- ▼ H3K27me3-marked TSS genes (5029 in PrEC, 4639 in LNCaP) are expressed only at basal levels, consistent with H3K27me3-mediated repression; repression occurs irrespective of DNA methylation level
- ▲ In cancer, more H3K27me3-marked TSS regions carry medium/high DNA methylation than in normal cells: 3330 (~71.7%) of LNCaP TSS regions vs 2866 (57.0%) of PrEC TSS regions 71.7% vs 57.0%
- ▼ Exonic, intronic, and intergenic H3K27me3-enriched regions are comparatively depleted of highly methylated DNA in LNCaP relative to PrEC, with a shift toward H3K27me3 binding unmethylated DNA in intergenic regions
- – Allele-specific differential methylation within H3K27me3-enriched regions at 762 (PrEC) and 1195 (LNCaP) heterozygous SNPs, i.e., both methylated and unmethylated alleles equally bound by the Polycomb mark; example SNP rs637481 shows high vs low methylation on the A and G alleles in PrEC 762 of 6472 het SNPs (PrEC); 1195 of 13,034 (LNCaP)
- ▲ The RCSD1 CpG island promoter retains similar H3K27me3 enrichment in both cell types while being unmethylated in PrEC and extensively hypermethylated in LNCaP — gain of DNA methylation in cancer without loss of H3K27me3
- – BisChIP-seq methylation calls are highly concordant with Infinium 450K methylation from native DNA, and plus/minus strand methylation calls are highly concordant, supporting technical validity; cross-linking does not affect bisulfite results r = 0.932 (PrEC), r = 0.925 (LNCaP) vs 450K; r = 0.94 / 0.91 strand concordance
- ▲ Modified Clark et al. (2006) 4-h bisulfite method with Microcon YM50 desalting gave the greatest yield of converted DNA; 22.5 μL of bisulfite-treated input inhibited library prep, and 14 PCR cycles gave adequate yield without overamplification
- correlation r = 0.932 (PrEC), r = 0.925 (LNCaP) (Concordance of BisChIP-seq methylation in H3K27me3-enriched regions with Infinium 450K array data from native DNA)
- correlation r = 0.94 (PrEC), r = 0.91 (LNCaP) (Plus- vs minus-strand CpG methylation concordance at individual CpG sites (minimum 10 reads))
- pvalue p < 1 × 10^-15 (χ2 test) (LNCaP vs PrEC difference in number of H3K27me3-marked TSS regions with medium/high DNA methylation (3330 vs 2866))
- count 38,403,614 reads (PrEC); 70,682,755 reads (LNCaP) (Total BisChIP-seq reads obtained per cell type)
- other 99.7% (PrEC); 99.8% (LNCaP) (Bisulfite conversion rate)
- count 2,482,996 (PrEC); 2,552,762 (LNCaP) (CpG sites interrogated after pooling strands; <1% of H3K27me3-enriched regions had insufficient coverage)
- count 6472 (PrEC) and 13,034 (LNCaP) clear heterozygous SNPs from 106,887 and 215,628 SNPs with ≥20 reads (SNPs available for allele-specific methylation analysis)
- count 762 (PrEC); 1195 (LNCaP) SNPs with allele-specific differential methylation (difference-in-proportions test, FDR < 0.05) (Allele-specific methylation within H3K27me3-enriched regions)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
The study compares genome-wide DNA methylation and H3K27me3 enrichment between two prostate cell lines (normal PrEC and cancer LNCaP) using a novel bisulfite-ChIP sequencing (BisChIP-seq) method. Region-level H3K27me3 enrichment was called with the ChromaBlocks algorithm under an FDR threshold; methylation concordance (between DNA strands, and between BisChIP-seq and an independent Infinium 450K array) was assessed with Pearson correlation coefficients; the difference in proportion of highly methylated TSS regions between cell lines was assessed with a chi-square test; and allele-specific differential methylation at heterozygous SNPs was assessed with a difference-in-proportions test under FDR control. Results are reported mainly as counts, percentages, correlation coefficients, and FDR-thresholded or inequality-bound p-values, without a specific statistical software package named.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| chi-square (χ2) test | comparison of the proportion of TSS regions with medium/high DNA methylation between LNCaP (cancer) and PrEC (normal) H3K27me3-enriched regions | 3330/4644 LNCaP TSS regions vs. 2866/5029 PrEC TSS regions | not stated |
| Pearson correlation (r) | concordance of CpG methylation calls between plus and minus DNA strands | CpG sites with a minimum of 10 reads (r = 0.94 PrEC, 0.91 LNCaP) | not stated |
| Pearson correlation (r) | concordance between BisChIP-seq methylation levels and Infinium 450K array methylation data | not explicitly stated (r = 0.932 PrEC, 0.925 LNCaP) | not stated |
| FDR-thresholded enrichment calling (ChromaBlocks) | genome-wide identification of H3K27me3-enriched regions in PrEC and LNCaP | 53,749 (PrEC) and 52,677 (LNCaP) enriched regions called at FDR < 0.001 | not stated |
| difference-in-proportions test | allele-specific differential methylation at heterozygous SNPs within H3K27me3-enriched regions | 6472 (PrEC) and 13,034 (LNCaP) clear heterozygous SNPs, FDR < 0.05 | not stated |
-
The difference in the proportion of highly methylated TSS regions between LNCaP and PrEC was assessed with a chi-square test on region counts.↳ Could also: A logistic regression or generalized linear model relating methylation category to cell type (and optionally other region-level covariates such as CpG density) — A regression-based approach would allow the group comparison to be adjusted for other region characteristics and would naturally extend to reporting an effect size (e.g., odds ratio) alongside the significance test.
-
FDR-based thresholds (FDR < 0.001 for region calling, FDR < 0.05 for allele-specific methylation) were used without naming the specific correction procedure.↳ Could also: Explicitly naming and reporting the FDR method (e.g., Benjamini-Hochberg) with per-site q-values — Stating the exact procedure would let readers evaluate its assumptions (e.g., independence or positive dependence among tests) and reproduce the multiple-testing adjustment.
-
Agreement between BisChIP-seq methylation calls and an independent method (strand-to-strand, and vs. the Infinium 450K array) was quantified using Pearson correlation coefficients.↳ Could also: Bland-Altman (limits-of-agreement) analysis or a concordance correlation coefficient (Lin's CCC) — These approaches directly quantify agreement and any systematic bias between two measurement methods, which a correlation coefficient alone does not fully capture since r can be high even when a consistent offset is present.
-
The comparison between normal and cancer states is based on one cell line per condition (PrEC vs. LNCaP), sequenced across multiple lanes.↳ Could also: Including additional independent normal and cancer cell lines or biological culture replicates, analyzed with replicate as a random effect — This would let variability between independent biological preparations be distinguished from the specific normal-vs-cancer contrast of interest.
-
Methylation levels are summarized as counts/percentages of regions falling into discrete methylation categories (Fig. 1D, Table 2) rather than as continuous distributions with a dispersion measure.↳ Could also: Reporting mean or median methylation with SD, IQR, or bootstrapped confidence intervals alongside the categorical breakdown — A continuous summary with a spread estimate would complement the categorical distribution and convey within-group variability more directly.
-
The chi-square result is reported as an inequality (p < 1 x 10^-15) rather than an exact p-value or test statistic.↳ Could also: Reporting the exact chi-square statistic, degrees of freedom, and exact p-value — Exact values support downstream use of the result, such as meta-analysis or recomputation of effect sizes, by other researchers.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
Strong core reproduction. All 7 raw SRA lanes (SRR309230-236) were independently re-aligned end-to-end and the deposited processed CpG calls were re-derived at r=0.9972 plus-strand in both LNCaP and PrEC (mean abs diff <0.008), with site counts within 0.25% and unique-read totals within ~1.7% (71,894,841 vs 70,682,755) — the data are fully derivable and nothing is fabrication-suspect.
The open items are all on our side, not the authors'. The minus-strand output recovered only ~60% of deposited sites (53,279 vs 87,915), the non-CpG conversion efficiency could not be isolated from the pipeline's .context files (we could only compute an overall ~92.6-95.2% C-retention rate against the paper's 99.7%/99.8% non-CpG figure), and the plus/minus join key produced r=0.012 on 569 sites — a broken pairing convention, correctly recorded as inconclusive rather than as a contradiction.
Severity is negligible where comparison was possible, so q6 is green; q7 is capped at yellow only because two internal QC claims and the ChromaBlocks peak calling (checked as deposited BED vs paper text only — 52,677 regions / 139.97 Mb) were never independently tested. Also worth an admin correction: the manifest's data_accession: GSE19726 is the wrong series (MeDIP-chip/expression arrays); the real data are GSE34403 → GSE30558 + GSE34340.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.