epiGBS2: Improvements and evaluation of highly multiplexed, epiGBS-based reduced representation bisulfite sequencing.
Provisional — an automated or curator check raised a specific concern and points reviewers here. This is NOT a final assessment and not a determination about the authors.
The main results reproduced, with only marginal, non-material deviations.
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
STRONG PARTIAL. epiGBS2 methods paper; Snakemake bisulfite-seq pipeline run on «our HPC» SLURM. KEY METHOD NOTE: the published pipeline demultiplexes RAW multiplexed reads with a barcode file published NOWHERE; the raw Zenodo deposit (92GB) is thus un-demultiplexable, so the SRA deposit PRJNA764918 (44 samples x Watson/Crick, already-demultiplexed 140bp) was staged as the pipeline's post-demultiplex 'clone-stacks' state and the pipeline run from trimming onward. BOTH BRANCHES NOW COMPLETE 44/44. REFERENCE BRANCH (TAIR10) reproduces the numeric claims essentially exactly: C05 mapping 53.1-68.3% all in 50-70% Col-0 highest; C08 3'-trim retention 99.9949% vs 99.99%; C10 methylation sites 1,827,216 vs 1,827,467 (0.014% diff). DE NOVO BRANCH reproduces mapping % (C06 50-54% per-accession, in 48-57%) and cluster mapping quality (C02 89.01% vs 87.69%), but has a single documented root-cause divergence: clustering on already-demultiplexed/trimmed SRA reads yields 29,112 clusters vs paper's 19,766 (C01 1.47x), which propagates to inflated de novo methylation sites (C09 1,738,348 vs 1,047,597) and high cluster edit distance (C04 13.89 vs 1.1) and low liftover overlap (C11 371,226 vs 928,070). C12/C13 baseline correlation partial (per-sample CG R2~0.85-0.91 strong; pooled ~0.53 - possible discrepancy, flagged). C16/C17 SNP: the freebayes SNP-calling pipeline reproduces and runs (575,780 biallelic SNPs/44 samples) and its calls overlap the 1001genomes baseline (thousands of genotype-matched true positives per divergent accession via rtg-vcfeval), but the paper's headline >90% precision / ~80% sensitivity are NOT reached with GQ-thresholded raw freebayes calls (reproduced precision ~0.56, sensitivity ~0.40 squash best-F) - the paper's figures require the authors' high-quality + both-strand call-set filtering (epiGBS-Benchmarking repo), which was not replicated; graded partial. NOTE: this room was REQUEUED after a «infra» janitor purge wiped all intermediates; the entire reference branch + SNP chain was regenerated from scratch (recipe recovered from prior-run transcripts) to close C16/C17. AUDIT FLAG: SRA library has 10 C24 samples vs paper's stated 8 (C15). NOT ATTEMPTED (documented): C14 de novo CHG correlation (de novo branch not regenerated this round), C18-C23 P. major (separate organism + external RRBS, out of primary scope).
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-30
- Rubric version
- not recorded
- Assessed by
- —
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetThe authors ask whether an improved epiGBS-based laboratory protocol and revised, user-friendly bioinformatics pipeline (epiGBS2) can accurately and efficiently determine de novo cytosine methylation and SNP variation in species with or without a reference genome, performing well against established benchmark datasets.
- ★ epiGBS2 provides a laboratory protocol and revised bioinformatics pipeline for de novo cytosine methylation and SNP calling in species with or without a reference genome resource
- ★ The bioinformatics pipeline is integrated into the snakemake workflow management system, making it modular, reproducible and flexible method
- ★ epiGBS2 implements bismark for alignment and methylation calling, replacing the legacy samtools mpileup-based approach method
- ★ Alignment files are preprocessed by double masking to enable SNP calling with freebayes (epifreebayes) method
- ★ The protocol allows flexible choice of restriction enzyme pairs, including a double digest, via a control nucleotide that still enables Watson/Crick strand differentiation mechanism
- Hemimethylated adapter pairs (instead of fully methylated adapters) reduce library preparation costs while nick translation with 5mC-dNTPs yields fully methylated adapters mechanism
- ★ A random 3-nucleotide UMI placed in the adapter enables computational identification and removal of PCR duplicates using stacks2 clone_filter method
- ★ Performance evaluation against Arabidopsis thaliana and great tit (Parus major) baseline datasets confirmed overall good performance of epiGBS2 finding
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| epiGBS2 reduced representation bisulfite sequencing (full pipeline) | Arabidopsis thaliana accessions | none | methylation calling and SNP calling performance compared to published benchmark data set | bismark, freebayes (epifreebayes) |
| epiGBS2 reduced representation bisulfite sequencing (full pipeline) | great tit (Parus major) | none | methylation calling and SNP calling performance compared to published benchmark data set | bismark, freebayes (epifreebayes) |
| PCR duplicate removal | epiGBS2 sequencing reads (Watson/Crick, UMI-tagged) | none | identification/removal of PCR clones based on identical fragment and UMI sequence | clone_filter (stacks2) |
| demultiplexing and Watson/Crick strand annotation | epiGBS2 sequencing reads | none | sample assignment and Watson/Crick strand classification via control nucleotide and barcode | process_radtags (stacks2) |
| de novo reference reconstruction | epiGBS2 reads from A. thaliana and great tit | none | consensus reference sequence clusters from deduplication, Watson/Crick pairing and identity clustering | pear, custom scripts |
| adapter trimming | demultiplexed, strand-annotated epiGBS2 reads | none | 3' adapter removal and minimum read length filtering (20 bp cutoff) | cutadapt |
| read mapping and DNA methylation calling | de novo clusters or pre-existing reference genome | none | alignment and cytosine methylation levels | bismark |
| SNP calling | double-masked bismark alignment files | none | single nucleotide polymorphism variants | freebayes (epifreebayes) |
- – Performance of critical epiGBS2 steps evaluated against A. thaliana and great tit baseline data sets confirmed overall good performance
- other 46% (proportion of BS-Seq data sets in the SRA generated from mouse)
- other 34% (proportion of BS-Seq data sets in the SRA generated from human)
- other 4% (proportion of BS-Seq data sets in the SRA generated from Arabidopsis thaliana)
- other 3-nucleotide UMI (length of unique molecular identifier used in adapter for PCR duplicate removal)
- other snakemake version 6.1.1 (version of workflow management system used to build the epiGBS2 pipeline)
- other minimum read length 20 bp (length cutoff used for filtering reads after adapter trimming)
- other 10 bp hard trim (additional bases removed from adapter-trimmed reads to discard UMI, barcode and control nucleotide)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This is a methods/tool paper describing epiGBS2, a laboratory protocol and snakemake-based bioinformatics pipeline for reduced-representation bisulfite sequencing. The manuscript text provided (introduction and methods sections) focuses on describing the wet-lab protocol and computational workflow (PCR duplicate removal, demultiplexing, de novo reference construction, adapter trimming, bismark-based alignment/methylation calling, and freebayes-based SNP calling) and states that pipeline performance was evaluated against baseline datasets from Arabidopsis thaliana and great tit (Parus major). The excerpt available does not include the Results section, so no specific inferential statistical test, p-value, or effect-size reporting is described in the text provided.
-
Pipeline performance is described as having been 'evaluated' against baseline datasets from A. thaliana and great tit without a specific quantitative statistic named in the available text.↳ Could also: Concordance metrics such as Pearson/Spearman correlation, concordance correlation coefficient, or Bland-Altman agreement analysis could also be used to summarize agreement between epiGBS2 outputs and baseline/reference calls. — These approaches quantify both the strength and the direction of any systematic bias between two measurement methods, which can complement a qualitative description of 'good performance.'
-
SNP calling performance and methylation calling performance are each described narratively as being benchmarked against reference datasets.↳ Could also: Standard classification-accuracy metrics (precision, recall/sensitivity, F1 score) relative to a gold-standard call set could also be reported. — These metrics are commonly used in bioinformatics tool benchmarking to give a standardized, comparable measure of calling accuracy across studies and tools.
-
The de novo reference construction and clustering steps involve tunable parameters (minimum/maximum cluster depth, identity percentage) described as adjustable by the user.↳ Could also: A sensitivity analysis or parameter sweep with a quantitative summary (e.g., variance in the number/quality of loci recovered across parameter settings) could also be used to characterize how outcomes depend on these settings. — This would provide readers with an explicit, quantitative picture of how parameter choices affect downstream results, complementing the qualitative guidance given in the text.
-
The comparison of the new de novo/reference branches to the previous 'legacy' epiGBS pipeline is described narratively as a methodological update.↳ Could also: A paired or matched-sample comparison (e.g., Wilcoxon signed-rank test or paired t-test on a relevant summary metric per sample) between legacy and updated pipeline outputs on the same samples could also be used. — A formal paired comparison would let readers see whether differences between pipeline versions are consistent across samples and how large they are, in addition to the qualitative description provided.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-35178872 (epiGBS2)
Paper: Gawehns et al. 2022, Mol Ecol Resour. "epiGBS2: Improvements and evaluation of highly multiplexed, epiGBS-based reduced representation bisulfite sequencing." DOI 10.1111/1755-0998.13597.
Nature of the paper: a methods/tool paper presenting the epiGBS2 Snakemake pipeline (https://github.com/nioo-knaw/epiGBS2). The "results" are an evaluation of the pipeline on two empirical datasets. This is the canonical P16 case: the repo IS the authors' own pipeline; reproducing = running that pipeline on the paper's own data with the described parameters.
Datasets the paper relies on
- PRJNA764918 (SRA) — A. thaliana epiGBS2, 1 multiplexed library, 44 barcoded samples (6 accessions). Raw also on Zenodo 5519370. Primary in-scope data.
- Zenodo 5878925 — small example dataset shipped to test the pipeline (smoke test).
- Zenodo 4764652 — pipeline code archive.
- GitHub MaartenPostuma/epiGBS-Benchmarking — benchmarking/validation scripts (minimap2 cluster validation, rtg vcfeval SNP benchmark, methylKit correlations).
- P. major (great tit) epiGBS2 + a prior RRBS dataset — used in §3.6. The RRBS comparison data accession is not the primary deposit; lower priority / likely out of scope.
- External baselines (NOT our data, used for correlation): 1001 Genomes SNPs, 1001 Epigenomes methylation, TAIR10 reference genome.
IN SCOPE (pipeline-derived, reproducible by running epiGBS2)
Ordered by tractability (80% floor = top items):
| id | result | branch | pipeline stage |
|---|---|---|---|
| C07/C08 | adapter content before / reads retained after trimming | both | cutadapt (preprocessing) |
| C01 | de novo clusters generated (19,766) | denovo | clustering (output_denovo) |
| C05/C06 | per-accession mapping % (ref 50-70%, denovo 48-57%) | both | bismark alignment reports |
| C09/C10/C11 | methylation sites de novo / ref / overlap | both | bismark methylation calling (CX reports) |
| C02/C03/C04 | de novo cluster → TAIR10 alignment (87.69%, 10,960 perfect, mean 1.1 mm) | denovo | minimap2 validation (benchmarking repo) |
| C12/C13/C14 | CG/CHG correlation vs 1001 epigenomes | both | methylKit + baseline |
| C16/C17 | SNP sensitivity vs 1001 Genomes | both | freebayes + rtg vcfeval |
OUT OF SCOPE (not pipeline-derived from our primary data, or external)
- Wet-lab protocol improvements (barcode design, RE choice, library prep) — manual.
- §3.6 P. major RRBS comparison (C18–C23): requires a separate prior RRBS dataset
- great-tit reference; reproduce only if the great-tit epiGBS2 data + RRBS data prove openly obtainable. Marked lower priority.
- Baseline datasets themselves (1001 Genomes/Epigenomes) — external references, used for comparison, not reproduced.
Plan
- Clone epiGBS2 + benchmarking repos on «infra»; pin commit.
- Download PRJNA764918 (44 samples) + barcode file to «infra» via front1.
- Smoke-test pipeline on Zenodo example data (fast, validates env).
- Run denovo + reference branches on A. thaliana data (SLURM, compute only).
- Extract: cluster count, mapping %, methylation site counts; validate clusters vs TAIR10.
- Compare to claims.tsv; grade in agreement.json + AUDIT.md.
Compute note: bisulfite alignment (bismark/bowtie2) on 44 paired-end 150bp libraries is the heavy step → SLURM on «our HPC», conda envs built on front1.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.