Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

epiGBS2: Improvements and evaluation of highly multiplexed, epiGBS-based reduced representation bisulfite sequencing.

Mol Ecol Resour · 2022
52/100 3/4
⚑ Flagged for review — a reproduced result did not match the reported value

Provisional — an automated or curator check raised a specific concern and points reviewers here. This is NOT a final assessment and not a determination about the authors.

Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
How its reproducibility compares
52/100
Reproducibility score
1.3 SD below mean
vs. all fields · 1173 studies
🎯 Scores higher than 13% of all assessed papers rank 1014 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

STRONG PARTIAL. epiGBS2 methods paper; Snakemake bisulfite-seq pipeline run on «our HPC» SLURM. KEY METHOD NOTE: the published pipeline demultiplexes RAW multiplexed reads with a barcode file published NOWHERE; the raw Zenodo deposit (92GB) is thus un-demultiplexable, so the SRA deposit PRJNA764918 (44 samples x Watson/Crick, already-demultiplexed 140bp) was staged as the pipeline's post-demultiplex 'clone-stacks' state and the pipeline run from trimming onward. BOTH BRANCHES NOW COMPLETE 44/44. REFERENCE BRANCH (TAIR10) reproduces the numeric claims essentially exactly: C05 mapping 53.1-68.3% all in 50-70% Col-0 highest; C08 3'-trim retention 99.9949% vs 99.99%; C10 methylation sites 1,827,216 vs 1,827,467 (0.014% diff). DE NOVO BRANCH reproduces mapping % (C06 50-54% per-accession, in 48-57%) and cluster mapping quality (C02 89.01% vs 87.69%), but has a single documented root-cause divergence: clustering on already-demultiplexed/trimmed SRA reads yields 29,112 clusters vs paper's 19,766 (C01 1.47x), which propagates to inflated de novo methylation sites (C09 1,738,348 vs 1,047,597) and high cluster edit distance (C04 13.89 vs 1.1) and low liftover overlap (C11 371,226 vs 928,070). C12/C13 baseline correlation partial (per-sample CG R2~0.85-0.91 strong; pooled ~0.53 - possible discrepancy, flagged). C16/C17 SNP: the freebayes SNP-calling pipeline reproduces and runs (575,780 biallelic SNPs/44 samples) and its calls overlap the 1001genomes baseline (thousands of genotype-matched true positives per divergent accession via rtg-vcfeval), but the paper's headline >90% precision / ~80% sensitivity are NOT reached with GQ-thresholded raw freebayes calls (reproduced precision ~0.56, sensitivity ~0.40 squash best-F) - the paper's figures require the authors' high-quality + both-strand call-set filtering (epiGBS-Benchmarking repo), which was not replicated; graded partial. NOTE: this room was REQUEUED after a «infra» janitor purge wiped all intermediates; the entire reference branch + SNP chain was regenerated from scratch (recipe recovered from prior-run transcripts) to close C16/C17. AUDIT FLAG: SRA library has 10 C24 samples vs paper's stated 8 (C15). NOT ATTEMPTED (documented): C14 de novo CHG correlation (de novo branch not regenerated this round), C18-C23 P. major (separate organism + external RRBS, out of primary scope).

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-30
Rubric version
not recorded
Assessed by
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The authors ask whether an improved epiGBS-based laboratory protocol and revised, user-friendly bioinformatics pipeline (epiGBS2) can accurately and efficiently determine de novo cytosine methylation and SNP variation in species with or without a reference genome, performing well against established benchmark datasets.

Core claims
  • epiGBS2 provides a laboratory protocol and revised bioinformatics pipeline for de novo cytosine methylation and SNP calling in species with or without a reference genome resource
  • The bioinformatics pipeline is integrated into the snakemake workflow management system, making it modular, reproducible and flexible method
  • epiGBS2 implements bismark for alignment and methylation calling, replacing the legacy samtools mpileup-based approach method
  • Alignment files are preprocessed by double masking to enable SNP calling with freebayes (epifreebayes) method
  • The protocol allows flexible choice of restriction enzyme pairs, including a double digest, via a control nucleotide that still enables Watson/Crick strand differentiation mechanism
  • Hemimethylated adapter pairs (instead of fully methylated adapters) reduce library preparation costs while nick translation with 5mC-dNTPs yields fully methylated adapters mechanism
  • A random 3-nucleotide UMI placed in the adapter enables computational identification and removal of PCR duplicates using stacks2 clone_filter method
  • Performance evaluation against Arabidopsis thaliana and great tit (Parus major) baseline datasets confirmed overall good performance of epiGBS2 finding
Experimental setups
Assay System Perturbation Readout Platform
epiGBS2 reduced representation bisulfite sequencing (full pipeline) Arabidopsis thaliana accessions none methylation calling and SNP calling performance compared to published benchmark data set bismark, freebayes (epifreebayes)
epiGBS2 reduced representation bisulfite sequencing (full pipeline) great tit (Parus major) none methylation calling and SNP calling performance compared to published benchmark data set bismark, freebayes (epifreebayes)
PCR duplicate removal epiGBS2 sequencing reads (Watson/Crick, UMI-tagged) none identification/removal of PCR clones based on identical fragment and UMI sequence clone_filter (stacks2)
demultiplexing and Watson/Crick strand annotation epiGBS2 sequencing reads none sample assignment and Watson/Crick strand classification via control nucleotide and barcode process_radtags (stacks2)
de novo reference reconstruction epiGBS2 reads from A. thaliana and great tit none consensus reference sequence clusters from deduplication, Watson/Crick pairing and identity clustering pear, custom scripts
adapter trimming demultiplexed, strand-annotated epiGBS2 reads none 3' adapter removal and minimum read length filtering (20 bp cutoff) cutadapt
read mapping and DNA methylation calling de novo clusters or pre-existing reference genome none alignment and cytosine methylation levels bismark
SNP calling double-masked bismark alignment files none single nucleotide polymorphism variants freebayes (epifreebayes)
Key results
  • Performance of critical epiGBS2 steps evaluated against A. thaliana and great tit baseline data sets confirmed overall good performance
Key statistics
  • other 46% (proportion of BS-Seq data sets in the SRA generated from mouse)
  • other 34% (proportion of BS-Seq data sets in the SRA generated from human)
  • other 4% (proportion of BS-Seq data sets in the SRA generated from Arabidopsis thaliana)
  • other 3-nucleotide UMI (length of unique molecular identifier used in adapter for PCR duplicate removal)
  • other snakemake version 6.1.1 (version of workflow management system used to build the epiGBS2 pipeline)
  • other minimum read length 20 bp (length cutoff used for filtering reads after adapter trimming)
  • other 10 bp hard trim (additional bases removed from adapter-trimmed reads to discard UMI, barcode and control nucleotide)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This is a methods/tool paper describing epiGBS2, a laboratory protocol and snakemake-based bioinformatics pipeline for reduced-representation bisulfite sequencing. The manuscript text provided (introduction and methods sections) focuses on describing the wet-lab protocol and computational workflow (PCR duplicate removal, demultiplexing, de novo reference construction, adapter trimming, bismark-based alignment/methylation calling, and freebayes-based SNP calling) and states that pipeline performance was evaluated against baseline datasets from Arabidopsis thaliana and great tit (Parus major). The excerpt available does not include the Results section, so no specific inferential statistical test, p-value, or effect-size reporting is described in the text provided.

Replicationunclear GroupsepiGBS2 pipeline output compared against baseline/benchmark datasets from A. thaliana and great tit (Parus major); specific comparison metrics not stated in the provided text Pairingunclear Randomization/blindingnot stated Dispersionnone
Approaches that could also have been used
  • Pipeline performance is described as having been 'evaluated' against baseline datasets from A. thaliana and great tit without a specific quantitative statistic named in the available text.
    Could also: Concordance metrics such as Pearson/Spearman correlation, concordance correlation coefficient, or Bland-Altman agreement analysis could also be used to summarize agreement between epiGBS2 outputs and baseline/reference calls. — These approaches quantify both the strength and the direction of any systematic bias between two measurement methods, which can complement a qualitative description of 'good performance.'
  • SNP calling performance and methylation calling performance are each described narratively as being benchmarked against reference datasets.
    Could also: Standard classification-accuracy metrics (precision, recall/sensitivity, F1 score) relative to a gold-standard call set could also be reported. — These metrics are commonly used in bioinformatics tool benchmarking to give a standardized, comparable measure of calling accuracy across studies and tools.
  • The de novo reference construction and clustering steps involve tunable parameters (minimum/maximum cluster depth, identity percentage) described as adjustable by the user.
    Could also: A sensitivity analysis or parameter sweep with a quantitative summary (e.g., variance in the number/quality of loci recovered across parameter settings) could also be used to characterize how outcomes depend on these settings. — This would provide readers with an explicit, quantitative picture of how parameter choices affect downstream results, complementing the qualitative guidance given in the text.
  • The comparison of the new de novo/reference branches to the previous 'legacy' epiGBS pipeline is described narratively as a methodological update.
    Could also: A paired or matched-sample comparison (e.g., Wilcoxon signed-rank test or paired t-test on a relevant summary metric per sample) between legacy and updated pipeline outputs on the same samples could also be used. — A formal paired comparison would let readers see whether differences between pipeline versions are consistent across samples and how large they are, in addition to the qualitative description provided.
Software: snakemake 6.1.1 · bismark · freebayes (epifreebayes) · cutadapt · stacks2 · conda

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-35178872 (epiGBS2)

Paper: Gawehns et al. 2022, Mol Ecol Resour. "epiGBS2: Improvements and evaluation of highly multiplexed, epiGBS-based reduced representation bisulfite sequencing." DOI 10.1111/1755-0998.13597.

Nature of the paper: a methods/tool paper presenting the epiGBS2 Snakemake pipeline (https://github.com/nioo-knaw/epiGBS2). The "results" are an evaluation of the pipeline on two empirical datasets. This is the canonical P16 case: the repo IS the authors' own pipeline; reproducing = running that pipeline on the paper's own data with the described parameters.

Datasets the paper relies on

  • PRJNA764918 (SRA) — A. thaliana epiGBS2, 1 multiplexed library, 44 barcoded samples (6 accessions). Raw also on Zenodo 5519370. Primary in-scope data.
  • Zenodo 5878925 — small example dataset shipped to test the pipeline (smoke test).
  • Zenodo 4764652 — pipeline code archive.
  • GitHub MaartenPostuma/epiGBS-Benchmarking — benchmarking/validation scripts (minimap2 cluster validation, rtg vcfeval SNP benchmark, methylKit correlations).
  • P. major (great tit) epiGBS2 + a prior RRBS dataset — used in §3.6. The RRBS comparison data accession is not the primary deposit; lower priority / likely out of scope.
  • External baselines (NOT our data, used for correlation): 1001 Genomes SNPs, 1001 Epigenomes methylation, TAIR10 reference genome.

IN SCOPE (pipeline-derived, reproducible by running epiGBS2)

Ordered by tractability (80% floor = top items):

id result branch pipeline stage
C07/C08 adapter content before / reads retained after trimming both cutadapt (preprocessing)
C01 de novo clusters generated (19,766) denovo clustering (output_denovo)
C05/C06 per-accession mapping % (ref 50-70%, denovo 48-57%) both bismark alignment reports
C09/C10/C11 methylation sites de novo / ref / overlap both bismark methylation calling (CX reports)
C02/C03/C04 de novo cluster → TAIR10 alignment (87.69%, 10,960 perfect, mean 1.1 mm) denovo minimap2 validation (benchmarking repo)
C12/C13/C14 CG/CHG correlation vs 1001 epigenomes both methylKit + baseline
C16/C17 SNP sensitivity vs 1001 Genomes both freebayes + rtg vcfeval

OUT OF SCOPE (not pipeline-derived from our primary data, or external)

  • Wet-lab protocol improvements (barcode design, RE choice, library prep) — manual.
  • §3.6 P. major RRBS comparison (C18–C23): requires a separate prior RRBS dataset
    • great-tit reference; reproduce only if the great-tit epiGBS2 data + RRBS data prove openly obtainable. Marked lower priority.
  • Baseline datasets themselves (1001 Genomes/Epigenomes) — external references, used for comparison, not reproduced.

Plan

  1. Clone epiGBS2 + benchmarking repos on «infra»; pin commit.
  2. Download PRJNA764918 (44 samples) + barcode file to «infra» via front1.
  3. Smoke-test pipeline on Zenodo example data (fast, validates env).
  4. Run denovo + reference branches on A. thaliana data (SLURM, compute only).
  5. Extract: cluster count, mapping %, methylation site counts; validate clusters vs TAIR10.
  6. Compare to claims.tsv; grade in agreement.json + AUDIT.md.

Compute note: bisulfite alignment (bismark/bowtie2) on 44 paired-end 150bp libraries is the heavy step → SLURM on «our HPC», conda envs built on front1.

Figures / tables: Fig 5Fig 6Fig 7Fig 9
C05
Reported
50-70% mapping, Col-0 best
Reproduced
44/44 samples 53.1-68.3%, all within 50-70%; Col-0 mean 66.6% highest
within tolerance
C06
Reported
48-57% mapping, de novo branch
Reproduced
44/44 samples; per-accession means 50.1-54.3% (Ler0 50.1, Gu0 52.0, C24 53.4, Col0 53.6, Cvi0 53.7, Ei2 54.3) all within 48-57%; global range 40.6-57.4% (40.6 = low-coverage outlier C24_29); denovo maps lower than ref branch as claimed
within tolerance
C07
Reported
up to 18% reads contain Illumina adapter before trimming
Reproduced
12.6-35.3% 'reads with adapters' across 176 cutadapt adapter-log entries (max 35.3%); NOT cleanly comparable - SRA reads are post-demultiplex and the paper's stat is on raw library reads
partial
C08
Reported
99.99% reads retained after 3' trim
Reproduced
99.9949% (16854494/16855354 pairs)
exact
C09
Reported
1047597 methylation sites, de novo branch
Reproduced
1738348 union covered cytosines (44/44 samples) = 1.66x reported; same entry-point cluster inflation as C01 (29112 vs 19766 clusters)
did not match
C10
Reported
1827467 methylation sites, reference branch
Reproduced
1827216 (union covered cytosines, 44 samples) = 99.986% match
within tolerance
C11
Reported
928070 overlapping methylation sites (de novo ∩ reference)
Reproduced
371226 (best-effort minimap2 liftover of denovo cluster sites to TAIR10 then intersect with ref union; 315842/1738348 denovo sites unliftable). Low concordance driven by entry-point divergence + simplified position-only liftover
did not match
C12
Reported
CG methylation R2>0.93 vs 1001 epigenomes (both branches)
Reproduced
ref branch per-sample CG R2~0.85-0.91 (strong, consistent); but authors' OWN pooled benchmark script reproduces only R2~0.53 - POSSIBLE DISCREPANCY
partial
C13
Reported
CHG methylation R2>0.79 (reference branch)
Reproduced
per-sample CHG R2~0.59-0.79 (reaches 0.79 only Cvi-0@30x); pooled ~0.49
partial
C16
Reported
>90% high-quality SNPs match 1001 Genomes baseline
Reproduced
epiGBS2 freebayes SNP pipeline REGENERATED + run («job» ref-align 44/44 -> 2258188 freebayes = 575780 biallelic SNPs/44 samples) and benchmarked with rtg-vcfeval vs 1001genomes per-accession reference VCFs restricted to each sample's covered regions («job»). Genotype-matched precision (squash, best-F GQ) per divergent accession: Cvi-0 0.638, Ei-2 0.535, Gu-0 0.516, Ler-0 0.556 (overall 0.563); even at most stringent GQ precision plateaus ~0.55-0.67, NOT >90%. Calls substantially overlap baseline (thousands of TP/accession) but the >90% figure is not reached with GQ-thresholded raw freebayes calls - the paper's >90% applies to HIGH-QUALITY + both-strand-filtered SNPs (authors' epiGBS-Benchmarking script), a call-set filter not replicated here.
partial
C17
Reported
~80% filtered SNPs detected in baseline
Reproduced
rtg-vcfeval sensitivity (squash, best-F) per divergent accession: Cvi-0 0.397, Ei-2 0.427, Gu-0 0.421, Ler-0 0.318 (overall 0.403); ~0.40-0.45 of baseline SNPs in covered regions recovered with raw freebayes calls, below the paper's ~80%. Same root cause as C16: authors' high-quality/both-strand call filtering not replicated. Col-0 excluded (it IS the TAIR10 reference -> ~no real SNPs vs itself, correctly ~empty baseline).
partial
C15
Reported
44 A. thaliana samples (8 C24,8 Col-0,8 Cvi-0,7 Ei-2,8 Gu-0,3 Ler-0=42)
Reproduced
44 deposited; composition C24:10,Col0:8,Cvi0:8,Ei2:7,Gu0:8,Ler0:3
partial
C01
Reported
19766 de novo clusters
Reproduced
29112 clusters (entry-point difference: SRA reads are already demultiplexed/trimmed; raw-read barcode file unpublished)
did not match
C02
Reported
87.69% clusters uniquely align to TAIR10
Reproduced
89.01% (25914/29112) uniquely aligning (minimap2 2.17)
within tolerance
C03
Reported
10960 clusters perfect match to TAIR10
Reproduced
9774 (NM:i:0)
partial
C04
Reported
1.1 mean mismatches vs TAIR10
Reproduced
13.89 (inflated by entry-point/bisulfite C-T in consensus; same root cause as C01)
did not match
C14
Reported
CHG methylation correlation de novo branch R2>0.8
Reproduced
NOT ATTEMPTED (deliberate deprioritization). Requires a full de-novo-branch pipeline re-run (~5h, the de novo «infra» outputs were lost to the janitor purge and only the reference branch was regenerated this round) plus cluster->TAIR10 liftover of de novo methyl sites. Low value: the de novo branch already shows a documented entry-point divergence (C01/C09/C11 mismatch), and the reference-branch CHG correlation C13 only reached partial (~0.59-0.79), so a de novo CHG R2>0.8 is unlikely to reproduce cleanly and would not change the overall verdict.
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

No assessment has been recorded yet.
🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.