Evaluation of core genome and whole genome multilocus sequence typing schemes for Campylobacter jejuni and Campylobacter coli outbreak detection i
The main results reproduced: recomputed values matched the published ones within tolerance.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Same input data as the authors
- ✓Reported values were directly comparable
- ✓Reported values are derivable from the shared data
- ✓The central claim held under reproduction
- 🟡A deviation arose in the data or preprocessing
- 🟡A deviation was attributed to the published material
- 🟡The deviation was non-trivial in magnitude
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
REPRODUCED (open-source pipeline). Lyve-SET v1.1.4f hqSNP within-outbreak pairwise SNP ranges (Table 1): 6/7 in-scope outbreaks graded = FIVE EXACT (1602VTDBR-1 0-1, 1302AKDBB-1 0-2 [C.coli], 1510WIDBR-1 0-3, 1612OHDBR-1 0, 2102NHDBR-1 0-1 — all identical to Table 1) + ONE PARTIAL (1509VTDBR-1 0-9 vs reported 0-5: 5 of 6 isolates cluster 0-5, one outlier 2015D-0144 at 9 SNPs; the excess is consistent with the phage/recombination masking the paper applied (Lyve-SET PHAST + PlasFlow) that we could NOT apply because the PHAST database service (phast.wishartlab.com) is permanently offline — so our unmasked counts are an upper bound). 0810PADBR-1 (deep-coverage, 15 isolates) still in final merge at finalization; matrix completes on «infra». VarScan params confirmed in logs = paper exactly (min coverage 20, min var freq 0.95, 100bp flanking; smalt). Reference per outbreak = Table-2 strain re-assembled with SPAdes 3.14.0. SEVEN repo/env fixes were needed to run this 2015 pipeline on a 2026 bioconda stack (missing config/LyveSET.conf; set_manage self-referential reference symlink; dead PHAST mask db; invalid --read_cleaner none; legacy samtools sort syntax; missing repo varscan.v2.3.7.jar; front-node TMPDIR absent on compute nodes) — all documented in AUDIT.md + the lyve-set kartei card. OUT OF SCOPE (declared, not attempted): cgMLST (1343 loci) + wgMLST (6623 loci) allele counts = BioNumerics v7.6.3 commercial; Baker's Gamma (Table S2). VERDICT: the paper's pipeline-derived open-source hqSNP results reproduce 1:1 (5/6 graded outbreaks exact; 1 partial fully explained by an unavailable masking database, itself a reproducibility finding).
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 92assessed: 2026-06-20 ⛓ f1621449835d
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-20
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-20no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetThe study tests whether whole-genome-sequencing-based subtyping methods (hqSNP, cgMLST, wgMLST) can differentiate outbreak-associated from sporadic Campylobacter jejuni and C. coli isolates in concordance with epidemiological data, and how well these three methods agree with one another.
- ★ cgMLST, wgMLST and hqSNP WGS-based analysis methods clustered C. jejuni and C. coli isolates in concordance with epidemiological data. finding
- ★ 68 of 73 sporadic isolates were differentiated from outbreak-associated isolates using all three methods (hqSNP, cgMLST, wgMLST). finding
- ★ cgMLST and wgMLST analyses of the isolates were highly correlated with each other (BGI, cophenetic correlation, linear regression R2 and Pearson correlation all >0.90). finding
- ★ Correlation between hqSNP analysis and the MLST-based methods was sometimes lower, with R2/Pearson between 0.60 and 0.86 and BGI/cophenetic correlation between 0.63 and 0.86 for some outbreak isolates. finding
- ★ cgMLST is sufficient for C. jejuni and C. coli outbreak detection and surveillance, while wgMLST or hqSNP can be used for further differentiation if needed. mechanism
- ★ Low average de novo coverage (<20x), low sequence length (<1.4 Mb) and low N50 values (<20 000) resulted in low core genome and whole genome allele calls; read quality, contig number and ambiguous base calls did not affect allele calling. finding
- The PulseNet cgMLST scheme contains 1343 C. jejuni/C. coli loci, and the wgMLST scheme adds 5280 further accessory loci plus 7-gene MLST loci for related Campylobacter species. resource
- ★ hqSNP analysis requires a priori selection of an appropriate reference genome and is more computationally intensive and less scalable than cgMLST/wgMLST allele-based approaches. mechanism
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| cgMLST/wgMLST allele calling and UPGMA cluster analysis | C. jejuni and C. coli isolates (n=315: 242 outbreak-associated, 73 sporadic) | none | allele differences, cluster concordance with epidemiology | BioNumerics v7.6.3 |
| hqSNP analysis (maximum-likelihood tree) | same 315 C. jejuni/C. coli isolates | none | SNP differences relative to reference genomes | Lyve-SET v1.1.4f with VarScan |
| Whole genome sequencing | C. jejuni and C. coli outbreak and sporadic isolates | none | sequence reads / de novo assemblies | Illumina sequencers (Nextera XT or DNA Prep kits); SPAdes v3.7.1/v3.14.0 |
| Pulsed-field gel electrophoresis (PFGE) | subset of C. jejuni isolates | none | SmaI/KpnI PFGE pattern combinations | BioNumerics v6.6.10 |
| In silico 7-gene MLST | C. jejuni/C. coli isolates | none | sequence type (ST) | BioNumerics v7.6.3 / PubMLST Campylobacter database |
| Sequence quality assessment | all isolate sequences plus additional low-quality sequences (short length n=6, low coverage n=15) | none (deliberately included low-quality sequences) | percentage core genome loci called, number of wgMLST alleles present vs. Q-score, coverage, length, N50, contigs, ambiguous bases | R Studio v4.1.1 with GGplot2 |
| Phylogenetic tree topology comparison (Baker's Gamma Index, cophenetic correlation) | cgMLST, wgMLST and hqSNP Newick trees from isolate dataset | none | BGI and cophenetic correlation coefficients | dendextend package in R v4.1.2 |
| Pairwise genetic distance linear regression | outbreak-related isolate pairwise distances (cgMLST/wgMLST vs hqSNP) | none | slope, y-intercept, R2, Pearson correlation coefficient | — |
- – 68/73 sporadic isolates were differentiated from outbreak-associated isolates by hqSNP, cgMLST and wgMLST. 68/73
- ▲ cgMLST and wgMLST showed high concordance with each other across BGI, cophenetic correlation, R2 and Pearson correlation. >0.90
- – hqSNP vs cgMLST/wgMLST linear regression R2 and Pearson correlation were lower for some comparisons. 0.60-0.86
- – hqSNP vs cgMLST/wgMLST BGI and cophenetic correlation were lower for some outbreak isolates. 0.63-0.86
- ▼ Low coverage (<20x), short sequence length (<1.4 Mb) and low N50 (<20 000) reduced core genome (<85% core called) and whole genome (<1300 present alleles) allele calling.
- – Analysed sequences met PulseNet QC thresholds: Q-score ≥32, length 1.59-2.12 Mb, average de novo coverage ≥29x, core loci allele calls present for 85-99% of loci (1142-1330 loci).
- count 315 isolates (237 C. jejuni + 5 C. coli outbreak-associated from 16 outbreaks; 69 C. jejuni + 4 C. coli sporadic) (study isolate composition)
- count 68/73 (sporadic isolates differentiated from outbreak isolates by all three methods)
- correlation >0.90 (BGI, cophenetic correlation coefficient, linear regression R2 and Pearson correlation between cgMLST and wgMLST)
- correlation 0.60-0.86 (linear regression R2 and Pearson correlation coefficients, hqSNP vs MLST-based methods)
- correlation 0.63-0.86 (BGI and cophenetic correlation coefficient, hqSNP vs MLST-based methods for some outbreak isolates)
- count 1343 loci (total number of cgMLST loci)
- count 6623 loci (total cgMLST loci plus accessory genome loci (wgMLST))
- other Q-score ≥32; sequence length 1.59-2.12 Mb; average de novo coverage ≥29x; core loci allele calls present for 85-99% (1142-1330 loci) (sequence quality metrics of isolates meeting PulseNet QC thresholds)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This study used a comparative/descriptive genomic epidemiology design rather than classical hypothesis testing: whole-genome sequences from 315 Campylobacter jejuni and C. coli isolates (outbreak-associated and sporadic) were subtyped by three WGS-based methods (hqSNP, cgMLST, wgMLST), clustered via UPGMA dendrograms, and the resulting phylogenies and pairwise genetic distances were compared using Baker's Gamma Index, cophenetic correlation coefficients, linear regression (R²), and Pearson correlation coefficients. Results were reported primarily as ranges of pairwise SNP/allele differences per outbreak, correlation/concordance statistics, and qualitative concordance with epidemiological data, without formal significance (p-value) testing.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Baker's Gamma Index (BGI) | comparison of phylogenetic tree topology between hqSNP, cgMLST and wgMLST dendrograms | outbreak/clade isolate sets listed in Table 1 | not stated |
| Cophenetic correlation coefficient | comparison of phylogenetic tree topology between hqSNP, cgMLST and wgMLST dendrograms | outbreak/clade isolate sets listed in Table 1 | not stated |
| Linear regression (y=mx+b, R²) | pairwise cgMLST/wgMLST allelic differences plotted against pairwise hqSNP differences for outbreak-related isolates | pairwise distances within each outbreak/clade listed in Table 1 | not stated |
| Pearson correlation coefficient | genetic distance correlations between cgMLST, wgMLST and hqSNP analysis methods | pairwise distances within each outbreak/clade listed in Table 1 | not stated |
| UPGMA cluster analysis (categorical similarity coefficient) | generation of cgMLST and wgMLST dendrograms from allele calls across all 315 isolates | 315 outbreak and sporadic isolates | not stated |
-
Pairwise genetic distances between methods (hqSNP vs cgMLST/wgMLST) were compared using Pearson correlation and linear regression R².↳ Could also: Spearman rank correlation — Since pairwise distance counts are non-negative, often skewed, and may include outlier isolates, a rank-based measure like Spearman's rho would also assess concordance without assuming a linear relationship or normally distributed residuals, which can be useful for these count-based genomic distance data.
-
Tree topology concordance was summarized with point estimates of Baker's Gamma Index and cophenetic correlation coefficients.↳ Could also: Bootstrap resampling to generate confidence intervals around BGI/cophenetic correlation values — Adding resampling-based interval estimates would also convey the precision/uncertainty of the topology-concordance statistics, complementing the single point-estimate values reported.
-
Pairwise hqSNP and allele differences within outbreaks/clades were reported as ranges (min–max) in Table 1.↳ Could also: Reporting mean and SD or IQR alongside the range — For clades with larger isolate counts, a central-tendency and spread measure (mean ± SD, or median with IQR) would also give readers a sense of the typical pairwise distance in addition to the extremes captured by the range.
-
Multiple correlation and regression comparisons were performed across 16 outbreaks/clades and three subtyping method pairs without a stated multiplicity adjustment.↳ Could also: A formal correction such as Bonferroni or Benjamini-Hochberg FDR when interpreting the collection of R²/correlation results as inferential tests — If these comparisons were treated as a family of statistical tests rather than purely descriptive summaries, a multiplicity correction would also help control the overall false-positive rate across the many pairwise comparisons.
-
Sequence-quality thresholds (coverage, length, N50) were evaluated by visual inspection of scatterplots against allele-calling completeness.↳ Could also: A formal regression-based or ROC-curve threshold analysis relating quality metrics to allele-calling failure — A quantitative threshold-selection method could also provide a statistically derived cutoff (e.g., via sensitivity/specificity trade-offs) to complement the visual/graphical determination of quality thresholds.
-
UPGMA was used to build cgMLST and wgMLST dendrograms for cluster analysis.↳ Could also: Neighbor-joining or maximum-likelihood phylogenetic reconstruction — These methods do not assume a constant rate of change across lineages (an assumption underlying UPGMA) and could also be used to cross-check clustering topology, particularly when isolates may have variable substitution rates.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
The open-source hqSNP results of Table 1 reproduce 1:1: 5 of 6 graded outbreaks are exact against public SRA data (PRJNA239251) with VarScan parameters matching the paper. The single partial (1509VTDBR-1, reproduced 0–9 vs reported 0–5) is fully explained on our/environmental side: the PHAST phage-masking database is permanently offline, so the authors' masking step could not be applied and our unmasked counts are an upper bound — 5/6 isolates still cluster at 0–5. The cgMLST/wgMLST and Baker's Gamma claims depend on commercial BioNumerics and are legitimately out of scope (q1/q2 limitation, not a defect). No fabrication concern; the central outbreak-detection claim holds.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.