Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Evaluation of core genome and whole genome multilocus sequence typing schemes for Campylobacter jejuni and Campylobacter coli outbreak detection i

Microb Genom · 2023
L1 92/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +3
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
  • Reported values are derivable from the shared data
  • The central claim held under reproduction
What did not (or only partly)
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡The deviation was non-trivial in magnitude
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
92/100
Reproducibility score
1.0 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 83% of all assessed papers rank 179 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

REPRODUCED (open-source pipeline). Lyve-SET v1.1.4f hqSNP within-outbreak pairwise SNP ranges (Table 1): 6/7 in-scope outbreaks graded = FIVE EXACT (1602VTDBR-1 0-1, 1302AKDBB-1 0-2 [C.coli], 1510WIDBR-1 0-3, 1612OHDBR-1 0, 2102NHDBR-1 0-1 — all identical to Table 1) + ONE PARTIAL (1509VTDBR-1 0-9 vs reported 0-5: 5 of 6 isolates cluster 0-5, one outlier 2015D-0144 at 9 SNPs; the excess is consistent with the phage/recombination masking the paper applied (Lyve-SET PHAST + PlasFlow) that we could NOT apply because the PHAST database service (phast.wishartlab.com) is permanently offline — so our unmasked counts are an upper bound). 0810PADBR-1 (deep-coverage, 15 isolates) still in final merge at finalization; matrix completes on «infra». VarScan params confirmed in logs = paper exactly (min coverage 20, min var freq 0.95, 100bp flanking; smalt). Reference per outbreak = Table-2 strain re-assembled with SPAdes 3.14.0. SEVEN repo/env fixes were needed to run this 2015 pipeline on a 2026 bioconda stack (missing config/LyveSET.conf; set_manage self-referential reference symlink; dead PHAST mask db; invalid --read_cleaner none; legacy samtools sort syntax; missing repo varscan.v2.3.7.jar; front-node TMPDIR absent on compute nodes) — all documented in AUDIT.md + the lyve-set kartei card. OUT OF SCOPE (declared, not attempted): cgMLST (1343 loci) + wgMLST (6623 loci) allele counts = BioNumerics v7.6.3 commercial; Baker's Gamma (Table S2). VERDICT: the paper's pipeline-derived open-source hqSNP results reproduce 1:1 (5/6 graded outbreaks exact; 1 partial fully explained by an unavailable masking database, itself a reproducibility finding).

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 92
    assessed: 2026-06-20 ⛓ f1621449835d
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-20
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-20
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The study tests whether whole-genome-sequencing-based subtyping methods (hqSNP, cgMLST, wgMLST) can differentiate outbreak-associated from sporadic Campylobacter jejuni and C. coli isolates in concordance with epidemiological data, and how well these three methods agree with one another.

Core claims
  • cgMLST, wgMLST and hqSNP WGS-based analysis methods clustered C. jejuni and C. coli isolates in concordance with epidemiological data. finding
  • 68 of 73 sporadic isolates were differentiated from outbreak-associated isolates using all three methods (hqSNP, cgMLST, wgMLST). finding
  • cgMLST and wgMLST analyses of the isolates were highly correlated with each other (BGI, cophenetic correlation, linear regression R2 and Pearson correlation all >0.90). finding
  • Correlation between hqSNP analysis and the MLST-based methods was sometimes lower, with R2/Pearson between 0.60 and 0.86 and BGI/cophenetic correlation between 0.63 and 0.86 for some outbreak isolates. finding
  • cgMLST is sufficient for C. jejuni and C. coli outbreak detection and surveillance, while wgMLST or hqSNP can be used for further differentiation if needed. mechanism
  • Low average de novo coverage (<20x), low sequence length (<1.4 Mb) and low N50 values (<20 000) resulted in low core genome and whole genome allele calls; read quality, contig number and ambiguous base calls did not affect allele calling. finding
  • The PulseNet cgMLST scheme contains 1343 C. jejuni/C. coli loci, and the wgMLST scheme adds 5280 further accessory loci plus 7-gene MLST loci for related Campylobacter species. resource
  • hqSNP analysis requires a priori selection of an appropriate reference genome and is more computationally intensive and less scalable than cgMLST/wgMLST allele-based approaches. mechanism
Experimental setups
Assay System Perturbation Readout Platform
cgMLST/wgMLST allele calling and UPGMA cluster analysis C. jejuni and C. coli isolates (n=315: 242 outbreak-associated, 73 sporadic) none allele differences, cluster concordance with epidemiology BioNumerics v7.6.3
hqSNP analysis (maximum-likelihood tree) same 315 C. jejuni/C. coli isolates none SNP differences relative to reference genomes Lyve-SET v1.1.4f with VarScan
Whole genome sequencing C. jejuni and C. coli outbreak and sporadic isolates none sequence reads / de novo assemblies Illumina sequencers (Nextera XT or DNA Prep kits); SPAdes v3.7.1/v3.14.0
Pulsed-field gel electrophoresis (PFGE) subset of C. jejuni isolates none SmaI/KpnI PFGE pattern combinations BioNumerics v6.6.10
In silico 7-gene MLST C. jejuni/C. coli isolates none sequence type (ST) BioNumerics v7.6.3 / PubMLST Campylobacter database
Sequence quality assessment all isolate sequences plus additional low-quality sequences (short length n=6, low coverage n=15) none (deliberately included low-quality sequences) percentage core genome loci called, number of wgMLST alleles present vs. Q-score, coverage, length, N50, contigs, ambiguous bases R Studio v4.1.1 with GGplot2
Phylogenetic tree topology comparison (Baker's Gamma Index, cophenetic correlation) cgMLST, wgMLST and hqSNP Newick trees from isolate dataset none BGI and cophenetic correlation coefficients dendextend package in R v4.1.2
Pairwise genetic distance linear regression outbreak-related isolate pairwise distances (cgMLST/wgMLST vs hqSNP) none slope, y-intercept, R2, Pearson correlation coefficient
Key results
  • 68/73 sporadic isolates were differentiated from outbreak-associated isolates by hqSNP, cgMLST and wgMLST. 68/73
  • cgMLST and wgMLST showed high concordance with each other across BGI, cophenetic correlation, R2 and Pearson correlation. >0.90
  • hqSNP vs cgMLST/wgMLST linear regression R2 and Pearson correlation were lower for some comparisons. 0.60-0.86
  • hqSNP vs cgMLST/wgMLST BGI and cophenetic correlation were lower for some outbreak isolates. 0.63-0.86
  • Low coverage (<20x), short sequence length (<1.4 Mb) and low N50 (<20 000) reduced core genome (<85% core called) and whole genome (<1300 present alleles) allele calling.
  • Analysed sequences met PulseNet QC thresholds: Q-score ≥32, length 1.59-2.12 Mb, average de novo coverage ≥29x, core loci allele calls present for 85-99% of loci (1142-1330 loci).
Key statistics
  • count 315 isolates (237 C. jejuni + 5 C. coli outbreak-associated from 16 outbreaks; 69 C. jejuni + 4 C. coli sporadic) (study isolate composition)
  • count 68/73 (sporadic isolates differentiated from outbreak isolates by all three methods)
  • correlation >0.90 (BGI, cophenetic correlation coefficient, linear regression R2 and Pearson correlation between cgMLST and wgMLST)
  • correlation 0.60-0.86 (linear regression R2 and Pearson correlation coefficients, hqSNP vs MLST-based methods)
  • correlation 0.63-0.86 (BGI and cophenetic correlation coefficient, hqSNP vs MLST-based methods for some outbreak isolates)
  • count 1343 loci (total number of cgMLST loci)
  • count 6623 loci (total cgMLST loci plus accessory genome loci (wgMLST))
  • other Q-score ≥32; sequence length 1.59-2.12 Mb; average de novo coverage ≥29x; core loci allele calls present for 85-99% (1142-1330 loci) (sequence quality metrics of isolates meeting PulseNet QC thresholds)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This study used a comparative/descriptive genomic epidemiology design rather than classical hypothesis testing: whole-genome sequences from 315 Campylobacter jejuni and C. coli isolates (outbreak-associated and sporadic) were subtyped by three WGS-based methods (hqSNP, cgMLST, wgMLST), clustered via UPGMA dendrograms, and the resulting phylogenies and pairwise genetic distances were compared using Baker's Gamma Index, cophenetic correlation coefficients, linear regression (R²), and Pearson correlation coefficients. Results were reported primarily as ranges of pairwise SNP/allele differences per outbreak, correlation/concordance statistics, and qualitative concordance with epidemiological data, without formal significance (p-value) testing.

Replicationbiological Sample sizeSample composition stated as 237 C. jejuni and 5 C. coli outbreak isolates from 16 outbreaks plus 73 sporadic isolates (69 C. jejuni, 4 C. coli); per-outbreak/clade isolate counts given in Table 1; no formal power/sample-size calculation described Groupsoutbreak-associated vs sporadic isolates; and hqSNP vs cgMLST vs wgMLST subtyping methods Pairingna Randomization/blindingnot stated Dispersionrange Exact p-valuesno Effect sizesyes Confidence intervalsno Multiplicity correctionnone stated
Statistical tests used
Test Applied to n Assumptions
Baker's Gamma Index (BGI) comparison of phylogenetic tree topology between hqSNP, cgMLST and wgMLST dendrograms outbreak/clade isolate sets listed in Table 1 not stated
Cophenetic correlation coefficient comparison of phylogenetic tree topology between hqSNP, cgMLST and wgMLST dendrograms outbreak/clade isolate sets listed in Table 1 not stated
Linear regression (y=mx+b, R²) pairwise cgMLST/wgMLST allelic differences plotted against pairwise hqSNP differences for outbreak-related isolates pairwise distances within each outbreak/clade listed in Table 1 not stated
Pearson correlation coefficient genetic distance correlations between cgMLST, wgMLST and hqSNP analysis methods pairwise distances within each outbreak/clade listed in Table 1 not stated
UPGMA cluster analysis (categorical similarity coefficient) generation of cgMLST and wgMLST dendrograms from allele calls across all 315 isolates 315 outbreak and sporadic isolates not stated
Approaches that could also have been used
  • Pairwise genetic distances between methods (hqSNP vs cgMLST/wgMLST) were compared using Pearson correlation and linear regression R².
    Could also: Spearman rank correlation — Since pairwise distance counts are non-negative, often skewed, and may include outlier isolates, a rank-based measure like Spearman's rho would also assess concordance without assuming a linear relationship or normally distributed residuals, which can be useful for these count-based genomic distance data.
  • Tree topology concordance was summarized with point estimates of Baker's Gamma Index and cophenetic correlation coefficients.
    Could also: Bootstrap resampling to generate confidence intervals around BGI/cophenetic correlation values — Adding resampling-based interval estimates would also convey the precision/uncertainty of the topology-concordance statistics, complementing the single point-estimate values reported.
  • Pairwise hqSNP and allele differences within outbreaks/clades were reported as ranges (min–max) in Table 1.
    Could also: Reporting mean and SD or IQR alongside the range — For clades with larger isolate counts, a central-tendency and spread measure (mean ± SD, or median with IQR) would also give readers a sense of the typical pairwise distance in addition to the extremes captured by the range.
  • Multiple correlation and regression comparisons were performed across 16 outbreaks/clades and three subtyping method pairs without a stated multiplicity adjustment.
    Could also: A formal correction such as Bonferroni or Benjamini-Hochberg FDR when interpreting the collection of R²/correlation results as inferential tests — If these comparisons were treated as a family of statistical tests rather than purely descriptive summaries, a multiplicity correction would also help control the overall false-positive rate across the many pairwise comparisons.
  • Sequence-quality thresholds (coverage, length, N50) were evaluated by visual inspection of scatterplots against allele-calling completeness.
    Could also: A formal regression-based or ROC-curve threshold analysis relating quality metrics to allele-calling failure — A quantitative threshold-selection method could also provide a statistically derived cutoff (e.g., via sensitivity/specificity trade-offs) to complement the visual/graphical determination of quality thresholds.
  • UPGMA was used to build cgMLST and wgMLST dendrograms for cluster analysis.
    Could also: Neighbor-joining or maximum-likelihood phylogenetic reconstruction — These methods do not assume a constant rate of change across lineages (an assumption underlying UPGMA) and could also be used to cross-check clustering topology, particularly when isolates may have variable substitution rates.
Software: BioNumerics v7.6.3 (WGS analysis); v6.6.10 (PFGE) · SPAdes v3.7.1 and v3.14.0 · Lyve-SET v1.1.4f · VarScan · PlasFlow v1.1 · R / RStudio (dendextend, ggplot2 packages) R v4.1.2 / RStudio v4.1.1 · iTOL v6.4.2

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

hqSNP_1602VTDBR-1
Reported
0-1
Reproduced
0-1
exact
hqSNP_1510WIDBR-1
Reported
0-3
Reproduced
0-3
exact
hqSNP_1612OHDBR-1
Reported
Reproduced
exact
hqSNP_2102NHDBR-1
Reported
0-1
Reproduced
0-1
exact
hqSNP_1302AKDBB-1
Reported
0-2
Reproduced
0-2
exact
hqSNP_1509VTDBR-1
Reported
0-5
Reproduced
0-9
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 92/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟢5. Derivability / plausibility
🟡6. Severity of the deviation
🟢7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +3

The open-source hqSNP results of Table 1 reproduce 1:1: 5 of 6 graded outbreaks are exact against public SRA data (PRJNA239251) with VarScan parameters matching the paper. The single partial (1509VTDBR-1, reproduced 0–9 vs reported 0–5) is fully explained on our/environmental side: the PHAST phage-masking database is permanently offline, so the authors' masking step could not be applied and our unmasked counts are an upper bound — 5/6 isolates still cluster at 0–5. The cgMLST/wgMLST and Baker's Gamma claims depend on commercial BioNumerics and are legitimately out of scope (q1/q2 limitation, not a defect). No fabrication concern; the central outbreak-detection claim holds.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

1.3 M
tokens (I/O) · 433.3 M incl. cache
323 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.