Genome-wide insights into population structure and host specificity of Campylobacter jejuni.
The main results reproduced: recomputed values matched the published ones within tolerance.
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
REPRODUCED (in-scope MLST pipeline, 1:1). Named code = tseemann/mlst (third-party tool, valid per P16). Ran SPAdes v3.11.1 (Python 3.7.12, default) -> mlst 2.35.0 (PubMLST C. jejuni/coli 7-gene scheme, DB snapshot 2026-03-11) on all 323 paired-end WGS runs of PRJNA648048, on «our HPC» SLURM/«infra». vs Supplementary Table S1: 321/321 RESOLVABLE STs EXACT (100%, ZERO true mismatches); 2 isolates not exactly resolved by mlst on our independent assembly (inexact alleles), giving 99.38% incl. unresolved. CC concordance 290/291 gradable = 99.66%; the few CC differences are all PubMLST database drift (paper ~2020 vs 2026 snapshot: ST-658 CC-177->CC-658; 2 paper-'unknown' STs now get a CC). The headline host-specificity claims reproduce qualitatively from the CC x host crosstab: CC-42/CC-61 cattle, CC-257/CC-353/CC-1034 chicken, CC-403 pig, CC-21/CC-45/CC-48 generalist. Host census reproduces Table S1 (Cattle97/Chicken101/Human96/Pig29 = 323) off by one chicken isolate (SRR10103068, single-end, different BioProject). DID NOT attempt (out of scope, separate pipelines): RAxML/BAPS phylogeny, pan-genome counts, pyseer GWAS, Canadian isolates. POSSIBLE-FABRICATION FLAG: the paper's MAIN TEXT host counts (Cattle 98 / Pig 28) contradict its own Table S1 (97/29); the data match Table S1 -> internal paper inconsistency, not a reproduction error. All grades provisional; a human reviewer signs off via AUDIT.md.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-25
- Rubric version
- not recorded
- Assessed by
- —
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetThe study tests whether identifiable genetic determinants in the core and accessory genome of Campylobacter jejuni underlie host-specific adaptation (to chicken, cattle, pig) versus host-generalist lifestyles, using genome-wide association across diverse strains.
- ★ Both core and accessory genome characteristics show strong association with distinct host animal species, indicating multiple independent adaptive trajectories rather than a single common evolutionary path finding
- ★ Host adaptation in C. jejuni is a long evolutionary, multifactorial process expressed through gene presence/absence (accessory genome) and allelic variation (core genome) finding
- ★ Host-specific allelic variants of genes including dnaE, rpoB, ftsX and pycB, involved in genome maintenance and metabolic pathways, distinguish lineages by host preference finding
- ★ Host-specific lineages (cattle, chicken) are phylogenetically closer to host-generalist lineages than to other host-specific lineages of the same host, rejecting a common evolutionary background hypothesis for host specificity finding
- ★ Host-generalist BAPS clusters have a broader/more diverse accessory gene pool and core genome population structure than host-specific clusters finding
- Pig-associated genomes carry accessory genes for type I and type II restriction-modification (RM) systems exclusive to pig hosts finding
- ★ A k-mer-based GWAS combined with bootstrapping/stratified random sampling can identify host-associated core and accessory genomic features across a population of 490 genomes method
- Cattle-associated CC-61 lineages carry a 9.9 kb accessory locus encoding a HicA-HicB toxin/antitoxin system, extracytoplasmic stress protein YafQ, and plasmid repair protein RepA finding
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Whole-genome sequencing and core/accessory genome analysis | 490 C. jejuni isolates from animal, human and environmental sources (Germany, Canada) | none | core gene count, accessory gene count, genome size | — |
| Core genome phylogenetics with BAPS clustering | 490 C. jejuni genomes | none | phylogenetic branches/clusters, association with clonal complex, lifestyle, geography | BAPS |
| k-mer based genome-wide association study (GWAS) | C. jejuni genomes stratified by host lifestyle (pig, cattle, chicken, host-generalist) | none | significant k-mers mapped to accessory genes and core-gene allelic variants associated with host | — |
| Recombination detection | Core genome alignment of 490 C. jejuni isolates | none | recombination events/profile across lineages | BRATNextGen; visualized in Phandango |
| t-SNE dimensionality reduction of accessory genome profiles | 490 C. jejuni genomes | none | clustering of accessory gene content by sample origin, BAPS cluster, lifestyle | — |
| Predicted amino acid sequence comparison / phylogenetic tree construction | Core gene loci (e.g. dnaE, ffh, rpoB, ftsX, dxs, tenI) across lineages | none | non-synonymous substitutions / allelic variant distribution by lifestyle | — |
- – 1,111 core genes identified, covering 60% of the average C. jejuni genome, plus 7,250 accessory genes across the dataset
- – 15 distinct phylogenetic branches/BAPS clusters identified among 490 genomes, concordant with known CC-lifestyle associations (e.g., CC-42/CC-61 cattle, CC-353 etc. chicken, CC-403 pig, CC-21/CC-45/CC-48 host-generalist)
- – 21,681 significant k-mers identified for pig-associated genomes, mapping to 49 accessory genes and 78 core genome allelic variants 21,681 k-mers; 49 genes; 78 variants
- – 66,491 significant k-mers identified for cattle-associated genomes, mapping to 71 accessory genes and 136 core gene variants 66,491 k-mers; 71 genes; 136 variants
- – 14 accessory genes found exclusively in pig-host genomes, including three type II RM system genes and one type I RM hsdR gene 14 genes
- – dnaE and ffh allelic variants identified as cattle-specific within a 9.7 kb ribosomal-complex-encoding locus of 9 adjacent core genes
- – Cattle-related BAPS cluster 4 more closely related to host-generalist BAPS cluster 6 than to other cattle-related cluster 10; similarly chicken clusters show closer ties to generalist clusters than to each other
- – Pig- and cattle-associated lineages (BAPS 11, BAPS 4) show low recombination rates with other lineages, suggesting lineage-specific recombination barriers, while host-generalist and some chicken lineages show frequent recombination exchange
- count 1,111 core genes (core genome size, 60% of average genome)
- count 7,250 accessory genes (total accessory gene content across dataset)
- mean 1,690,635 bp (average C. jejuni genome size)
- count 15 phylogenetic branches/BAPS clusters (population structure clustering)
- count 21,681 k-mers; 49 accessory genes; 78 core variants (pig-associated GWAS hits (BAPS cluster 11, CC-403))
- count 66,491 k-mers; 71 accessory genes; 136 core variants (cattle-associated GWAS hits (BAPS clusters 4 and 10))
- count 14 accessory genes exclusive to pig hosts (pig-specific accessory gene content)
- count 16 accessory genes in 9.9 kb region (CC-61 (BAPS cluster 10) cattle-associated locus (NCTC13261_01705–01720))
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
The study used a population-genomics approach on 490 Campylobacter jejuni genomes: core-genome phylogenetic reconstruction with BAPS clustering to define lineages, a stratified random sampling/bootstrapping scheme to assemble host-origin groups, and a k-mer-based genome-wide association ('consensus GWAS approach') to link core and accessory genome features with host lifestyle. Recombination was assessed separately with BRATNextGen, and population/accessory-genome structure was visualized with phylogenetic trees, minimum spanning trees, and t-SNE plots. Associations are reported as p-values and k-mer/gene frequencies per lifestyle group (Tables 1–2, Figure S2) rather than through classical two-group hypothesis tests.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| k-mer-based genome-wide association ('consensus GWAS approach') | identifying core/accessory genome features associated with host lifestyle (pig, cattle, chicken, host-generalist) | 490 genomes; per-lifestyle counts as tabulated (e.g. 26 pig, 56 cattle, 90 chicken, 255 host-generalist, 63 other) | not stated |
| BAPS (Bayesian Analysis of Population Structure) clustering | assignment of core-genome phylogenetic branches to 15 BAPS clusters (Fig. 1) | 490 core genomes | not stated |
| Bootstrapping based on stratified random sampling | construction of the genome set/groups used for the GWAS | 490 genomes | not stated |
| Recombination detection (BRATNextGen) | identifying significant recombination events across the core genome alignment (Fig. 3) | 490 isolates | not stated |
| t-SNE dimensionality reduction | visualizing accessory genome profile clustering by origin, BAPS cluster, and lifestyle (Fig. 2b–d) | 490 genomes | na |
-
Host-associated genomic features were identified via a k-mer-based GWAS across genomes with underlying clonal population structure.↳ Could also: Bacterial GWAS tools using linear mixed models that explicitly incorporate phylogenetic relatedness as a random effect (e.g., pyseer, treeWAS) — These approaches can help separate host-association signals from associations that arise simply from shared ancestry between related lineages, a known consideration in bacterial GWAS.
-
Balance across host-origin groups was achieved through bootstrapping based on stratified random sampling.↳ Could also: A formal sample-size/power assessment for each lifestyle group — This would quantify the statistical power available for the smaller groups (e.g., pig, n=26) to detect lifestyle-associated variants, complementing the resampling strategy already used.
-
Population structure was defined using BAPS clustering of the core-genome alignment.↳ Could also: Complementary clustering methods such as fastBAPS, hierBAPS, or ancestry-based approaches (e.g., ADMIXTURE/STRUCTURE) — Cross-checking cluster boundaries with an independent method can corroborate lineage assignments, and some of these tools scale differently on large genomic datasets.
-
Recombination events were identified using BRATNextGen.↳ Could also: Gubbins or ClonalFrameML — These tools use alternative statistical models for detecting recombinant regions and are often shown alongside BRATNextGen results as a complementary cross-check in bacterial phylogenomics.
-
Accessory genome relationships were visualized with t-SNE.↳ Could also: UMAP or classical PCA/MDS shown alongside t-SNE — Different dimensionality-reduction methods preserve different aspects of structure (e.g., UMAP often better preserves global distances, PCA offers directly interpretable linear axes), so presenting more than one can reinforce confidence in the observed groupings.
-
Significance of k-mer/gene associations is conveyed via p-values and presence/absence frequency counts.↳ Could also: Reporting effect sizes (e.g., odds ratios or allele-frequency differences) with confidence intervals, alongside an explicitly stated multiple-testing correction threshold — This would let readers assess both the magnitude and precision of each association and clarify how false-positive risk was controlled given the very large number of k-mers tested.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
scope.md — pmid-33990625
Paper: Epping L, et al. Genome-wide insights into population structure and host specificity of Campylobacter jejuni. Sci Rep 2021. PMID 33990625 / PMC8121833 / DOI 10.1038/s41598-021-89683-6.
Named code: https://github.com/tseemann/mlst (third-party tool — valid per BRIEF P16). Data: NCBI BioProject PRJNA648048 (Illumina WGS raw reads, German isolates).
Reported pipeline (Methods)
- Assembly: SPAdes v3.11.1, default settings.
- MLST: BLAST-based tool
mlst(tseemann/mlst) against the C. jejuni/coli PubMLST 7-gene scheme (aspA, glnA, gltA, glyA, pgm, tkt, uncA) → sequence type (ST)- clonal complex (CC).
- Downstream (separate tools, see below): core-genome alignment + RAxML phylogeny, BAPS clustering (15 branches), pan-genome (Roary-class: 1,111 core / 7,250 accessory genes), pyseer k-mer GWAS for host association.
IN SCOPE (matches the named repo tseemann/mlst)
- MLST sequence typing of the German isolates: per-isolate ST and CC, derived by
mlston SPAdes assemblies. Ground truth = Supplementary Table S1 (per-isolate source + ST for all 490 genomes). - Host distribution of clonal complexes — the paper's headline host-specificity claims are at the CC level (CC-42 & CC-61 → cattle; CC-257/CC-353/CC-1034 → chicken; CC-403 → pig; CC-21/CC-45/CC-48 → host-generalists). These CCs come directly from MLST output → reproducible by tabulating our ST→CC assignments against host source.
- Sample/host census of PRJNA648048 (N per host) vs the paper's reported N.
OUT OF SCOPE (different pipelines, not the named repo; not attempted here)
- RAxML core-genome phylogeny / the 15 BAPS branches (Fig 1) — separate alignment +
RAxML + BAPS stack; heavy, not
mlst. - Pan-genome counts (1,111 core / 7,250 accessory genes) — Roary-class, not
mlst. - pyseer k-mer GWAS host-association statistics — separate tool.
- Wet-lab / phenotype work. These may be revisited after the MLST core is reproduced (80/20 is a floor), but the primary, directly-pinnable reproduction tied to the named code is the MLST typing (1–3).
Data note
- PRJNA648048 = 323 runs (paper says 324 German isolates → off by one).
ENA
isolation_source: chicken 101 / cattle 97 / human 96 / pig 29 (=323) vs paper chicken 102 / cattle 98 / human 96 / pig 28 (=324). The 166 Canadian isolates of the 490 total are from external/other collections, not in this BioProject → MLST repro is on the 323 German runs. - No assemblies deposited (NCBI assembly search for PRJNA648048 → 0) → must assemble from reads with SPAdes (faithful to Methods).
Repro plan («our HPC» / «infra»)
Per-isolate, self-cleaning to bound peak disk (118 GB total reads):
download fastq → SPAdes 3.11.1 (default) → mlst → keep tiny TSV (run, ST, alleles, CC)
- small contigs checksum; delete reads + SPAdes working dir. Aggregate ST/CC table, join host source, compare to Table S1.
Caveat (database drift): mlst's bundled PubMLST profile DB is a 2026 snapshot; the
paper typed in ~2020. STs already in PubMLST then will match; isolates the paper called
"novel ST" may now have assigned ST numbers. Recorded as a provisional grade, not a
mismatch where drift explains it.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.