Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Genome-wide insights into population structure and host specificity of Campylobacter jejuni.

Sci Rep · 2021
94/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
How its reproducibility compares
94/100
Reproducibility score
1.1 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 87% of all assessed papers rank 133 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

REPRODUCED (in-scope MLST pipeline, 1:1). Named code = tseemann/mlst (third-party tool, valid per P16). Ran SPAdes v3.11.1 (Python 3.7.12, default) -> mlst 2.35.0 (PubMLST C. jejuni/coli 7-gene scheme, DB snapshot 2026-03-11) on all 323 paired-end WGS runs of PRJNA648048, on «our HPC» SLURM/«infra». vs Supplementary Table S1: 321/321 RESOLVABLE STs EXACT (100%, ZERO true mismatches); 2 isolates not exactly resolved by mlst on our independent assembly (inexact alleles), giving 99.38% incl. unresolved. CC concordance 290/291 gradable = 99.66%; the few CC differences are all PubMLST database drift (paper ~2020 vs 2026 snapshot: ST-658 CC-177->CC-658; 2 paper-'unknown' STs now get a CC). The headline host-specificity claims reproduce qualitatively from the CC x host crosstab: CC-42/CC-61 cattle, CC-257/CC-353/CC-1034 chicken, CC-403 pig, CC-21/CC-45/CC-48 generalist. Host census reproduces Table S1 (Cattle97/Chicken101/Human96/Pig29 = 323) off by one chicken isolate (SRR10103068, single-end, different BioProject). DID NOT attempt (out of scope, separate pipelines): RAxML/BAPS phylogeny, pan-genome counts, pyseer GWAS, Canadian isolates. POSSIBLE-FABRICATION FLAG: the paper's MAIN TEXT host counts (Cattle 98 / Pig 28) contradict its own Table S1 (97/29); the data match Table S1 -> internal paper inconsistency, not a reproduction error. All grades provisional; a human reviewer signs off via AUDIT.md.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-25
Rubric version
not recorded
Assessed by
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The study tests whether identifiable genetic determinants in the core and accessory genome of Campylobacter jejuni underlie host-specific adaptation (to chicken, cattle, pig) versus host-generalist lifestyles, using genome-wide association across diverse strains.

Core claims
  • Both core and accessory genome characteristics show strong association with distinct host animal species, indicating multiple independent adaptive trajectories rather than a single common evolutionary path finding
  • Host adaptation in C. jejuni is a long evolutionary, multifactorial process expressed through gene presence/absence (accessory genome) and allelic variation (core genome) finding
  • Host-specific allelic variants of genes including dnaE, rpoB, ftsX and pycB, involved in genome maintenance and metabolic pathways, distinguish lineages by host preference finding
  • Host-specific lineages (cattle, chicken) are phylogenetically closer to host-generalist lineages than to other host-specific lineages of the same host, rejecting a common evolutionary background hypothesis for host specificity finding
  • Host-generalist BAPS clusters have a broader/more diverse accessory gene pool and core genome population structure than host-specific clusters finding
  • Pig-associated genomes carry accessory genes for type I and type II restriction-modification (RM) systems exclusive to pig hosts finding
  • A k-mer-based GWAS combined with bootstrapping/stratified random sampling can identify host-associated core and accessory genomic features across a population of 490 genomes method
  • Cattle-associated CC-61 lineages carry a 9.9 kb accessory locus encoding a HicA-HicB toxin/antitoxin system, extracytoplasmic stress protein YafQ, and plasmid repair protein RepA finding
Experimental setups
Assay System Perturbation Readout Platform
Whole-genome sequencing and core/accessory genome analysis 490 C. jejuni isolates from animal, human and environmental sources (Germany, Canada) none core gene count, accessory gene count, genome size
Core genome phylogenetics with BAPS clustering 490 C. jejuni genomes none phylogenetic branches/clusters, association with clonal complex, lifestyle, geography BAPS
k-mer based genome-wide association study (GWAS) C. jejuni genomes stratified by host lifestyle (pig, cattle, chicken, host-generalist) none significant k-mers mapped to accessory genes and core-gene allelic variants associated with host
Recombination detection Core genome alignment of 490 C. jejuni isolates none recombination events/profile across lineages BRATNextGen; visualized in Phandango
t-SNE dimensionality reduction of accessory genome profiles 490 C. jejuni genomes none clustering of accessory gene content by sample origin, BAPS cluster, lifestyle
Predicted amino acid sequence comparison / phylogenetic tree construction Core gene loci (e.g. dnaE, ffh, rpoB, ftsX, dxs, tenI) across lineages none non-synonymous substitutions / allelic variant distribution by lifestyle
Key results
  • 1,111 core genes identified, covering 60% of the average C. jejuni genome, plus 7,250 accessory genes across the dataset
  • 15 distinct phylogenetic branches/BAPS clusters identified among 490 genomes, concordant with known CC-lifestyle associations (e.g., CC-42/CC-61 cattle, CC-353 etc. chicken, CC-403 pig, CC-21/CC-45/CC-48 host-generalist)
  • 21,681 significant k-mers identified for pig-associated genomes, mapping to 49 accessory genes and 78 core genome allelic variants 21,681 k-mers; 49 genes; 78 variants
  • 66,491 significant k-mers identified for cattle-associated genomes, mapping to 71 accessory genes and 136 core gene variants 66,491 k-mers; 71 genes; 136 variants
  • 14 accessory genes found exclusively in pig-host genomes, including three type II RM system genes and one type I RM hsdR gene 14 genes
  • dnaE and ffh allelic variants identified as cattle-specific within a 9.7 kb ribosomal-complex-encoding locus of 9 adjacent core genes
  • Cattle-related BAPS cluster 4 more closely related to host-generalist BAPS cluster 6 than to other cattle-related cluster 10; similarly chicken clusters show closer ties to generalist clusters than to each other
  • Pig- and cattle-associated lineages (BAPS 11, BAPS 4) show low recombination rates with other lineages, suggesting lineage-specific recombination barriers, while host-generalist and some chicken lineages show frequent recombination exchange
Key statistics
  • count 1,111 core genes (core genome size, 60% of average genome)
  • count 7,250 accessory genes (total accessory gene content across dataset)
  • mean 1,690,635 bp (average C. jejuni genome size)
  • count 15 phylogenetic branches/BAPS clusters (population structure clustering)
  • count 21,681 k-mers; 49 accessory genes; 78 core variants (pig-associated GWAS hits (BAPS cluster 11, CC-403))
  • count 66,491 k-mers; 71 accessory genes; 136 core variants (cattle-associated GWAS hits (BAPS clusters 4 and 10))
  • count 14 accessory genes exclusive to pig hosts (pig-specific accessory gene content)
  • count 16 accessory genes in 9.9 kb region (CC-61 (BAPS cluster 10) cattle-associated locus (NCTC13261_01705–01720))

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The study used a population-genomics approach on 490 Campylobacter jejuni genomes: core-genome phylogenetic reconstruction with BAPS clustering to define lineages, a stratified random sampling/bootstrapping scheme to assemble host-origin groups, and a k-mer-based genome-wide association ('consensus GWAS approach') to link core and accessory genome features with host lifestyle. Recombination was assessed separately with BRATNextGen, and population/accessory-genome structure was visualized with phylogenetic trees, minimum spanning trees, and t-SNE plots. Associations are reported as p-values and k-mer/gene frequencies per lifestyle group (Tables 1–2, Figure S2) rather than through classical two-group hypothesis tests.

Replicationbiological Sample size490 total genomes from diverse origins (Germany and Canada); stratified random sampling with bootstrapping used to compose groups for GWAS; no formal power calculation stated Groupshost lifestyle groups derived from BAPS cluster/clonal complex (pig, cattle, chicken, host-generalist, other/control) Pairingna Randomization/blindingnot stated Dispersionnone Confidence intervalsno
Statistical tests used
Test Applied to n Assumptions
k-mer-based genome-wide association ('consensus GWAS approach') identifying core/accessory genome features associated with host lifestyle (pig, cattle, chicken, host-generalist) 490 genomes; per-lifestyle counts as tabulated (e.g. 26 pig, 56 cattle, 90 chicken, 255 host-generalist, 63 other) not stated
BAPS (Bayesian Analysis of Population Structure) clustering assignment of core-genome phylogenetic branches to 15 BAPS clusters (Fig. 1) 490 core genomes not stated
Bootstrapping based on stratified random sampling construction of the genome set/groups used for the GWAS 490 genomes not stated
Recombination detection (BRATNextGen) identifying significant recombination events across the core genome alignment (Fig. 3) 490 isolates not stated
t-SNE dimensionality reduction visualizing accessory genome profile clustering by origin, BAPS cluster, and lifestyle (Fig. 2b–d) 490 genomes na
Approaches that could also have been used
  • Host-associated genomic features were identified via a k-mer-based GWAS across genomes with underlying clonal population structure.
    Could also: Bacterial GWAS tools using linear mixed models that explicitly incorporate phylogenetic relatedness as a random effect (e.g., pyseer, treeWAS) — These approaches can help separate host-association signals from associations that arise simply from shared ancestry between related lineages, a known consideration in bacterial GWAS.
  • Balance across host-origin groups was achieved through bootstrapping based on stratified random sampling.
    Could also: A formal sample-size/power assessment for each lifestyle group — This would quantify the statistical power available for the smaller groups (e.g., pig, n=26) to detect lifestyle-associated variants, complementing the resampling strategy already used.
  • Population structure was defined using BAPS clustering of the core-genome alignment.
    Could also: Complementary clustering methods such as fastBAPS, hierBAPS, or ancestry-based approaches (e.g., ADMIXTURE/STRUCTURE) — Cross-checking cluster boundaries with an independent method can corroborate lineage assignments, and some of these tools scale differently on large genomic datasets.
  • Recombination events were identified using BRATNextGen.
    Could also: Gubbins or ClonalFrameML — These tools use alternative statistical models for detecting recombinant regions and are often shown alongside BRATNextGen results as a complementary cross-check in bacterial phylogenomics.
  • Accessory genome relationships were visualized with t-SNE.
    Could also: UMAP or classical PCA/MDS shown alongside t-SNE — Different dimensionality-reduction methods preserve different aspects of structure (e.g., UMAP often better preserves global distances, PCA offers directly interpretable linear axes), so presenting more than one can reinforce confidence in the observed groupings.
  • Significance of k-mer/gene associations is conveyed via p-values and presence/absence frequency counts.
    Could also: Reporting effect sizes (e.g., odds ratios or allele-frequency differences) with confidence intervals, alongside an explicitly stated multiple-testing correction threshold — This would let readers assess both the magnitude and precision of each association and clarify how false-positive risk was controlled given the very large number of k-mers tested.
Software: BAPS (population clustering) · BRATNextGen (recombination detection) · Phandango (visualization) · t-SNE

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

scope.md — pmid-33990625

Paper: Epping L, et al. Genome-wide insights into population structure and host specificity of Campylobacter jejuni. Sci Rep 2021. PMID 33990625 / PMC8121833 / DOI 10.1038/s41598-021-89683-6.

Named code: https://github.com/tseemann/mlst (third-party tool — valid per BRIEF P16). Data: NCBI BioProject PRJNA648048 (Illumina WGS raw reads, German isolates).

Reported pipeline (Methods)

  • Assembly: SPAdes v3.11.1, default settings.
  • MLST: BLAST-based tool mlst (tseemann/mlst) against the C. jejuni/coli PubMLST 7-gene scheme (aspA, glnA, gltA, glyA, pgm, tkt, uncA) → sequence type (ST)
    • clonal complex (CC).
  • Downstream (separate tools, see below): core-genome alignment + RAxML phylogeny, BAPS clustering (15 branches), pan-genome (Roary-class: 1,111 core / 7,250 accessory genes), pyseer k-mer GWAS for host association.

IN SCOPE (matches the named repo tseemann/mlst)

  1. MLST sequence typing of the German isolates: per-isolate ST and CC, derived by mlst on SPAdes assemblies. Ground truth = Supplementary Table S1 (per-isolate source + ST for all 490 genomes).
  2. Host distribution of clonal complexes — the paper's headline host-specificity claims are at the CC level (CC-42 & CC-61 → cattle; CC-257/CC-353/CC-1034 → chicken; CC-403 → pig; CC-21/CC-45/CC-48 → host-generalists). These CCs come directly from MLST output → reproducible by tabulating our ST→CC assignments against host source.
  3. Sample/host census of PRJNA648048 (N per host) vs the paper's reported N.

OUT OF SCOPE (different pipelines, not the named repo; not attempted here)

  • RAxML core-genome phylogeny / the 15 BAPS branches (Fig 1) — separate alignment + RAxML + BAPS stack; heavy, not mlst.
  • Pan-genome counts (1,111 core / 7,250 accessory genes) — Roary-class, not mlst.
  • pyseer k-mer GWAS host-association statistics — separate tool.
  • Wet-lab / phenotype work. These may be revisited after the MLST core is reproduced (80/20 is a floor), but the primary, directly-pinnable reproduction tied to the named code is the MLST typing (1–3).

Data note

  • PRJNA648048 = 323 runs (paper says 324 German isolates → off by one). ENA isolation_source: chicken 101 / cattle 97 / human 96 / pig 29 (=323) vs paper chicken 102 / cattle 98 / human 96 / pig 28 (=324). The 166 Canadian isolates of the 490 total are from external/other collections, not in this BioProject → MLST repro is on the 323 German runs.
  • No assemblies deposited (NCBI assembly search for PRJNA648048 → 0) → must assemble from reads with SPAdes (faithful to Methods).

Repro plan («our HPC» / «infra»)

Per-isolate, self-cleaning to bound peak disk (118 GB total reads): download fastq → SPAdes 3.11.1 (default) → mlst → keep tiny TSV (run, ST, alleles, CC)

  • small contigs checksum; delete reads + SPAdes working dir. Aggregate ST/CC table, join host source, compare to Table S1.

Caveat (database drift): mlst's bundled PubMLST profile DB is a 2026 snapshot; the paper typed in ~2020. STs already in PubMLST then will match; isolates the paper called "novel ST" may now have assigned ST numbers. Recorded as a provisional grade, not a mismatch where drift explains it.

Figures / tables: Table
C1
Reported
Per-isolate ST in Suppl. Table S1
Reproduced
321/321 resolvable ST exact (100%); 2/323 unresolved; 99.38% incl. unresolved; 0 true mismatches
within tolerance
C2
Reported
Per-isolate CC in Suppl. Table S1
Reproduced
290/291 gradable CC match (99.66%); remaining diffs explained by PubMLST database drift
within tolerance
C3
Reported
CC-42, CC-61 associated with cattle
Reproduced
CC-42 75% cattle (6/8); CC-61 71% cattle (15/21), 0 chicken
exact
C4
Reported
CC-257, CC-353, CC-1034 associated with chicken
Reproduced
CC-353 7/11 chicken (0 cattle/pig); CC-1034 5/6; CC-257 chicken-plurality
exact
C5
Reported
CC-403 associated with pig
Reproduced
CC-403 80% pig (20/25), 0 chicken
exact
C6
Reported
CC-21, CC-45, CC-48 are host generalists
Reproduced
all three spread across cattle/chicken/human/pig, no single-host dominance
exact
C7
Reported
324 German isolates (Table S1: Cattle 97 / Chicken 102 / Human 96 / Pig 29)
Reproduced
323 typed (Cattle 97 / Chicken 101 / Human 96 / Pig 29); off-by-one = SRR10103068 single-end in PRJNA562653
within tolerance

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

No assessment has been recorded yet.
🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.