Genomic approaches used to investigate an atypical outbreak of Salmonella Adjame.
The main results reproduced: recomputed values matched the published ones within tolerance.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Same input data as the authors
- ✓Reported values were directly comparable
- ✓No relevant deviation in data/preprocessing
- ✓No authors-side cause for any deviation
- ✓Reported values are derivable from the shared data
- ✓Any deviation was negligible
- ✓The central claim held under reproduction
- ✓Overall, the reproduction was clean
- Every checked point held up.
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
DESCRIBED WELL ENOUGH -> 1:1. The paper's headline computational result (Achtman 7-gene MLST of the S. Adjame outbreak, typed with MOST) reproduces EXACTLY. On the 3 representative NCTC isolates the paper names with SRA accessions, SRR5583198 & SRR6237100 = ST3929 and SRR6190984 = ST4023 (a single-locus dnaN variant), recovering BOTH reported STs. Confirmed two independent ways: (Route A) skesa+mlst (current PubMLST senterica_achtman_2 DB) gives the ST numbers directly; (Route B) MOST, the paper's own Py2.7 tool, gives the identical underlying allele sequences (it can't print the ST number only because its 2016 bundled database predates purE allele 731 / dnaN allele 651). SISTR independently confirms serovar Adjame with antigenic formula 13,23:r:1,6 in-silico. NOT ATTEMPTED (out-of-scope hard 20%, honestly disclosed): the SnapperDB reference-mapped SNP cluster distances (max 449/692/410; within 0/0/0-9), the Enterobase cgMLST allele differences (max 101/99/49; within 0/0/0-2), and the SnapperDB SNP addresses (Table 2) -- these depend on a specific reference genome + SnapperDB params and the Enterobase-hosted cgMLST scheme that are not fully specified in the paper. No fabrication concerns: every reported value we checked is derivable from the deposited reads with the designated tool.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 100assessed: 2026-06-16 ⛓ 71b72fdeebd6
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-16
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-16no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusCan whole genome sequencing-based clustering (SNP versus cgMLST) appropriately characterize an atypical outbreak of the rare serovar Salmonella Adjame, and can cgMLST allele-based typing serve as a standardizable format for international outbreak comparison and alerting?
- ★ The S. Adjame outbreak produced a heterogeneous phylogeny with multiple temporally/geographically linked sub-clusters, atypical of a point-source Salmonella outbreak and consistent with contamination from an endemic mixed-strain source (imported South Asian herbs/spices). finding
- ★ cgMLST allele-based clustering correlated with and was comparable to SNP-derived phylogenetic analysis in defining clusters for this outbreak. finding
- ★ Incorporating fixed SNP or allelic differences into a case definition may not always be appropriate for heterogeneous outbreaks or rare serovars involving small case numbers. finding
- ★ cgMLST (cgST) is recommended for international comparison and alerting of WGS data via platforms such as EPIS, pending further validation. method
- ★ Imported herbs or spices from South Asia were the suspected vehicle of infection, though backward tracing identified no common source. finding
- SnapperDB reference-based high-quality SNP pipeline and Enterobase cgMLST V2 (3002 loci) were used and compared as parallel typing approaches. method
- All outbreak strains belonged to a single EBurst Group (EBG421) comprising two sequence types, ST4023 and ST3929. finding
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Whole genome sequencing (Illumina) | Salmonella Adjame clinical isolates (human; UK, Ireland, Denmark, Cote d'Ivoire reference) | none | FASTQ sequence reads / genome sequence | Illumina HiSeq 2500, rapid run mode, 2x100 bp |
| Reference-based SNP typing (SnapperDB pipeline) | 31 S. Adjame isolates (20 UK 2017, 6 historical UK, 1 reference, 2 Irish, 1 Danish, 1 Cote d'Ivoire) | none | high-quality SNP positions, SNP distances, maximum-likelihood phylogeny/clusters | BWA mem, Samtools, GATK2, SnapperDB, RAxML v8.2.8; reference SRR5583198 assembled with SPAdes v3.8.0 |
| cgMLST allele-based typing | S. Adjame isolates | none | core genome sequence type (cgST) over 3002 loci, minimum spanning tree, allele differences | Enterobase (SPAdes assembly, cgMLST V2, MSTreeV2 algorithm) |
| 7-gene MLST / sequence typing | S. Adjame isolates | none | ST and EBurst Group assignment | MOST (Metric Orientated Sequence Typer), Achtman MLST database |
| Phenotypic serology (serotyping) | S. Adjame isolates (first three of each ST) | none | antigenic profile (13,23:r:1,6) confirming serovar | White-Kauffman-Le Minor scheme |
| Epidemiological investigation (trawling/targeted questionnaire) | 14 human outbreak cases in England | none | demographics, clinical info, food/travel exposures | Stata-13, Microsoft Excel 2010 |
| Food trace-back investigation | Commonly reported food items/spice brands and South Asian grocers | none | supply chain provenance / country of origin | — |
- – 14 cases met the outbreak case definition, 11 confirmed S. Adjame in London plus one each in South East, East of England, South West England 14 cases
- – SNP clustering revealed three genetically distinct clusters plus single outlying strains; 13/14 cases fell into the Red and Green clusters with one strain unrelated 3 clusters
- – Large SNP distances between clusters indicate marked heterogeneity atypical of point-source outbreak max 449 SNPs (cluster 1-2), 692 SNPs (1-3), 410 SNPs (2-3)
- – Red cluster comprised eight strains (late June/early July 2017) with min 0, max 9 SNP distance; blue and green clusters each had zero SNP distance internally min 0, max 9 SNPs
- – cgMLST clusters were comparable to SNP-derived phylogenetic clusters
- – All strains fell into one EBurst Group (EBG421) comprising ST4023 and ST3929 1 EBG, 2 STs
- – Four EU countries (Italy, Germany, Ireland, Denmark) reported S. Adjame cases in response to EPIS alert; suspected link to South Asian groceries
- – Backward tracing found multiple possible sources with no common supplier; spices frequently sourced from South Asian spice farms outside Europe
- count 14 cases (cases meeting outbreak case definition)
- mean median age 66.5 years (range 3–85) (age of outbreak cases)
- count 31 S. Adjame SNP-typed (20 UK 2017, 1 reference, 6 historical UK, 2 Irish, 1 Danish, 1 Cote d'Ivoire)
- count 449 / 692 / 410 SNPs (maximum SNP distances between clusters 1-2, 1-3, 2-3)
- count 3002 genes/loci (cgMLST V2 scheme loci present in >98% of 3144 reference genomes)
- count 10/11 exposed to brand X, 8/11 to Y, 7/11 to Z (likelihood of spice brand exposure among cases)
- count pepper and turmeric 9/10 (most commonly reported herb/spice exposures)
- count 8 grocer A, 4 grocer B, 3 grocer C (number of cases purchasing from each South Asian grocer)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This outbreak investigation report combined descriptive epidemiology with comparative whole-genome sequencing (WGS) genomics for 14 cases of Salmonella Adjame. Epidemiological data from trawling and targeted questionnaires were summarised using descriptive statistics (counts, proportions, medians with ranges) in Stata-13 and Excel; no inferential hypothesis tests were applied. Genomic relatedness was assessed through two parallel approaches: a high-quality SNP pipeline (SnapperDB) with single-linkage hierarchical clustering and RAxML maximum-likelihood phylogenetics, and an allele-based cgMLST scheme (Enterobase/MSTreeV2 minimum spanning tree). The primary analytic goal was to compare the concordance of SNP- and cgMLST-derived cluster assignments rather than to test any pre-specified statistical hypothesis.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Descriptive statistics (counts, proportions, median with range) | Demographic and exposure data for 14 outbreak cases (age, illness duration, food exposures, hospitalisations) | 14 cases; sub-denominators vary (e.g. 10/11, 9/10) due to missing questionnaire responses | na |
| Maximum-likelihood phylogeny (RAxML v8.2.8, GTR+GAMMA, 1000 bootstrap replicates) | SNP-based phylogenetic tree of 31 S. Adjame isolates mapped against a multi-contig reference (SRR5583198) | 31 isolates (20 UK 2017, 6 historical UK, 2 Irish, 1 Danish, 1 reference Cote d'Ivoire strain) | not stated |
| Single-linkage hierarchical clustering of pairwise SNP distances (SnapperDB SNP address) | Assignment of isolates to SNP-defined cluster groups (Clusters 1–3 / Blue, Green, Red) | 31 isolates | not stated |
| cgMLST minimum spanning tree (MSTreeV2 algorithm, Enterobase, 3002-locus scheme) | Allele-based clustering and cgST assignment of the same isolate set for comparison with SNP clusters | 31 isolates (subset with passing QC assembly) | not stated |
-
Food exposures were summarised as the number/proportion of cases reporting each item (e.g. pepper 9/10, coriander 8/10), with no comparison group↳ Could also: A case-control design with matched community controls could have been used, enabling calculation of odds ratios (with 95% CIs) for each food exposure via conditional logistic regression or Fisher's exact test — With only 14 cases, matched case-control analysis (even with 2:1 or 3:1 matching) would allow estimation of the relative odds of exposure for each food vehicle and formal ranking of suspected sources, which descriptive proportions alone cannot provide
-
Central tendency and spread for age and illness duration were reported as median and range↳ Could also: Interquartile range (IQR) alongside the median would also be a standard summary for small, potentially skewed continuous data — IQR is less sensitive to extreme values than full range and is the conventional companion to the median in non-parametric reporting, giving a more stable picture of spread when n is small
-
Phylogenetic inference used maximum likelihood (RAxML, GTR+GAMMA, 1000 bootstraps) from SNP pseudosequences↳ Could also: Bayesian inference (e.g., BEAST or MrBayes) could also have been applied to the same SNP alignment — Bayesian methods provide posterior probability support values and can additionally estimate substitution rates and divergence dates, which would be informative for assessing whether outbreak strains share a common recent ancestor or represent repeated introductions from an endemic source
-
SNP-based clustering used single-linkage hierarchical clustering to define SNP address groups↳ Could also: Complete-linkage or average-linkage (UPGMA) hierarchical clustering, or density-based clustering (e.g. DBSCAN on pairwise SNP distances), could also have been applied — Single-linkage clustering is susceptible to chaining artefacts, where one intermediate isolate can join otherwise distant groups; complete or average linkage tends to produce more compact, internally coherent clusters, which is relevant here given the reported heterogeneity within the red cluster (0–9 SNPs)
-
Cluster concordance between SNP and cgMLST methods was assessed qualitatively by visual comparison of the two tree/network topologies↳ Could also: A quantitative measure of clustering agreement such as the Adjusted Rand Index, or a Mantel test comparing pairwise SNP-distance and allele-difference matrices, could also have been used — A numeric concordance measure would provide an objective, reproducible summary of how well the two typing methods agree across all isolate pairs, going beyond visual inspection of tree topology
-
Exposure data were analysed with Stata but only descriptive analyses are reported; no attack rates or relative risk estimates are presented↳ Could also: If denominator data (number of people buying from each grocer or consuming each food) were available, food-specific attack rates or risk ratios could have been calculated; even without controls, a source-attribution model weighting by exposure frequency could have been applied — Presenting exposure-stratified attack rates or risk ratios alongside crude proportions would allow formal ranking of the most likely vehicles and quantify uncertainty around those estimates, which is particularly valuable when trace-back does not identify a single common supplier
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-30648934
Paper: Chattaway et al. 2019, Microb Genom 5(3):e000248. "Genomic approaches used to investigate an atypical outbreak of Salmonella Adjame." Designated code: MOST (Metric Orientated Sequence Typer), PHE, https://github.com/phe-bioinformatics/MOST — MLST-from-reads using the Salmonella enterica Achtman 7-gene scheme (aroC,dnaN,hemD,hisD,purE,sucA,thrA). Data: Illumina reads under SRA. Paper names 3 representative NCTC strains with SRA run accessions (the cleanest 1:1 targets):
- SRR5583198 (NCTC 14246)
- SRR6237100 (NCTC 14247)
- SRR6190984 (NCTC 14248)
Reported computational results & in/out of scope
| Result | Method/pipeline | Scope |
|---|---|---|
| MLST sequence types ST4023 and ST3929 (Achtman scheme via MOST) | MOST on raw reads | IN — primary target |
| Confirmed serovar S. Adjame, antigenic formula 13,23:r:1,6 | phenotypic serology + WGS confirmation | partial/out — formula is wet-lab (WKLM scheme); WGS serovar predictable in-silico (SISTR/SeqSero) as a cross-check only |
| SNP distances between clusters (max 449 / 692 / 410 SNPs; within-cluster 0, 0, 0–9) | PHE SnapperDB reference-mapped SNP pipeline | OUT — hard 20%: needs the exact reference + SnapperDB params not fully specified; not attempted |
| cgMLST allele differences (max 101 / 99 / 49 alleles; within 0/0/0–2) | Enterobase cgMLST (3002-locus scheme) | OUT — hard 20%: Enterobase-hosted scheme/calling; not attempted |
| SNP addresses (Table 2, e.g. 1.8.8.8.8.8.8) | SnapperDB hierarchical SNP address | OUT — derivative of the SnapperDB pipeline above |
| Sequencing: 2×100 bp HiSeq 2500; coverage >30 added to DB | wet-lab / QC | OUT — descriptive |
Reproduction strategy (in-scope target)
The paper's tool MOST ships its own Salmonella scheme, but the bundled
profiles.txt maxes at ST3360 and lacks ST4023/ST3929 (these were assigned
later in the live PHE/Enterobase DB). Faithful plan:
- MOST (paper's tool): run on each of the 3 SRA runs → 7-gene Achtman allele profile + the ST MOST resolves against its bundled DB.
- Resolve allele profile → ST against the current PubMLST/Enterobase senterica profile table → confirm it equals ST4023 / ST3929.
- Independent cross-check (third-party, equally valid per brief P16):
assemble each run (skesa) +
mlst(Torsten Seemann; current senterica DB) → ST number directly.
Agreement decision is on the ST assignment of the 3 representative isolates.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
The paper's headline computational claim — Achtman 7-gene MLST of the S. Adjame outbreak typed with MOST — reproduces exactly: SRR5583198 & SRR6237100 = ST3929 and SRR6190984 = ST4023, recovering both reported STs, with serovar Adjame (13,23:r:1,6) confirmed in-silico by SISTR. It is cross-validated by two independent routes, with the MOST tool's older DB difference reconciled at the allele-sequence level (an expected version artifact, not a discrepancy). The only gaps are the SnapperDB-SNP and Enterobase-cgMLST distance matrices, honestly disclosed as out-of-scope because their reference/params are underspecified by the authors — these do not undermine the central conclusion. No fabrication concern.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.