Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Genomic approaches used to investigate an atypical outbreak of Salmonella Adjame.

Microb Genom · 2019
L1 100/100 PQI 97
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Every question reproduced
-1 pts
From: “every question reproduced”
Total score -7
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
  • No relevant deviation in data/preprocessing
  • No authors-side cause for any deviation
  • Reported values are derivable from the shared data
  • Any deviation was negligible
  • The central claim held under reproduction
  • Overall, the reproduction was clean
What did not (or only partly)
  • Every checked point held up.
How its reproducibility compares
100/100
Reproducibility score
1.5 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 95% of all assessed papers rank 1 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

DESCRIBED WELL ENOUGH -> 1:1. The paper's headline computational result (Achtman 7-gene MLST of the S. Adjame outbreak, typed with MOST) reproduces EXACTLY. On the 3 representative NCTC isolates the paper names with SRA accessions, SRR5583198 & SRR6237100 = ST3929 and SRR6190984 = ST4023 (a single-locus dnaN variant), recovering BOTH reported STs. Confirmed two independent ways: (Route A) skesa+mlst (current PubMLST senterica_achtman_2 DB) gives the ST numbers directly; (Route B) MOST, the paper's own Py2.7 tool, gives the identical underlying allele sequences (it can't print the ST number only because its 2016 bundled database predates purE allele 731 / dnaN allele 651). SISTR independently confirms serovar Adjame with antigenic formula 13,23:r:1,6 in-silico. NOT ATTEMPTED (out-of-scope hard 20%, honestly disclosed): the SnapperDB reference-mapped SNP cluster distances (max 449/692/410; within 0/0/0-9), the Enterobase cgMLST allele differences (max 101/99/49; within 0/0/0-2), and the SnapperDB SNP addresses (Table 2) -- these depend on a specific reference genome + SnapperDB params and the Enterobase-hosted cgMLST scheme that are not fully specified in the paper. No fabrication concerns: every reported value we checked is derivable from the deposited reads with the designated tool.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 100
    assessed: 2026-06-16 ⛓ 71b72fdeebd6
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-16
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-16
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

Can whole genome sequencing-based clustering (SNP versus cgMLST) appropriately characterize an atypical outbreak of the rare serovar Salmonella Adjame, and can cgMLST allele-based typing serve as a standardizable format for international outbreak comparison and alerting?

Core claims
  • The S. Adjame outbreak produced a heterogeneous phylogeny with multiple temporally/geographically linked sub-clusters, atypical of a point-source Salmonella outbreak and consistent with contamination from an endemic mixed-strain source (imported South Asian herbs/spices). finding
  • cgMLST allele-based clustering correlated with and was comparable to SNP-derived phylogenetic analysis in defining clusters for this outbreak. finding
  • Incorporating fixed SNP or allelic differences into a case definition may not always be appropriate for heterogeneous outbreaks or rare serovars involving small case numbers. finding
  • cgMLST (cgST) is recommended for international comparison and alerting of WGS data via platforms such as EPIS, pending further validation. method
  • Imported herbs or spices from South Asia were the suspected vehicle of infection, though backward tracing identified no common source. finding
  • SnapperDB reference-based high-quality SNP pipeline and Enterobase cgMLST V2 (3002 loci) were used and compared as parallel typing approaches. method
  • All outbreak strains belonged to a single EBurst Group (EBG421) comprising two sequence types, ST4023 and ST3929. finding
Experimental setups
Assay System Perturbation Readout Platform
Whole genome sequencing (Illumina) Salmonella Adjame clinical isolates (human; UK, Ireland, Denmark, Cote d'Ivoire reference) none FASTQ sequence reads / genome sequence Illumina HiSeq 2500, rapid run mode, 2x100 bp
Reference-based SNP typing (SnapperDB pipeline) 31 S. Adjame isolates (20 UK 2017, 6 historical UK, 1 reference, 2 Irish, 1 Danish, 1 Cote d'Ivoire) none high-quality SNP positions, SNP distances, maximum-likelihood phylogeny/clusters BWA mem, Samtools, GATK2, SnapperDB, RAxML v8.2.8; reference SRR5583198 assembled with SPAdes v3.8.0
cgMLST allele-based typing S. Adjame isolates none core genome sequence type (cgST) over 3002 loci, minimum spanning tree, allele differences Enterobase (SPAdes assembly, cgMLST V2, MSTreeV2 algorithm)
7-gene MLST / sequence typing S. Adjame isolates none ST and EBurst Group assignment MOST (Metric Orientated Sequence Typer), Achtman MLST database
Phenotypic serology (serotyping) S. Adjame isolates (first three of each ST) none antigenic profile (13,23:r:1,6) confirming serovar White-Kauffman-Le Minor scheme
Epidemiological investigation (trawling/targeted questionnaire) 14 human outbreak cases in England none demographics, clinical info, food/travel exposures Stata-13, Microsoft Excel 2010
Food trace-back investigation Commonly reported food items/spice brands and South Asian grocers none supply chain provenance / country of origin
Key results
  • 14 cases met the outbreak case definition, 11 confirmed S. Adjame in London plus one each in South East, East of England, South West England 14 cases
  • SNP clustering revealed three genetically distinct clusters plus single outlying strains; 13/14 cases fell into the Red and Green clusters with one strain unrelated 3 clusters
  • Large SNP distances between clusters indicate marked heterogeneity atypical of point-source outbreak max 449 SNPs (cluster 1-2), 692 SNPs (1-3), 410 SNPs (2-3)
  • Red cluster comprised eight strains (late June/early July 2017) with min 0, max 9 SNP distance; blue and green clusters each had zero SNP distance internally min 0, max 9 SNPs
  • cgMLST clusters were comparable to SNP-derived phylogenetic clusters
  • All strains fell into one EBurst Group (EBG421) comprising ST4023 and ST3929 1 EBG, 2 STs
  • Four EU countries (Italy, Germany, Ireland, Denmark) reported S. Adjame cases in response to EPIS alert; suspected link to South Asian groceries
  • Backward tracing found multiple possible sources with no common supplier; spices frequently sourced from South Asian spice farms outside Europe
Key statistics
  • count 14 cases (cases meeting outbreak case definition)
  • mean median age 66.5 years (range 3–85) (age of outbreak cases)
  • count 31 S. Adjame SNP-typed (20 UK 2017, 1 reference, 6 historical UK, 2 Irish, 1 Danish, 1 Cote d'Ivoire)
  • count 449 / 692 / 410 SNPs (maximum SNP distances between clusters 1-2, 1-3, 2-3)
  • count 3002 genes/loci (cgMLST V2 scheme loci present in >98% of 3144 reference genomes)
  • count 10/11 exposed to brand X, 8/11 to Y, 7/11 to Z (likelihood of spice brand exposure among cases)
  • count pepper and turmeric 9/10 (most commonly reported herb/spice exposures)
  • count 8 grocer A, 4 grocer B, 3 grocer C (number of cases purchasing from each South Asian grocer)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This outbreak investigation report combined descriptive epidemiology with comparative whole-genome sequencing (WGS) genomics for 14 cases of Salmonella Adjame. Epidemiological data from trawling and targeted questionnaires were summarised using descriptive statistics (counts, proportions, medians with ranges) in Stata-13 and Excel; no inferential hypothesis tests were applied. Genomic relatedness was assessed through two parallel approaches: a high-quality SNP pipeline (SnapperDB) with single-linkage hierarchical clustering and RAxML maximum-likelihood phylogenetics, and an allele-based cgMLST scheme (Enterobase/MSTreeV2 minimum spanning tree). The primary analytic goal was to compare the concordance of SNP- and cgMLST-derived cluster assignments rather than to test any pre-specified statistical hypothesis.

Replicationunclear Sample size14 epidemiological cases defined by clinical + time criteria; 31 isolates selected for WGS (all viable historical + outbreak + international strains available); no formal sample-size or power calculation stated GroupsOutbreak cases vs. historical UK isolates vs. international (Irish, Danish, Cote d'Ivoire) isolates; SNP-derived clusters vs. cgMLST-derived clusters Pairingna Randomization/blindingnot stated Dispersionrange Effect sizesno Confidence intervalsno
Statistical tests used
Test Applied to n Assumptions
Descriptive statistics (counts, proportions, median with range) Demographic and exposure data for 14 outbreak cases (age, illness duration, food exposures, hospitalisations) 14 cases; sub-denominators vary (e.g. 10/11, 9/10) due to missing questionnaire responses na
Maximum-likelihood phylogeny (RAxML v8.2.8, GTR+GAMMA, 1000 bootstrap replicates) SNP-based phylogenetic tree of 31 S. Adjame isolates mapped against a multi-contig reference (SRR5583198) 31 isolates (20 UK 2017, 6 historical UK, 2 Irish, 1 Danish, 1 reference Cote d'Ivoire strain) not stated
Single-linkage hierarchical clustering of pairwise SNP distances (SnapperDB SNP address) Assignment of isolates to SNP-defined cluster groups (Clusters 1–3 / Blue, Green, Red) 31 isolates not stated
cgMLST minimum spanning tree (MSTreeV2 algorithm, Enterobase, 3002-locus scheme) Allele-based clustering and cgST assignment of the same isolate set for comparison with SNP clusters 31 isolates (subset with passing QC assembly) not stated
Approaches that could also have been used
  • Food exposures were summarised as the number/proportion of cases reporting each item (e.g. pepper 9/10, coriander 8/10), with no comparison group
    Could also: A case-control design with matched community controls could have been used, enabling calculation of odds ratios (with 95% CIs) for each food exposure via conditional logistic regression or Fisher's exact test — With only 14 cases, matched case-control analysis (even with 2:1 or 3:1 matching) would allow estimation of the relative odds of exposure for each food vehicle and formal ranking of suspected sources, which descriptive proportions alone cannot provide
  • Central tendency and spread for age and illness duration were reported as median and range
    Could also: Interquartile range (IQR) alongside the median would also be a standard summary for small, potentially skewed continuous data — IQR is less sensitive to extreme values than full range and is the conventional companion to the median in non-parametric reporting, giving a more stable picture of spread when n is small
  • Phylogenetic inference used maximum likelihood (RAxML, GTR+GAMMA, 1000 bootstraps) from SNP pseudosequences
    Could also: Bayesian inference (e.g., BEAST or MrBayes) could also have been applied to the same SNP alignment — Bayesian methods provide posterior probability support values and can additionally estimate substitution rates and divergence dates, which would be informative for assessing whether outbreak strains share a common recent ancestor or represent repeated introductions from an endemic source
  • SNP-based clustering used single-linkage hierarchical clustering to define SNP address groups
    Could also: Complete-linkage or average-linkage (UPGMA) hierarchical clustering, or density-based clustering (e.g. DBSCAN on pairwise SNP distances), could also have been applied — Single-linkage clustering is susceptible to chaining artefacts, where one intermediate isolate can join otherwise distant groups; complete or average linkage tends to produce more compact, internally coherent clusters, which is relevant here given the reported heterogeneity within the red cluster (0–9 SNPs)
  • Cluster concordance between SNP and cgMLST methods was assessed qualitatively by visual comparison of the two tree/network topologies
    Could also: A quantitative measure of clustering agreement such as the Adjusted Rand Index, or a Mantel test comparing pairwise SNP-distance and allele-difference matrices, could also have been used — A numeric concordance measure would provide an objective, reproducible summary of how well the two typing methods agree across all isolate pairs, going beyond visual inspection of tree topology
  • Exposure data were analysed with Stata but only descriptive analyses are reported; no attack rates or relative risk estimates are presented
    Could also: If denominator data (number of people buying from each grocer or consuming each food) were available, food-specific attack rates or risk ratios could have been calculated; even without controls, a source-attribution model weighting by exposure frequency could have been applied — Presenting exposure-stratified attack rates or risk ratios alongside crude proportions would allow formal ranking of the most likely vehicles and quantify uncertainty around those estimates, which is particularly valuable when trace-back does not identify a single common supplier
Software: Stata 13 (StataCorp) 13 · Microsoft Excel 2010 · SnapperDB (PHE SNP pipeline) · Enterobase (cgMLST / MSTreeV2) · RAxML 8.2.8 · SPAdes 3.8.0 · BWA mem · Samtools · GATK2 · Trimmomatic · MOST (Metric Orientated Sequence Typer)

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
15
Impact: medium
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

PRJNA248792 BioProject in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet
SRR6193063 ENA in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
SRR6233875 ENA in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
SRR6233881 ENA in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
SRR6233939 ENA in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
SRR6234003 ENA in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
SRR6237100 ENA in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-30648934

Paper: Chattaway et al. 2019, Microb Genom 5(3):e000248. "Genomic approaches used to investigate an atypical outbreak of Salmonella Adjame." Designated code: MOST (Metric Orientated Sequence Typer), PHE, https://github.com/phe-bioinformatics/MOST — MLST-from-reads using the Salmonella enterica Achtman 7-gene scheme (aroC,dnaN,hemD,hisD,purE,sucA,thrA). Data: Illumina reads under SRA. Paper names 3 representative NCTC strains with SRA run accessions (the cleanest 1:1 targets):

  • SRR5583198 (NCTC 14246)
  • SRR6237100 (NCTC 14247)
  • SRR6190984 (NCTC 14248)

Reported computational results & in/out of scope

Result Method/pipeline Scope
MLST sequence types ST4023 and ST3929 (Achtman scheme via MOST) MOST on raw reads IN — primary target
Confirmed serovar S. Adjame, antigenic formula 13,23:r:1,6 phenotypic serology + WGS confirmation partial/out — formula is wet-lab (WKLM scheme); WGS serovar predictable in-silico (SISTR/SeqSero) as a cross-check only
SNP distances between clusters (max 449 / 692 / 410 SNPs; within-cluster 0, 0, 0–9) PHE SnapperDB reference-mapped SNP pipeline OUT — hard 20%: needs the exact reference + SnapperDB params not fully specified; not attempted
cgMLST allele differences (max 101 / 99 / 49 alleles; within 0/0/0–2) Enterobase cgMLST (3002-locus scheme) OUT — hard 20%: Enterobase-hosted scheme/calling; not attempted
SNP addresses (Table 2, e.g. 1.8.8.8.8.8.8) SnapperDB hierarchical SNP address OUT — derivative of the SnapperDB pipeline above
Sequencing: 2×100 bp HiSeq 2500; coverage >30 added to DB wet-lab / QC OUT — descriptive

Reproduction strategy (in-scope target)

The paper's tool MOST ships its own Salmonella scheme, but the bundled profiles.txt maxes at ST3360 and lacks ST4023/ST3929 (these were assigned later in the live PHE/Enterobase DB). Faithful plan:

  1. MOST (paper's tool): run on each of the 3 SRA runs → 7-gene Achtman allele profile + the ST MOST resolves against its bundled DB.
  2. Resolve allele profile → ST against the current PubMLST/Enterobase senterica profile table → confirm it equals ST4023 / ST3929.
  3. Independent cross-check (third-party, equally valid per brief P16): assemble each run (skesa) + mlst (Torsten Seemann; current senterica DB) → ST number directly.

Agreement decision is on the ST assignment of the 3 representative isolates.

C1_ST_SRR5583198
Reported
ST4023 or ST3929 (Achtman MLST)
Reproduced
ST3929 [aroC127 dnaN287 hemD289 hisD94 purE731 sucA199 thrA126]
exact
C2_ST_SRR6237100
Reported
ST4023 or ST3929 (Achtman MLST)
Reproduced
ST3929 [same profile]
exact
C3_ST_SRR6190984
Reported
ST4023 or ST3929 (Achtman MLST)
Reproduced
ST4023 [aroC127 dnaN651 hemD289 hisD94 purE731 sucA199 thrA126; SLV of ST3929]
exact
C4_two_STs
Reported
outbreak comprised ST4023 and ST3929
Reproduced
both recovered across the 3 named reps (ST3929 x2, ST4023 x1)
exact
C5_serovar
Reported
S. Adjame, antigenic formula 13,23:r:1,6
Reproduced
SISTR: serovar=Adjame, antigen 13,23:r:1,6 (O13,23;H1 r;H2 1,6) for all 3
exact
C6_MOST_faithful
Reported
typing performed with MOST (PHE)
Reproduced
MOST (paper's own tool) run on all 3 reps recovers the IDENTICAL 7-locus allele sequences as Route A; reports NOVEL/SLV/DLV only because its 2016 bundled DB lacks purE731/dnaN651 (PURE*84+SNP==purE731, DNAN*287+SNP==dnaN651)
exact

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 100/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟢3. Location of the main deviation
🟢4. Cause of the deviation
🟢5. Derivability / plausibility
🟢6. Severity of the deviation
🟢7. Core claim
🟢8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Every question reproduced
-1 pts
From: “every question reproduced”
Total score -7

The paper's headline computational claim — Achtman 7-gene MLST of the S. Adjame outbreak typed with MOST — reproduces exactly: SRR5583198 & SRR6237100 = ST3929 and SRR6190984 = ST4023, recovering both reported STs, with serovar Adjame (13,23:r:1,6) confirmed in-silico by SISTR. It is cross-validated by two independent routes, with the MOST tool's older DB difference reconciled at the allele-sequence level (an expected version artifact, not a discrepancy). The only gaps are the SnapperDB-SNP and Enterobase-cgMLST distance matrices, honestly disclosed as out-of-scope because their reference/params are underspecified by the authors — these do not undermine the central conclusion. No fabrication concern.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

160 k
tokens (I/O) · 14.3 M incl. cache
38 min
runtime · 0.69 CPU-h
7.5 GB
peak RAM
3
HPC jobs
hummel
machine