Acquisition and loss of CTX-M plasmids in Shigella species associated with MSM transmission in the UK.
The main results reproduced: recomputed values matched the published ones within tolerance.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Same input data as the authors
- ✓Reported values were directly comparable
- ✓Reported values are derivable from the shared data
- ✓Any deviation was negligible
- ✓The central claim held under reproduction
- 🟡A deviation arose in the data or preprocessing
- 🟡A deviation was attributed to the published material
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
REPRODUCED. The paper is described well enough that its central pipeline-derived genomic claims reproduce 1:1 from its OWN deposited data using standard third-party tools (abricate/PlasmidFinder/blastn/seqkit/shovill) -- a valid P16 reproduction. The brief's pinned code (Porechop) is a Nanopore adapter-trimmer and the brief's pinned datum (SRX1766927) is Illumina: a text-mining mismatch, and the Nanopore reads needed for Porechop were never deposited, so the pinned tool cannot reproduce any value (documented, not fatal). HEADLINE REPRODUCED EXACTLY: blaCTX-M-27 sits on the S. sonnei 893916 IncFII plasmid (100%) and is ABSENT from the S. flexneri 888048 complete assembly, yet PRESENT in 888048's Illumina reads (shovill contig00250, 100%) -- precisely the 'acquisition and loss' of the CTX-M plasmid the title describes. Replicon types (IncFII 67-83kb, IncB/O/K/Z 86-103kb), pINV (~220kb), chromosome (~4.7Mb) and MDR status all match. NEW vs prior attempt: the p183660 identity reproduced EXACTLY at 99.73% (paper 99.7-99.9%) using KX008967. Only pKSR100 identity is graded partial (top HSPs 98.0-98.6% inside the band, coverage-weighted mean ~1pt low; the paper does not define its identity/cover metric). Not attempted: Porechop (no Nanopore reads), phylogenetics (context isolates not enumerated), de novo of 598080/607387 draft WGS.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 86assessed: 2026-06-18 ⛓ f5f23359df31
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-24
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetThe study investigates where bla_CTX-M-27 and other antimicrobial resistance determinants are located within the genomes/plasmids of MSM-associated Shigella isolates, using combined long- and short-read sequencing to resolve complete plasmid structures and mobile genetic element context.
- ★ bla_CTX-M-27 is located on IncFII pKSR100-like plasmids, flanked by IS26 and IS903B finding
- ★ All S. sonnei isolates harboured Tn7/Int2 chromosomal integrons, whereas S. flexneri 3a contained the Shigella Resistance Locus (SRL) finding
- ★ All four strains harboured IncFII pKSR100-like plasmids (67-83 kbp) finding
- ★ bla_CTX-M-27 was lost in the S. flexneri 3a isolate during storage between Illumina and Nanopore sequencing finding
- ★ IncFII AMR regions were mosaic and likely reorganised by IS26 mechanism
- ★ Three of four plasmids contained azithromycin-resistance genes erm(B) and mph(A), and one harboured the pKSR100 integron finding
- ★ All S. sonnei isolates possessed a large IncB/O/K/Z plasmid, two of which carried aph(3')-Ib/aph(6)-Id/sul2 and tet(A) finding
- ★ Hybrid long-read (Nanopore) and short-read (Illumina) sequencing/assembly (Flye plus Pilon/Racon polishing) resolves complete plasmid sequences better than short-read-only approaches method
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Illumina short-read whole genome sequencing | S. sonnei (n=3) and S. flexneri 3a (n=1) isolates from MSM patients | none | genome sequence reads for assembly and AMR gene detection | HiSeq 2500, Nextera XP library kit |
| Oxford Nanopore long-read whole genome sequencing | same four Shigella isolates | none | long reads for complete genome/plasmid assembly | MinION, FLO-MIN106 R9.4.1 flow cell, SQK-RBK004 kit |
| De novo genome assembly and polishing | Shigella isolate genomes | none | assembly contiguity, contig number, N50, plasmid recovery | Flye v2.7.1, Unicycler v0.4.8, Pilon v1.23, Racon v1.4.13, QUAST v5.0.2 |
| BLASTn plasmid comparison | assembled Shigella IncFII and IncB/O/K/Z plasmids vs reference plasmids pKSR100 and p183660 | none | percent nucleotide identity and query coverage | blastn v2.10.1, BRIG |
| In silico AMR gene detection | Shigella genome assemblies | none | presence/absence and count of acquired resistance genes and mutations | ResFinder-3.2, CARD Resistance Gene Identifier |
| Multi-Locus Sequence Typing (MLST) | Shigella isolates | none | sequence type based on 7 housekeeping genes | CGE E. coli MLST scheme #1 |
| Maximum-likelihood core-genome SNP phylogenetics | S. sonnei (198 additional England isolates) and S. flexneri (49 additional isolates) | none | phylogenetic clustering, MSM clade/lineage assignment | SnapperDB v0.2.6, BWA-MEM, GATK v2.6.5, IQ-Tree v2.0.6 |
| Virulence factor detection | Shigella genome assemblies | none | presence/absence of virulence genes on chromosome and pINV plasmid | VirulenceFinder (CGE) |
- – bla_CTX-M-27 located on IncFII plasmids flanked by IS26 and IS903B
- – IncFII pKSR100-like plasmids ranged 67-83 kbp across the four isolates 67-83 kbp
- ▼ bla_CTX-M-27 was lost from the S. flexneri 3a isolate between Illumina and Nanopore sequencing during storage
- – 3 of 4 plasmids carried erm(B) and mph(A) azithromycin resistance genes 3/4
- – All S. sonnei isolates carried a large IncB/O/K/Z plasmid; 2 of 3 carried aph(3')-Ib/aph(6)-Id/sul2 and tet(A) 2/3
- – Flye produced more contiguous assemblies than Unicycler for all isolates
- – Acquired resistance genes detected ranged 7-11 (ResFinder) and 47-60 (CARD) per isolate 7-11 (ResFinder), 47-60 (CARD)
- – Predicted IS elements per genome ranged 504-588, representing an estimated 38-48 distinct IS types 504-588
- count 67-83 kbp (size range of IncFII pKSR100-like plasmids across isolates)
- count ~220 kbp (size of the virulence plasmid pINV present in all isolates)
- count 7-11 (acquired resistance genes detected by ResFinder per isolate)
- count 47-60 (resistance genes/mutations detected by CARD (Perfect and Strict hits) per isolate)
- count 504-588 (predicted total insertion sequence (IS) elements per genome)
- count 38-48 (estimated number of distinct IS types per genome across isolates)
- count 3 of 4 (plasmids carrying azithromycin-resistance genes erm(B) and mph(A))
- other 30x theoretical coverage (target Nanopore read coverage of the ~4.7 Mb Shigella genome after Filtlong filtering)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This is a descriptive comparative genomics study of four Shigella isolates; no inferential hypothesis tests were applied. The primary analytical methods were de novo long-read genome assembly (Flye, with Illumina polishing), in silico resistance and virulence gene identification (ResFinder, CARD, VirulenceFinder), BLAST-based plasmid comparison, and maximum-likelihood phylogenetics (IQ-Tree, GTR+ASC model) built from SNP alignments generated via SnapperDB. Results were reported descriptively using phylogenetic bootstrap support values, percent nucleotide identity, assembly contiguity metrics, and presence/absence of resistance genes; no traditional inferential statistics or p-values were produced.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Maximum-likelihood phylogeny (IQ-Tree v2.0.6, GTR+ASC model, 1000 ultrafast bootstrap replicates) | S. sonnei (n=201) and S. flexneri 3a (n=50) phylogenetic context trees | 201 S. sonnei isolates; 50 S. flexneri 3a isolates | not stated |
| BLAST nucleotide identity comparison (blastn v2.10.1, default parameters) | Plasmid comparisons to pKSR100 and p183660 references; replicon typing via PlasmidFinder (>95% identity, >60% query coverage) | — | na |
| SNP calling with quality filters (GATK v2.6.5; MQ>30, minimum depth>10, variant ratio>0.9) | Core SNP alignment used as input for maximum-likelihood phylogenetic trees | — | not stated |
| Single-linkage hierarchical clustering (SnapperDB v0.2.6) | SNP Address assignment for all context isolates; cluster representative selection | — | not stated |
| Assembly comparison by contiguity metrics (total contig number, N50) | Flye v2.7.1 vs Unicycler v0.4.8 assembler selection across four isolates | 4 isolates | na |
-
Phylogenetic branch support was assessed using 1000 ultrafast bootstrap approximations in IQ-Tree↳ Could also: Bayesian phylogenetic inference (e.g., MrBayes or BEAST) with posterior probability support values could also be applied — Bayesian posterior probabilities carry a different probabilistic interpretation than bootstrap values; BEAST additionally enables molecular-clock-dated phylogenies, allowing estimation of divergence times relevant to outbreak reconstruction
-
Assembler selection (Flye over Unicycler) was based on informal comparison of contiguity metrics (N50, total contig count) across four isolates↳ Could also: Reference-based accuracy metrics such as QUAST genome fraction, Merqury k-mer completeness, or per-base error rate could also be used to complement contiguity evaluation — Contiguity metrics alone do not capture base-level accuracy; adding accuracy metrics would give a more complete picture of assembly quality when choosing among assemblers
-
Plasmid relatedness was visualized and assessed using BLAST percent nucleotide identity and BRIG ring diagrams↳ Could also: Whole-sequence similarity statistics such as average nucleotide identity (ANI) or Mash distance could also quantify plasmid relatedness numerically — ANI and Mash provide compact, single-value summaries of overall sequence similarity that complement local-alignment visualizations and facilitate systematic comparison across larger plasmid sets
-
Context isolates for phylogenetic trees were drawn from single-linkage hierarchical clusters by representative sampling across time frames↳ Could also: Maximum-diversity sampling, structured random sampling stratified by year and geography, or explicit rarefaction could also be used to select context genomes — An explicit sampling strategy reduces potential ascertainment bias in the phylogenetic context and makes the representativeness of the tree more straightforward to evaluate and reproduce
-
MSM transmission status was inferred from metadata criteria (male sex, adult age, no reported foreign travel) rather than direct epidemiological confirmation↳ Could also: Phylogenetic transmission cluster analysis (e.g., pairwise SNP distance thresholds or TransPhylo) combined with metadata could also be used to support or refine transmission inference — Integrating genomic clustering with epidemiological metadata provides an additional, independent line of evidence for transmission linkage beyond metadata-based proxies alone
-
Resistance gene presence/absence was determined using fixed identity and coverage thresholds applied independently by each tool (ResFinder ≥90% identity/≥80% length; PlasmidFinder ≥95% identity/≥60% coverage)↳ Could also: Reporting the full distribution of per-gene identity and coverage values, or using a unified probabilistic resistance-calling framework, could also characterize detection confidence — Fixed thresholds applied independently across multiple databases may classify borderline matches differently; reporting underlying identity and coverage values makes detection decisions more transparent and reproducible
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Reproduction scope — pmid-34427554
Paper: Locke RK, Greig DR, Jenkins C, Dallman TJ, Cowley LA. Acquisition and loss of CTX-M plasmids in Shigella species associated with MSM transmission in the UK. Microb Genom 2021. PMID 34427554 · PMC8549364 · DOI 10.1099/mgen.0.000644
Study design. Hybrid (Illumina + Oxford Nanopore) WGS of 4 MDR Shigella isolates from MSM-associated cases in London: 3 S. sonnei (598080, 607387, 893916) + 1 S. flexneri 3a (888048). Core thesis: blaCTX-M-27 sits on a transmissible IncFII plasmid (pKSR100-like) that is independently acquired and lost across the MSM transmission network.
Pinned artifacts vs. what is actually reproducible
- Code pinned in brief:
github.com/rrwick/Porechop— a Nanopore adapter trimmer; ONE step in a ~30-tool pipeline (Guppy→Porechop→Filtlong→Flye→ Pilon/Racon→…). It is a third-party tool, valid per brief rule P16. - Data pinned in brief:
SRA:SRX1766927= S. sonnei isolate 183660, Illumina HiSeq, BioProject PRJNA315192 (PHE surveillance). This is the source of the p183660 reference plasmid the paper compares its IncFII plasmids to (99.7–99.9 % identity), NOT one of the 4 study isolates. - Tool/data mismatch (documented, not fatal): Porechop trims Nanopore adapters; SRX1766927 is Illumina. Running Porechop on SRX1766927 is biologically meaningless. The paper's Nanopore reads were never deposited (only Illumina SRRs + final assemblies), so Porechop cannot be used to reproduce any reported value. We therefore reproduce the paper's pipeline-derived genomic claims by applying standard third-party tools (abricate/ResFinder/PlasmidFinder/blastn/shovill) to the paper's own deposited genomes and Illumina reads — fully valid under P16.
IN SCOPE (pipeline-derived, attempted)
| # | Reported result | Pipeline | How we reproduce |
|---|---|---|---|
| C1 | blaCTX-M-27 present in S. sonnei IncFII plasmid | ResFinder/CARD | abricate on MW396858 |
| C2 | blaCTX-M-27 LOST from S. flexneri 3a 888048 assembly | ResFinder on assembly | abricate on CP066809+MW3968xx |
| C3 | CTX-M-27 detected by Illumina in 888048 (basis of "loss") | read assembly + ResFinder | shovill(SRR11096691)+abricate |
| C4 | IncFII replicon on CTX-M plasmids (67–83 kbp) | PlasmidFinder | abricate plasmidfinder + seqkit sizes |
| C5 | IncB/O/K/Z plasmid (86–103 kbp) | PlasmidFinder | abricate plasmidfinder + seqkit |
| C6 | pINV virulence plasmid ~220 kbp | size | seqkit on MW396859/MW396862 |
| C7 | IncFII identity to pKSR100 98–99.5 % | blastn | blastn vs LN624486 |
| C8 | All isolates MDR (≥3 antimicrobial classes) | ResFinder | abricate gene catalogue |
| C9 | Chromosome ~4.7 Mb | size | seqkit on CP066809/CP066810 |
OUT OF SCOPE (not pipeline-reproducible / not attempted)
- Porechop run on the paper's Nanopore reads — reads not deposited (drop of
this sub-result:
data_unavailablefor the Nanopore layer). - IncFII identity to p183660 — p183660 has no standalone GenBank accession; would require de novo assembly of SRX1766927 then plasmid extraction (partial attempt possible, lower priority).
- Phylogenetics (50-isolate S. flexneri tree, 201-isolate S. sonnei tree): the 198/49 context isolates are PHE-internal accessions, not enumerated as a reusable set → out of scope.
- Wet-lab / epidemiological / MSM-network interpretation — manual, out of scope.
- Full hybrid de novo assembly (Flye) — Nanopore reads not public.
Compute
All on «our HPC» («infra») via «host» ssh. Env: «infra» _shared_envs/gbs-typing
(abricate 1.4.0; DBs resfinder/plasmidfinder/card/ncbi 2026-Apr-3; blastn 2.16;
shovill; skesa; seqkit). SLURM partition std. Work dir:
«path».
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
The paper's central genomic claim — acquisition and loss of the CTX-M-27 IncFII plasmid — reproduces 1:1 from the authors' own deposited genomes: blaCTX-M-27 is 100%id/100%cov on the S. sonnei 893916 plasmid (MW396858) and absent from all S. flexneri 888048 sequences, with matching replicon types and within-tolerance plasmid/chromosome sizes. The deviations are minor and on our methodology (pKSR100 identity 97.25% vs reported 98-99.5%, a blast-weighting artefact) plus data-availability gaps (Nanopore reads not deposited; the Porechop code_url is a text-mining mismatch). One claim (C3, Illumina-read detection in 888048) is still pending, so the run is solid-but-preliminary rather than fully closed — no fabrication concern.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.