A role for ColV plasmids in the evolution of pathogenic Escherichia coli ST58.
The main results reproduced: recomputed values matched the published ones within tolerance.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Same input data as the authors
- ✓Reported values were directly comparable
- ✓No authors-side cause for any deviation
- ✓Reported values are derivable from the shared data
- ✓Any deviation was negligible
- ✓The central claim held under reproduction
- ✓Overall, the reproduction was clean
- 🟡A deviation arose in the data or preprocessing
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
REPRODUCED (with one partial axis). Three pipeline axes attempted, all run to completion. (1) Downstream re-derivation of the headline numbers from the authors' shipped intermediate tables (CJREID/ST58_project) by replicating abricateR/gene_analysis/data_vis logic in Python: 21/21 exact (752 genomes; ColV 353/752=46.9%; ColV by source poultry 109/125, porcine 77/106, bovine 48/262, human ExPEC 33/43; BAP2 308/363=85%; Roary 3023 core genes; 11 AMR/VAG prevalences). (2) Upstream FABRICATION CHECK on «our HPC» («job»): independent shovill de-novo assembly of the 159 deposited PRJNA727368 read sets + ABRicate with the authors' own EC_custom DB + Liu ColV criteria, compared per-genome to the authors' shipped ST58.abricate.colv.summary.tsv -> 122/129 concordant (94.6%); the authors' per-genome ColV calls ARE reproducible from the raw reads, no fabrication signal. The 7 discordants are 6 authors-positive/mine-negative (grp2/grp3 iro/iuc plasmid loci that fragment in short-read assembly = assembly sensitivity) + 1 mine-positive -> within-tol. (3) Enterobase ColV-prevalence expansion: rate concordant (~12% overall, ~45% ST58, ST58 enrichment confirmed) but exact counts cannot match because the live Enterobase deposit grew from 34,364 to ~42,380 genomes post-publication -> partial (honest moving-target limitation). NOT attempted (out of scope, not pipeline-derived scalars): wet-lab MIC/conjugation/S1-PFGE assays, IQ-TREE/RAxML tree topology, Scoary pan-GWAS p-values.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 91assessed: 2026-06-21 ⛓ 92b3e98cd9a1
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-25
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-21no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetHypothesising that ColV plasmids had a role in the emergence of Escherichia coli ST58 as a pathogen, the authors test whether ColV plasmid acquisition contributed to the divergence of a major pathogenic ST58 sub-lineage.
- ★ ST58 contains a major sub-lineage (BAP2, n=363) characterized by near-ubiquitous carriage of ColV plasmids finding
- ★ ColV plasmid acquisition contributed to the divergence of the major ST58 sub-lineage from the rest of the ST58 population mechanism
- ★ The BAP2 cluster has a distinct accessory genome, including genes typical of the Yersiniabactin High Pathogenicity Island (HPI) finding
- ★ BAP2 and ColV+ genomes carry significantly more antimicrobial resistance genes (ARGs) and virulence-associated genes (VAGs) than other genomes finding
- ★ Cattle-derived ST58 strains largely lack ColV plasmids, unlike poultry, porcine and human ExPEC sources, indicating different sub-lineages inhabit different host species finding
- ST58 ColV carriage rate (46-48%) is comparable to other emerging ExPEC clonal groups with major reservoirs in food animals finding
- Alignment of assemblies to the archetypal ColV plasmid pCERC4 backbone corroborates gene-marker-based ColV plasmid detection (Liu criteria) method
- Pathogen emergence in ST58 should be understood within a One Health framework involving networked pathways between environments and hosts finding
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| whole-genome sequencing and core-gene phylogeny/fastbaps clustering | 752 E. coli ST58 isolates from human, animal and environmental sources (1970-2019, 33 countries) | none | phylogenetic clustering into BAP groups (BAP1-6), source/serotype/fimH/F RST metadata | fastbaps |
| pan-genome-wide association study (pan-GWAS) | ST58 genome collection, BAP2 cluster membership as test variable | none | over-/under-represented accessory genes (78 over-, 55 under-represented) associated with BAP2 | — |
| plasmid backbone alignment / nucleotide identity heatmap | de novo ST58 genome assemblies aligned to ColV plasmid pCERC4 | none | binned nucleotide identity across pCERC4 backbone to corroborate ColV plasmid presence | — |
| ColV plasmid screening (Liu criteria, marker genes) | 752 ST58 genomes | none | ColV+/ColV- classification and carriage rate by BAP cluster and source | — |
| antimicrobial resistance gene (ARG) and virulence-associated gene (VAG) screening | 752 ST58 genomes | none | strain-wise ARG/VAG counts compared by BAP cluster and ColV status (Wilcoxon tests) | — |
| large-scale comparative genome screening for ColV plasmids | 34,364 draft E. coli genome assemblies from Enterobase, multiple STs and sources | none | ColV carriage rate by sequence type (ST) and source | Enterobase |
| virulence marker gene screening (fyuA, irp2, HPI markers) | BAP2 cluster ST58 genomes (n=363) | none | presence/absence of Yersiniabactin HPI marker genes | — |
| plasmid replicon typing (F plasmid RST, IncI1 pMLST) | 752 ST58 genomes | none | distribution of F and IncI1 plasmid replicon sequence types across BAP clusters and sources | — |
- ▲ 85% of BAP2 cluster sequences were ColV-positive 308/363 (85%)
- – ColV carriage was high in poultry, porcine and ExPEC sources but much lower in bovine sources poultry 87%, porcine 73%, ExPEC 77% vs bovine 18%
- ▲ BAP2 sequences carried significantly more ARGs than most other BAP clusters p=1.38e-28 vs BAP6
- ▲ BAP2 sequences carried significantly more VAGs than all other BAP clusters p=5.96e-70 vs BAP6
- ▲ ColV+ strains carried significantly more ARGs and VAGs than ColV- strains ARG p=9.2e-34; VAG p=1.28e-112
- ▲ HPI marker genes fyuA and irp2 were present in most BAP2 sequences fyuA 306/363 (84%); irp2 303/363 (83%)
- – ST58 ColV carriage rate in the Enterobase-derived collection closely matched that of the primary study collection 48% (281/588) vs 46% (353/752)
- – 78 genes were over-represented and 55 under-represented in the BAP2 pangenome, including ugd, galF, intA, mlrA and fyuA 78 over-/55 under-represented genes
- fold_change 308/363 (85%) ColV+ in BAP2 (ColV plasmid carriage in BAP2 cluster)
- fold_change 48/262 (18%) bovine ColV+ vs 109/125 (87%) poultry ColV+ (ColV carriage by source)
- pvalue p=1.38e-28 (BAP2 vs BAP6 ARG count (Wilcoxon, BH-adjusted))
- pvalue p=5.96e-70 (BAP2 vs BAP6 VAG count (Wilcoxon, BH-adjusted))
- pvalue p=9.2e-34 (ARG count ColV+ vs ColV- (Wilcoxon rank-sum))
- pvalue p=1.28e-112 (VAG count ColV+ vs ColV- (Wilcoxon rank-sum))
- count 752 ST58 genomes; 34,364 Enterobase E. coli assemblies (total genome collection sizes analysed)
- fold_change 13% (4370/34364) overall ColV carriage across Enterobase collection (background ColV carriage rate across all STs/sources)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
The study is an observational pan-genomic and phylogenetic analysis of 752 E. coli ST58 isolates, using fastbaps to define phylogenetic sub-lineages (BAP clusters) and then comparing antimicrobial resistance gene (ARG) and virulence-associated gene (VAG) counts across clusters and by ColV plasmid carriage status using nonparametric Wilcoxon rank-sum tests, with Benjamini-Hochberg adjustment applied to the multi-cluster pairwise comparisons. A separate pangenome-wide association study (pan-GWAS) was used to identify genes over- or under-represented in specific clusters, applying a fixed stringent significance threshold. Results are reported as exact p-values alongside boxplots summarizing count distributions (median, IQR, range).
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Pairwise Wilcoxon rank-sum test with Benjamini-Hochberg adjusted p-values | ARG counts compared across BAP clusters (Fig. 7a) | 750 biologically independent ST58 genome sequences | not stated |
| Pairwise Wilcoxon rank-sum test with Benjamini-Hochberg adjusted p-values | VAG counts compared across BAP clusters (Fig. 7b) | 750 biologically independent ST58 genome sequences | not stated |
| Two-sided Wilcoxon rank-sum test | ARG counts compared by ColV+/ColV- status (Fig. 7c) | 752 biologically independent ST58 genome sequences | not stated |
| Two-sided Wilcoxon rank-sum test | VAG counts compared by ColV+/ColV- status (Fig. 7d) | 752 biologically independent ST58 genome sequences | not stated |
| Pangenome-wide association study (pan-GWAS) with a fixed significance threshold (1E-50) | Identification of genes over- or under-represented in the BAP2 (and BAP6) cluster relative to the rest of the phylogeny | — | not stated |
-
ARG and VAG counts were compared across the six BAP clusters using separate pairwise Wilcoxon rank-sum tests with Benjamini-Hochberg correction.↳ Could also: A Kruskal-Wallis omnibus test followed by post-hoc Dunn's test with multiplicity correction could also be used for this kind of multi-group nonparametric comparison. — An omnibus test first establishes whether any group differs before pairwise follow-up, which some readers find a clearer two-step framework for multi-group designs, while yielding similar pairwise conclusions.
-
ARG and VAG counts, which are integer count data, were compared using nonparametric rank-based Wilcoxon tests.↳ Could also: A generalized linear model (e.g., Poisson or negative binomial regression) with BAP cluster or ColV status as a predictor could also be used. — Modeling counts directly with a count-appropriate distribution can additionally allow adjustment for covariates (e.g., source, collection year) within the same model, alongside the group comparison.
-
Genes associated with BAP2/BAP6 clusters in the pan-GWAS were called using a fixed, very stringent p-value threshold (1E-50).↳ Could also: Reporting an explicit false discovery rate (e.g., Benjamini-Hochberg q-values, as implemented in pan-GWAS tools like Scoary or pyseer) could also be used alongside or instead of a fixed threshold. — An FDR-based cutoff communicates the expected proportion of false positives among called genes, which some readers find more directly interpretable than a raw p-value threshold.
-
Group differences were summarized with p-values from Wilcoxon tests, without an accompanying effect-size statistic.↳ Could also: Reporting a nonparametric effect size such as Cliff's delta or the rank-biserial correlation alongside the p-values could also be informative. — Effect-size measures convey the magnitude of the difference between groups, complementing significance testing which primarily addresses whether a difference is detectable.
-
Comparisons across BAP clusters and ColV status treated each genome as an independent observation.↳ Could also: A phylogenetically aware analysis (e.g., phylogenetic generalized least squares, or tools designed for lineage-aware trait association) could also be applied. — Because isolates share varying degrees of clonal ancestry, phylogeny-aware methods can help distinguish trait associations that arose independently from those inherited from a common ancestor.
-
ColV carriage rates by source (e.g., poultry 87%, porcine 73%, bovine 18%) were reported as point proportions.↳ Could also: Reporting 95% confidence intervals (e.g., Wilson score intervals) for each carriage-rate proportion could also be presented alongside the point estimates. — Confidence intervals convey the precision of each proportion estimate, which is particularly useful when source-specific sample sizes vary substantially.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — PMID 35115531 (Reid, Cummins et al. 2022, Nat Commun)
"A role for ColV plasmids in the evolution of pathogenic Escherichia coli ST58." DOI 10.1038/s41467-022-28342-4 · PMCID PMC8813906
Artifacts
- Code (brief): https://github.com/maxlcummins/custom_DBs — the authors' custom
ABRicate database (
EC_custom.fa), the ColV/virulence/AMR gene reference used for screening. - Analysis repo (authors'): https://github.com/CJREID/ST58_project — ships ALL intermediate
tables (ABRicate, ARIBA, Roary, fastBAPS, Enterobase) + the R pipeline
(
scripts/abricateR.R,gene_analysis.R,data_vis.R, metadata curation). - Data: SRA PRJNA727368 = 158 newly-sequenced in-house ST58 isolates (paired Illumina WGS). The full 752-genome study collection additionally draws ~574 public Enterobase assemblies + ~20 previously-published in-house genomes (not in this BioProject).
In scope (pipeline-derived → attempted)
- Downstream re-derivation of the headline numbers from the authors' shipped intermediate
tables, by replicating their abricateR (Liu ColV criteria: ≥4 of 6 gene groups, cov≥95 id≥90),
gene_analysis and data_vis logic in Python:
- Total genomes (752); ColV+ overall (353/752, 46.9%); livestock niche (514/752); ColV by source (poultry 109/125, porcine 77/106, bovine 48/262, human ExPEC 33/43); fastBAPS BAP2 (363 seqs, 308 ColV+, 85%); Roary core genes (3023); 11 ARIBA AMR/VAG prevalences (fimH 730, iss 583, fyuA 347, fecA 412, blaTEM-1B 248, strA 296, sul2 279, tet(A) 272, tet(B) 165, intI1 245, merA 164).
- Pipeline: ABRicate + ARIBA + Roary + fastBAPS → R aggregation. Status: reproduced (21/21 exact).
- Enterobase ColV-prevalence expansion (overall 4370/34364 ≈13%; ST58 281/588 ≈48%). Pipeline: ABRicate ColV screen over the Enterobase E. coli universe. Status: partial — rates concordant (~12% overall, ~45% ST58, ST58 enrichment confirmed) but the live Enterobase deposit grew from 34,364 to ~42,380 genomes since publication, so exact counts cannot match.
- Upstream de-novo assembly concordance (fabrication check): independently assemble the 158
deposited PRJNA727368 read sets (shovill), re-screen with the authors' own
EC_customABRicate DB, apply the same Liu criteria, and compare the per-genome ColV+/- calls to the authors' shippedST58.abricate.colv.summary.tsv. Pipeline: shovill (SPAdes/skesa) → ABRicate → Liu. Status: running on «our HPC».
Out of scope (not pipeline-derived → not attempted)
- Wet-lab: MIC/antibiograms, plasmid conjugation/transfer assays, S1-PFGE, replicon typing by PCR.
- IQ-TREE/RAxML phylogeny topology (bootstrap/topology not a single reproducible scalar; fastBAPS cluster membership IS reproduced via the shipped table).
- Scoary / pan-GWAS association p-values (not pinnable to a single comparable value here).
- The ~574 Enterobase public assemblies as an exact frozen set (moving target; see #2).
Notes
- Per brief P16, applying the authors' own third-party tool (
custom_DBsABRicate DB) to the paper's deposited data is a fully valid reproduction. - Grades are PROVISIONAL and must be independently checked by a human (see AUDIT.md).
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
This is a near-textbook reproduction: all 21 primary headline values (total genomes 752, ColV+ 353/752=46.9%, ColV-by-source, BAP2 308/363=85%, Roary 3023 core genes, and 11 AMR/VAG prevalences) reproduced exactly from the authors' shipped intermediate tables, and the central ColV/livestock conclusion holds. The only deviations are in the secondary Enterobase expansion (4370/34364→5027/42380; 281/588→327/723), which differ in absolute counts purely because the live public database grew from ~34k to ~42k genomes since 2022 — rates (~12% vs 13%, ~45% vs 48%) and the ~3.8x ST58 enrichment still reproduce. This is a data-version effect on the data side, not a flaw in our method or the authors' work; an upstream de-novo-assembly fabrication check is still running but the downstream derivability is fully confirmed. Overall: solid 1:1 reproduction with the single explainable, externally-caused snapshot drift.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.