Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

A role for ColV plasmids in the evolution of pathogenic Escherichia coli ST58.

Nat Commun · 2022
L1 95/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score -5
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
  • No authors-side cause for any deviation
  • Reported values are derivable from the shared data
  • Any deviation was negligible
  • The central claim held under reproduction
  • Overall, the reproduction was clean
What did not (or only partly)
  • 🟡A deviation arose in the data or preprocessing
How its reproducibility compares
95/100
Reproducibility score
1.2 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 89% of all assessed papers rank 105 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

REPRODUCED (with one partial axis). Three pipeline axes attempted, all run to completion. (1) Downstream re-derivation of the headline numbers from the authors' shipped intermediate tables (CJREID/ST58_project) by replicating abricateR/gene_analysis/data_vis logic in Python: 21/21 exact (752 genomes; ColV 353/752=46.9%; ColV by source poultry 109/125, porcine 77/106, bovine 48/262, human ExPEC 33/43; BAP2 308/363=85%; Roary 3023 core genes; 11 AMR/VAG prevalences). (2) Upstream FABRICATION CHECK on «our HPC» («job»): independent shovill de-novo assembly of the 159 deposited PRJNA727368 read sets + ABRicate with the authors' own EC_custom DB + Liu ColV criteria, compared per-genome to the authors' shipped ST58.abricate.colv.summary.tsv -> 122/129 concordant (94.6%); the authors' per-genome ColV calls ARE reproducible from the raw reads, no fabrication signal. The 7 discordants are 6 authors-positive/mine-negative (grp2/grp3 iro/iuc plasmid loci that fragment in short-read assembly = assembly sensitivity) + 1 mine-positive -> within-tol. (3) Enterobase ColV-prevalence expansion: rate concordant (~12% overall, ~45% ST58, ST58 enrichment confirmed) but exact counts cannot match because the live Enterobase deposit grew from 34,364 to ~42,380 genomes post-publication -> partial (honest moving-target limitation). NOT attempted (out of scope, not pipeline-derived scalars): wet-lab MIC/conjugation/S1-PFGE assays, IQ-TREE/RAxML tree topology, Scoary pan-GWAS p-values.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 91
    assessed: 2026-06-21 ⛓ 92b3e98cd9a1
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-25
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-21
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

Hypothesising that ColV plasmids had a role in the emergence of Escherichia coli ST58 as a pathogen, the authors test whether ColV plasmid acquisition contributed to the divergence of a major pathogenic ST58 sub-lineage.

Core claims
  • ST58 contains a major sub-lineage (BAP2, n=363) characterized by near-ubiquitous carriage of ColV plasmids finding
  • ColV plasmid acquisition contributed to the divergence of the major ST58 sub-lineage from the rest of the ST58 population mechanism
  • The BAP2 cluster has a distinct accessory genome, including genes typical of the Yersiniabactin High Pathogenicity Island (HPI) finding
  • BAP2 and ColV+ genomes carry significantly more antimicrobial resistance genes (ARGs) and virulence-associated genes (VAGs) than other genomes finding
  • Cattle-derived ST58 strains largely lack ColV plasmids, unlike poultry, porcine and human ExPEC sources, indicating different sub-lineages inhabit different host species finding
  • ST58 ColV carriage rate (46-48%) is comparable to other emerging ExPEC clonal groups with major reservoirs in food animals finding
  • Alignment of assemblies to the archetypal ColV plasmid pCERC4 backbone corroborates gene-marker-based ColV plasmid detection (Liu criteria) method
  • Pathogen emergence in ST58 should be understood within a One Health framework involving networked pathways between environments and hosts finding
Experimental setups
Assay System Perturbation Readout Platform
whole-genome sequencing and core-gene phylogeny/fastbaps clustering 752 E. coli ST58 isolates from human, animal and environmental sources (1970-2019, 33 countries) none phylogenetic clustering into BAP groups (BAP1-6), source/serotype/fimH/F RST metadata fastbaps
pan-genome-wide association study (pan-GWAS) ST58 genome collection, BAP2 cluster membership as test variable none over-/under-represented accessory genes (78 over-, 55 under-represented) associated with BAP2
plasmid backbone alignment / nucleotide identity heatmap de novo ST58 genome assemblies aligned to ColV plasmid pCERC4 none binned nucleotide identity across pCERC4 backbone to corroborate ColV plasmid presence
ColV plasmid screening (Liu criteria, marker genes) 752 ST58 genomes none ColV+/ColV- classification and carriage rate by BAP cluster and source
antimicrobial resistance gene (ARG) and virulence-associated gene (VAG) screening 752 ST58 genomes none strain-wise ARG/VAG counts compared by BAP cluster and ColV status (Wilcoxon tests)
large-scale comparative genome screening for ColV plasmids 34,364 draft E. coli genome assemblies from Enterobase, multiple STs and sources none ColV carriage rate by sequence type (ST) and source Enterobase
virulence marker gene screening (fyuA, irp2, HPI markers) BAP2 cluster ST58 genomes (n=363) none presence/absence of Yersiniabactin HPI marker genes
plasmid replicon typing (F plasmid RST, IncI1 pMLST) 752 ST58 genomes none distribution of F and IncI1 plasmid replicon sequence types across BAP clusters and sources
Key results
  • 85% of BAP2 cluster sequences were ColV-positive 308/363 (85%)
  • ColV carriage was high in poultry, porcine and ExPEC sources but much lower in bovine sources poultry 87%, porcine 73%, ExPEC 77% vs bovine 18%
  • BAP2 sequences carried significantly more ARGs than most other BAP clusters p=1.38e-28 vs BAP6
  • BAP2 sequences carried significantly more VAGs than all other BAP clusters p=5.96e-70 vs BAP6
  • ColV+ strains carried significantly more ARGs and VAGs than ColV- strains ARG p=9.2e-34; VAG p=1.28e-112
  • HPI marker genes fyuA and irp2 were present in most BAP2 sequences fyuA 306/363 (84%); irp2 303/363 (83%)
  • ST58 ColV carriage rate in the Enterobase-derived collection closely matched that of the primary study collection 48% (281/588) vs 46% (353/752)
  • 78 genes were over-represented and 55 under-represented in the BAP2 pangenome, including ugd, galF, intA, mlrA and fyuA 78 over-/55 under-represented genes
Key statistics
  • fold_change 308/363 (85%) ColV+ in BAP2 (ColV plasmid carriage in BAP2 cluster)
  • fold_change 48/262 (18%) bovine ColV+ vs 109/125 (87%) poultry ColV+ (ColV carriage by source)
  • pvalue p=1.38e-28 (BAP2 vs BAP6 ARG count (Wilcoxon, BH-adjusted))
  • pvalue p=5.96e-70 (BAP2 vs BAP6 VAG count (Wilcoxon, BH-adjusted))
  • pvalue p=9.2e-34 (ARG count ColV+ vs ColV- (Wilcoxon rank-sum))
  • pvalue p=1.28e-112 (VAG count ColV+ vs ColV- (Wilcoxon rank-sum))
  • count 752 ST58 genomes; 34,364 Enterobase E. coli assemblies (total genome collection sizes analysed)
  • fold_change 13% (4370/34364) overall ColV carriage across Enterobase collection (background ColV carriage rate across all STs/sources)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The study is an observational pan-genomic and phylogenetic analysis of 752 E. coli ST58 isolates, using fastbaps to define phylogenetic sub-lineages (BAP clusters) and then comparing antimicrobial resistance gene (ARG) and virulence-associated gene (VAG) counts across clusters and by ColV plasmid carriage status using nonparametric Wilcoxon rank-sum tests, with Benjamini-Hochberg adjustment applied to the multi-cluster pairwise comparisons. A separate pangenome-wide association study (pan-GWAS) was used to identify genes over- or under-represented in specific clusters, applying a fixed stringent significance threshold. Results are reported as exact p-values alongside boxplots summarizing count distributions (median, IQR, range).

Replicationbiological Sample sizeSample sizes were stated per comparison as the number of biologically independent genome sequences (e.g., n=750 for BAP cluster comparisons, n=752 for ColV status comparisons); no formal power or sample-size calculation was described GroupsBAP phylogenetic clusters (BAP1-6) and ColV plasmid presence/absence, compared for ARG and VAG gene counts Pairingunpaired Randomization/blindingna DispersionIQR Exact p-valuesyes Effect sizesno Confidence intervalsno Multiplicity correctionBenjamini-Hochberg FDR adjustment (pairwise Wilcoxon tests); a stringent fixed p-value threshold (1E-50) for pan-GWAS gene associations
Statistical tests used
Test Applied to n Assumptions
Pairwise Wilcoxon rank-sum test with Benjamini-Hochberg adjusted p-values ARG counts compared across BAP clusters (Fig. 7a) 750 biologically independent ST58 genome sequences not stated
Pairwise Wilcoxon rank-sum test with Benjamini-Hochberg adjusted p-values VAG counts compared across BAP clusters (Fig. 7b) 750 biologically independent ST58 genome sequences not stated
Two-sided Wilcoxon rank-sum test ARG counts compared by ColV+/ColV- status (Fig. 7c) 752 biologically independent ST58 genome sequences not stated
Two-sided Wilcoxon rank-sum test VAG counts compared by ColV+/ColV- status (Fig. 7d) 752 biologically independent ST58 genome sequences not stated
Pangenome-wide association study (pan-GWAS) with a fixed significance threshold (1E-50) Identification of genes over- or under-represented in the BAP2 (and BAP6) cluster relative to the rest of the phylogeny not stated
Approaches that could also have been used
  • ARG and VAG counts were compared across the six BAP clusters using separate pairwise Wilcoxon rank-sum tests with Benjamini-Hochberg correction.
    Could also: A Kruskal-Wallis omnibus test followed by post-hoc Dunn's test with multiplicity correction could also be used for this kind of multi-group nonparametric comparison. — An omnibus test first establishes whether any group differs before pairwise follow-up, which some readers find a clearer two-step framework for multi-group designs, while yielding similar pairwise conclusions.
  • ARG and VAG counts, which are integer count data, were compared using nonparametric rank-based Wilcoxon tests.
    Could also: A generalized linear model (e.g., Poisson or negative binomial regression) with BAP cluster or ColV status as a predictor could also be used. — Modeling counts directly with a count-appropriate distribution can additionally allow adjustment for covariates (e.g., source, collection year) within the same model, alongside the group comparison.
  • Genes associated with BAP2/BAP6 clusters in the pan-GWAS were called using a fixed, very stringent p-value threshold (1E-50).
    Could also: Reporting an explicit false discovery rate (e.g., Benjamini-Hochberg q-values, as implemented in pan-GWAS tools like Scoary or pyseer) could also be used alongside or instead of a fixed threshold. — An FDR-based cutoff communicates the expected proportion of false positives among called genes, which some readers find more directly interpretable than a raw p-value threshold.
  • Group differences were summarized with p-values from Wilcoxon tests, without an accompanying effect-size statistic.
    Could also: Reporting a nonparametric effect size such as Cliff's delta or the rank-biserial correlation alongside the p-values could also be informative. — Effect-size measures convey the magnitude of the difference between groups, complementing significance testing which primarily addresses whether a difference is detectable.
  • Comparisons across BAP clusters and ColV status treated each genome as an independent observation.
    Could also: A phylogenetically aware analysis (e.g., phylogenetic generalized least squares, or tools designed for lineage-aware trait association) could also be applied. — Because isolates share varying degrees of clonal ancestry, phylogeny-aware methods can help distinguish trait associations that arose independently from those inherited from a common ancestor.
  • ColV carriage rates by source (e.g., poultry 87%, porcine 73%, bovine 18%) were reported as point proportions.
    Could also: Reporting 95% confidence intervals (e.g., Wilson score intervals) for each carriage-rate proportion could also be presented alongside the point estimates. — Confidence intervals convey the precision of each proportion estimate, which is particularly useful when source-specific sample sizes vary substantially.
Software: fastbaps

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — PMID 35115531 (Reid, Cummins et al. 2022, Nat Commun)

"A role for ColV plasmids in the evolution of pathogenic Escherichia coli ST58." DOI 10.1038/s41467-022-28342-4 · PMCID PMC8813906

Artifacts

  • Code (brief): https://github.com/maxlcummins/custom_DBs — the authors' custom ABRicate database (EC_custom.fa), the ColV/virulence/AMR gene reference used for screening.
  • Analysis repo (authors'): https://github.com/CJREID/ST58_project — ships ALL intermediate tables (ABRicate, ARIBA, Roary, fastBAPS, Enterobase) + the R pipeline (scripts/abricateR.R, gene_analysis.R, data_vis.R, metadata curation).
  • Data: SRA PRJNA727368 = 158 newly-sequenced in-house ST58 isolates (paired Illumina WGS). The full 752-genome study collection additionally draws ~574 public Enterobase assemblies + ~20 previously-published in-house genomes (not in this BioProject).

In scope (pipeline-derived → attempted)

  1. Downstream re-derivation of the headline numbers from the authors' shipped intermediate tables, by replicating their abricateR (Liu ColV criteria: ≥4 of 6 gene groups, cov≥95 id≥90), gene_analysis and data_vis logic in Python:
    • Total genomes (752); ColV+ overall (353/752, 46.9%); livestock niche (514/752); ColV by source (poultry 109/125, porcine 77/106, bovine 48/262, human ExPEC 33/43); fastBAPS BAP2 (363 seqs, 308 ColV+, 85%); Roary core genes (3023); 11 ARIBA AMR/VAG prevalences (fimH 730, iss 583, fyuA 347, fecA 412, blaTEM-1B 248, strA 296, sul2 279, tet(A) 272, tet(B) 165, intI1 245, merA 164).
    • Pipeline: ABRicate + ARIBA + Roary + fastBAPS → R aggregation. Status: reproduced (21/21 exact).
  2. Enterobase ColV-prevalence expansion (overall 4370/34364 ≈13%; ST58 281/588 ≈48%). Pipeline: ABRicate ColV screen over the Enterobase E. coli universe. Status: partial — rates concordant (~12% overall, ~45% ST58, ST58 enrichment confirmed) but the live Enterobase deposit grew from 34,364 to ~42,380 genomes since publication, so exact counts cannot match.
  3. Upstream de-novo assembly concordance (fabrication check): independently assemble the 158 deposited PRJNA727368 read sets (shovill), re-screen with the authors' own EC_custom ABRicate DB, apply the same Liu criteria, and compare the per-genome ColV+/- calls to the authors' shipped ST58.abricate.colv.summary.tsv. Pipeline: shovill (SPAdes/skesa) → ABRicate → Liu. Status: running on «our HPC».

Out of scope (not pipeline-derived → not attempted)

  • Wet-lab: MIC/antibiograms, plasmid conjugation/transfer assays, S1-PFGE, replicon typing by PCR.
  • IQ-TREE/RAxML phylogeny topology (bootstrap/topology not a single reproducible scalar; fastBAPS cluster membership IS reproduced via the shipped table).
  • Scoary / pan-GWAS association p-values (not pinnable to a single comparable value here).
  • The ~574 Enterobase public assemblies as an exact frozen set (moving target; see #2).

Notes

  • Per brief P16, applying the authors' own third-party tool (custom_DBs ABRicate DB) to the paper's deposited data is a fully valid reproduction.
  • Grades are PROVISIONAL and must be independently checked by a human (see AUDIT.md).
total_genomes
Reported
752
Reproduced
752
exact
colv_pos
Reported
353/752 (46.9%)
Reproduced
353/752 (46.9%)
exact
livestock_niche
Reported
514/752 (68%)
Reproduced
514/752 (68.4%)
exact
colv_poultry
Reported
109/125 (87%)
Reproduced
109/125 (87%)
exact
colv_porcine
Reported
77/106 (73%)
Reproduced
77/106 (73%)
exact
colv_bovine
Reported
48/262 (18%)
Reproduced
48/262 (18%)
exact
colv_expec
Reported
33/43 (77%)
Reproduced
33/43 (77%)
exact
bap2
Reported
363 seqs, 308 ColV (85%)
Reproduced
363, 308 (85%)
exact
core_genes
Reported
3023
Reproduced
3023
exact
prev_fimH
Reported
730 (97%)
Reproduced
730/752 (97%)
exact
prev_iss
Reported
583 (78%)
Reproduced
583/752 (78%)
exact
prev_fyuA
Reported
347 (46%)
Reproduced
347/752 (46%)
exact
prev_fecA
Reported
412 (55%)
Reproduced
412/752 (55%)
exact
prev_blaTEM1B
Reported
248 (33%)
Reproduced
248/752 (33%)
exact
prev_strA
Reported
296 (39%)
Reproduced
296/752 (39%)
exact
prev_sul2
Reported
279 (37%)
Reproduced
279/752 (37%)
exact
prev_tetA
Reported
272 (36%)
Reproduced
272/752 (36%)
exact
prev_tetB
Reported
165 (22%)
Reproduced
165/752 (22%)
exact
prev_intI1
Reported
245 (33%)
Reproduced
245/752 (33%)
exact
prev_merA
Reported
164 (22%)
Reproduced
164/752 (22%)
exact
upstream_colv_concordance
Reported
shipped per-genome ColV calls reproducible from raw reads (PRJNA727368)
Reproduced
122/129 concordant (94.6%); authors ColV+ 86 vs mine 81; confusion TP80/FP1/FN6/TN42
within tolerance
enterobase_overall
Reported
4370/34364 (13%)
Reproduced
5027/42380 (11.9%)
partial
enterobase_st58
Reported
281/588 (48%)
Reproduced
327/723 (45.2%)
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 95/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟡3. Location of the main deviation
🟢4. Cause of the deviation
🟢5. Derivability / plausibility
🟢6. Severity of the deviation
🟢7. Core claim
🟢8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score -5

This is a near-textbook reproduction: all 21 primary headline values (total genomes 752, ColV+ 353/752=46.9%, ColV-by-source, BAP2 308/363=85%, Roary 3023 core genes, and 11 AMR/VAG prevalences) reproduced exactly from the authors' shipped intermediate tables, and the central ColV/livestock conclusion holds. The only deviations are in the secondary Enterobase expansion (4370/34364→5027/42380; 281/588→327/723), which differ in absolute counts purely because the live public database grew from ~34k to ~42k genomes since 2022 — rates (~12% vs 13%, ~45% vs 48%) and the ~3.8x ST58 enrichment still reproduce. This is a data-version effect on the data side, not a flaw in our method or the authors' work; an upstream de-novo-assembly fabrication check is still running but the downstream derivability is fully confirmed. Overall: solid 1:1 reproduction with the single explainable, externally-caused snapshot drift.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

335.1 k
tokens (I/O) · 22.1 M incl. cache
53 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.