A step forward for Shiga toxin-producing Escherichia coli identification and characterization in raw milk using long-read metagenomics.
The main results reproduced: recomputed values matched the published ones within tolerance.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Same input data as the authors
- ✓No relevant deviation in data/preprocessing
- ✓No authors-side cause for any deviation
- ✓Any deviation was negligible
- 🟡Reported values were only indirectly comparable
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
REPRODUCED (partial, strong). Re-ran the paper's pipeline on its own data PRJNA835223 (31 ONT MinION runs) on «our HPC» SLURM. EXACT: C1 rasusa determinism (out bases = cov x genome; same seed byte-identical). WITHIN-TOL: C2 read/base retention (reads 8.40-57.46% vs 8.49-57.5%; bases mean 73.64 vs 73.85%), C3 Kraken2 classification (n=30: 88.07-99.67% vs 88.73-99.77%), C4a uncontaminated E.coli (min 0.03% exact), C5 Flye assembly length (6.14 vs 5.92 Mb). PARTIAL: C4b/C8 contaminated E.coli % run ~10pp high (Kraken DB build 2019 vs 2017, same direction/regime); C6 co-localization 15/21 strict (18/21 with >=2 markers) vs 19/21 (assembly fragmentation, no Strainberry step). C7 min-coverage titration running. All numeric grades PROVISIONAL (human audit). Wet-lab claims (qdPCR Tables 2-3, enrichment optimization, cfu detection limits) OUT OF SCOPE. No fabrication concerns: every reproduced value is derivable from the shipped data via the named tools.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 100assessed: 2026-06-19 ⛓ 6ba0c1a5b03b
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-30
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetThe paper tests whether long-read (MinION) sequencing metagenomics can identify and characterize eae-positive Shiga toxin-producing E. coli (STEC) directly from artificially contaminated raw cow's milk without requiring a laborious bacterial isolation step.
- ★ Long-read metagenomics enables isolation-independent identification and characterization of eae-positive STEC directly from raw milk. finding
- ★ Short-read sequencing metagenomics produces poor assembly contiguity for STEC due to high mobile genetic element content, limiting co-localization of stx and eae on the same contig. finding
- ★ The stx-prophage can integrate at variable chromosomal sites, placing stx and eae genes up to 2.1 Mb apart, which requires long-read assembly to resolve on a single contig. mechanism
- ★ Optimized enrichment conditions (acriflavine-supplemented buffered peptone water at 37°C) improve recovery/detectability of STEC in raw milk prior to sequencing. method
- ★ STECmetadetector, a freely available Snakemake pipeline, automates STEC read classification, assembly, and virulence/serotype/MLST characterization from long-read metagenomics data. resource
- ★ In silico subsampling of ONT reads to defined genome coverages was used to determine the minimum coverage needed for reliable co-assembly of stx and eae on the same contig. method
- ★ An eae-positive STEC O26 strain was successfully identified in raw milk artificially contaminated at levels as low as 5 c.f.u. ml-1 after enrichment. finding
- Multiple long-read assemblers (Flye, Raven, Canu) were compared for assembly quality during method development. method
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| ONT long-read sequencing and de novo assembly (Flye/Raven/Canu) with quast/genial-abricate/VFDB virulence screening | pure culture of 10 eae-positive STEC strains | in silico subsampling to defined genome coverages (3x-70x) | assembly contiguity and co-localization of stx and eae on the same contig | MinION |
| real-time PCR (qPCR) | enriched raw cow's milk samples (5 fresh field samples) | enrichment temperature (37 or 41.5°C) in buffered peptone water | presence of stx1, stx2, eae, and wecA/cdgR markers | CFX96 real-time detection system (BioRad) |
| artificial contamination and enrichment culture | STEC-negative raw cow's milk spiked with STEC O26 strains (4712-O26, 6423-O26) | inoculation at ~10, 100, 1000 c.f.u. ml-1 with/without acriflavine, 37 or 41.5°C incubation | bacterial recovery/growth enabling downstream detection | — |
| quantitative digital PCR (qdPCR) | DNA from enriched artificially contaminated raw milk | none (quantification of prior contamination/enrichment conditions) | quantification of total E. coli (wecA) and inoculated STEC O26 (wzxO26) | Fluidigm BioMark system, qdPCR 37K IFC microfluidic chips |
| long-read metagenomic sequencing | DNA extracted from enriched artificially contaminated raw milk | STEC contamination level and enrichment condition | raw sequencing reads for downstream taxonomic classification and assembly | MinION Mk1B/Mk1C, R9.4.1 FLO-MIN106 flow cell |
| metagenomic bioinformatics pipeline (kraken2 classification, Flye metagenome assembly, abricate serotyping/virulence typing, mlst, Checkm, Strainberry) | E. coli-assigned reads extracted from raw milk metagenome | none | serotype, virulence gene content, MLST, assembly completeness/contamination, strain heterogeneity | STECmetadetector (Snakemake pipeline) |
- – An eae-positive STEC O26 strain was directly identified from raw milk enriched in acriflavine-supplemented BPW at 37°C, down to an artificial contamination level of 5 c.f.u. ml-1. 5 c.f.u. ml-1
- – Two eae-positive STEC O26 (4712-O26, 6423-O26) strains used for contamination both had prior-characterized stx/eae genomic distance of 1.9 Mb. 1.9 Mb
- – ONT reads from 10 STEC strains were subsampled across a coverage gradient to empirically determine the minimum coverage sufficient for stx and eae to assemble onto the same contig.
- other 5 c.f.u. ml-1 (lowest raw milk contamination level at which eae-positive STEC O26 was successfully identified after enrichment)
- other stx/eae distance = 1.9 Mb (genomic distance between stx and eae genes in the two inoculated O26:H11 strains (Table 1))
- other stx/eae distance up to 2.1 Mb (previously reported maximum distance between stx-phage integration site and eae across STEC strains generally)
- other N50 range 9087-22 759 bp (read N50 values of ONT sequencing data from 10 STEC strains used for coverage subsampling)
- other coverage levels tested: 3x, 5x, 10x, 15x, 20x, 25x, 30x, 35x, 40x, 50x, 60x, 70x (genome coverages used for in silico subsampling with rasusa)
- count 10 eae-positive STEC strains (strains used for in silico minimum coverage determination)
- count inoculation levels ~10^1, 10^2, 10^3 c.f.u. ml-1 (actual 1.33x10^1 to 1.78x10^3 c.f.u. ml-1) (artificial STEC O26 spike-in levels in raw milk)
- count 31 SRA BioSample accessions deposited (raw sequence data deposition under BioProject PRJNA835223)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This is a methods-development study describing optimization of raw milk enrichment conditions, a coverage-titration experiment across sequencing assemblers, and construction of a bioinformatics pipeline (STECmetadetector) for STEC identification from long-read metagenomics data. Results are reported descriptively (e.g., contig assembly metrics, c.f.u. ranges from triplicate plate counts, presence/absence of virulence genes on the same contig) rather than through formal inferential hypothesis testing, and no dedicated statistics section or p-values appear in the text provided.
-
Genome assembly quality across different sequencing coverages and three assemblers (Flye, Raven, Canu) was compared descriptively using QUAST metrics and presence of stx/eae on the same contig.↳ Could also: A formal statistical comparison (e.g., a mixed-effects or ANOVA-type model treating assembler and coverage as factors) on contiguity metrics such as N50 or contig count — This would allow quantifying whether observed differences between assemblers or coverage levels exceed what could be expected from sampling variation, complementing the descriptive threshold-based approach used here.
-
Spike-in bacterial concentrations were reported as ranges derived from triplicate plate counts (e.g., 1.25–1.78×10^3 c.f.u. ml⁻¹).↳ Could also: Reporting the mean with SD or a 95% confidence interval alongside the range — Mean ± SD/CI conveys both central tendency and variability in a standardized way and facilitates comparison across conditions or with other studies, whereas a range alone reports only the extremes observed.
-
Enrichment conditions (temperature, acriflavine supplementation) and STEC detectability were compared without a stated formal hypothesis test.↳ Could also: A chi-square or Fisher's exact test (for detection/non-detection outcomes) or ANOVA (for continuous quantification outcomes such as qdPCR counts) across enrichment conditions — Such tests would provide a formal statistical basis for concluding whether detection rates or quantitative recovery differ between enrichment conditions, in addition to the descriptive comparison presented.
-
STEC quantification in enriched raw milk was performed using quantitative digital PCR (qdPCR) via the Fluidigm BioMark system.↳ Could also: Explicitly reporting Poisson-based confidence intervals for the digital PCR concentration estimates — Digital PCR quantification is inherently based on Poisson statistics, and stating the associated confidence interval would communicate the precision of the concentration estimate in addition to the point estimate.
-
Multiple experimental factors (temperature, acriflavine, spike level, coverage, assembler) were varied and compared without an explicit multiplicity-correction method reported.↳ Could also: A false-discovery-rate method such as Benjamini-Hochberg, or a Bonferroni correction, if multiple pairwise statistical tests were to be performed across these conditions — When many conditions or comparisons are evaluated statistically, a multiplicity correction helps control the overall false-positive rate across the family of tests.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-36748417
Paper: Jaudou S, Deneke C, Tran ML, et al. (2022) A step forward for Shiga toxin-producing Escherichia coli identification and characterization in raw milk using long-read metagenomics. Microb Genom 8(12):000911. PMID 36748417 / PMC9836091 / DOI 10.1099/mgen.0.000911.
Data: SRA BioProject PRJNA835223 — 31 Oxford Nanopore MinION metagenomic
sequencing runs of raw cow milk (uncontaminated controls + artificially
STEC-contaminated, enriched in BPW ± acriflavine at 37 °C / 41.5 °C). Confirmed
via ENA filereport: 31 runs, 36.4 M reads, 38.8 Gbp total. (See data/ena_runs.tsv.)
Authors' pipeline: STECmetadetector — a Snakemake pipeline
(gitlab.com/bfr_bioinformatics/STECmetadetector). The BRIEF's code link points
to rasusa (mbhall88/rasusa), one named component used for coverage
down-sampling. Per the study rules a third-party tool applied to the paper's own
data is an equally valid reproduction target.
Pipeline steps (from Methods)
| Step | Tool (version) | Key params |
|---|---|---|
| Basecall / demux | Guppy 4.4.2+ | (NOT in pipeline; reads deposited already basecalled) |
| Adapter/barcode trim | Porechop 0.2.4 | defaults |
| Quality filter | NanoFilt 2.8.0 | length ≥ 1 kb, q-score ≥ 7 |
| Read QC | NanoPlot 1.39 | — |
| Taxonomic classification | Kraken2 2.1.2 | Minikraken DB 8 GB (2017-10-18) |
| E. coli read extraction | KrakenTools 1.2 | — |
| Pre-assembly gene detection | Minimap2 2.24 | vs CGE stx/eae/O-group DB |
| Down-sampling | rasusa 0.6.0 | target coverage 3×…70× |
| Assembly | Flye 2.8.1 / 2.9 | --meta (metagenome) |
| (assembler comparison) | Raven 1.2.2/1.7.0, Canu 2.1.1 | — |
| Assembly QC | QUAST 5.0.2 | — |
| Virulence genes / co-loc | GENIAL 1.0 (abricate 0.8.7), VFDB 2020-05-29 | — |
| Serotype / virulence | abricate 1.0.1 | — |
| MLST | mlst 2.19 | ecoli scheme |
| Assembly contamination | CheckM 1.1.3 | — |
| Strain separation | Strainberry 1.1 | — |
IN SCOPE (pipeline-derived, computational — attempt to reproduce)
| ID | Reported result | Pipeline | Tractability |
|---|---|---|---|
| C1 | rasusa deterministically down-samples an ONT run to a target coverage (output bases ≈ coverage × genome size) | rasusa 0.6.0 | HIGH — named tool, deterministic, cheap |
| C2 | Read filtering retains 8.49–57.5 % of reads; ≥42.73 % of bases (mean 73.85 ± 11.88 %) across 31 runs | Porechop + NanoFilt | HIGH — per-run, deterministic |
| C3 | 88.73–99.77 % of filtered reads taxonomically assigned (Kraken2) | Kraken2 + Minikraken8GB | MED — needs DB download |
| C4 | E. coli read proportion: uncontaminated 0.03–46.24 % (mean 5.73); contaminated 13.9–80.62 % (mean 64.46 ± 18.1) | Kraken2/KrakenTools | MED |
| C5 | Mean total Flye assembly length 5.92 Mb (4.96–6.4 Mb) for n=21 contaminated samples | Flye --meta |
MED-HEAVY |
| C6 | wzx-O26 + stx + eae co-localized on one contig in 19/21 Flye assemblies | Flye + abricate | HEAVY |
| C7 | Minimum coverage ~35× (Flye) for reliable stx/eae co-localization (down-sampling series) | rasusa + Flye | HEAVY |
| C8 | Fig 4a: E. coli read proportion vs contamination level (63.75/67.93/72.53 %) | Kraken2 | MED |
OUT OF SCOPE (wet-lab / not pipeline-derived — not attempted)
- qdPCR quantification (Tables 2, 3) — laboratory measurement.
- Enrichment optimization (temperature, acriflavine effect on growth) — wet-lab.
- Detection-limit c.f.u./ml claims — depend on inoculation, wet-lab.
- DNA extraction / real-time PCR screening — wet-lab.
Strategy
Reach the cheap, deterministic floor first (C1 rasusa, C2 NanoFilt retention),
then push into Kraken2 (C3/C4) and Flye assembly (C5/C6) as compute allows.
All heavy compute on «our HPC» SLURM; data on «infra». No --mem flag («infra» rule).
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
This is an explicitly preliminary reproduction: agreement.json is not-run-yet, completed_utc is null, and only DS1 (31 ONT runs in PRJNA835223) has been checked, matching the paper 1:1. The substantive claims (C1–C8: read/base retention, Kraken2 E.coli proportions, Flye assembly length, stx/eae co-localization, ~35x coverage threshold) were scoped but not computed, so derivability and the central conclusion are only partly established — no deviation is on the authors' or data side; the data is fully public and the count is exact. Severity is negligible for what was run, but overall the reproduction is incomplete (compute pending), warranting yellow on q5/q7/q8 rather than green or red.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.