Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

A step forward for Shiga toxin-producing Escherichia coli identification and characterization in raw milk using long-read metagenomics.

Microb Genom · 2022
L1 75/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +5
✓ What held up
  • Same input data as the authors
  • No relevant deviation in data/preprocessing
  • No authors-side cause for any deviation
  • Any deviation was negligible
What did not (or only partly)
  • 🟡Reported values were only indirectly comparable
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
75/100
Reproducibility score
at the mean
vs. all fields · 1173 studies
🎯 Scores higher than 45% of all assessed papers rank 612 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

REPRODUCED (partial, strong). Re-ran the paper's pipeline on its own data PRJNA835223 (31 ONT MinION runs) on «our HPC» SLURM. EXACT: C1 rasusa determinism (out bases = cov x genome; same seed byte-identical). WITHIN-TOL: C2 read/base retention (reads 8.40-57.46% vs 8.49-57.5%; bases mean 73.64 vs 73.85%), C3 Kraken2 classification (n=30: 88.07-99.67% vs 88.73-99.77%), C4a uncontaminated E.coli (min 0.03% exact), C5 Flye assembly length (6.14 vs 5.92 Mb). PARTIAL: C4b/C8 contaminated E.coli % run ~10pp high (Kraken DB build 2019 vs 2017, same direction/regime); C6 co-localization 15/21 strict (18/21 with >=2 markers) vs 19/21 (assembly fragmentation, no Strainberry step). C7 min-coverage titration running. All numeric grades PROVISIONAL (human audit). Wet-lab claims (qdPCR Tables 2-3, enrichment optimization, cfu detection limits) OUT OF SCOPE. No fabrication concerns: every reproduced value is derivable from the shipped data via the named tools.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 100
    assessed: 2026-06-19 ⛓ 6ba0c1a5b03b
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-30
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The paper tests whether long-read (MinION) sequencing metagenomics can identify and characterize eae-positive Shiga toxin-producing E. coli (STEC) directly from artificially contaminated raw cow's milk without requiring a laborious bacterial isolation step.

Core claims
  • Long-read metagenomics enables isolation-independent identification and characterization of eae-positive STEC directly from raw milk. finding
  • Short-read sequencing metagenomics produces poor assembly contiguity for STEC due to high mobile genetic element content, limiting co-localization of stx and eae on the same contig. finding
  • The stx-prophage can integrate at variable chromosomal sites, placing stx and eae genes up to 2.1 Mb apart, which requires long-read assembly to resolve on a single contig. mechanism
  • Optimized enrichment conditions (acriflavine-supplemented buffered peptone water at 37°C) improve recovery/detectability of STEC in raw milk prior to sequencing. method
  • STECmetadetector, a freely available Snakemake pipeline, automates STEC read classification, assembly, and virulence/serotype/MLST characterization from long-read metagenomics data. resource
  • In silico subsampling of ONT reads to defined genome coverages was used to determine the minimum coverage needed for reliable co-assembly of stx and eae on the same contig. method
  • An eae-positive STEC O26 strain was successfully identified in raw milk artificially contaminated at levels as low as 5 c.f.u. ml-1 after enrichment. finding
  • Multiple long-read assemblers (Flye, Raven, Canu) were compared for assembly quality during method development. method
Experimental setups
Assay System Perturbation Readout Platform
ONT long-read sequencing and de novo assembly (Flye/Raven/Canu) with quast/genial-abricate/VFDB virulence screening pure culture of 10 eae-positive STEC strains in silico subsampling to defined genome coverages (3x-70x) assembly contiguity and co-localization of stx and eae on the same contig MinION
real-time PCR (qPCR) enriched raw cow's milk samples (5 fresh field samples) enrichment temperature (37 or 41.5°C) in buffered peptone water presence of stx1, stx2, eae, and wecA/cdgR markers CFX96 real-time detection system (BioRad)
artificial contamination and enrichment culture STEC-negative raw cow's milk spiked with STEC O26 strains (4712-O26, 6423-O26) inoculation at ~10, 100, 1000 c.f.u. ml-1 with/without acriflavine, 37 or 41.5°C incubation bacterial recovery/growth enabling downstream detection
quantitative digital PCR (qdPCR) DNA from enriched artificially contaminated raw milk none (quantification of prior contamination/enrichment conditions) quantification of total E. coli (wecA) and inoculated STEC O26 (wzxO26) Fluidigm BioMark system, qdPCR 37K IFC microfluidic chips
long-read metagenomic sequencing DNA extracted from enriched artificially contaminated raw milk STEC contamination level and enrichment condition raw sequencing reads for downstream taxonomic classification and assembly MinION Mk1B/Mk1C, R9.4.1 FLO-MIN106 flow cell
metagenomic bioinformatics pipeline (kraken2 classification, Flye metagenome assembly, abricate serotyping/virulence typing, mlst, Checkm, Strainberry) E. coli-assigned reads extracted from raw milk metagenome none serotype, virulence gene content, MLST, assembly completeness/contamination, strain heterogeneity STECmetadetector (Snakemake pipeline)
Key results
  • An eae-positive STEC O26 strain was directly identified from raw milk enriched in acriflavine-supplemented BPW at 37°C, down to an artificial contamination level of 5 c.f.u. ml-1. 5 c.f.u. ml-1
  • Two eae-positive STEC O26 (4712-O26, 6423-O26) strains used for contamination both had prior-characterized stx/eae genomic distance of 1.9 Mb. 1.9 Mb
  • ONT reads from 10 STEC strains were subsampled across a coverage gradient to empirically determine the minimum coverage sufficient for stx and eae to assemble onto the same contig.
Key statistics
  • other 5 c.f.u. ml-1 (lowest raw milk contamination level at which eae-positive STEC O26 was successfully identified after enrichment)
  • other stx/eae distance = 1.9 Mb (genomic distance between stx and eae genes in the two inoculated O26:H11 strains (Table 1))
  • other stx/eae distance up to 2.1 Mb (previously reported maximum distance between stx-phage integration site and eae across STEC strains generally)
  • other N50 range 9087-22 759 bp (read N50 values of ONT sequencing data from 10 STEC strains used for coverage subsampling)
  • other coverage levels tested: 3x, 5x, 10x, 15x, 20x, 25x, 30x, 35x, 40x, 50x, 60x, 70x (genome coverages used for in silico subsampling with rasusa)
  • count 10 eae-positive STEC strains (strains used for in silico minimum coverage determination)
  • count inoculation levels ~10^1, 10^2, 10^3 c.f.u. ml-1 (actual 1.33x10^1 to 1.78x10^3 c.f.u. ml-1) (artificial STEC O26 spike-in levels in raw milk)
  • count 31 SRA BioSample accessions deposited (raw sequence data deposition under BioProject PRJNA835223)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This is a methods-development study describing optimization of raw milk enrichment conditions, a coverage-titration experiment across sequencing assemblers, and construction of a bioinformatics pipeline (STECmetadetector) for STEC identification from long-read metagenomics data. Results are reported descriptively (e.g., contig assembly metrics, c.f.u. ranges from triplicate plate counts, presence/absence of virulence genes on the same contig) rather than through formal inferential hypothesis testing, and no dedicated statistics section or p-values appear in the text provided.

Replicationunclear Sample sizeArtificial contamination experiments were 'performed in triplicate'; spike-in levels were determined by triplicate plate counting; no formal sample-size or power calculation is described Groupsenrichment temperature (37 vs 41.5 °C), acriflavine supplementation vs none, spike-in concentration levels, and sequencing coverage/assembler combinations (Flye, Raven, Canu) for genome assembly quality Pairingunclear Randomization/blindingnot stated Dispersionrange Exact p-valuesno Effect sizesno Confidence intervalsno
Approaches that could also have been used
  • Genome assembly quality across different sequencing coverages and three assemblers (Flye, Raven, Canu) was compared descriptively using QUAST metrics and presence of stx/eae on the same contig.
    Could also: A formal statistical comparison (e.g., a mixed-effects or ANOVA-type model treating assembler and coverage as factors) on contiguity metrics such as N50 or contig count — This would allow quantifying whether observed differences between assemblers or coverage levels exceed what could be expected from sampling variation, complementing the descriptive threshold-based approach used here.
  • Spike-in bacterial concentrations were reported as ranges derived from triplicate plate counts (e.g., 1.25–1.78×10^3 c.f.u. ml⁻¹).
    Could also: Reporting the mean with SD or a 95% confidence interval alongside the range — Mean ± SD/CI conveys both central tendency and variability in a standardized way and facilitates comparison across conditions or with other studies, whereas a range alone reports only the extremes observed.
  • Enrichment conditions (temperature, acriflavine supplementation) and STEC detectability were compared without a stated formal hypothesis test.
    Could also: A chi-square or Fisher's exact test (for detection/non-detection outcomes) or ANOVA (for continuous quantification outcomes such as qdPCR counts) across enrichment conditions — Such tests would provide a formal statistical basis for concluding whether detection rates or quantitative recovery differ between enrichment conditions, in addition to the descriptive comparison presented.
  • STEC quantification in enriched raw milk was performed using quantitative digital PCR (qdPCR) via the Fluidigm BioMark system.
    Could also: Explicitly reporting Poisson-based confidence intervals for the digital PCR concentration estimates — Digital PCR quantification is inherently based on Poisson statistics, and stating the associated confidence interval would communicate the precision of the concentration estimate in addition to the point estimate.
  • Multiple experimental factors (temperature, acriflavine, spike level, coverage, assembler) were varied and compared without an explicit multiplicity-correction method reported.
    Could also: A false-discovery-rate method such as Benjamini-Hochberg, or a Bonferroni correction, if multiple pairwise statistical tests were to be performed across these conditions — When many conditions or comparisons are evaluated statistically, a multiplicity correction helps control the overall false-positive rate across the family of tests.
Software: QUAST 5.0.2 · GENIAL/abricate 1.0 / 0.8.7 · rasusa 0.6.0 · Flye 2.8.1-b1676 / 2.9-b1768 · Fluidigm digital PCR analysis software 4.1.2 · R (utility scripts, packages listed in Table S4)

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-36748417

Paper: Jaudou S, Deneke C, Tran ML, et al. (2022) A step forward for Shiga toxin-producing Escherichia coli identification and characterization in raw milk using long-read metagenomics. Microb Genom 8(12):000911. PMID 36748417 / PMC9836091 / DOI 10.1099/mgen.0.000911.

Data: SRA BioProject PRJNA835223 — 31 Oxford Nanopore MinION metagenomic sequencing runs of raw cow milk (uncontaminated controls + artificially STEC-contaminated, enriched in BPW ± acriflavine at 37 °C / 41.5 °C). Confirmed via ENA filereport: 31 runs, 36.4 M reads, 38.8 Gbp total. (See data/ena_runs.tsv.)

Authors' pipeline: STECmetadetector — a Snakemake pipeline (gitlab.com/bfr_bioinformatics/STECmetadetector). The BRIEF's code link points to rasusa (mbhall88/rasusa), one named component used for coverage down-sampling. Per the study rules a third-party tool applied to the paper's own data is an equally valid reproduction target.

Pipeline steps (from Methods)

Step Tool (version) Key params
Basecall / demux Guppy 4.4.2+ (NOT in pipeline; reads deposited already basecalled)
Adapter/barcode trim Porechop 0.2.4 defaults
Quality filter NanoFilt 2.8.0 length ≥ 1 kb, q-score ≥ 7
Read QC NanoPlot 1.39
Taxonomic classification Kraken2 2.1.2 Minikraken DB 8 GB (2017-10-18)
E. coli read extraction KrakenTools 1.2
Pre-assembly gene detection Minimap2 2.24 vs CGE stx/eae/O-group DB
Down-sampling rasusa 0.6.0 target coverage 3×…70×
Assembly Flye 2.8.1 / 2.9 --meta (metagenome)
(assembler comparison) Raven 1.2.2/1.7.0, Canu 2.1.1
Assembly QC QUAST 5.0.2
Virulence genes / co-loc GENIAL 1.0 (abricate 0.8.7), VFDB 2020-05-29
Serotype / virulence abricate 1.0.1
MLST mlst 2.19 ecoli scheme
Assembly contamination CheckM 1.1.3
Strain separation Strainberry 1.1

IN SCOPE (pipeline-derived, computational — attempt to reproduce)

ID Reported result Pipeline Tractability
C1 rasusa deterministically down-samples an ONT run to a target coverage (output bases ≈ coverage × genome size) rasusa 0.6.0 HIGH — named tool, deterministic, cheap
C2 Read filtering retains 8.49–57.5 % of reads; ≥42.73 % of bases (mean 73.85 ± 11.88 %) across 31 runs Porechop + NanoFilt HIGH — per-run, deterministic
C3 88.73–99.77 % of filtered reads taxonomically assigned (Kraken2) Kraken2 + Minikraken8GB MED — needs DB download
C4 E. coli read proportion: uncontaminated 0.03–46.24 % (mean 5.73); contaminated 13.9–80.62 % (mean 64.46 ± 18.1) Kraken2/KrakenTools MED
C5 Mean total Flye assembly length 5.92 Mb (4.96–6.4 Mb) for n=21 contaminated samples Flye --meta MED-HEAVY
C6 wzx-O26 + stx + eae co-localized on one contig in 19/21 Flye assemblies Flye + abricate HEAVY
C7 Minimum coverage ~35× (Flye) for reliable stx/eae co-localization (down-sampling series) rasusa + Flye HEAVY
C8 Fig 4a: E. coli read proportion vs contamination level (63.75/67.93/72.53 %) Kraken2 MED

OUT OF SCOPE (wet-lab / not pipeline-derived — not attempted)

  • qdPCR quantification (Tables 2, 3) — laboratory measurement.
  • Enrichment optimization (temperature, acriflavine effect on growth) — wet-lab.
  • Detection-limit c.f.u./ml claims — depend on inoculation, wet-lab.
  • DNA extraction / real-time PCR screening — wet-lab.

Strategy

Reach the cheap, deterministic floor first (C1 rasusa, C2 NanoFilt retention), then push into Kraken2 (C3/C4) and Flye assembly (C5/C6) as compute allows. All heavy compute on «our HPC» SLURM; data on «infra». No --mem flag («infra» rule).

Figures / tables: Fig 1Fig 4a
DS1
Reported
31 MinION metagenomic runs (PRJNA835223)
Reproduced
31 runs (ENA filereport)
exact
C1
Reported
rasusa deterministic down-sampling; out bases ~= cov x genome
Reproduced
cov35x5Mb -> 175,001,227 bp (+0.0007%); same seed byte-identical
exact
C2a
Reported
8.49-57.5 % reads retained (31 runs)
Reproduced
8.40-57.46 % (n=31)
within tolerance
C2b
Reported
>42.73 % bases; mean 73.85+/-11.88 %
Reproduced
42.34-94.79 %; mean 73.64+/-11.82 % (n=31)
within tolerance
C3
Reported
88.73-99.77 % filtered reads classified (Kraken2, n=30)
Reproduced
88.07-99.67 % (n=30); 40.71-99.67 % (n=31)
within tolerance
C4a
Reported
uncontaminated E.coli 0.03-46.24 % mean 5.73 (n=10)
Reproduced
0.03-52.73 % mean 6.61 (n=10)
within tolerance
C4b
Reported
contaminated E.coli 13.9-80.62 % mean 64.46 (n=21)
Reproduced
16.57-92.08 % mean 74.85 (n=21)
partial
C5
Reported
mean total Flye assembly length 5.92 Mb (4.96-6.4, n=21)
Reproduced
mean 6.14 Mb (5.90-7.02, n=21)
within tolerance
C6
Reported
wzx-O26+stx+eae co-localized on one contig 19/21
Reproduced
15/21 strict; 18/21 with >=2 of 3
partial
C7
Reported
min ~35x coverage for reliable co-localization (8/10 strains)
Reproduced
coverage titration running («job»)
partial
C8
Reported
Fig 4a E.coli proportion by cfu 63.75/67.93/72.53 % (5/50/500)
Reproduced
75.57/78.60/87.12 % (monotonic increase reproduced)
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 75/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟡2. Endpoint comparability
🟢3. Location of the main deviation
🟢4. Cause of the deviation
🟡5. Derivability / plausibility
🟢6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +5

This is an explicitly preliminary reproduction: agreement.json is not-run-yet, completed_utc is null, and only DS1 (31 ONT runs in PRJNA835223) has been checked, matching the paper 1:1. The substantive claims (C1–C8: read/base retention, Kraken2 E.coli proportions, Flye assembly length, stx/eae co-localization, ~35x coverage threshold) were scoped but not computed, so derivability and the central conclusion are only partly established — no deviation is on the authors' or data side; the data is fully public and the count is exact. Severity is negligible for what was run, but overall the reproduction is incomplete (compute pending), warranting yellow on q5/q7/q8 rather than green or red.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

57.5 k
tokens (I/O) · 2.7 M incl. cache
8 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.