Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Comparative genomics of dairy-associated Staphylococcus aureus from selected sub-Saharan African regions reveals milk as reservoir for human-and animal-d

Front Microbiol · 2022
L1 89/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Every question reproduced
-1 pts
From: “every question reproduced”
Total score -7
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
  • No relevant deviation in data/preprocessing
  • No authors-side cause for any deviation
  • Reported values are derivable from the shared data
  • Any deviation was negligible
  • The central claim held under reproduction
  • Overall, the reproduction was clean
What did not (or only partly)
  • Every checked point held up.
How its reproducibility compares
89/100
Reproducibility score
0.8 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 77% of all assessed papers rank 246 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

1:1 REPRODUCED. Third-party-tool-on-paper's-data (BRIEF P16): re-ran sanger-pathogens/assembly-stats v1.0.1 + GenBank .gbff feature parsing + tseemann/mlst v2.35.0 on all 20 deposited ILS S. aureus assemblies (PRJNA310553 -> 20 GCA accessions) on «our HPC» SLURM «job», compared to Table 1. 115/120 cells exact: genome length 20/20, contigs 20/20, CDS 20/20, genes 20/20 all EXACT; GC% 19/20 exact + 1 within 0.1 rounding (ILS-37). MLST 16/20 exact; the 4 non-exact MLST cells (ILS-12/49/71 unassignable, ILS-37 ST130->ST3602) are PubMLST S. aureus database evolution since 2022, NOT data discrepancies - the deterministic columns from the same assemblies all match. No fabrication signal: every reported Table 1 number is derivable from the deposited data. NOT attempted (out of scope): the pathogenwatch/cgMLST 956-genome phylogeny + clades (external web service), phenotypic AMR, yersiniabactin screen over external genomes.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 89
    assessed: 2026-06-21 ⛓ 4f4987873e8f
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-21
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-21
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The study aimed to use whole genome sequencing to characterize the population structure and genetic composition of sub-Saharan African dairy-associated S. aureus strains, in order to determine their relationship to human and animal strains and the public health risks they pose.

Core claims
  • Milk serves as a reservoir for both human- and animal-derived S. aureus strains in sub-Saharan Africa finding
  • A putative novel siderophore (yersiniabactin-like) operon was identified in multiple strains forming a distinct animal-associated clade with strains of global origin finding
  • The putative siderophore shares high genetic identity with that of Streptococcus equi, suggesting possible horizontal gene transfer mechanism
  • Dairy-derived African S. aureus strains harbor virulence genes (pvl, scn, various enterotoxins, leucocidins) and antibiotic resistance genes, indicating public health risk requiring a One Health approach finding
  • cgMLST-based WGS analysis provides high-resolution insight into S. aureus population structure compared to conventional MLST method
  • PCR primers were designed targeting four core genes of the putative yersiniabactin siderophore operon for strain screening method
  • Growth under iron-limited conditions was compared between yersiniabactin operon-positive and -negative strains to assess siderophore functional relevance method
Experimental setups
Assay System Perturbation Readout Platform
Whole genome sequencing / de novo assembly 20 S. aureus ILS strains from camel, cow, and goat milk (Côte d'Ivoire, Kenya, Somalia) none genome sequence, size, GC content, gene/CDS counts, contigs Illumina MiSeq 2x150bp, Nextera XT library prep, CLC Genomics Workbench v11
cgMLST phylogenetic analysis 956 S. aureus genomes (20 ILS + 936 global reference/African genomes) none clade assignment, phylogenetic relationships WGSA.net / pathogen.watch
Comparative genomics / singleton (unique CDS) analysis Reduced set of 79 S. aureus reference and African strains none presumably unique CDS per strain/group EDGAR v2.3
Virulence factor gene screening S. aureus ILS genome assemblies none presence of virulence factor genes (bidirectional best hit, 70% identity cutoff) VFDB with blastp
Antimicrobial resistance gene and plasmid replicon screening S. aureus ILS genome assemblies none presence of ARGs and plasmid replicons Mobtyper, ResFinder v4.1, Center for Genomic Epidemiology tools
Antibiotic disk diffusion sensitivity testing 20 S. aureus ILS strains bacitracin, ciprofloxacin, cefoxitin, penicillin, teicoplanin, fosfomycin, tetracycline disks zone of inhibition / antibiotic sensitivity classification CLSI disk diffusion method
Yersiniabactin operon BLAST screening and PCR validation 928 S. aureus genomes; subset of strains for PCR none presence/absence of yersiniabactin siderophore operon genes (amp ligase, non-ribosomal peptide synthetase, polyketide synthase/synthetase) CLC Genomics Workbench BLAST database; PCR with agarose gel electrophoresis
Growth curve analysis under iron limitation S. aureus strains, yersiniabactin operon-positive (n=11) vs negative (n=18) iron-chelated (Chelex-100) deferrated TSB (DTSB) vs normal TSB lag phase (LPD), maximum growth rate (MGR), area under curve (AUC) from OD600 growth curves Eon microtiter plate reader (BioTek); R package 'opm'; GraphPad Prism v9.3.1
Key results
  • The 20 ILS genomes had sizes of 2.66-2.82 Mb, GC content of 32.6-32.7%, and encoded 2,676-2,960 genes, assembled into 21-134 contigs
  • cgMLST of 956 S. aureus genomes identified seven clades, with five major clades represented by reference strains (USA300/COL/Newman/NCTC8325, TW20/JKD6008, Mu50/N315/JH1, HO50960412, MRSA252)
  • African human isolates (n=33) were distributed among all clades except clade 1, with the majority clustering in clades 6 and 7 n=33 total, n=12 in clades 6/7
  • A putative novel siderophore operon was detected in multiple strains forming a distinct animal-related clade containing strains of global origin
  • The putative siderophore showed high genetic identity to the siderophore of Streptococcus equi
  • Dairy isolates carried virulence genes including pvl, scn (human immune evasion factor), enterotoxins, leucocidins, and various antibiotic resistance genes
  • Yersiniabactin operon-positive (n=11) and negative (n=18) strains were compared for growth under iron limitation
Key statistics
  • count 20 (number of African S. aureus dairy isolates (ILS strains) sequenced)
  • count 956 (total genomes analyzed in cgMLST (20 ILS + 936 global/reference genomes))
  • count seven clades (number of clades identified in cgMLST phylogenetic tree)
  • count clade1=186, clade2=161, clade3=213, clade4=208, clade5=111 (isolate counts within the five major cgMLST clades)
  • count 33 (number of African human-derived S. aureus isolates included in phylogenetic comparison)
  • count 11 positive, 18 negative (strains classified by yersiniabactin siderophore operon presence for growth comparison assay)
  • other 2.66-2.82 Mb (genome size range of the 20 ILS strains)
  • other 32.6%-32.7% (G+C content range of the 20 ILS strains)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The study is primarily descriptive comparative genomics: whole-genome sequences of 20 African dairy-associated S. aureus strains were compared to a global collection via cgMLST phylogenetics, and genomes were screened for virulence, resistance, and siderophore-operon genes using BLAST-based homology searches, without an inferential statistical framework applied to these genomic comparisons. Phenotypic data (antibiotic disk diffusion, iron-limited growth curves) were also generated. For one specific comparison — growth parameters (lag phase, maximum growth rate, area under the curve) derived from OD600 growth curves — yersiniabactin operon-positive (n=11) versus negative (n=18) strains were compared using the nonparametric Mann-Whitney U rank-sum test, run in GraphPad Prism (v9.3.1) on values standardized relative to iron-replete (TSB) growth.

Replicationmixed Sample sizeGrowth comparison based on 11 yersiniabactin-positive vs 18 yersiniabactin-negative strains, each strain tested in three independent biological experiments performed in duplicate; no formal power/sample-size calculation stated. Antibiotic sensitivity assays performed in duplicate on two separate occasions. GroupsYersiniabactin siderophore operon-positive vs negative S. aureus strains (growth under iron limitation) Pairingunpaired Randomization/blindingnot stated Dispersionunclear
Statistical tests used
Test Applied to n Assumptions
Mann-Whitney U rank-sum test Standardized DTSB (iron-limited) growth parameters (lag phase duration, maximum growth rate, area under the curve) compared between yersiniabactin operon-positive and negative S. aureus strains n=11 (operon-positive) vs n=18 (operon-negative) strains, each strain assessed in 3 independent biological experiments performed in duplicate not stated
Approaches that could also have been used
  • Growth parameter differences between operon-positive and operon-negative strains were assessed with the Mann-Whitney U test.
    Could also: An unpaired two-tailed Student's t-test (if parameter distributions are approximately normal) or a linear mixed-effects model incorporating strain and biological-replicate as random effects — A parametric test can offer more statistical power when normality holds, and a mixed-effects model would formally account for the nested structure of three biological replicates run in duplicate per strain rather than treating replicate-derived values as fully independent.
  • Three growth parameters (lag phase, maximum growth rate, area under the curve) were each tested separately for the same group comparison.
    Could also: A multiplicity correction such as Bonferroni or Benjamini-Hochberg FDR applied across the three parameter tests — Since all three parameters test the same underlying group comparison, a family-wise or false-discovery correction would help contextualize the joint chance of at least one false-positive result across the parameter set.
  • Growth curves were reduced to three summary parameters (LPD, MGR, AUC) before statistical comparison.
    Could also: Functional data analysis or nonlinear mixed-model fitting of the full OD600-versus-time trajectories — Modeling the entire growth curve can reveal timing or shape differences between groups that summary parameters alone might not fully capture.
  • The Mann-Whitney U comparison reports (implicitly, per the methods) a significance test without an accompanying effect-size statistic.
    Could also: A rank-biserial correlation or Hodges-Lehmann estimator alongside the U test — An effect-size measure would convey the magnitude of the difference between operon-positive and operon-negative strains, complementing the significance result, which is particularly informative given the modest group sizes (11 vs 18).
  • Genomic features such as yersiniabactin operon presence are described across clades and geographic/host origins in narrative and tabular form without a formal test of association.
    Could also: Fisher's exact test or chi-square test of independence between operon presence and categorical variables (e.g., clade, host species, country) — A formal association test, with Fisher's exact test being well suited to small cell counts, could quantify whether the observed distribution patterns depart from what would be expected by chance.
  • Antibiotic sensitivity testing was performed in duplicate on two separate occasions, with results presumably reported as categorical susceptibility calls.
    Could also: Reporting inter-assay agreement statistics (e.g., percent agreement or a concordance measure) between the two testing occasions — A concordance metric would communicate the reproducibility of the disk-diffusion results across the two independent testing occasions.
Software: GraphPad Prism 9.3.1 · R package 'opm' · CLC Genomics Workbench 11

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Figures / tables: Table
C1-length
Reported
Table 1 genome length (bp), 20 ILS genomes
Reproduced
20/20 exact (assembly-stats sum)
exact
C2-contigs
Reported
Table 1 No. of contigs, 20 genomes
Reproduced
20/20 exact (assembly-stats n)
exact
C3-gc
Reported
Table 1 G+C content (%), 20 genomes
Reproduced
19/20 exact, 1 within 0.1 (ILS-37 32.7 vs 32.8 rounding)
within tolerance
C4-cds
Reported
Table 1 CDS count, 20 genomes
Reproduced
20/20 exact (GenBank .gbff CDS features)
exact
C5-genes
Reported
Table 1 Genes count, 20 genomes
Reproduced
20/20 exact (GenBank .gbff gene features)
exact
C6-mlst
Reported
Table 1 MLST sequence type, 20 genomes
Reproduced
16/20 exact; 3 unassignable (novel/partial allele) + 1 mismatch ILS-37 ST130->ST3602, all PubMLST DB-version drift since 2022
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 89/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟢3. Location of the main deviation
🟢4. Cause of the deviation
🟢5. Derivability / plausibility
🟢6. Severity of the deviation
🟢7. Core claim
🟢8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Every question reproduced
-1 pts
From: “every question reproduced”
Total score -7

1:1 reproduction. Re-running assembly-stats, GenBank .gbff parsing, and tseemann/mlst on all 20 deposited ILS assemblies (PRJNA310553) reproduces Table 1 with 115/120 cells exact. The only deviations are one GC% rounding (ILS-37 32.7 vs 32.8, Δ0.1) and 4 MLST cells explained by PubMLST database drift since 2022 — both on the technical/expected side, not the authors' or a data-availability defect. Every reported value is derivable from the deposited data with no fabrication signal; the broader phylogeny/cgMLST/reservoir conclusions were legitimately out of scope (external web service).

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

156.1 k
tokens (I/O) · 8.7 M incl. cache
45 min
runtime · 0.01 CPU-h
0.3 GB
peak RAM
1
HPC jobs
hummel
machine