Comparative Genomics of Borderline Oxacillin-Resistant Staphylococcus aureus Detected during a Pseudo-outbreak of Methicillin-Resistant S. aureus in a Neonatal
The main results reproduced, with only marginal, non-material deviations.
- ✓Any deviation was negligible
- 🟡Could not use the authors’ exact input data
- 🟡Reported values were only indirectly comparable
- 🟡A deviation arose in the data or preprocessing
- 🟡A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
PARTIAL - strong core reproduction. Applied third-party tools (mlst, AMRFinderPlus, Prokka+Roary, snp-dists, rhierbaps) to the paper's own 101 deposited S. aureus assemblies (PRJNA695316; no raw reads). EXACT: C2 class split (33 BORSA/59 MSSA/9 MRSA) and C7 mecA-in-9-MRSA-only/mecC-absent. WITHIN-TOL: C3 Roary core genome 1865 vs 1859 (0.3%) and C5 dominant-ST fraction 57.6% (exact percentage; 11 vs 12 distinct STs). PARTIAL: C1 (n=101 exact, CheckM not run), C6 (ST97=5/ST398=5 exact but ST27/ST72 BORSA singletons not recovered - likely PubMLST DB version drift, 2 BORSA untyped), C9/C10 SNP distances (clonal clusters + near-identical MSSA/BORSA pair confirmed but absolute counts lower because we used Roary core-GENE SNPs vs the paper's Snippy whole-genome SNPs). MISMATCH: C4 hierBAPS 8 vs 5 lineages (param/alignment-sensitive). NOT REPRODUCIBLE/out-of-scope: the headline Random Forest AUROC 0.902 (C11/C12) - feature matrices not deposited and one feature is a wet-lab phenotype; all wet-lab phenotyping (susceptibility, PBP2a, beta-lactamase assays, chromogenic agar Tables 1-3); the assembly step itself (no reads). Conclusion: the paper's comparative-genomics backbone is well-described and reproduces 1:1 from the deposited data; the ML headline cannot be independently verified from what was deposited.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 63assessed: 2026-06-21 ⛓ fcd4d5c61c57
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-21
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetFollowing atypical MRSA screening culture results in a NICU, the study asks whether isolates flagged as MRSA are true mecA-mediated MRSA or a distinct borderline oxacillin-resistant S. aureus (BORSA) phenotype, and whether genomic features can distinguish BORSA from MSSA and explain inconsistent detection by commercial MRSA screening agars.
- ★ Of 42 isolates flagged as MRSA by screening agar, only 9 were PBP2a- and mecA-positive true MRSA, while the remaining 33 were mecA-negative and largely met criteria for BORSA finding
- ★ BORSA isolates identified in the NICU were phylogenetically diverse and did not represent clonal expansion or shared gene content, though two NICU strains showed infrequent clonal recurrence over 8 months finding
- ★ Six genomic features (substitutions/truncations in PBP2, PBP4, and GdpP, plus beta-lactamase hyperproduction) were used to build a random forest classifier distinguishing BORSA from MSSA method
- ★ The random forest classifier robustly predicted the BORSA phenotype in an external validation cohort spanning two continents (AUC = 0.902) finding
- ★ Commercial MRSA screening agars vary substantially in sensitivity and specificity for detecting MRSA versus BORSA, leading to misclassification of BORSA as MRSA finding
- ★ Beta-lactamase hyperproduction alone is insufficient to explain borderline oxacillin resistance in many isolates, implicating additional PBP and GdpP-related mechanisms mechanism
- Beta-lactamase inhibitor potentiation effect (amoxicillin vs. amoxicillin-clavulanate MIC shift) was significantly greater in BORSA than MSSA isolates finding
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| MALDI-TOF mass spectrometry | clinical S. aureus isolates (NICU screening + comparator blood isolates) | none | species confirmation | MALDI-TOF MS |
| PBP2a lateral flow immunoassay and mecA/mecC PCR | clinical S. aureus isolates | none | PBP2a protein presence and mecA/mecC gene detection | — |
| Cefoxitin and oxacillin susceptibility testing (disk diffusion and gradient diffusion on 2% NaCl agar) | clinical S. aureus isolates | none | zone size / MIC for resistance classification | Mueller-Hinton agar, CLSI methods |
| Growth assay on commercial MRSA screening agars (Spectra, MRSA Select II, BBL CHROMagar MRSA II, chromID MRSA, nonchromogenic MRSA screen, HardyCHROM) | clinical S. aureus isolates | none | colony growth abundance and pigmentation | 6 commercial chromogenic/nonchromogenic agars |
| Beta-lactamase detection (disk diffusion penicillin zone edge test, nitrocefin/Cefinase test) | clinical S. aureus isolates | none | beta-lactamase production status | Cefinase test |
| Beta-lactamase inhibitor potentiation testing (amoxicillin vs. amoxicillin-clavulanate gradient diffusion) | clinical S. aureus isolates | beta-lactamase inhibitor (clavulanate) addition | fold-change in MIC | gradient diffusion strips |
| Whole-genome sequencing / comparative genomics | 42 NICU isolates + 60 comparator blood isolates | none | phylogenetic relatedness, gene content, mutations in pbp2/pbp4/gdpP | WGS |
| Random forest classification | genomic feature set from sequenced isolates (training + 2-continent validation cohort) | none | predicted BORSA vs. MSSA phenotype | — |
- ▼ Only 9 of 42 NICU screening isolates were PBP2a/mecA-positive true MRSA 9/42
- ▲ 33 of 42 isolates were oxacillin resistant by 2% NaCl gradient diffusion 33/42
- – 24 of 33 non-MRSA investigated isolates and 9 of 60 comparator isolates met BORSA criteria 24/33; 9/60
- ▲ Random forest classifier validated across a two-continent cohort AUC = 0.902
- – MRSA-selective agars (BBL CHROMagar MRSA II, MRSA Select II, chromID MRSA) showed high specificity but lower sensitivity for BORSA detection 89% sensitivity, 100% specificity for MRSA
- – Nonchromogenic and Spectra MRSA agars showed high sensitivity but reduced specificity 100% sensitivity; 86% and 60% specificity, respectively
- ▲ Beta-lactamase inhibitor effect (MIC shift) differed significantly between BORSA and MSSA isolates P = 0.0030
- – BORSA isolates were genomically diverse rather than clonal, with rare exceptions
- pvalue P = 0.0030 (Mann-Whitney test comparing beta-lactamase inhibitor MIC-shift effect between BORSA and MSSA isolates)
- other AUC = 0.902 (Validation performance of random forest classifier predicting BORSA phenotype across a two-continent cohort)
- count 9/42 (NICU isolates confirmed as true MRSA (PBP2a+/mecA+))
- count 33/42 (NICU isolates oxacillin resistant by 2% NaCl gradient diffusion)
- count 24/33 investigated isolates; 9/60 comparator isolates (Isolates meeting BORSA case definition)
- other 89% sensitivity / 100% specificity (BBL CHROMagar MRSA II, MRSA Select II, chromID MRSA performance for MRSA detection across all 102 isolates)
- other 100% sensitivity / 86% and 60% specificity (Nonchromogenic MRSA screen agar and Spectra MRSA agar performance, respectively)
- count 6/33 BORSA vs. 6/60 MSSA (Isolates exhibiting a 4-fold difference in lactamase inhibitor effect)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This retrospective observational study phenotypically and genomically characterized 42 suspected MRSA NICU isolates alongside 60 comparator blood isolates (102 total). The primary analytical contribution was a random forest classifier trained on 6 genomic features (substitutions/truncations in PBP2, PBP4, GdpP, and beta-lactamase hyperproduction) to distinguish BORSA from MSSA, evaluated by AUC on a validation set. A Mann-Whitney U test was used to compare the beta-lactamase inhibitor effect between BORSA and MSSA groups, and diagnostic performance (sensitivity and specificity) was reported descriptively for each of six MRSA screening agars.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Mann-Whitney U test (two-group comparison) | Comparison of log2 fold change in MICs (amoxicillin vs. amoxicillin-clavulanic acid) between BORSA and MSSA isolates; Fig. S1 | 33 BORSA vs. 60 MSSA | not stated |
| Random forest classifier (machine learning model, not a classical hypothesis test); performance reported as AUC | Genomic feature-based classification of BORSA vs. MSSA across cohorts spanning two continents; validation AUC = 0.902 | — | not stated |
| Sensitivity and specificity (diagnostic accuracy calculation) | Analytical performance of six MRSA screening agars for MRSA and BORSA detection; Table 3 | 102 isolates total | na |
-
Sensitivity and specificity for each screening agar were reported as single point estimates without confidence intervals, with small true-positive denominators (e.g., 8 or 9 confirmed MRSA isolates)↳ Could also: Wilson score or Clopper-Pearson exact binomial 95% confidence intervals around each sensitivity and specificity estimate could also be reported — With very small numerators, point estimates alone carry substantial uncertainty; CIs would quantify that uncertainty and allow readers to assess whether apparent differences in agar performance are distinguishable from chance
-
Six MRSA screening agars were each evaluated on the same 102 isolates, producing separate sensitivity/specificity point estimates that are compared informally↳ Could also: McNemar's test (or its extension for multiple paired proportions) could also formally compare sensitivities or specificities across agars, since the same isolates were tested on all agars — The paired structure of the data (same isolate evaluated on every agar) is a natural match for paired-proportion tests, which would provide p-values and effect estimates for pairwise agar differences rather than relying on visually non-overlapping percentages
-
The Mann-Whitney U test was used to compare the beta-lactamase inhibitor effect (log2 MIC fold change) between BORSA and MSSA↳ Could also: A two-sample t-test or Welch's t-test on the log2-transformed fold changes could also be applied if approximate normality holds, together with a reported mean difference and 95% CI — Log-transformed ratio data often approaches normality; a parametric test would additionally yield a mean difference and confidence interval that directly quantify the magnitude of the effect, complementing the p-value
-
Dispersion around the beta-lactamase inhibitor effect in Fig. S1 was shown as standard deviation (SD)↳ Could also: Median with interquartile range (IQR), or 95% confidence intervals around group means, could also convey spread — MIC and fold-change distributions are often right-skewed; IQR or bootstrapped CIs may more faithfully represent central tendency and spread for such data, and the IQR is a natural companion to the non-parametric Mann-Whitney test used
-
A random forest classifier was used to predict the BORSA phenotype; performance was summarized with AUC on a validation set↳ Could also: Penalized logistic regression (e.g., LASSO or elastic net) could also be applied as a companion or alternative classifier on the same 6-feature set — With a small feature set and modest sample size, penalized logistic regression yields odds ratios and a transparent probability model directly interpretable as effect sizes; it would complement the random forest by providing feature-level inferential estimates
-
Phylogenetic diversity among BORSA isolates was described qualitatively as 'phylogenetically diverse and not representative of clonal expansion'↳ Could also: A quantitative measure of population structure such as the index of association (IA) or a permutation-based clustering test could also be reported alongside the phylogeny — Formal statistical tests of clonal structure provide an objective, reproducible basis for statements about phylogenetic diversity, allowing readers to assess whether the observed spread differs significantly from a random distribution of genotypes
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
scope.md — pmid-35038924
Paper: Sawhney, Ransom, Wallace, Reich, Dantas, Burnham. Comparative Genomics of Borderline Oxacillin-Resistant Staphylococcus aureus Detected during a Pseudo-outbreak of MRSA in a Neonatal Intensive Care Unit. mBio 2022. PMID 35038924 · PMC8764539 · DOI 10.1128/mbio.03196-21.
Code: https://github.com/sanjsawhney/BORSA_RFC (commit 8ee78886, pushed 2021-12-02,
MIT, default branch main). Repo contains ONLY LICENSE + README.md; all R code is
embedded inline in the README. No input data CSVs and no WGS-pipeline scripts are shipped.
Data: BioProject PRJNA695316. Deposited as 101 genome assemblies (GenBank GCA
accessions, scaffold level) — NOT raw reads (NCBI SRA count = 0, ENA read_run = 0).
Strain names encode phenotype: WUSa_oxa_BOR### (BORSA), _MS### (MSSA), _MR### (MRSA).
Observed class split: 33 BORSA + 59 MSSA + 9 MRSA = 101 (matches paper "101 high-quality
assemblies", "33 high-quality BORSA assemblies for public use").
What the paper reports (candidate claims)
| ID | Result | Reported value | Location |
|---|---|---|---|
| C1 | High-quality assemblies | 101 (1 MSSA dropped, low cov) | Results / Methods |
| C2 | Class composition | 33 BORSA, 9 MRSA, 59 MSSA | Results |
| C3 | Core genome (Roary) | 1,859 genes shared >99% isolates @ >95% id | Methods/Results |
| C4 | hierBAPS lineages | 5 lineages | "WGS reveals…" |
| C5 | MLST sequence types in BORSA | 12 STs; ST398/ST15/ST97/ST8 = 57.6% of BORSA | Results |
| C6 | ST counts | ST97 = 5, ST398 = 5; ST27 = 1, ST72 = 1 | Results |
| C7 | AMR / mecA | mecA in 9 MRSA only; mecC in 0 | Table 1 / Results |
| C8 | Clone definition / SNP distance | clones within 30 WG-SNPs / ≥99.999% ANI | Methods |
| C9 | SNP distances within cluster | isolates 316/318/334 "10 to 17 SNPs apart" | Fig 3B legend |
| C10 | Near-identical pair | isolates 343 (MSSA) / 344 (BORSA) "10 SNP" apart | Results |
| C11 | RFC performance | AUROC 0.902 ± 0.009 over 100 iter; 91.9% acc | "A sparse RFC…" |
| C12 | RFC 6 features | GdpP trunc, PBP2 trunc, β-lactamase-inhibitor effect, GdpP I52V, PBP2 A285P, PBP4 T189S | Fig 5B |
IN SCOPE (pipeline-derived, reproducible from the deposited 101 assemblies)
Third-party bioinformatics tools applied to the paper's own deposited assemblies (BRIEF P16: equally valid). All deterministic or near-deterministic given assemblies:
- C2 (class composition) — derive from deposited strain labels. Already done, control-plane. MATCH.
- C5/C6 (MLST) —
mlst(PubMLST S. aureus scheme) on each assembly → ST per isolate → ST distribution among the 33 BORSA. Deterministic. QUICK WIN. - C7 (AMR / mecA / mecC) —
AMRFinderPlus(paper: AMRFinder 3.8.4, >90% id) on each assembly → presence/absence of mecA, mecC, blaZ. Deterministic. QUICK WIN. - C1 (assembly QC) — re-run
CheckM/quaston the deposited assemblies to confirm99% completeness / <1% contamination and count = 101. (Cannot reproduce the assembly step — no reads — but can confirm QC of the deposited product.)
- C3 (Roary core genome 1,859) —
Prokkaannotate all 101 →Roary -i 95→ count core genes present in >99% isolates. Heavier; param-sensitive (expect within-tol). - C4 (hierBAPS 5 lineages) — from Roary core alignment →
rhierbaps. Stretch. - C9/C10 (SNP distances) —
snippy/snippy-core(contigs mode) orsnp-distson the core alignment for the named isolate sets. Stretch; reference-dependent.
PARTIAL / BLOCKED
- C11/C12 (Random Forest Classifier, AUROC 0.902) — the repo R code is present and
runnable, BUT its input feature matrices are NOT deposited (
RFC_metadata_noClass.csv,reduced_data_corr_withClass.csv,reduced_data_6features.csvall absent). 5 of the 6 final features are sequence-derivable (GdpP/PBP2/PBP4 mutations) but "β-lactamase inhibitor effect" (clav_effect) is a wet-lab phenotype (clavulanat
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.