Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Comparative Genomics of Borderline Oxacillin-Resistant Staphylococcus aureus Detected during a Pseudo-outbreak of Methicillin-Resistant S. aureus in a Neonatal

mBio · 2022
L1 63/100 3/4
Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
✓ What held up
  • Any deviation was negligible
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🟡Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
63/100
Reproducibility score
0.6 SD below mean
vs. all fields · 1173 studies
🎯 Scores higher than 24% of all assessed papers rank 875 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

PARTIAL - strong core reproduction. Applied third-party tools (mlst, AMRFinderPlus, Prokka+Roary, snp-dists, rhierbaps) to the paper's own 101 deposited S. aureus assemblies (PRJNA695316; no raw reads). EXACT: C2 class split (33 BORSA/59 MSSA/9 MRSA) and C7 mecA-in-9-MRSA-only/mecC-absent. WITHIN-TOL: C3 Roary core genome 1865 vs 1859 (0.3%) and C5 dominant-ST fraction 57.6% (exact percentage; 11 vs 12 distinct STs). PARTIAL: C1 (n=101 exact, CheckM not run), C6 (ST97=5/ST398=5 exact but ST27/ST72 BORSA singletons not recovered - likely PubMLST DB version drift, 2 BORSA untyped), C9/C10 SNP distances (clonal clusters + near-identical MSSA/BORSA pair confirmed but absolute counts lower because we used Roary core-GENE SNPs vs the paper's Snippy whole-genome SNPs). MISMATCH: C4 hierBAPS 8 vs 5 lineages (param/alignment-sensitive). NOT REPRODUCIBLE/out-of-scope: the headline Random Forest AUROC 0.902 (C11/C12) - feature matrices not deposited and one feature is a wet-lab phenotype; all wet-lab phenotyping (susceptibility, PBP2a, beta-lactamase assays, chromogenic agar Tables 1-3); the assembly step itself (no reads). Conclusion: the paper's comparative-genomics backbone is well-described and reproduces 1:1 from the deposited data; the ML headline cannot be independently verified from what was deposited.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 63
    assessed: 2026-06-21 ⛓ fcd4d5c61c57
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-21
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

Following atypical MRSA screening culture results in a NICU, the study asks whether isolates flagged as MRSA are true mecA-mediated MRSA or a distinct borderline oxacillin-resistant S. aureus (BORSA) phenotype, and whether genomic features can distinguish BORSA from MSSA and explain inconsistent detection by commercial MRSA screening agars.

Core claims
  • Of 42 isolates flagged as MRSA by screening agar, only 9 were PBP2a- and mecA-positive true MRSA, while the remaining 33 were mecA-negative and largely met criteria for BORSA finding
  • BORSA isolates identified in the NICU were phylogenetically diverse and did not represent clonal expansion or shared gene content, though two NICU strains showed infrequent clonal recurrence over 8 months finding
  • Six genomic features (substitutions/truncations in PBP2, PBP4, and GdpP, plus beta-lactamase hyperproduction) were used to build a random forest classifier distinguishing BORSA from MSSA method
  • The random forest classifier robustly predicted the BORSA phenotype in an external validation cohort spanning two continents (AUC = 0.902) finding
  • Commercial MRSA screening agars vary substantially in sensitivity and specificity for detecting MRSA versus BORSA, leading to misclassification of BORSA as MRSA finding
  • Beta-lactamase hyperproduction alone is insufficient to explain borderline oxacillin resistance in many isolates, implicating additional PBP and GdpP-related mechanisms mechanism
  • Beta-lactamase inhibitor potentiation effect (amoxicillin vs. amoxicillin-clavulanate MIC shift) was significantly greater in BORSA than MSSA isolates finding
Experimental setups
Assay System Perturbation Readout Platform
MALDI-TOF mass spectrometry clinical S. aureus isolates (NICU screening + comparator blood isolates) none species confirmation MALDI-TOF MS
PBP2a lateral flow immunoassay and mecA/mecC PCR clinical S. aureus isolates none PBP2a protein presence and mecA/mecC gene detection
Cefoxitin and oxacillin susceptibility testing (disk diffusion and gradient diffusion on 2% NaCl agar) clinical S. aureus isolates none zone size / MIC for resistance classification Mueller-Hinton agar, CLSI methods
Growth assay on commercial MRSA screening agars (Spectra, MRSA Select II, BBL CHROMagar MRSA II, chromID MRSA, nonchromogenic MRSA screen, HardyCHROM) clinical S. aureus isolates none colony growth abundance and pigmentation 6 commercial chromogenic/nonchromogenic agars
Beta-lactamase detection (disk diffusion penicillin zone edge test, nitrocefin/Cefinase test) clinical S. aureus isolates none beta-lactamase production status Cefinase test
Beta-lactamase inhibitor potentiation testing (amoxicillin vs. amoxicillin-clavulanate gradient diffusion) clinical S. aureus isolates beta-lactamase inhibitor (clavulanate) addition fold-change in MIC gradient diffusion strips
Whole-genome sequencing / comparative genomics 42 NICU isolates + 60 comparator blood isolates none phylogenetic relatedness, gene content, mutations in pbp2/pbp4/gdpP WGS
Random forest classification genomic feature set from sequenced isolates (training + 2-continent validation cohort) none predicted BORSA vs. MSSA phenotype
Key results
  • Only 9 of 42 NICU screening isolates were PBP2a/mecA-positive true MRSA 9/42
  • 33 of 42 isolates were oxacillin resistant by 2% NaCl gradient diffusion 33/42
  • 24 of 33 non-MRSA investigated isolates and 9 of 60 comparator isolates met BORSA criteria 24/33; 9/60
  • Random forest classifier validated across a two-continent cohort AUC = 0.902
  • MRSA-selective agars (BBL CHROMagar MRSA II, MRSA Select II, chromID MRSA) showed high specificity but lower sensitivity for BORSA detection 89% sensitivity, 100% specificity for MRSA
  • Nonchromogenic and Spectra MRSA agars showed high sensitivity but reduced specificity 100% sensitivity; 86% and 60% specificity, respectively
  • Beta-lactamase inhibitor effect (MIC shift) differed significantly between BORSA and MSSA isolates P = 0.0030
  • BORSA isolates were genomically diverse rather than clonal, with rare exceptions
Key statistics
  • pvalue P = 0.0030 (Mann-Whitney test comparing beta-lactamase inhibitor MIC-shift effect between BORSA and MSSA isolates)
  • other AUC = 0.902 (Validation performance of random forest classifier predicting BORSA phenotype across a two-continent cohort)
  • count 9/42 (NICU isolates confirmed as true MRSA (PBP2a+/mecA+))
  • count 33/42 (NICU isolates oxacillin resistant by 2% NaCl gradient diffusion)
  • count 24/33 investigated isolates; 9/60 comparator isolates (Isolates meeting BORSA case definition)
  • other 89% sensitivity / 100% specificity (BBL CHROMagar MRSA II, MRSA Select II, chromID MRSA performance for MRSA detection across all 102 isolates)
  • other 100% sensitivity / 86% and 60% specificity (Nonchromogenic MRSA screen agar and Spectra MRSA agar performance, respectively)
  • count 6/33 BORSA vs. 6/60 MSSA (Isolates exhibiting a 4-fold difference in lactamase inhibitor effect)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This retrospective observational study phenotypically and genomically characterized 42 suspected MRSA NICU isolates alongside 60 comparator blood isolates (102 total). The primary analytical contribution was a random forest classifier trained on 6 genomic features (substitutions/truncations in PBP2, PBP4, GdpP, and beta-lactamase hyperproduction) to distinguish BORSA from MSSA, evaluated by AUC on a validation set. A Mann-Whitney U test was used to compare the beta-lactamase inhibitor effect between BORSA and MSSA groups, and diagnostic performance (sensitivity and specificity) was reported descriptively for each of six MRSA screening agars.

Replicationunclear Sample size42 NICU surveillance isolates and 60 comparator blood isolates; no formal power calculation described GroupsMRSA vs. BORSA vs. MSSA; BORSA vs. MSSA for beta-lactamase inhibitor effect Pairingunpaired Randomization/blindingnot stated DispersionSD Exact p-valuesyes Effect sizesno Confidence intervalsno Multiplicity correctionnone stated
Statistical tests used
Test Applied to n Assumptions
Mann-Whitney U test (two-group comparison) Comparison of log2 fold change in MICs (amoxicillin vs. amoxicillin-clavulanic acid) between BORSA and MSSA isolates; Fig. S1 33 BORSA vs. 60 MSSA not stated
Random forest classifier (machine learning model, not a classical hypothesis test); performance reported as AUC Genomic feature-based classification of BORSA vs. MSSA across cohorts spanning two continents; validation AUC = 0.902 not stated
Sensitivity and specificity (diagnostic accuracy calculation) Analytical performance of six MRSA screening agars for MRSA and BORSA detection; Table 3 102 isolates total na
Approaches that could also have been used
  • Sensitivity and specificity for each screening agar were reported as single point estimates without confidence intervals, with small true-positive denominators (e.g., 8 or 9 confirmed MRSA isolates)
    Could also: Wilson score or Clopper-Pearson exact binomial 95% confidence intervals around each sensitivity and specificity estimate could also be reported — With very small numerators, point estimates alone carry substantial uncertainty; CIs would quantify that uncertainty and allow readers to assess whether apparent differences in agar performance are distinguishable from chance
  • Six MRSA screening agars were each evaluated on the same 102 isolates, producing separate sensitivity/specificity point estimates that are compared informally
    Could also: McNemar's test (or its extension for multiple paired proportions) could also formally compare sensitivities or specificities across agars, since the same isolates were tested on all agars — The paired structure of the data (same isolate evaluated on every agar) is a natural match for paired-proportion tests, which would provide p-values and effect estimates for pairwise agar differences rather than relying on visually non-overlapping percentages
  • The Mann-Whitney U test was used to compare the beta-lactamase inhibitor effect (log2 MIC fold change) between BORSA and MSSA
    Could also: A two-sample t-test or Welch's t-test on the log2-transformed fold changes could also be applied if approximate normality holds, together with a reported mean difference and 95% CI — Log-transformed ratio data often approaches normality; a parametric test would additionally yield a mean difference and confidence interval that directly quantify the magnitude of the effect, complementing the p-value
  • Dispersion around the beta-lactamase inhibitor effect in Fig. S1 was shown as standard deviation (SD)
    Could also: Median with interquartile range (IQR), or 95% confidence intervals around group means, could also convey spread — MIC and fold-change distributions are often right-skewed; IQR or bootstrapped CIs may more faithfully represent central tendency and spread for such data, and the IQR is a natural companion to the non-parametric Mann-Whitney test used
  • A random forest classifier was used to predict the BORSA phenotype; performance was summarized with AUC on a validation set
    Could also: Penalized logistic regression (e.g., LASSO or elastic net) could also be applied as a companion or alternative classifier on the same 6-feature set — With a small feature set and modest sample size, penalized logistic regression yields odds ratios and a transparent probability model directly interpretable as effect sizes; it would complement the random forest by providing feature-level inferential estimates
  • Phylogenetic diversity among BORSA isolates was described qualitatively as 'phylogenetically diverse and not representative of clonal expansion'
    Could also: A quantitative measure of population structure such as the index of association (IA) or a permutation-based clustering test could also be reported alongside the phylogeny — Formal statistical tests of clonal structure provide an objective, reproducible basis for statements about phylogenetic diversity, allowing readers to assess whether the observed spread differs significantly from a random distribution of genotypes
Software: Not stated in provided text

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

scope.md — pmid-35038924

Paper: Sawhney, Ransom, Wallace, Reich, Dantas, Burnham. Comparative Genomics of Borderline Oxacillin-Resistant Staphylococcus aureus Detected during a Pseudo-outbreak of MRSA in a Neonatal Intensive Care Unit. mBio 2022. PMID 35038924 · PMC8764539 · DOI 10.1128/mbio.03196-21.

Code: https://github.com/sanjsawhney/BORSA_RFC (commit 8ee78886, pushed 2021-12-02, MIT, default branch main). Repo contains ONLY LICENSE + README.md; all R code is embedded inline in the README. No input data CSVs and no WGS-pipeline scripts are shipped.

Data: BioProject PRJNA695316. Deposited as 101 genome assemblies (GenBank GCA accessions, scaffold level) — NOT raw reads (NCBI SRA count = 0, ENA read_run = 0). Strain names encode phenotype: WUSa_oxa_BOR### (BORSA), _MS### (MSSA), _MR### (MRSA). Observed class split: 33 BORSA + 59 MSSA + 9 MRSA = 101 (matches paper "101 high-quality assemblies", "33 high-quality BORSA assemblies for public use").


What the paper reports (candidate claims)

ID Result Reported value Location
C1 High-quality assemblies 101 (1 MSSA dropped, low cov) Results / Methods
C2 Class composition 33 BORSA, 9 MRSA, 59 MSSA Results
C3 Core genome (Roary) 1,859 genes shared >99% isolates @ >95% id Methods/Results
C4 hierBAPS lineages 5 lineages "WGS reveals…"
C5 MLST sequence types in BORSA 12 STs; ST398/ST15/ST97/ST8 = 57.6% of BORSA Results
C6 ST counts ST97 = 5, ST398 = 5; ST27 = 1, ST72 = 1 Results
C7 AMR / mecA mecA in 9 MRSA only; mecC in 0 Table 1 / Results
C8 Clone definition / SNP distance clones within 30 WG-SNPs / ≥99.999% ANI Methods
C9 SNP distances within cluster isolates 316/318/334 "10 to 17 SNPs apart" Fig 3B legend
C10 Near-identical pair isolates 343 (MSSA) / 344 (BORSA) "10 SNP" apart Results
C11 RFC performance AUROC 0.902 ± 0.009 over 100 iter; 91.9% acc "A sparse RFC…"
C12 RFC 6 features GdpP trunc, PBP2 trunc, β-lactamase-inhibitor effect, GdpP I52V, PBP2 A285P, PBP4 T189S Fig 5B

IN SCOPE (pipeline-derived, reproducible from the deposited 101 assemblies)

Third-party bioinformatics tools applied to the paper's own deposited assemblies (BRIEF P16: equally valid). All deterministic or near-deterministic given assemblies:

  • C2 (class composition) — derive from deposited strain labels. Already done, control-plane. MATCH.
  • C5/C6 (MLST)mlst (PubMLST S. aureus scheme) on each assembly → ST per isolate → ST distribution among the 33 BORSA. Deterministic. QUICK WIN.
  • C7 (AMR / mecA / mecC)AMRFinderPlus (paper: AMRFinder 3.8.4, >90% id) on each assembly → presence/absence of mecA, mecC, blaZ. Deterministic. QUICK WIN.
  • C1 (assembly QC) — re-run CheckM/quast on the deposited assemblies to confirm

    99% completeness / <1% contamination and count = 101. (Cannot reproduce the assembly step — no reads — but can confirm QC of the deposited product.)

  • C3 (Roary core genome 1,859)Prokka annotate all 101 → Roary -i 95 → count core genes present in >99% isolates. Heavier; param-sensitive (expect within-tol).
  • C4 (hierBAPS 5 lineages) — from Roary core alignment → rhierbaps. Stretch.
  • C9/C10 (SNP distances)snippy/snippy-core (contigs mode) or snp-dists on the core alignment for the named isolate sets. Stretch; reference-dependent.

PARTIAL / BLOCKED

  • C11/C12 (Random Forest Classifier, AUROC 0.902) — the repo R code is present and runnable, BUT its input feature matrices are NOT deposited (RFC_metadata_noClass.csv, reduced_data_corr_withClass.csv, reduced_data_6features.csv all absent). 5 of the 6 final features are sequence-derivable (GdpP/PBP2/PBP4 mutations) but "β-lactamase inhibitor effect" (clav_effect) is a wet-lab phenotype (clavulanat
Figures / tables: TableFig 3B
C2
Reported
33 BORSA, 9 MRSA, 59 MSSA (101 total)
Reproduced
33 BORSA, 9 MRSA, 59 MSSA (101)
exact
C7
Reported
mecA in 9 MRSA only; mecC in 0
Reproduced
mecA in 9 MRSA only; mecC in 0; absent from all BORSA+MSSA
exact
C3
Reported
1859 core genes (Roary, >99% isolates, >95% id)
Reproduced
1865 core genes (Prokka 1.14.6 + Roary -i 95)
within tolerance
C5
Reported
12 MLST STs in BORSA; ST398/15/97/8 = 57.6%
Reproduced
57.6% (19/33) EXACT; 11 named + 2 untyped STs
within tolerance
C1
Reported
101 high-quality assemblies
Reproduced
101 assemblies, 2.65-2.90 Mb (median 2.74)
partial
C6
Reported
ST97=5, ST398=5, ST27=1, ST72=1 (BORSA)
Reproduced
ST97=5, ST398=5, ST27=0, ST72=0
partial
C9
Reported
316/318/334 10-17 SNPs apart
Reproduced
0-5 core-gene SNPs (clonal cluster confirmed)
partial
C10
Reported
343(MSSA)/344(BORSA) 10 SNP apart
Reproduced
3 core-gene SNPs (near-identical pair confirmed)
partial
C4
Reported
5 hierBAPS lineages, >half in lineage 1
Reproduced
8 level-1 lineages (largest 28/101)
did not match
C11
Reported
RFC AUROC 0.902 +/- 0.009; 91.9% acc
Reproduced
NOT ATTEMPTED - input feature matrices not deposited; one feature is wet-lab
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 63/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟡2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟡5. Derivability / plausibility
🟢6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

245.1 k
tokens (I/O) · 16.2 M incl. cache
152 min
runtime · 14.36 CPU-h
4.4 GB
peak RAM
2
HPC jobs
hummel
machine