Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Molecular Epidemiology of Invasive Group B Streptococcus in South Africa, 2019-2020.

J Infect Dis · 2025
L1 No data access 2/4
Why this verdict

The main result did not reproduce in this reproduction attempt. Where our recomputation produced values that differ from the published ones, those discrepancies are listed below. This is a single automated attempt — not peer review and not a finding of error or misconduct — and differences can also arise from data access, undocumented parameters or the computing environment. The verdict can be contested via “report an error”.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Main result did not reproduce
Decisive
From: Q5 · Derivability / plausibility 🔴
Main result did not reproduce
Decisive
From: Q7 · Core claim 🔴
Main result did not reproduce
Decisive
From: Q8 · Severity of the miss (overall human judgment) 🔴
✓ What held up
  • Nothing in this column.
What did not (or only partly)
  • 🔴Could not use the authors’ exact input data
  • 🔴Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🔴A deviation was attributed to the published material
  • 🔴Reported values were not (fully) derivable from the shared data
  • 🔴The deviation was non-trivial in magnitude
  • 🔴The central claim did not (fully) hold under reproduction
  • 🔴Overall, the reproduction showed a material discrepancy
No data access Data access not granted

This paper has a computational component, but its primary data is legally or ethically access-restricted — identifiable patient cohorts, rare-disease genomes, or controlled-access biobanks that cannot be openly shared. The reproduction therefore could not be attempted. That is a neutral verdict: it does not mean the result is wrong or that the authors fell short — only that, for legitimate privacy reasons, it cannot be independently checked from public data. We deliberately do NOT assign a 0–100 score here, because a low number would wrongly read as a failed reproduction.

Reproduction agent’s raw note

DROP (data_unavailable). Described well enough and CODE is fully available: github.com/shaze/GS (public, master 0a2b49fa5eee392d9fbc7c97885e284d8cf8a188, pushed 2025-04-02), a containerized Nextflow wrap of the CDC GBS typer (VelvetOptimiser assembly; GBS_Serotyper/MLST/GBS_Res_Typer/GBS_Surface_Typer/PBP->MIC), plus kSNP3 + RAxML for the SNP phylogeny. All seven in-scope results (serotype %, 5 CCs, AMR-gene %, penicillin MIC, surface proteins incl. hvgA-in-CC17, 435/658 QC, core-SNP tree) are pipeline-derived and clearly specified. BUT the data are not obtainable: the paper's Data Availability statement cites ENA PRJEB78571, and that accession resolves only to an empty project metadata record -- 0 records across read_run/read_experiment/analysis/assembly/wgs_set in ENA (filereport AND advanced search) and 0 in NCBI (SRA esearch=0, no SRA/biosample elinks), despite being public since 2024-08-01 (not indexing lag). The 658 isolates' WGS reads are the sole input to every result, so none can be reproduced 1:1; substituting unrelated GBS genomes would not reproduce this cohort's distributions, so per the brief no fabricated proxy was attempted. NOT ATTEMPTED: any «our HPC» compute (no input to run), the kSNP3/RAxML phylogeny, and PBP->MIC penicillin calling -- all blocked by the same missing data. INTEGRITY FLAG for human review: stated-vs-actual data-availability discrepancy (data declared public at ENA but not retrievable there); see AUDIT.md and data/data.json.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment
    assessed: 2026-06-15 ⛓ 00a973ec6102
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-15
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

To characterize invasive Group B Streptococcus (GBS) isolates from individuals of all ages in South Africa during 2019–2020, describing antimicrobial resistance, distribution of capsular serotypes and surface proteins, and genomic lineages, in order to inform antibiotic treatment and the expected coverage of polysaccharide and protein-based vaccines under development.

Core claims
  • β-lactam antibiotics remain appropriate for treatment of invasive GBS in South Africa, as only 1 isolate showed reduced penicillin susceptibility. finding
  • Polysaccharide and protein-based GBS vaccines under development are expected to provide good coverage in this setting, as all isolates fell within the 6 vaccine-targeted serotypes and carried targeted surface protein determinants. finding
  • Serotype III was the most common invasive GBS serotype (42.8%), followed by Ia (27.9%). finding
  • Moderate erythromycin (16.1%) and clindamycin (3.8%) resistance and high tetracycline resistance (91.5%) were observed among invasive GBS isolates. finding
  • All isolates carried at least 1 pilus gene cluster, 1 of 4 alpha/Rib family determinants, and 98% harbored a serine-rich repeat protein gene; hvgA was found exclusively in CC17 isolates. finding
  • Isolates belonged to 5 clonal complexes (CC1, CC8/10, CC17, CC19, CC23) characterized via whole-genome sequencing and MLST. finding
  • An adapted CDC GBS bioinformatics pipeline containerized in Nextflow was used for serotyping, MLST, resistance, and surface protein determination from WGS data. method
  • Phenotypic and in silico serotyping methods were highly concordant (99.7%). finding
Experimental setups
Assay System Perturbation Readout Platform
Phenotypic capsular serotyping (latex agglutination) 658 invasive GBS isolates, patients of all ages, South Africa none GBS serotype (Ia, Ib, II-IX) Immulex Group B Streptococcus antisera (SSI Diagnostica)
Antimicrobial susceptibility testing (disc diffusion and broth microdilution/MIC) 658 invasive GBS isolates antibiotic exposure (erythromycin, clindamycin, chloramphenicol, tetracycline, rifampicin, vancomycin, cotrimoxazole, penicillin) susceptible/intermediate/resistant and MIC Sensititre Streptococcus STP6F AST panels (ThermoFisher Scientific), CLSI guidelines
Whole-genome sequencing (WGS) and de novo assembly 661 viable invasive GBS isolates (658 with WGS results) none serotype, MLST/clonal complex, resistance genes, surface protein genes Illumina NextSeq 550, Nextera DNA Flex Library Prep, 2×150-bp paired-end
Multilocus sequence typing (MLST) 658 invasive GBS isolates none sequence types and clonal complexes (7 housekeeping genes) SRST2, PubMLST GBS database, goeBURST/PHYLOViZ v2.0
In silico antimicrobial resistance gene detection 658 invasive GBS isolates none presence of ermB, ermA/TR, ermC, mef, tet, gyrA/gyrB, parC/parE, pbp1a/pbp2x mutations CDC GBS bioinformatics pipeline (Nextflow)
In silico surface protein gene detection 658 invasive GBS isolates none presence/absence of alpha/Rib/Alp1/Alp2-3, hvgA, PI-1/PI-2a/PI-2b, srr1/srr2 CDC GBS bioinformatics pipeline
SNP-based maximum-likelihood phylogenetic analysis 435/658 high-quality GBS genomes none core SNP phylogeny kSNP v3, RAxML (GTR+gamma, 100 bootstraps), iTOL v6
MALDI-TOF mass spectrometry confirmation isolates not morphologically appearing as GBS none species confirmation as GBS MALDI-TOF MS
Key results
  • Serotype III was the most common serotype among invasive isolates 42.8% (281/656)
  • Serotype Ia was the second most common serotype 27.9% (183/656)
  • LOD cases were more often caused by serotype III than EOD cases 68.0% (124/183) LOD vs 38.6% (95/249) EOD
  • Only 1 isolate exhibited reduced penicillin susceptibility (MIC 0.25 µg/mL); 99.8% susceptible 657/658 (99.8%) susceptible
  • Phenotypic tetracycline resistance was very high 91.5% of isolates
  • Phenotypic erythromycin resistance observed at moderate level 16.1% of isolates
  • hvgA was found exclusively in CC17 isolates
  • tetM accounted for the majority of tetracycline resistance 95.8%
Key statistics
  • count 1748 invasive GBS cases reported (Total cases reported through GERMS-SA, 2019-2020)
  • count 658 isolates with both phenotypic and WGS results characterized (Isolates analyzed in study)
  • pvalue P < .001 (Serotype III more common in LOD (68.0%) vs EOD (38.6%))
  • pvalue P < .001 (Serotype Ia in EOD (30.9%) vs LOD (23.5%))
  • fold_change tetM 95.8% of tetracycline resistance (tetracycline resistance gene distribution)
  • count ermTR 34.9% and mefA/E 30.1% (most common genes among erythromycin-resistant isolates)
  • count ermB 32.0% (predominant gene in clindamycin-resistant isolates)
  • count serotype concordance 99.7% (647/649) (phenotypic vs in silico serotyping agreement)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This national laboratory-based surveillance study characterized 658 invasive GBS isolates from South Africa (2019–2020) using descriptive statistics (counts and percentages) for serotype distribution, antimicrobial susceptibility, surface protein genes, and clonal complexes. The primary inferential test was the chi-square (χ²) test to evaluate associations between serotype distribution and age group or sex, with a P < .05 significance threshold. Phylogenetic relationships were inferred by maximum-likelihood using RAxML with bootstrap support and Felsenstein ascertainment-bias correction. Results were reported as proportions with threshold-based P values.

Replicationbiological Sample sizeSample size determined by national surveillance catchment; no formal power calculation described. 1,748 total cases reported; 661 viable isolates characterized; 658 with WGS results analyzed. GroupsSix serotypes (Ia, Ib, II, III, IV, V) compared across six age groups (EOD, LOD, 3 mo–<5 y, 5–<18 y, 18–45 y, >45 y) and by sex; clonal complexes cross-tabulated with serotypes and surface protein profiles Pairingunpaired Randomization/blindingnot stated Dispersionnone Exact p-valuesno Effect sizesno Confidence intervalsno Multiplicity correctionnone stated
Statistical tests used
Test Applied to n Assumptions
Pearson chi-square (χ²) test Association between serotype distribution and age group (e.g., serotype III in LOD 68.0% vs EOD 38.6%, P < .001; serotype Ia in EOD 30.9% vs LOD 23.5%, P < .001) 656 serotyped isolates; specific pairwise comparisons used subsets (e.g., EOD n=249, LOD n=183) not stated
Pearson chi-square (χ²) test Association between serotype distribution and sex 656 serotyped isolates not stated
Maximum-likelihood phylogenetic inference (RAxML, GTR + gamma model, 100 bootstraps, Felsenstein ascertainment-bias correction) SNP-based phylogenetic tree of quality-filtered WGS assemblies 435 (filtered from 658 for ≤150 contigs and N50 ≥30,000) stated
goeBURST algorithm (minimum spanning tree, sharing ≥6/7 MLST loci) via PHYLOViZ 2.0 Clonal complex assignment of all sequenced isolates 658 stated
Approaches that could also have been used
  • Multiple pairwise serotype-by-age-group proportional comparisons were reported (e.g., serotype III in LOD vs EOD; serotype Ia in EOD vs LOD) without adjustment for multiplicity
    Could also: Apply Bonferroni correction or Benjamini-Hochberg FDR adjustment across the family of pairwise tests, or use a single omnibus χ² test with planned post-hoc contrasts — Multiplicity correction reduces the probability of spurious significant findings when many comparisons are made; FDR methods are less conservative than Bonferroni for larger comparison families common in surveillance studies
  • Key proportions (e.g., serotype prevalence, resistance rates) were reported as point estimates only, without confidence intervals
    Could also: Report 95% Wilson or Clopper-Pearson confidence intervals alongside each proportion — Confidence intervals convey both the precision of an estimate and its epidemiological magnitude, which is particularly informative for surveillance data where proportions are the primary quantities of interest and sample sizes vary across subgroups
  • Fisher's exact test was not mentioned; the χ² test was used for all categorical comparisons including strata with small cell counts (e.g., serotype IV: n=15)
    Could also: Fisher's exact test (or its r×c generalization) for comparisons where expected cell frequencies fall below 5 — The χ² large-sample approximation can be unreliable for sparse cells; Fisher's exact test does not rely on this approximation and is a standard alternative when expected counts are small
  • Concordance between phenotypic and genotypic serotyping/resistance classification was summarized as a percentage agreement (99.7% concordance)
    Could also: Cohen's kappa statistic with 95% CI to quantify and test agreement between the two classification methods — Kappa quantifies agreement beyond chance, providing a more nuanced measure of method concordance than simple percent agreement, and its CI expresses uncertainty in that estimate
  • Phylogenetic inference was performed on a filtered subset of 435/658 (66%) isolates meeting assembly-quality thresholds, with the remainder excluded
    Could also: A reference-based SNP caller (e.g., Snippy against a curated GBS reference) to enable inclusion of lower-quality assemblies, or a Bayesian approach (BEAST/MrBayes) for posterior branch-support estimates — Broader inclusion would reduce potential selection bias from assembly-quality filtering; Bayesian methods provide posterior probabilities rather than bootstrap percentages and can incorporate a molecular clock when sampling dates are available
  • Association between resistance genotype (gene presence) and resistance phenotype was assessed descriptively by concordance percentage, without formal regression modelling
    Could also: Logistic regression or a generalized linear model with serotype/clonal complex as covariates to evaluate independent predictors of phenotypic resistance — Regression modelling would allow simultaneous adjustment for serotype, clonal complex, and patient factors (age group, year), helping distinguish strain-level from host-level contributors to resistance patterns
Software: R 4.1.2 · kSNP 3 · RAxML (standard-RAxML) · PHYLOViZ 2.0 · iTol (Interactive Tree of Life) 6 · SRST2

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
5
Impact: low
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

scope.md — pmid-39737783

Paper: Molecular Epidemiology of Invasive Group B Streptococcus in South Africa, 2019–2020. Ntozini et al., J Infect Dis 2025. PMID 39737783 · PMCID PMC11998550 · DOI 10.1093/infdis/jiae633.

What the paper does (computational summary)

National laboratory surveillance collected invasive GBS isolates (2019–2020). 658 isolates were whole-genome sequenced (Illumina NextSeq 550, 2×150 bp). The authors ran the CDC GBS bioinformatics pipeline (BenJamesMetcalf) wrapped as a containerized Nextflow workflow (github.com/shaze/GS, strepB.nf) for QC

  • de novo Velvet assembly, then derived serotype, MLST/clonal complex, and antibiotic-resistance determinants (PBP mutations → penicillin MIC; macrolide/lincosamide erm/mef; tetracycline tet; surface proteins: pilus clusters, alpha/Rib family, Srr, hvgA). Core-SNP phylogeny on the 435/658 high-quality genomes (≤150 contigs, N50 ≥30 000) via kSNP v3 (k=19) and a RAxML GTR+Γ ML tree (100 bootstraps), visualized in iTOL.

In scope (pipeline-derived, would be attempted)

# Result Pipeline Reported
C1 Serotype distribution (n=658) shaze/GS GBS_Serotyper III 42.8%, Ia 27.9%, V 11.9%, II 8.4%, Ib 6.7%, IV 2.3%
C2 Clonal complexes (MLST) shaze/GS MLST + PubMLST CC 5 CCs: CC1, CC8/10, CC17, CC19, CC23
C3 AMR genes among resistant isolates GBS_Res_Typer ermTR 34.9%, mefA/E 29.2% (ery-R); ermB 32.0% (clinda-R); tetM 95.5% (tet-R)
C4 Penicillin: PBP-derived MIC PBP-Gene_Typer → Target2MIC 1 isolate reduced susceptibility (MIC 0.25 µg/mL)
C5 Surface proteins GBS_Surface_Typer all ≥1 of 3 pilus clusters; all ≥1 of 4 alpha/Rib; 98% Srr; hvgA only in CC17
C6 High-quality genome count Velvet QC 435/658 (66%) ≤150 contigs & N50 ≥30 000
C7 Core-SNP ML phylogeny kSNP3 + RAxML tree structure / CC clustering (Fig)

Out of scope (not a pipeline result)

  • Phenotypic serotyping & antimicrobial susceptibility testing (wet-lab, CLSI).
  • Epidemiological case counts (1748 cases; surveillance, not pipeline).
  • DNA extraction / library prep / sequencing (wet-lab).

Reproducibility surface

  • Code: AVAILABLE. github.com/shaze/GS public, not archived, master 0a2b49fa5eee392d9fbc7c97885e284d8cf8a188 (pushed 2025-04-02). Nextflow + Docker, runnable in principle. Upstream CDC tool: github.com/BenJamesMetcalf. Phylo: github.com/stamatak/standard-RAxML.
  • Expected results: IDENTIFIABLE. Specific %s in abstract/Results/Tables.
  • Data: NOT OBTAINABLE — blocks every in-scope result. See evidence below.

DROP decision: data_unavailable

The Data Availability statement says (verbatim): "Data supporting this study are available at the European Nucleotide Archive at https://www.ebi.ac.uk/ena/browser/view/PRJEB78571." The project record exists and is public (first-public 2024-08-01), but it contains no retrievable sequence data: every ENA result type returns 0 records and NCBI has no linked SRA. The 658 isolates' reads — the sole input to every in-scope result — are not downloadable, so none of C1–C7 can be reproduced. Substituting other GBS genomes would not reproduce this paper's South-African-isolate distributions, so per HARD RULE 6 the honest outcome is a drop (data_unavailable), not a fabricated proxy result. See data/data.json for the machine-checkable evidence and AUDIT.md for the data-availability-discrepancy (possible-fabrication) note.

Figures / tables: Figure
C1
Reported
Serotypes (n=658): III 42.8%, Ia 27.9%, V 11.9%, II 8.4%, Ib 6.7%, IV 2.3%
Reproduced
not reproduced — data unavailable
partial
C2
Reported
5 clonal complexes: CC1, CC8/10, CC17, CC19, CC23
Reproduced
not reproduced — data unavailable
partial
C3
Reported
AMR genes: ermTR 34.9%, mefA/E ~29-30%, ermB 32.0%, tetM ~95.5-95.8%
Reproduced
not reproduced — data unavailable
partial
C4
Reported
Penicillin reduced susceptibility: 1 isolate, MIC 0.25 ug/mL
Reproduced
not reproduced — data unavailable
partial
C5
Reported
Surface proteins: all >=1 pilus + >=1 alpha/Rib; Srr 98%; hvgA only in CC17
Reproduced
not reproduced — data unavailable
partial
C6
Reported
435/658 (66%) high-quality genomes (<=150 contigs, N50>=30000)
Reproduced
not reproduced — data unavailable
partial
C7
Reported
Core-SNP ML phylogeny: kSNP3 k=19 + RAxML GTR+gamma, 100 bootstraps
Reproduced
not reproduced — data unavailable
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 6/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🔴1. Data identity
🔴2. Endpoint comparability
🟡3. Location of the main deviation
🔴4. Cause of the deviation
🔴5. Derivability / plausibility
🔴6. Severity of the deviation
🔴7. Core claim
🔴8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Main result did not reproduce
Decisive
From: Q5 · Derivability / plausibility 🔴
Main result did not reproduce
Decisive
From: Q7 · Core claim 🔴
Main result did not reproduce
Decisive
From: Q8 · Severity of the miss (overall human judgment) 🔴

Drop — data_unavailable. The code (github.com/shaze/GS) is fully public and the seven in-scope results are clearly pipeline-derived, but the cited accession ENA PRJEB78571 resolves to an empty project record (0 records across ENA read_run/experiment/analysis/assembly and NCBI SRA, public since 2024-08-01), so none of C1–C7 could be reproduced 1:1. The blocker sits on the authors'/data-availability side: the deposit declared public is not retrievable, making every reported value non-derivable from the shared data. This is a genuine stated-vs-actual availability integrity flag for human confirmation — not evidence the numbers themselves are wrong, but they are unverifiable as deposited.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

120.8 k
tokens (I/O) · 6.2 M incl. cache
11 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.