Invasive bacterial disease trends and characterization of group B streptococcal isolates among young infants in southern Mozambique, 2001-2015.
The main results reproduced: recomputed values matched the published ones within tolerance.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Same input data as the authors
- ✓Reported values were directly comparable
- ✓No relevant deviation in data/preprocessing
- ✓No authors-side cause for any deviation
- ✓Reported values are derivable from the shared data
- ✓Any deviation was negligible
- ✓The central claim held under reproduction
- ✓Overall, the reproduction was clean
- Every checked point held up.
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Described well enough and 1:1. The paper's genomic typing of 35 GBS isolates (SRA PRJNA407943) was reproduced with a third-party in-silico toolchain on «our HPC» (shovill/skesa assembly -> mlst sagalactiae + GBS-SBG serotyper + abricate/ResFinder), independent of the authors' own CDC StrepLab GBS_Scripts_Reference pipeline (legacy SRST2 stack, deliberately not run; P16 third-party reproduction). All serotype, MLST sequence-type, clonal-complex and resistance-gene claims reproduce EXACTLY (14/15 claims exact): serotype III 33/35, V 1, Ia 1; ST17 24, ST109 7, ST866 1, ST1089 1, V=ST1, Ia=ST23, all III=CC17; tetM 35/35, mef 6, ermTR 1 (on the serotype-V isolate). The serotype x ST cross-tab matches Table 3 cell-for-cell. The only non-exact item is the CC17 surface-protein/pilus panel (hvgA/srr2/rib/PI-1/PI-2b): graded partial because the generic VFDB DB does not cleanly resolve those CC17-specific markers (it confirms shared GBS virulence genes); typing them would need the authors' GBS_Surface DB and was deliberately skipped as the optional last 20%. NOT attempted (out of scope, not pipeline-derived): invasive-disease incidence trends, penicillin/antibiotic MIC phenotypes, latex serotyping, clinical metadata. No fabrication concerns: every reproduced value is independently derivable from the public reads.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 97assessed: 2026-06-16 ⛓ f59fcb82ef35
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-16
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-16no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusWhat are the trends in invasive bacterial disease among infants <90 days in rural southern Mozambique during 2001–2015, with a focus on group B streptococcal (GBS) disease burden and the clinical and microbiological strain characteristics of circulating GBS isolates?
- ★ A notable young infant GBS disease burden persisted during 2001–2015 despite significant declines in overall IBD, neonatal mortality, and stillbirth rates. finding
- ★ By 2015, GBS had become the leading cause of young infant IBD at 2.7 per 1,000 live births. finding
- ★ Most GBS isolates were highly related serotype III strains belonging to ST17 or ST109, representing a well-established clone. finding
- ★ All ST109 isolates carried a PBP2x G398A substitution associated with elevated penicillin MIC, a first-step mutation toward reduced penicillin susceptibility within a well-known virulent lineage. mechanism
- GBS isolates were characterized by serotyping (multiplex PCR), antimicrobial susceptibility testing, and whole genome sequencing including MLST and SNP analysis. method
- Findings underscore the need for non-antibiotic GBS prevention strategies such as maternal vaccination. finding
- Whole genome sequences of the GBS isolates are publicly available (NCBI SRA BioProject PRJNA407943) as a resource. resource
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Demographic surveillance / vital statistics analysis | Manhiça district population, rural southern Mozambique (DSS catchment) | none | annual live births, stillbirths, neonatal mortality rate, admission rate, causes of death | Manhiça DSS; verbal autopsy (WHO model); FoxPro v2.6 |
| Invasive bacterial disease surveillance with blood/CSF bacterial culture | Infants <90 days admitted to Manhiça District Hospital | none | microbiologically-confirmed IBD (positive blood or CSF culture), pathogen identity | Pedibact pediatric blood culture bottle, BACTEC 9050 (Becton-Dickinson) |
| GBS phenotypic identification | Recovered GBS bacterial isolates | none | beta-hemolysis, catalase, bacitracin resistance, Lancefield group B antigen | Bio RAD PASTOREX STREP latex agglutination |
| Serotyping by multiplex PCR | Stored GBS isolates (CDC Streptococcus Lab) | none | capsular serotype | — |
| Antimicrobial susceptibility testing (broth microdilution) | Stored GBS isolates | antibiotics (e.g., penicillin) | minimum inhibitory concentration (MIC) using CLSI breakpoints | — |
| Whole genome sequencing with MLST and SNP analysis | GBS isolates (CSF isolate preferred when both available) | none | serotype deduction, antimicrobial resistance determinants (PBP2x substitution), sequence types, surface protein/virulence factor presence, core genome SNPs | Cutadapt v1.8.1, VelvetOptimiser v2.2.5/VelvetK, kSNP3.0; GBS_Scripts_Reference pipeline |
- – 437 IBD cases identified, including 57 GBS cases 437 IBD; 57 GBS
- – GBS was the leading cause of young infant IBD in 2015 2.7 per 1,000 live births
- – Significant declines in overall IBD, neonatal mortality, and stillbirth rates, but no significant decline for GBS P<0.0001 (overall); GBS P=0.17
- – Among 35 GBS isolates tested, 31 were highly related serotype III isolates within ST17 or ST109 31/35 (88.6%); ST17 68.6%, ST109 20.0%
- ▲ All seven ST109 isolates had elevated penicillin MIC associated with PBP2x substitution G398A 7 isolates (21.9%); MIC ≥0.12 μg/mL
- ▼ Stillbirth rate declined over study period 38.6 to 5.6 per 1,000 births
- ▼ Neonatal mortality rate declined 35.8 (2001) to 16.7 (2013) per 1,000 live births
- ▲ Infants with IBD had higher in-hospital mortality than those without IBD 11.8% (51/437) vs 5.5% (216/3956)
- count 437 IBD cases including 57 GBS cases (Total IBD and GBS cases identified 2001–2015)
- fold_change 2.7 per 1,000 live births (GBS incidence in 2015, leading cause of young infant IBD)
- pvalue P<0.0001 (Declines in overall IBD, neonatal mortality, stillbirth rates)
- pvalue P=0.17 (No significant decline in GBS rate trend)
- count 31/35 (88.6%) serotype III; ST17 68.6%, ST109 20.0% (GBS isolates available for testing)
- count 7 (21.9%) ST109 isolates with elevated penicillin MIC (≥0.12 μg/mL) (PBP2x G398A substitution)
- count 11.8% (51/437) IBD vs 5.5% (216/3956) non-IBD died; P<0.0001 (In-hospital mortality comparison)
- count 47,651 live births and 993 stillbirths (Reported within DSS catchment 2001–2015)
Statistical methods review
Model: opusA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This is a descriptive epidemiologic study using long-term demographic and invasive bacterial disease (IBD) surveillance data (2001–2015) from a rural Mozambican district, supplemented by molecular characterization of GBS isolates. Annual incidence, mortality, stillbirth, and admission rates were computed using live births (or total births) as denominators; temporal trends in rates were assessed with Poisson regression, trends in proportions with the Cochran-Armitage test, and group proportions compared with chi-square or Fisher's exact tests. Results were reported largely as counts, percentages, rates per 1,000 live births, medians with IQR, and P-values, with isolates further characterized by serotyping, MIC testing, MLST, and whole genome SNP analysis.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Poisson regression (trend in rates) | Trends of IBD incidence, neonatal mortality, stillbirth, and admission rates over 2001–2015 (Fig 2); reported P<0.0001 for several and P=0.17 for GBS | 47,651 live births and 993 stillbirths over the period; 437 IBD cases including 57 GBS | not stated |
| Cochran-Armitage trend test (trend in proportions) | Trend in proportion of young infant deaths occurring at a health facility (P=0.66 for trend) | — | not stated |
| Chi-square or Fisher's exact test | Comparison of proportions, e.g., culture collection by age (78.5% vs 90.7%, P<0.0001) and in-hospital mortality in IBD vs non-IBD (11.8% vs 5.5%, P<0.0001) | e.g., 437 IBD vs 3,956 non-IBD admissions | not stated |
-
Temporal trends in rates were assessed with Poisson regression.↳ Could also: A negative binomial regression model could also be used, and incidence rate ratios with 95% confidence intervals could accompany the trend P-values. — Negative binomial models accommodate overdispersion common in count data, and reporting rate ratios with confidence intervals would convey the magnitude and precision of the trend in addition to its statistical significance.
-
Annual rates and trends were summarized primarily with P-values and point estimates.↳ Could also: Reporting 95% confidence intervals around each annual rate and around the trend estimate would also be standard. — Confidence intervals communicate the uncertainty around each estimate, which is especially informative given the smaller counts in early years and for the GBS subgroup.
-
P-values were largely reported as thresholds (e.g., P<0.0001).↳ Could also: Exact P-values could also be reported. — Exact values let readers gauge how far results sit from conventional cutoffs and support any later meta-analytic use.
-
Several proportion comparisons were conducted without a stated multiplicity adjustment.↳ Could also: A family-wise or false-discovery-rate correction (e.g., Bonferroni or Benjamini-Hochberg) could also be applied when many comparisons are made. — Such corrections control the chance of false-positive findings across a family of tests, which can be helpful when many rates and proportions are examined together.
-
Trends were modeled with calendar year as the predictor across an expanding catchment area.↳ Could also: A model offsetting for person-time or live births and including terms for the catchment expansion could also be used. — Explicitly accounting for the phased population expansion would help distinguish underlying epidemiologic trends from changes driven by surveillance coverage.
-
Group differences (e.g., EOD vs LOD characteristics) were described with proportions and chi-square/Fisher tests.↳ Could also: Effect measures such as risk ratios or odds ratios with confidence intervals could also be presented alongside the tests. — Effect sizes with intervals quantify the strength of associations rather than only indicating whether a difference reached significance.
Result convergence & founder nodes
Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.
-
Infants with invasive bacterial disease had higher in-hospital mortality (11.8%) compared to those without IBD (5.5%)other human infant manhica mozambique up 2018×1papers★ This paper is the founder (earliest)
-
GBS was the leading cause of invasive bacterial disease among young infants in Manhiça, Mozambique in 2015 with an incidence of 2.7 per 1,000 live birthsother human infant manhica mozambique 2018×1papers★ This paper is the founder (earliest)
-
GBS invasive bacterial disease incidence showed no significant decline over 2001-2015 (P=0.17) despite significant declines in overall IBD incidence (P<0.0001)other human infant manhica mozambique none 2018×1papers★ This paper is the founder (earliest)
-
Neonatal mortality rate declined from 35.8 to 16.7 per 1,000 live births between 2001 and 2013 in Manhiça districtother manhica district mozambique population down 2018×1papers★ This paper is the founder (earliest)
-
Stillbirth rate declined significantly from 38.6 to 5.6 per 1,000 births over 2001-2015 in Manhiça districtother manhica district mozambique population down 2018×1papers★ This paper is the founder (earliest)
-
All ST109 GBS isolates (21.9% of collection) carried PBP2x G398A substitution and had elevated penicillin MIC (>=0.12 ug/mL)other streptococcus-agalactiae isolate manhica mozambique up 2018×1papers★ This paper is the founder (earliest)
-
88.6% of GBS isolates were serotype III, predominantly sequence type ST17 (68.6%) and ST109 (20.0%), indicating clonal dominanceWGS streptococcus-agalactiae isolate manhica mozambique 2018×1papers★ This paper is the founder (earliest)
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-29351318
Paper: Sigaúque et al. 2018, PLoS One 13(1):e0191193. "Invasive bacterial disease trends and characterization of group B streptococcal isolates among young infants in southern Mozambique, 2001–2015."
Code artifact: https://github.com/BenJamesMetcalf/GBS_Scripts_Reference — CDC
StrepLab GBS WGS-typing pipeline (SRST2-style: bowtie2-indexed gene DBs for
serotype/resistance/surface genes + an MLST allele profile set; Perl/Shell
wrappers StrepLab-JanOw_GBS-Typer.sh). This is a third-party tool (P16): applying
it to the paper's own data is a valid reproduction.
Data: SRA BioProject PRJNA407943 — exactly 35 paired-end Illumina runs (SRR6050832–SRR6050866), one per GBS isolate. Matches the paper's "35 GBS isolates available for characterization." Public, downloadable from ENA FTP.
In scope (pipeline-derived genomic typing of the 35 isolates → Table 3)
These are produced by a bioinformatic pipeline from the WGS reads and are what we reproduce:
| # | Result | Reported (Table 3 / Results) |
|---|---|---|
| C1 | Serotype distribution | III 33/35 (94.3%); V 1 (2.9%); Ia 1 (2.9%) |
| C2 | MLST sequence types | ST17 24 (68.6%), ST109 7 (20.0%), ST866 1 (2.9%), ST1089 1 (2.9%); plus ST1 (the serotype-V isolate) and ST23 (the serotype-Ia isolate) |
| C3 | Clonal complex | all 33 serotype-III isolates = CC17 |
| C4 | Tetracycline resistance gene | all 35 carry tetM |
| C5 | Macrolide/lincosamide genes | mef-positive in 6 isolates; ermTR-positive in 1 (serotype V) |
| C6 | Surface protein / pilus genes (CC17) | serotype-III isolates uniformly hvgA, srr2, rib, PI-1, PI-2b |
Reproduction strategy
Heavy compute on «our HPC» (SLURM, «infra»). Faithful-but-feasible third-party toolchain rather than fighting the original pipeline's pinned legacy SRST2 / samtools-0.1.18 / bowtie2-2.1 stack (80/20 rule):
- Download 35 ENA runs → assemble each (shovill/skesa).
- C2/C3 ST + CC:
mlst(Seemann)sagalactiaescheme → ST per isolate. - C1 serotype: GBS-SBG (swainechen) in-silico capsular serotyper on assemblies.
- C4/C5 resistance genes:
abricateResFinder DB → tet(M), erm, mef per isolate. - C6 surface/pilus genes: best-effort (abricate vs VFDB / authors' GBS_Surface DB). Treated as the optional "last 20%".
Out of scope (NOT pipeline-derived → not attempted)
- Invasive-disease incidence trends/rates 2001–2015 (epidemiologic surveillance).
- Penicillin/antibiotic MIC phenotypes (wet-lab broth microdilution).
- Latex-agglutination serotyping, case ascertainment, clinical metadata.
Drop-risk notes
Data + code both resolve and are public → eligible. Main risk is env/run-time (legacy pipeline deps) — mitigated by the modern third-party toolchain above.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
This is a clean 1:1 reproduction: the paper's genomic typing of 35 GBS isolates (SRA PRJNA407943) reproduced exactly on 14/15 claims using an independent toolchain — serotype III 33/35 (94.3%), ST17 24, ST109 7, all serotype-III = CC17, tetM 35/35, mef 6, ermTR 1 — matching Table 3 cell-for-cell. The only non-exact item (C6, the CC17 surface/pilus markers) is on our side: a deliberate optional skip because the generic VFDB DB cannot resolve those CC17-specific markers without the authors' GBS_Surface DB. No fabrication concern; every value is independently derivable from the public reads, and the central characterization conclusion holds fully.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.