Whole genome sequencing reveals possible host species adaptation of Streptococcus dysgalactiae.
Part of the results reproduced; minor but material deviations remained.
- ✓Reported values were directly comparable
- 🟡Could not use the authors’ exact input data
- 🟡A deviation arose in the data or preprocessing
- 🟡A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The deviation was non-trivial in magnitude
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
▸Reproduction agent’s raw note
IN PROGRESS. Reproducing shovill de-novo assembly + Prokka + MLST + fastANI on ENA PRJEB42928 (124 paired-end SDSD WGS runs). Targets: avg genome size 2.04 Mb, ~1990 CDS, 14 MLST STs, mean ANI 99.0%. Env building on «our HPC»; downloads starting.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessmentassessed: 2026-06-18 ⛓ 485ad2f3fcfb
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-18
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19no human curator yet
- Last updated
- 2026-07-31
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusCan whole genome sequencing of Streptococcus dysgalactiae (SD) from multiple host species reveal virulence factors, phylogenetic relationships, and genetic factors contributing to host adaptation, particularly for the understudied subspecies dysgalactiae (SDSD) from sheep and cattle?
- ★ SDSD constitutes a distinct taxonomic entity within S. dysgalactiae, with a mean intra-subspecies average nucleotide identity of 99%. finding
- ★ Phylogenetic clustering of SD isolates extends beyond the two-subspecies division and reflects adaptive evolution into multiple host-associated lineages. finding
- ★ No single gene or genetic region was uniquely associated with host species among bovine vs ovine SDSD, despite separate pangenome clustering. finding
- ★ Lineage-specific virulence factors exist and several reside near hotspots for integration of mobile genetic elements, implicating horizontal genetic transfer in niche specialization. mechanism
- ★ A previously unidentified emm-like (M-protein) gene and a novel Srr-like glycoprotein operon were detected in SDSD. finding
- ★ The glycosylation gene gtfC exists in two allelic variants whose distribution correlates with bovine vs ovine host of origin. finding
- Intact bacteriophages are more prevalent in SDSD whereas Integrative Conjugative Elements (ICEs) are more prevalent in human-associated SDSE. finding
- This is the first comprehensive genomic characterization of SDSD and the first to include ovine isolates, providing a sequenced genome resource. resource
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Whole genome sequencing | S. dysgalactiae isolates from cows and sheep (SDSD) | none | genome sequence, genome size, CDS count | — |
| Whole genome sequencing | SDSE isolates from human, dog, horse, pig | none | genome sequence | — |
| Multilocus sequence typing (MLST) | SDSD isolates from cows and sheep | none | sequence types / MLST profiles | — |
| Pangenome analysis | 156 SD isolates (78 SDSD + 78 SDSE) | none | gene clusters, core/pangenome size, single orthologue genes | — |
| Phylogenetic analysis | 78 SDSD and 78 SDSE genomes | none | phylogenetic tree from 752 single orthologue gene clusters, bootstrap support | — |
| Average nucleotide identity (fastANI) | 156 SD whole genome sequences | none | pairwise ANI values / heatmap | fastANI algorithm |
| Virulence gene profiling / in silico screening | SDSD and SDSE genomes by host species | none | presence/percentage of virulence factors | — |
| Mobile genetic element screening | SDSD and human-associated SDSE genomes | none | bacteriophage, ICE, and associated virulence/resistance gene counts | — |
- – Mean intra-subspecies ANI for SDSD was 99% with all pairwise comparisons >98% 99% (>98%)
- – gtfC allele A in 50/55 bovine isolates and allele B in 20/22 ovine isolates, the two variants 88% similar 88% similarity
- ▼ Intact bacteriophages in 81% (63/78) of SDSD vs 40% (14/35) of human SDSE 81% vs 40%
- ▲ ICEs averaged 2.5 per genome in human SDSE vs 1.6 in SDSD 2.5 vs 1.6 per genome
- – 17 genetic loci (40 genes) unique and ubiquitous to SDSD; 73 genes in 19 regions specific to human SDSE 40 genes / 73 genes
- – 14 different MLST profiles among SDSD including 5 novel profiles; bovine 13 types, ovine 4 types 14 profiles
- – Pangenome of 156 isolates: 6464 gene clusters, core-genome 871 clusters (13.5%); SDSD more closed (alpha 0.97) than SDSE (alpha 0.78) alpha 0.97 vs 0.78
- – emm-like gene found in SDSD with high homology to SDSE/S. pyogenes M-proteins; CDC primers had 4 forward and 1 reverse mismatch
- other 99.0% (SDSD), 97.9% (SDSE), 96.0% (between subsp) (average pairwise ANI within and between subspecies)
- count 63/78 (81%) SDSD vs 14/35 (40%) human SDSE harbored bacteriophage (bacteriophage prevalence)
- pvalue p < 0.0001 (difference in bacteriophage prevalence between subspecies)
- pvalue p < 0.0001 (difference in ICE carriage (2.5 vs 1.6 per genome))
- mean 2.04 MB average genome size; 1990 average CDS (SDSD genome statistics for cow and sheep isolates)
- count 752 single orthologue gene clusters (basis for phylogenetic reconstruction)
- count 31 of 78 SDSD harbored demA; 48 contained MIG; 29 harbored MAG (virulence factor carriage in SDSD)
- other pangenome size 9137 (Chao1) / 8669 (binomial mixture); core 871 (13.5%); Heaps alpha 0.81 (pangenome estimation across both subspecies)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This comparative genomics study analyzed 156 whole-genome sequences of Streptococcus dysgalactiae (78 SDSD, 78 SDSE) from multiple host species. Core analyses comprised phylogenetic reconstruction from 752 single-copy orthologue genes with 100 bootstrap replicates, pairwise average nucleotide identity (ANI) via fastANI, and pangenome characterization using the Chao1 index, Heaps' law, and a binomial mixture model. Prevalence differences for mobile genetic elements between subspecies were tested for significance (p < 0.0001 reported), but the specific statistical tests were not named in the text.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Phylogenetic bootstrap resampling (100 replicates) | SDSD phylogenetic tree (Fig. 3) and combined SDSD/SDSE midpoint-rooted tree (Fig. 4); alignment of 752 single-orthologue genes | 78 SDSD + 78 SDSE = 156 isolates | not stated |
| Unnamed significance test (p < 0.0001 reported) | Comparison of bacteriophage prevalence between SDSD (63/78, 81%) and human-associated SDSE (14/35, 40%) | 78 SDSD vs 35 human SDSE | not stated |
| Unnamed significance test (p < 0.0001 reported) | Comparison of ICE count per genome between human SDSE (mean 2.5, range 0–4) and SDSD (mean 1.6, range 1–4) | 78 SDSD vs 35 human SDSE | not stated |
| Chao1 nonparametric richness estimator | Pangenome size estimation across all 156 genomes (estimated size 9137 gene clusters from 6464 observed) | 156 genomes | na |
| Heaps' law (power-law regression for pangenome openness, alpha parameter) | Pangenome openness assessment for SD overall (alpha 0.81), SDSD (alpha 0.97), and SDSE (alpha 0.78) | 156 total; 78 SDSD; 78 SDSE | not stated |
| Binomial mixture model | Alternative pangenome size estimate (8669 gene clusters) and core-genome size estimate (871 gene clusters, 13.5% of total) | 156 genomes | not stated |
-
Significance testing for bacteriophage and ICE differences was reported with p < 0.0001 but the specific test was not named↳ Could also: Fisher's exact test (for the binary prevalence comparison 63/78 vs 14/35) or a Mann-Whitney U test (for the per-genome ICE count comparison) could each be explicitly named and reported — Naming the test allows readers to verify that its assumptions were met (independence, distributional form for count data) and enables exact reproducibility
-
Virulence factor presence/absence in Table 1 was compared across seven host groups and approximately 35 genes without any multiplicity correction↳ Could also: A Benjamini-Hochberg FDR correction applied across the family of group-by-gene comparisons would also be a standard approach in multi-gene presence/absence studies — With a large number of implicit comparisons across host groups and virulence genes, FDR control is a widely used strategy for managing the expected number of false discoveries
-
The gtfC allele–host association was presented descriptively (50/55 bovine vs 20/22 ovine carrying the respective allele)↳ Could also: A formal test of association (Fisher's exact test) with an effect size measure such as an odds ratio or phi coefficient would also quantify and test this concordance — Formal testing and an effect size allow readers to gauge how much of the host grouping is explained by allele type and to distinguish the observed pattern from chance variation
-
Phylogenetic branch support was assessed with 100 bootstrap replicates↳ Could also: Bayesian posterior probabilities (e.g., via MrBayes or IQ-TREE's ultrafast bootstrap) are also widely used to quantify phylogenetic branch support — Bayesian posteriors and classical bootstrap values capture different aspects of phylogenetic uncertainty; reporting both or choosing explicitly between them is common practice in large-genome phylogenomics
-
Genome size and CDS count dispersion was reported as mean ± a single value (implied SD) without confidence intervals or distributional context↳ Could also: 95% confidence intervals around group means, or distributional plots (e.g., violin or box plots), would also convey spread — With n = 55 bovine and n = 22 ovine isolates, CIs or distributional summaries facilitate direct assessment of overlap between groups
-
Pangenome openness was characterized using a single Heaps' law alpha estimate per group↳ Could also: Bayesian nonparametric pangenome models (e.g., the Infinitely Many Genes model) could also estimate pangenome openness and provide credible intervals on both size and openness — Bayesian approaches yield uncertainty bounds on pangenome estimates, whereas a single Heaps' law alpha conveys a point estimate with no associated precision
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-34462475
Paper: Porcellato et al. 2021, Sci Rep — "Whole genome sequencing reveals possible host species adaptation of Streptococcus dysgalactiae." DOI 10.1038/s41598-021-96710-z.
Named code artifact (P16, third-party tool): Shovill (https://github.com/tseemann/shovill) — a de-novo assembly pipeline wrapper (SPAdes/SKESA). Applying this tool to the paper's own deposited reads is a valid reproduction per the brief.
Data: ENA BioProject PRJEB42928 = the SDSD (subsp. dysgalactiae, bovine+ovine) genomes. (Companion PRJEB43000 = SDSE; out of our accession scope.)
Pipeline described in Methods
- Trimming: Trimmomatic
- De-novo assembly: Shovill pipeline
- Contig filter: drop contigs < 1000 bp and coverage < 3
- Annotation: Prokka
- MLST: CGE "MLST 2.0" (we use Seemann's
mlst+ PubMLSTsdysgalactiaescheme — equivalent) - ANI: fastANI
- ML tree: Geneious V10 (Jukes–Cantor) — proprietary GUI, OUT of scope
- Pangenome: all-vs-all blastp + R package micropan (Chao1)
In scope (pipeline-derived, attempted)
| id | reported result | tool | location |
|---|---|---|---|
| C1 | average genome size ≈ 2.04 Mb (bovine 2.04±0.1, ovine 2.02±0.05) | Shovill assembly | Results, "Genomic features" |
| C2 | average 1990 CDS (bovine 1993±95, ovine 1992±40) | Prokka annotation | Results |
| C3 | 14 MLST profiles among SDSD, 5 novel; bovine 13 STs, ovine 4 STs | MLST | Results |
| C4 | within-SDSD pairwise ANI > 98 %; mean SDSD ANI 99.0 % | fastANI | Results |
In scope but heavier (attempt after core)
| id | reported result | tool |
|---|---|---|
| C5 | SDSD pangenome 3550 gene clusters (subsp.-level), more-closed (alpha 0.97) | micropan blastp |
| C6 | virulence-gene host-association percentages (Table 1) | VF DB screen |
Out of scope (not pipeline / not attempted)
- ML phylogenetic tree (Geneious, proprietary GUI).
- Prophage / ICE counts (PHASTER web service / manual) — not a deposited reproducible pipeline.
- Wet-lab isolation, antimicrobial phenotypes.
- SDSE-only numbers (PRJEB43000, different accession).
Dataset N note (profiling)
Paper Methods: 60 newly sequenced SDSD (37 bovine + 23 ovine). Full phylo set = 78 SDSD (60 new + 18 public). ENA PRJEB42928 actually holds 124 paired-end WGS runs / 124 samples (Stdys001–Stdys221 with gaps + B02/B149/B237/MA201) → deposit contains MORE than the 60 "newly sequenced," recorded as N mismatch in dataset_profile.json.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.