Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Whole genome sequencing reveals possible host species adaptation of Streptococcus dysgalactiae.

Sci Rep · 2021
not yet assessed 2/4
Why this verdict

Part of the results reproduced; minor but material deviations remained.

Reproduced on the brainbox compute brainarbeit.com
✓ What held up
  • Reported values were directly comparable
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
Reproduction agent’s raw note

IN PROGRESS. Reproducing shovill de-novo assembly + Prokka + MLST + fastANI on ENA PRJEB42928 (124 paired-end SDSD WGS runs). Targets: avg genome size 2.04 Mb, ~1990 CDS, 14 MLST STs, mean ANI 99.0%. Env building on «our HPC»; downloads starting.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment
    assessed: 2026-06-18 ⛓ 485ad2f3fcfb
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-18
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19
no human curator yet
Last updated
2026-07-31

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

Can whole genome sequencing of Streptococcus dysgalactiae (SD) from multiple host species reveal virulence factors, phylogenetic relationships, and genetic factors contributing to host adaptation, particularly for the understudied subspecies dysgalactiae (SDSD) from sheep and cattle?

Core claims
  • SDSD constitutes a distinct taxonomic entity within S. dysgalactiae, with a mean intra-subspecies average nucleotide identity of 99%. finding
  • Phylogenetic clustering of SD isolates extends beyond the two-subspecies division and reflects adaptive evolution into multiple host-associated lineages. finding
  • No single gene or genetic region was uniquely associated with host species among bovine vs ovine SDSD, despite separate pangenome clustering. finding
  • Lineage-specific virulence factors exist and several reside near hotspots for integration of mobile genetic elements, implicating horizontal genetic transfer in niche specialization. mechanism
  • A previously unidentified emm-like (M-protein) gene and a novel Srr-like glycoprotein operon were detected in SDSD. finding
  • The glycosylation gene gtfC exists in two allelic variants whose distribution correlates with bovine vs ovine host of origin. finding
  • Intact bacteriophages are more prevalent in SDSD whereas Integrative Conjugative Elements (ICEs) are more prevalent in human-associated SDSE. finding
  • This is the first comprehensive genomic characterization of SDSD and the first to include ovine isolates, providing a sequenced genome resource. resource
Experimental setups
Assay System Perturbation Readout Platform
Whole genome sequencing S. dysgalactiae isolates from cows and sheep (SDSD) none genome sequence, genome size, CDS count
Whole genome sequencing SDSE isolates from human, dog, horse, pig none genome sequence
Multilocus sequence typing (MLST) SDSD isolates from cows and sheep none sequence types / MLST profiles
Pangenome analysis 156 SD isolates (78 SDSD + 78 SDSE) none gene clusters, core/pangenome size, single orthologue genes
Phylogenetic analysis 78 SDSD and 78 SDSE genomes none phylogenetic tree from 752 single orthologue gene clusters, bootstrap support
Average nucleotide identity (fastANI) 156 SD whole genome sequences none pairwise ANI values / heatmap fastANI algorithm
Virulence gene profiling / in silico screening SDSD and SDSE genomes by host species none presence/percentage of virulence factors
Mobile genetic element screening SDSD and human-associated SDSE genomes none bacteriophage, ICE, and associated virulence/resistance gene counts
Key results
  • Mean intra-subspecies ANI for SDSD was 99% with all pairwise comparisons >98% 99% (>98%)
  • gtfC allele A in 50/55 bovine isolates and allele B in 20/22 ovine isolates, the two variants 88% similar 88% similarity
  • Intact bacteriophages in 81% (63/78) of SDSD vs 40% (14/35) of human SDSE 81% vs 40%
  • ICEs averaged 2.5 per genome in human SDSE vs 1.6 in SDSD 2.5 vs 1.6 per genome
  • 17 genetic loci (40 genes) unique and ubiquitous to SDSD; 73 genes in 19 regions specific to human SDSE 40 genes / 73 genes
  • 14 different MLST profiles among SDSD including 5 novel profiles; bovine 13 types, ovine 4 types 14 profiles
  • Pangenome of 156 isolates: 6464 gene clusters, core-genome 871 clusters (13.5%); SDSD more closed (alpha 0.97) than SDSE (alpha 0.78) alpha 0.97 vs 0.78
  • emm-like gene found in SDSD with high homology to SDSE/S. pyogenes M-proteins; CDC primers had 4 forward and 1 reverse mismatch
Key statistics
  • other 99.0% (SDSD), 97.9% (SDSE), 96.0% (between subsp) (average pairwise ANI within and between subspecies)
  • count 63/78 (81%) SDSD vs 14/35 (40%) human SDSE harbored bacteriophage (bacteriophage prevalence)
  • pvalue p < 0.0001 (difference in bacteriophage prevalence between subspecies)
  • pvalue p < 0.0001 (difference in ICE carriage (2.5 vs 1.6 per genome))
  • mean 2.04 MB average genome size; 1990 average CDS (SDSD genome statistics for cow and sheep isolates)
  • count 752 single orthologue gene clusters (basis for phylogenetic reconstruction)
  • count 31 of 78 SDSD harbored demA; 48 contained MIG; 29 harbored MAG (virulence factor carriage in SDSD)
  • other pangenome size 9137 (Chao1) / 8669 (binomial mixture); core 871 (13.5%); Heaps alpha 0.81 (pangenome estimation across both subspecies)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This comparative genomics study analyzed 156 whole-genome sequences of Streptococcus dysgalactiae (78 SDSD, 78 SDSE) from multiple host species. Core analyses comprised phylogenetic reconstruction from 752 single-copy orthologue genes with 100 bootstrap replicates, pairwise average nucleotide identity (ANI) via fastANI, and pangenome characterization using the Chao1 index, Heaps' law, and a binomial mixture model. Prevalence differences for mobile genetic elements between subspecies were tested for significance (p < 0.0001 reported), but the specific statistical tests were not named in the text.

Replicationbiological Sample size78 SDSD (55 bovine, 22 ovine, 1 human) and 78 SDSE (35 human, 20 horse, 8 dog, 7 swine, 5 fish); 60 SDSD and 40 SDSE newly sequenced; remainder retrieved from public databases; no formal power calculation stated GroupsSDSD vs SDSE; within SDSD: bovine vs ovine; within SDSE: human vs horse vs dog vs swine vs fish Pairingunpaired Randomization/blindingnot stated Dispersionmixed Exact p-valuesno Effect sizesno Confidence intervalsno Multiplicity correctionnone stated
Statistical tests used
Test Applied to n Assumptions
Phylogenetic bootstrap resampling (100 replicates) SDSD phylogenetic tree (Fig. 3) and combined SDSD/SDSE midpoint-rooted tree (Fig. 4); alignment of 752 single-orthologue genes 78 SDSD + 78 SDSE = 156 isolates not stated
Unnamed significance test (p < 0.0001 reported) Comparison of bacteriophage prevalence between SDSD (63/78, 81%) and human-associated SDSE (14/35, 40%) 78 SDSD vs 35 human SDSE not stated
Unnamed significance test (p < 0.0001 reported) Comparison of ICE count per genome between human SDSE (mean 2.5, range 0–4) and SDSD (mean 1.6, range 1–4) 78 SDSD vs 35 human SDSE not stated
Chao1 nonparametric richness estimator Pangenome size estimation across all 156 genomes (estimated size 9137 gene clusters from 6464 observed) 156 genomes na
Heaps' law (power-law regression for pangenome openness, alpha parameter) Pangenome openness assessment for SD overall (alpha 0.81), SDSD (alpha 0.97), and SDSE (alpha 0.78) 156 total; 78 SDSD; 78 SDSE not stated
Binomial mixture model Alternative pangenome size estimate (8669 gene clusters) and core-genome size estimate (871 gene clusters, 13.5% of total) 156 genomes not stated
Approaches that could also have been used
  • Significance testing for bacteriophage and ICE differences was reported with p < 0.0001 but the specific test was not named
    Could also: Fisher's exact test (for the binary prevalence comparison 63/78 vs 14/35) or a Mann-Whitney U test (for the per-genome ICE count comparison) could each be explicitly named and reported — Naming the test allows readers to verify that its assumptions were met (independence, distributional form for count data) and enables exact reproducibility
  • Virulence factor presence/absence in Table 1 was compared across seven host groups and approximately 35 genes without any multiplicity correction
    Could also: A Benjamini-Hochberg FDR correction applied across the family of group-by-gene comparisons would also be a standard approach in multi-gene presence/absence studies — With a large number of implicit comparisons across host groups and virulence genes, FDR control is a widely used strategy for managing the expected number of false discoveries
  • The gtfC allele–host association was presented descriptively (50/55 bovine vs 20/22 ovine carrying the respective allele)
    Could also: A formal test of association (Fisher's exact test) with an effect size measure such as an odds ratio or phi coefficient would also quantify and test this concordance — Formal testing and an effect size allow readers to gauge how much of the host grouping is explained by allele type and to distinguish the observed pattern from chance variation
  • Phylogenetic branch support was assessed with 100 bootstrap replicates
    Could also: Bayesian posterior probabilities (e.g., via MrBayes or IQ-TREE's ultrafast bootstrap) are also widely used to quantify phylogenetic branch support — Bayesian posteriors and classical bootstrap values capture different aspects of phylogenetic uncertainty; reporting both or choosing explicitly between them is common practice in large-genome phylogenomics
  • Genome size and CDS count dispersion was reported as mean ± a single value (implied SD) without confidence intervals or distributional context
    Could also: 95% confidence intervals around group means, or distributional plots (e.g., violin or box plots), would also convey spread — With n = 55 bovine and n = 22 ovine isolates, CIs or distributional summaries facilitate direct assessment of overlap between groups
  • Pangenome openness was characterized using a single Heaps' law alpha estimate per group
    Could also: Bayesian nonparametric pangenome models (e.g., the Infinitely Many Genes model) could also estimate pangenome openness and provide credible intervals on both size and openness — Bayesian approaches yield uncertainty bounds on pangenome estimates, whereas a single Heaps' law alpha conveys a point estimate with no associated precision
Software: fastANI · Geneious (protein alignment with BLOSUM62 cost matrix)

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-34462475

Paper: Porcellato et al. 2021, Sci Rep — "Whole genome sequencing reveals possible host species adaptation of Streptococcus dysgalactiae." DOI 10.1038/s41598-021-96710-z.

Named code artifact (P16, third-party tool): Shovill (https://github.com/tseemann/shovill) — a de-novo assembly pipeline wrapper (SPAdes/SKESA). Applying this tool to the paper's own deposited reads is a valid reproduction per the brief.

Data: ENA BioProject PRJEB42928 = the SDSD (subsp. dysgalactiae, bovine+ovine) genomes. (Companion PRJEB43000 = SDSE; out of our accession scope.)

Pipeline described in Methods

  1. Trimming: Trimmomatic
  2. De-novo assembly: Shovill pipeline
  3. Contig filter: drop contigs < 1000 bp and coverage < 3
  4. Annotation: Prokka
  5. MLST: CGE "MLST 2.0" (we use Seemann's mlst + PubMLST sdysgalactiae scheme — equivalent)
  6. ANI: fastANI
  7. ML tree: Geneious V10 (Jukes–Cantor) — proprietary GUI, OUT of scope
  8. Pangenome: all-vs-all blastp + R package micropan (Chao1)

In scope (pipeline-derived, attempted)

id reported result tool location
C1 average genome size ≈ 2.04 Mb (bovine 2.04±0.1, ovine 2.02±0.05) Shovill assembly Results, "Genomic features"
C2 average 1990 CDS (bovine 1993±95, ovine 1992±40) Prokka annotation Results
C3 14 MLST profiles among SDSD, 5 novel; bovine 13 STs, ovine 4 STs MLST Results
C4 within-SDSD pairwise ANI > 98 %; mean SDSD ANI 99.0 % fastANI Results

In scope but heavier (attempt after core)

id reported result tool
C5 SDSD pangenome 3550 gene clusters (subsp.-level), more-closed (alpha 0.97) micropan blastp
C6 virulence-gene host-association percentages (Table 1) VF DB screen

Out of scope (not pipeline / not attempted)

  • ML phylogenetic tree (Geneious, proprietary GUI).
  • Prophage / ICE counts (PHASTER web service / manual) — not a deposited reproducible pipeline.
  • Wet-lab isolation, antimicrobial phenotypes.
  • SDSE-only numbers (PRJEB43000, different accession).

Dataset N note (profiling)

Paper Methods: 60 newly sequenced SDSD (37 bovine + 23 ovine). Full phylo set = 78 SDSD (60 new + 18 public). ENA PRJEB42928 actually holds 124 paired-end WGS runs / 124 samples (Stdys001–Stdys221 with gaps + B02/B149/B237/MA201) → deposit contains MORE than the 60 "newly sequenced," recorded as N mismatch in dataset_profile.json.

C1
Reported
avg genome 2.04 Mb (bovine 2.04±0.1; ovine 2.02±0.05)
Reproduced
C2
Reported
avg 1990 CDS (bovine 1993±95; ovine 1992±40)
Reproduced
C3
Reported
14 MLST profiles, 5 novel (bovine 13, ovine 4 STs)
Reproduced
C4
Reported
within-SDSD ANI mean 99.0%, all >98%
Reproduced

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 56/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟢2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

71.9 k
tokens (I/O) · 3.5 M incl. cache
8 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.