Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Population structure analysis of Salmonella serovar Muenchen to redefine geno-serotyping using genome indexing approaches.

Front Microbiol · 2026
L1 54/100 3/4
⚑ Flagged for review — a reproduced result did not match the reported value

Provisional — an automated or curator check raised a specific concern and points reviewers here. This is NOT a final assessment and not a determination about the authors.

Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +5
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
  • The central claim held under reproduction
What did not (or only partly)
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
54/100
Reproducibility score
1.1 SD below mean
vs. all fields · 1173 studies
🎯 Scores higher than 14% of all assessed papers rank 997 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough to reproduce the per-sample serotyping fields of Table 2's worked example SRR5235480 with the standalone third-party tools the paper cites (SeqSero2 v1.3.1, SISTR v1.0.2, MLST) run on the paper's own SRA data (P16). Result is a partial 1:1: the robust geno-serotyping signals match EXACTLY (ST 112; H antigens d/1,2; serogroup C2-C3; dominant serovar Muenchen). The contamination-sensitive fields differ (SeqSero2 O group 6,8 vs 3,10 => Muenchen vs Stormont; SISTR core-genome cgMLST Muenchen vs Anatum; QC PASS vs WARNING) -- consistent with the paper's own framing of SRR5235480 as a MIXED multiple-serovar sample, where O-antigen, cgMLST core match and QC are exactly the labile quantities sensitive to assembler (skesa vs MEGAHIT-in-bettercallsal) and SeqSero2 mode. SeqSero2 mode a independently flagged inter-serotype contamination=yes, corroborating the mixed-sample premise. The paper's central qualitative claim for this row (conflicting multi-serovar calls across methods) is itself reproduced. NOT attempted: the bettercallsal genome-index multi-call itself (needs the authors' unpinned PDG 1,963-genome snapshot from 2025-02-20; out of scope per 80/20) and cohort-level results (98%/92.87% concordance, pangenome 3159/586/8139, RAxML tree). No fabrication concern: all reported values are plausible tool outputs; divergences are explained by documented assembler/version/mode sensitivity.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 54
    assessed: 2026-06-16 ⛓ 7e28a75afba7
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-16
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-16
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

The study hypothesizes that combining DNA sketching-based serotyping via bettercallsal with the validated geno-serotyping tool SeqSero2 provides a robust framework for Salmonella serovar characterization—maximizing information retrieval and improving serovar resolution while preserving historical serovar definitions through antigen profiles—using the polyphyletic serovar S. Muenchen as a test case.

Core claims
  • Integrating genome-indexing (bettercallsal, DNA sketching + genome proximity) with SeqSero2 yields complementary serovar calls that improve discrimination of genomically distinct but antigenically similar serovars while retaining historical nomenclature finding
  • bettercallsal, leveraging the NCBI Pathogen Detection database, enhances serovar resolution by incorporating genome proximity-based calls method
  • S. Muenchen is a polyphyletic serovar (same O and H antigens but derived from multiple genetically distinct ancestors per MLST phylogeny) mechanism
  • A dual-tool strategy (SeqSero2 + bettercallsal) plus cgMLST/pangenome conflict resolution improves source attribution accuracy in outbreak investigations finding
  • WGS of isolates from mixed cultures can pass QC yet yield false 'mixed' serovar calls, indicating a need for enhanced QC tools to detect mixtures finding
  • A curated dataset of 1,963 S. Muenchen plus 27 outgroup SRRs (1,990 total) was assembled from NCBI Pathogen Detection for comparative serotyping resource
Experimental setups
Assay System Perturbation Readout Platform
Genome-indexing serovar calling (DNA sketching, assembly-free) 1,990 Salmonella SRRs (Muenchen + outgroups), Illumina paired-end reads none serovar calls / genome proximity predictions bettercallsal v0.7.0 (with fastMLST, KMA aligner, MEGAHIT assembler)
In-silico antigen-based serotyping 1,990 Salmonella SRRs and MEGAHIT-assembled genomes none O/H antigen serovar prediction SeqSero2 v1.3.1 (k-mer mode)
cgMLST phylogenetic clustering MEGAHIT-assembled genomes of 1,990 SRRs none core genome MLST profiles, genome relatedness clusters SISTR CLI v1.0.2; EnteroBase Salmonella.cgMLSTv2 scheme; GrapeTree/Microreact
7-gene MLST sequence typing genome assemblies of Salmonella SRRs none sequence types (ST) fastMLST / Achtman 7-gene scheme (within bettercallsal pipeline)
Pangenome analysis (core/shell/cloud gene classification) MEGAHIT-assembled Salmonella genomes none gene presence/absence matrix, accessory genome diversity Prokka v1.14.6 + PIRATE (Roary also evaluated)
Phylogenetic clustering of fliC alleles and rfb gene cluster Salmonella genome sequences none phylogenetic tree of antigen-encoding loci RAxML v8.2.9
Key results
  • bettercallsal and NCBI computed type calls agreed for the Muenchen dataset 92.87% (1,823/1,963 concordant; 140 conflicts)
  • bettercallsal and locally-run SeqSero2 concordantly called Muenchen 98% (1,933/1,963)
  • Among 140 conflicts, bettercallsal called Valdosta (8:a:1,2) instead of Muenchen n=90
  • bettercallsal called Umbadah (1,3,19:d:1,2) for a subset of conflicting SRRs n=19
  • Subset of conflicting SRRs returned no serovar call / incomplete antigen formula (I -:-:-) n=32
  • bettercallsal called Newport (8:e,h:1,2) for conflicting SRRs n=1 (plus 1 as dual call)
  • SRRs received multiple/dual serovar calls by bettercallsal (flagged WARNING QC) n=11
Key statistics
  • count 1,963 (S. Muenchen (6,8:d:1,2, serogroup C2-C3) genomes/SRRs in dataset)
  • count 1,990 (total SRRs analyzed (Muenchen + 27 outgroups))
  • other 92.87% (agreement between bettercallsal and NCBI computed type)
  • other 98% (1,933) (concordance between bettercallsal and locally-run SeqSero2)
  • count 140 (SRRs with conflicts between bettercallsal and computed type)
  • count 27 (outgroup SRRs included)
  • count ~730,000 (Salmonella genomes in NCBI Pathogen Detection as of Feb 20, 2025)
  • count >2,600 (identified S. enterica serovars in White-Kauffmann-Le Minor scheme)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This is a bioinformatics methods-comparison study evaluating four in-silico serotyping/genomic-typing tools (bettercallsal, SeqSero2, SISTR cgMLST, PIRATE pangenome) on 1,990 Salmonella Muenchen whole-genome sequences from NCBI Pathogen Detection. The primary analytical approach is tool-concordance counting (percentage agreement between pairwise tool calls), supplemented by maximum-likelihood phylogenetic inference of antigen-encoding genes and pangenome gene-prevalence categorisation. Results are reported as raw counts and percent agreement; no inferential hypothesis tests or p-values are presented.

Replicationunclear Sample size1,963 Salmonella Muenchen SRRs downloaded from NCBI PD plus 27 outgroup SRRs; no a-priori power calculation described; sample size determined by NCBI PD database availability GroupsMultiple in-silico serotyping tools compared against each other and against NCBI computed_type labels; concordant vs. discordant serovar calls Pairingpaired Randomization/blindingnot stated Dispersionnone Exact p-valuesno Effect sizesno Confidence intervalsno Multiplicity correctionnone stated
Statistical tests used
Test Applied to n Assumptions
Percent pairwise concordance (raw count of matching serovar calls between tool pairs) Table 3 — all four tool-pair combinations across 1,963 Muenchen SRRs 1,963 SRRs (Muenchen) + 27 outgroups = 1,990 total not stated
Maximum-likelihood phylogenetic inference (RAxML v8.2.9) Comparison and phylogenetic clustering of fliC alleles and rfb gene cluster not stated not stated
cgMLST hierarchical clustering (GrapeTree / Microreact visualisation) SISTR cgMLST profiles clustered for all 1,990 SRRs 1,990 SRRs not stated
Pangenome gene-prevalence thresholding (PIRATE: core >95%, shell 15–95%, cloud <15%) Pangenome analysis of MEGAHIT-assembled genomes as part of conflict resolution not stated (subset of 1,990 assemblies) not stated
Approaches that could also have been used
  • Tool agreement is quantified as simple percent concordance (e.g., 92.87%, 98%)
    Could also: Cohen's kappa (κ) or Gwet's AC1 could also quantify inter-tool agreement, with 95% confidence intervals — Chance-corrected agreement coefficients account for agreement expected by random coincidence of call frequencies, and CIs communicate uncertainty in the concordance estimate — both are standard in diagnostic-accuracy and typing-tool benchmarking studies
  • Maximum-likelihood phylogeny (RAxML) was used for fliC alleles and rfb gene cluster without a stated substitution model selection step or bootstrap support threshold
    Could also: ModelTest-NG or IQ-TREE's built-in ModelFinder could be used to select the best-fit substitution model before ML inference; bootstrap or ultrafast-bootstrap support values could be reported — Explicit model selection and branch-support values are standard practice in phylogenetics and allow readers to assess confidence in topological conclusions
  • cgMLST clustering was visualised in GrapeTree/Microreact without a formal cluster-validity or distance-threshold criterion stated
    Could also: A defined allele-distance threshold (e.g., HC threshold in EnteroBase) or quantitative cluster-validation indices (e.g., silhouette score, Davies–Bouldin index) could also be used to objectively delineate clusters — Explicit thresholds make cluster boundaries reproducible and comparable across studies; quantitative indices allow evaluation of cluster cohesion independent of visualisation
  • Pangenome gene categories were defined by fixed prevalence thresholds (core >95%, shell 15–95%, cloud <15%)
    Could also: Alternative threshold schemes (e.g., strict core 99–100%, soft core 95–99%; or data-driven breakpoints from the gene-frequency histogram) are also widely used — Threshold choice affects the size and interpreted biological meaning of each gene category; reporting sensitivity to threshold values or following a community-standard scheme aids comparability with other Salmonella pangenome studies
  • Discordant calls between tools were resolved through qualitative integration of cgMLST, pangenome, and antigen-gene evidence without a pre-specified decision rule
    Could also: A formal majority-vote or weighted-ensemble rule, or a pre-registered conflict-resolution algorithm, could also be stated explicitly before analysis — A transparent, pre-specified adjudication rule reduces subjectivity and makes the framework reproducible when applied to new datasets or other serovars
  • Tool performance is characterised only by pairwise percent agreement; no reference-standard sensitivity/specificity analysis is performed
    Could also: If a curated ground-truth serovar label set were available (e.g., wet-lab-confirmed phenotypic serotyping results), sensitivity, specificity, positive/negative predictive value, and their 95% CIs (e.g., Wilson score interval) could also be calculated — Agreement metrics measure consistency between tools but not accuracy relative to a known truth; accuracy metrics are the standard in diagnostic tool validation and would complement the concordance analysis
Software: bettercallsal 0.7.0 · SeqSero2 1.3.1 (seen in Table 2 header; v2.0 also referenced) · SISTR 1.0.2 · RAxML 8.2.9 · PIRATE · Prokka 1.14.6 · MEGAHIT · fastMLST · KMA aligner · GrapeTree · Microreact

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
0
Impact: low
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-41743541

Paper: Ramachandran P, Konganti K, Windsor AM, Grim CJ, Pradhan AK. "Population structure analysis of Salmonella serovar Muenchen to redefine geno-serotyping using genome indexing approaches." Front Microbiol 2025/2026. PMID 41743541 · PMCID PMC12929376 · DOI 10.3389/fmicb.2025.1681711.

Pipeline/tool (P16 — third-party tool on the paper's data is equally valid):

  • bettercallsal v0.7.0 (CFSAN-Biostatistics, github.com/CFSAN-Biostatistics/bettercallsal, default branch master; current release v1.2.0) — assembly-free genome-indexing serovar caller. Uses KMA + MEGAHIT + fastMLST + salmon over an NCBI Pathogen Detection (PDG) genome index.
  • Independent serotyping tools the paper reports per sample: SeqSero2 v1.3.1 (antigen-based serotyping from reads) and SISTR v1.0.2 (cgMLST + antigen + serogroup + QC from an assembly).

In scope (attempted)

Concrete, low-hanging, per-sample reproducible values from Table 2 ("SRRs with multiple serovar calls identified by bettercallsal"), row SRR5235480 (the worked example). These are reproducible with the standalone third-party tools run directly on the paper's own data (SRA SRR5235480), no custom DB needed:

Reported field (Table 2, SRR5235480) Reported value Tool to reproduce
O group 3,10 SeqSero2 v1.3.1
fliC (H1) d SeqSero2 v1.3.1
fljB (H2) 1,2 SeqSero2 v1.3.1
SeqSero2 serovar Stormont SeqSero2 v1.3.1
ST 112 mlst (Achtman senterica) / SISTR
Serovar overall (SISTR serovar) Muenchen|Virginia|Manhattan|Yovokome SISTR v1.0.2
SISTR serovar_antigen Muenchen|Virginia|Manhattan|Yovokome SISTR v1.0.2
SISTR serovar_cgmlst Anatum SISTR v1.0.2
qc_status WARNING SISTR v1.0.2
serogroup C2-C3 SISTR v1.0.2

Out of scope / not attempted (the hard ~20%)

  • bcs_serotype1/2 = Muenchen / Anatum (the bettercallsal multi-call itself). Reproducing this requires the exact bettercallsal genome index built from the authors' PDG snapshot (1,963 Muenchen genomes downloaded 2025-02-20, embedded in ~730k-genome PDG release). The bettercallsal_db build downloads the full NCBI Pathogen Detection Salmonella set — large, time-consuming, and the specific snapshot is not pinned in the repo. Per the 80/20 rule this is skipped; the SISTR serovar_cgmlst=Anatum and SISTR serovar=Muenchen|... we DO reproduce are the same underlying signals that drive the bettercallsal Muenchen+Anatum call.
  • Cohort-level results (98% / 92.87% concordance over 1,963 SRRs; pangenome 3,159 core / 586 shell / 8,139 cloud genes; RAxML phylogeny). These need the full 1,963-genome cohort + custom DB; out of scope for a few-clear-datapoints repro.
  • All wet-lab / database-curation steps — non-pipeline, out of scope.

Data

  • SRA SRR5235480, paired-end Illumina, 841,548 read pairs (ENA filereport). md5: _1 9f169848a19d209cd7e232818a8f9c1d, _2 638adf2d4d4d2b95cf6e35e8a66f7c4a.
  • Downloaded + processed entirely on «infra» inside the «our HPC» job.

Reproduction approach

Single SLURM job (run.sbatch, «job»): download SRR5235480 from ENA → SeqSero2 v1.3.1 (mode a, paired reads) → skesa assembly → SISTR v1.0.2 (--qc) → mlst. Consolidate to results.json. Compare to Table 2 row.

Figures / tables: Table
C1
Reported
ST 112
Reproduced
ST 112
exact
C2
Reported
fliC(H1)=d
Reproduced
d
exact
C3
Reported
fljB(H2)=1,2
Reproduced
1,2
exact
C4
Reported
serogroup C2-C3
Reproduced
C2-C3
exact
C5
Reported
SISTR serovar Muenchen|Virginia|Manhattan|Yovokome
Reproduced
Muenchen|Virginia
partial
C6
Reported
SeqSero2 O group 3,10
Reproduced
6,8
did not match
C7
Reported
SeqSero2 serovar Stormont
Reproduced
Muenchen
did not match
C8
Reported
SISTR serovar_cgmlst Anatum
Reproduced
Muenchen
did not match
C9
Reported
qc_status WARNING
Reproduced
PASS
did not match
C10
Reported
bettercallsal bcs_serotype1/2 = Muenchen, Anatum
Reproduced
not attempted
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 54/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟢7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +5

For Table 2's worked example SRR5235480, the robust serotyping signals reproduced exactly (ST 112; H1=d; H2=1,2; serogroup C2-C3; dominant serovar Muenchen) directly from the paper's own SRA data. The divergent fields (O group 6,8 vs 3,10 → Muenchen vs Stormont; cgMLST Muenchen vs Anatum; QC PASS vs WARNING) are the labile quantities on a known mixed/contaminated sample and track our self-chosen assembler (skesa vs the paper's MEGAHIT-in-bettercallsal), so the deviation sits on our methodology, not an authors' defect. No fabrication concern — every reported value is a plausible tool output — and the row's central premise (conflicting multi-serovar calls) is itself reproduced. Overall a solid partial reproduction with explainable deviations.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

148 k
tokens (I/O) · 11 M incl. cache
23 min
runtime · 0.97 CPU-h
13.6 GB
peak RAM
2
HPC jobs
hummel
machine