Draft Genome Sequences of Antimicrobial-Resistant Shigella Clinical Isolates from Pakistan.
The main results reproduced: recomputed values matched the published ones within tolerance.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Same input data as the authors
- ✓Reported values were directly comparable
- ✓No relevant deviation in data/preprocessing
- ✓No authors-side cause for any deviation
- ✓Reported values are derivable from the shared data
- ✓Any deviation was negligible
- ✓The central claim held under reproduction
- ✓Overall, the reproduction was clean
- Every checked point held up.
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Described well enough -> reproduced ~1:1. MRA genome announcement (Lomonaco 2019); pipeline = Shovill 0.9 de novo assembly of 3 Shigella isolates from SRA PRJNA342326, plus in-silico MLST and ResFinder AMR typing. 24/25 graded claims reproduce exact or within-tolerance: all 3 N50, all 3 GC%, and all 3 MLST sequence types (ST245/ST245/ST152) are bit-EXACT; total lengths within 0.03%; contig counts off by <=3; AMR gene counts exact for 2/3 isolates (CFSAN059651 9 vs 8, attributable to a newer ResFinder DB); the qualitative shared-AMR-gene claim (dfrA1, sul2, CTX-M, tet in all 3) reproduces. shovill version matched exactly (0.9.0) though conda pulled SPAdes 4.3.0 vs the original ~3.12 -- the agreement is nonetheless near-perfect, strong evidence the Table 1 values are genuine pipeline output. NOT attempted (out of scope): wet-lab culture/AST/DNA-prep/sequencing, NCBI PGAP gene annotation (external service), and registry accession numbers.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 92assessed: 2026-06-16 ⛓ 650f5bb43846
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-16
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-18no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opus- ★ Draft genome sequences are reported for three multidrug-resistant Shigella clinical isolates from Pakistan (two S. flexneri and one S. sonnei). resource
- ★ All three isolates were resistant to ampicillin, cefazolin, ceftriaxone, cefotaxime, trimethoprim-sulfamethoxazole, and tetracycline. finding
- ★ The two S. flexneri isolates belonged to sequence type ST245 and the S. sonnei isolate belonged to ST152. finding
- ★ Isolates CFSAN059650, CFSAN059651, and CFSAN059652 harbored 10, 8, and 7 known AMR genes, respectively. finding
- ★ All Shigella isolates carried drfA1, sul2, a CTX-M gene, and a tet gene, conferring resistance to β-lactams, aminoglycosides, trimethoprim, sulfonamides, and tetracyclines. finding
- The genomic data provide a comparative genetic context for AMR in Shigella spp. useful for infectious disease epidemiology and public health monitoring. resource
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| antimicrobial susceptibility testing (disk diffusion) | Shigella flexneri and Shigella sonnei clinical isolates from Pakistan | none | antimicrobial resistance/susceptibility profile | API 20E system and disk diffusion method (bioMérieux) |
| antimicrobial susceptibility testing (broth microdilution, confirmatory) | Shigella clinical isolates | none | antimicrobial resistance confirmation | Vitek 2 system (bioMérieux) and conventional broth microdilution |
| whole-genome sequencing | three MDR Shigella isolates (CFSAN059650, CFSAN059651, CFSAN059652) | none | genome sequence (reads, contigs, genome size, GC content) | Illumina MiSeq, MiSeq reagent kit v2 (2 × 250-bp paired-end); DNeasy blood and tissue kit (Qiagen); Nextera XT DNA library kit (Illumina) |
| in silico multilocus sequence typing (MLST) | S. flexneri and S. sonnei draft genomes | none | sequence type (ST) | ResFinder |
| AMR gene identification (in silico) | S. flexneri and S. sonnei draft genomes | none | AMR genes with >99% identity to reference sequences | ResFinder |
| de novo genome assembly and annotation (bioinformatics) | Shigella sequencing reads | none | genome assemblies and annotations | Shovill 0.9 (GalaxyTrakr), Kraken, NCBI Prokaryotic Genome Annotation Pipeline |
- – CFSAN059650 (S. flexneri) genome: 356 contigs, 4,629,951 bp total length, N50 33,749 bp, GC 50.55%, coverage 52×. 4,629,951 bp; 52×
- – CFSAN059651 (S. flexneri) genome: 359 contigs, 4,620,515 bp total length, N50 31,916 bp, GC 50.45%, coverage 69×. 4,620,515 bp; 69×
- – CFSAN059652 (S. sonnei) genome: 384 contigs, 4,528,736 bp total length, N50 25,093 bp, GC 50.85%, coverage 71×. 4,528,736 bp; 71×
- – CFSAN059651 was additionally resistant to cefepime, aztreonam, ampicillin-sulbactam, and chloramphenicol; carried a phenicol resistance gene.
- – CFSAN059650 and CFSAN059652 carried additional fluoroquinolone resistance genes; showed intermediate resistance to ciprofloxacin/ampicillin-sulbactam and aztreonam/ceftazidime, respectively.
- – S. flexneri isolates typed as ST245 and S. sonnei isolate as ST152.
- – Isolates harbored 10, 8, and 7 known AMR genes (CFSAN059650, CFSAN059651, CFSAN059652). 10, 8, 7 genes
- count 356 contigs (CFSAN059650 assembly contig count)
- count 4,629,951 bp (CFSAN059650 total genome length)
- count 4,620,515 bp (CFSAN059651 total genome length)
- count 4,528,736 bp (CFSAN059652 total genome length)
- other 50.55%, 50.45%, 50.85% (GC content of CFSAN059650, CFSAN059651, CFSAN059652)
- count 10, 8, 7 (number of known AMR genes per isolate)
- other >99% identity (ResFinder threshold for AMR gene identification vs reference)
- count >50× coverage; Q score >26 (minimum sequence quality thresholds set)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This is a genomic resource announcement reporting draft whole-genome sequences for three multidrug-resistant Shigella clinical isolates. No inferential statistical tests were performed; the paper presents per-isolate descriptive assembly metrics (contig count, total length, N50, GC content, coverage) and qualitative/threshold-based bioinformatics outputs (in silico MLST sequence typing, AMR gene identification at >99% nucleotide identity). Antimicrobial susceptibility was characterized phenotypically by disk diffusion confirmed with broth microdilution per CLSI standards, with results reported as categorical interpretive categories (resistant/intermediate/susceptible).
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Disk diffusion (Kirby-Bauer), confirmed by broth microdilution | Antimicrobial susceptibility characterization of all three isolates | 3 isolates | na |
| In silico multilocus sequence typing (MLST) via ResFinder | Sequence type assignment for all three genome assemblies | 3 genome assemblies | na |
| AMR gene identification with >99% nucleotide identity threshold (ResFinder) | Detection of acquired AMR genes in all three isolates | 3 genome assemblies | na |
| Metagenomic classification for contamination screening (Kraken) | Sequence integrity evaluation of all three assemblies | 3 sequencing runs | na |
-
Short-read Illumina sequencing produced draft assemblies of 356–384 contigs per isolate↳ Could also: Long-read sequencing (e.g., Oxford Nanopore MinION, PacBio) alone or in a hybrid short+long-read assembly could also have been used — Long-read or hybrid assemblies can yield closed or near-closed chromosomes and resolve repetitive regions, which may capture mobile genetic elements (plasmids, integrons) carrying AMR genes that are fragmented across contigs in short-read-only assemblies
-
Sequence typing was performed with 7-locus MLST via ResFinder↳ Could also: Core-genome MLST (cgMLST) or whole-genome SNP-based phylogenetic analysis could also have been applied — cgMLST and SNP phylogenetics provide higher discriminatory resolution than 7-locus MLST, which can be valuable for contextualizing isolates within global Shigella population structure and identifying outbreak clusters
-
AMR genes were identified using a single tool (ResFinder) with a >99% nucleotide identity threshold↳ Could also: Additional tools such as AMRFinderPlus (NCBI), CARD/RGI, or ARIBA could also have been applied, with varied identity/coverage thresholds — Different databases cover different gene families and allelic variants; using complementary tools or a relaxed identity threshold can improve sensitivity for detecting divergent AMR gene homologs and provide cross-validation of findings
-
Quality control used Kraken for contamination screening and coverage/Q-score thresholds↳ Could also: Tools such as FastQC/MultiQC for read-level QC and CheckM for assembly completeness/contamination estimation could also have been used — Complementary QC layers (per-base quality profiles, adapter content, genome completeness percentages) provide additional characterization of assembly quality that can help readers assess data reliability
-
The three isolates are described individually without comparative phylogenetic placement↳ Could also: A maximum-likelihood or Bayesian phylogenetic tree incorporating publicly available reference genomes from global collections could also have been constructed — Phylogenetic contextualization would allow readers to assess whether these Pakistani isolates belong to recognized international lineages of MDR Shigella, strengthening the epidemiological interpretation of the AMR gene profiles reported
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-31346012
Paper: Draft Genome Sequences of Antimicrobial-Resistant Shigella Clinical Isolates from Pakistan. Lomonaco et al., Microbiol Resour Announc 2019. PMID 31346012 · PMCID PMC6658682 · DOI 10.1128/mra.00500-19.
This is a Microbiology Resource Announcement (MRA): it reports draft genome assemblies of 3 multidrug-resistant Shigella isolates and a small assembly-stats table (Table 1) plus typing/AMR-gene results derived from standard bioinformatic tools. There are no figures or large analyses — Table 1 + the typing sentences ARE the reproducible computational output.
Pipeline described in Methods (verbatim essentials)
- De novo genome assemblies created with Shovill 0.9 (https://github.com/tseemann/shovill), via GalaxyTrakr. "trim reads" option selected; minimum contig length set to 500 bp; default parameters otherwise.
- Read QC gates: min avg coverage >50×, R1/R2 Q scores >26 (a sequencing gate, not a recomputable pipeline output).
- Contamination checked with Kraken (QC; no numeric claim → not a primary target).
- Annotation by NCBI PGAP (external NCBI service → OUT of scope).
- ResFinder used for in-silico MLST (sequence types) and AMR-gene detection at
99% identity.
Data
- BioProject PRJNA342326; 3 Illumina MiSeq 2×250 paired-end runs:
- SRR8836963 — CFSAN059650 — S. flexneri — SAMN10086692
- SRR8836950 — CFSAN059651 — S. flexneri — SAMN10086664
- SRR8837012 — CFSAN059652 — S. sonnei — SAMN10086663
- Public on ENA/SRA (fastq.gz available), no restriction.
In scope (pipeline-derived, attempted)
For each of the 3 isolates, from the Shovill 0.9 assembly (Table 1):
- No. of contigs (minlen 500)
- Total length (bp)
- N50 (bp)
- GC content (%) Plus, derivable from the runs/metadata:
- Genome coverage (×) = total bases / genome size (sanity: matches 52/69/71 from ENA metadata)
- Coverage read count = reads used (≈ 2× ENA read pairs) Typing (third-party tools on the paper's data — equally valid per P16):
- MLST sequence type (ST245 / ST245 / ST152) via
mlst(PubMLST E. coli Achtman scheme; ResFinder-equivalent) - No. of AMR genes at >99% identity (10 / 8 / 7) via ResFinder/abricate
- Shared AMR genes (dfrA1, sul2, a CTX-M gene, a tet gene) — qualitative
Out of scope (not attempted)
- Wet-lab: culture, susceptibility testing (API 20E, Vitek, broth microdilution), DNA extraction, library prep, sequencing. (Reported phenotypes, not pipeline output.)
- NCBI PGAP gene annotation (external NCBI service, not locally reproducible).
- WGS/BioSample accession numbers (registry IDs, not computed).
Tooling plan («our HPC» SLURM, conda envs on «infra»)
- env_asm:
shovill=0.9,sra-tools— download fastq (ENA) + assemble. - env_typing:
mlst,abricate(bundles resfinder DB),assembly-stats,seqkit. - Caveat: assembly stats depend on the exact SPAdes version pulled by shovill 0.9 and thread count → expect small (within-tol) differences in contig count/N50, not bit-identical output. AMR-gene counts depend on ResFinder DB version (drift likely).
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
This MRA genome announcement reproduces essentially 1:1 from the deposited SRA data: all three N50, GC% and MLST sequence types are bit-exact, total lengths agree within 0.03%, and the shared AMR-gene claim holds. The only deviations — contig counts off by ≤3 and one AMR count off by one — are attributable to technical tool/DB version drift (SPAdes 4.3.0 vs ~3.12; newer ResFinder DB), not to authors or data availability. No fabrication concern: the near-exact agreement is strong positive evidence the Table 1 values are genuine pipeline output.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.