Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Draft Genome Sequences of Antimicrobial-Resistant Shigella Clinical Isolates from Pakistan.

Microbiol Resour Announc · 2019
L1 92/100 PQI 87
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Every question reproduced
-1 pts
From: “every question reproduced”
Total score -7
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
  • No relevant deviation in data/preprocessing
  • No authors-side cause for any deviation
  • Reported values are derivable from the shared data
  • Any deviation was negligible
  • The central claim held under reproduction
  • Overall, the reproduction was clean
What did not (or only partly)
  • Every checked point held up.
How its reproducibility compares
92/100
Reproducibility score
1.0 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 83% of all assessed papers rank 179 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough -> reproduced ~1:1. MRA genome announcement (Lomonaco 2019); pipeline = Shovill 0.9 de novo assembly of 3 Shigella isolates from SRA PRJNA342326, plus in-silico MLST and ResFinder AMR typing. 24/25 graded claims reproduce exact or within-tolerance: all 3 N50, all 3 GC%, and all 3 MLST sequence types (ST245/ST245/ST152) are bit-EXACT; total lengths within 0.03%; contig counts off by <=3; AMR gene counts exact for 2/3 isolates (CFSAN059651 9 vs 8, attributable to a newer ResFinder DB); the qualitative shared-AMR-gene claim (dfrA1, sul2, CTX-M, tet in all 3) reproduces. shovill version matched exactly (0.9.0) though conda pulled SPAdes 4.3.0 vs the original ~3.12 -- the agreement is nonetheless near-perfect, strong evidence the Table 1 values are genuine pipeline output. NOT attempted (out of scope): wet-lab culture/AST/DNA-prep/sequencing, NCBI PGAP gene annotation (external service), and registry accession numbers.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 92
    assessed: 2026-06-16 ⛓ 650f5bb43846
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-16
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-18
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Core claims
  • Draft genome sequences are reported for three multidrug-resistant Shigella clinical isolates from Pakistan (two S. flexneri and one S. sonnei). resource
  • All three isolates were resistant to ampicillin, cefazolin, ceftriaxone, cefotaxime, trimethoprim-sulfamethoxazole, and tetracycline. finding
  • The two S. flexneri isolates belonged to sequence type ST245 and the S. sonnei isolate belonged to ST152. finding
  • Isolates CFSAN059650, CFSAN059651, and CFSAN059652 harbored 10, 8, and 7 known AMR genes, respectively. finding
  • All Shigella isolates carried drfA1, sul2, a CTX-M gene, and a tet gene, conferring resistance to β-lactams, aminoglycosides, trimethoprim, sulfonamides, and tetracyclines. finding
  • The genomic data provide a comparative genetic context for AMR in Shigella spp. useful for infectious disease epidemiology and public health monitoring. resource
Experimental setups
Assay System Perturbation Readout Platform
antimicrobial susceptibility testing (disk diffusion) Shigella flexneri and Shigella sonnei clinical isolates from Pakistan none antimicrobial resistance/susceptibility profile API 20E system and disk diffusion method (bioMérieux)
antimicrobial susceptibility testing (broth microdilution, confirmatory) Shigella clinical isolates none antimicrobial resistance confirmation Vitek 2 system (bioMérieux) and conventional broth microdilution
whole-genome sequencing three MDR Shigella isolates (CFSAN059650, CFSAN059651, CFSAN059652) none genome sequence (reads, contigs, genome size, GC content) Illumina MiSeq, MiSeq reagent kit v2 (2 × 250-bp paired-end); DNeasy blood and tissue kit (Qiagen); Nextera XT DNA library kit (Illumina)
in silico multilocus sequence typing (MLST) S. flexneri and S. sonnei draft genomes none sequence type (ST) ResFinder
AMR gene identification (in silico) S. flexneri and S. sonnei draft genomes none AMR genes with >99% identity to reference sequences ResFinder
de novo genome assembly and annotation (bioinformatics) Shigella sequencing reads none genome assemblies and annotations Shovill 0.9 (GalaxyTrakr), Kraken, NCBI Prokaryotic Genome Annotation Pipeline
Key results
  • CFSAN059650 (S. flexneri) genome: 356 contigs, 4,629,951 bp total length, N50 33,749 bp, GC 50.55%, coverage 52×. 4,629,951 bp; 52×
  • CFSAN059651 (S. flexneri) genome: 359 contigs, 4,620,515 bp total length, N50 31,916 bp, GC 50.45%, coverage 69×. 4,620,515 bp; 69×
  • CFSAN059652 (S. sonnei) genome: 384 contigs, 4,528,736 bp total length, N50 25,093 bp, GC 50.85%, coverage 71×. 4,528,736 bp; 71×
  • CFSAN059651 was additionally resistant to cefepime, aztreonam, ampicillin-sulbactam, and chloramphenicol; carried a phenicol resistance gene.
  • CFSAN059650 and CFSAN059652 carried additional fluoroquinolone resistance genes; showed intermediate resistance to ciprofloxacin/ampicillin-sulbactam and aztreonam/ceftazidime, respectively.
  • S. flexneri isolates typed as ST245 and S. sonnei isolate as ST152.
  • Isolates harbored 10, 8, and 7 known AMR genes (CFSAN059650, CFSAN059651, CFSAN059652). 10, 8, 7 genes
Key statistics
  • count 356 contigs (CFSAN059650 assembly contig count)
  • count 4,629,951 bp (CFSAN059650 total genome length)
  • count 4,620,515 bp (CFSAN059651 total genome length)
  • count 4,528,736 bp (CFSAN059652 total genome length)
  • other 50.55%, 50.45%, 50.85% (GC content of CFSAN059650, CFSAN059651, CFSAN059652)
  • count 10, 8, 7 (number of known AMR genes per isolate)
  • other >99% identity (ResFinder threshold for AMR gene identification vs reference)
  • count >50× coverage; Q score >26 (minimum sequence quality thresholds set)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This is a genomic resource announcement reporting draft whole-genome sequences for three multidrug-resistant Shigella clinical isolates. No inferential statistical tests were performed; the paper presents per-isolate descriptive assembly metrics (contig count, total length, N50, GC content, coverage) and qualitative/threshold-based bioinformatics outputs (in silico MLST sequence typing, AMR gene identification at >99% nucleotide identity). Antimicrobial susceptibility was characterized phenotypically by disk diffusion confirmed with broth microdilution per CLSI standards, with results reported as categorical interpretive categories (resistant/intermediate/susceptible).

Replicationunclear Sample sizeThree convenience clinical isolates described individually; no power calculation or sample-size justification applicable to this resource announcement format GroupsNo inferential group comparisons; three individual clinical isolates described in parallel Pairingna Randomization/blindingnot stated Dispersionnone
Statistical tests used
Test Applied to n Assumptions
Disk diffusion (Kirby-Bauer), confirmed by broth microdilution Antimicrobial susceptibility characterization of all three isolates 3 isolates na
In silico multilocus sequence typing (MLST) via ResFinder Sequence type assignment for all three genome assemblies 3 genome assemblies na
AMR gene identification with >99% nucleotide identity threshold (ResFinder) Detection of acquired AMR genes in all three isolates 3 genome assemblies na
Metagenomic classification for contamination screening (Kraken) Sequence integrity evaluation of all three assemblies 3 sequencing runs na
Approaches that could also have been used
  • Short-read Illumina sequencing produced draft assemblies of 356–384 contigs per isolate
    Could also: Long-read sequencing (e.g., Oxford Nanopore MinION, PacBio) alone or in a hybrid short+long-read assembly could also have been used — Long-read or hybrid assemblies can yield closed or near-closed chromosomes and resolve repetitive regions, which may capture mobile genetic elements (plasmids, integrons) carrying AMR genes that are fragmented across contigs in short-read-only assemblies
  • Sequence typing was performed with 7-locus MLST via ResFinder
    Could also: Core-genome MLST (cgMLST) or whole-genome SNP-based phylogenetic analysis could also have been applied — cgMLST and SNP phylogenetics provide higher discriminatory resolution than 7-locus MLST, which can be valuable for contextualizing isolates within global Shigella population structure and identifying outbreak clusters
  • AMR genes were identified using a single tool (ResFinder) with a >99% nucleotide identity threshold
    Could also: Additional tools such as AMRFinderPlus (NCBI), CARD/RGI, or ARIBA could also have been applied, with varied identity/coverage thresholds — Different databases cover different gene families and allelic variants; using complementary tools or a relaxed identity threshold can improve sensitivity for detecting divergent AMR gene homologs and provide cross-validation of findings
  • Quality control used Kraken for contamination screening and coverage/Q-score thresholds
    Could also: Tools such as FastQC/MultiQC for read-level QC and CheckM for assembly completeness/contamination estimation could also have been used — Complementary QC layers (per-base quality profiles, adapter content, genome completeness percentages) provide additional characterization of assembly quality that can help readers assess data reliability
  • The three isolates are described individually without comparative phylogenetic placement
    Could also: A maximum-likelihood or Bayesian phylogenetic tree incorporating publicly available reference genomes from global collections could also have been constructed — Phylogenetic contextualization would allow readers to assess whether these Pakistani isolates belong to recognized international lineages of MDR Shigella, strengthening the epidemiological interpretation of the AMR gene profiles reported
Software: Shovill 0.9 · GalaxyTrakr · Kraken · ResFinder (MLST + AMR gene identification) · NCBI Prokaryotic Genome Annotation Pipeline (PGAP) · Vitek 2 (bioMérieux) · API 20E (bioMérieux)

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
2
Impact: low
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

PRJNA342326 BioProject in Table (http://semanticscience.org/resource/SIO_000419)
no other assessed paper uses this yet
SAMN10086663 BioSamples in Table (http://semanticscience.org/resource/SIO_000419)
no other assessed paper uses this yet
SAMN10086664 BioSamples in Table (http://semanticscience.org/resource/SIO_000419)
no other assessed paper uses this yet
SAMN10086692 BioSamples in Table (http://semanticscience.org/resource/SIO_000419)
no other assessed paper uses this yet
SRR8836950 ENA in Table (http://semanticscience.org/resource/SIO_000419)
no other assessed paper uses this yet
SRR8836963 ENA in Table (http://semanticscience.org/resource/SIO_000419)
no other assessed paper uses this yet
SRR8837012 ENA in Table (http://semanticscience.org/resource/SIO_000419)
no other assessed paper uses this yet

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-31346012

Paper: Draft Genome Sequences of Antimicrobial-Resistant Shigella Clinical Isolates from Pakistan. Lomonaco et al., Microbiol Resour Announc 2019. PMID 31346012 · PMCID PMC6658682 · DOI 10.1128/mra.00500-19.

This is a Microbiology Resource Announcement (MRA): it reports draft genome assemblies of 3 multidrug-resistant Shigella isolates and a small assembly-stats table (Table 1) plus typing/AMR-gene results derived from standard bioinformatic tools. There are no figures or large analyses — Table 1 + the typing sentences ARE the reproducible computational output.

Pipeline described in Methods (verbatim essentials)

  • De novo genome assemblies created with Shovill 0.9 (https://github.com/tseemann/shovill), via GalaxyTrakr. "trim reads" option selected; minimum contig length set to 500 bp; default parameters otherwise.
  • Read QC gates: min avg coverage >50×, R1/R2 Q scores >26 (a sequencing gate, not a recomputable pipeline output).
  • Contamination checked with Kraken (QC; no numeric claim → not a primary target).
  • Annotation by NCBI PGAP (external NCBI service → OUT of scope).
  • ResFinder used for in-silico MLST (sequence types) and AMR-gene detection at

    99% identity.

Data

  • BioProject PRJNA342326; 3 Illumina MiSeq 2×250 paired-end runs:
    • SRR8836963 — CFSAN059650 — S. flexneri — SAMN10086692
    • SRR8836950 — CFSAN059651 — S. flexneri — SAMN10086664
    • SRR8837012 — CFSAN059652 — S. sonnei — SAMN10086663
  • Public on ENA/SRA (fastq.gz available), no restriction.

In scope (pipeline-derived, attempted)

For each of the 3 isolates, from the Shovill 0.9 assembly (Table 1):

  1. No. of contigs (minlen 500)
  2. Total length (bp)
  3. N50 (bp)
  4. GC content (%) Plus, derivable from the runs/metadata:
  5. Genome coverage (×) = total bases / genome size (sanity: matches 52/69/71 from ENA metadata)
  6. Coverage read count = reads used (≈ 2× ENA read pairs) Typing (third-party tools on the paper's data — equally valid per P16):
  7. MLST sequence type (ST245 / ST245 / ST152) via mlst (PubMLST E. coli Achtman scheme; ResFinder-equivalent)
  8. No. of AMR genes at >99% identity (10 / 8 / 7) via ResFinder/abricate
  9. Shared AMR genes (dfrA1, sul2, a CTX-M gene, a tet gene) — qualitative

Out of scope (not attempted)

  • Wet-lab: culture, susceptibility testing (API 20E, Vitek, broth microdilution), DNA extraction, library prep, sequencing. (Reported phenotypes, not pipeline output.)
  • NCBI PGAP gene annotation (external NCBI service, not locally reproducible).
  • WGS/BioSample accession numbers (registry IDs, not computed).

Tooling plan («our HPC» SLURM, conda envs on «infra»)

  • env_asm: shovill=0.9, sra-tools — download fastq (ENA) + assemble.
  • env_typing: mlst, abricate (bundles resfinder DB), assembly-stats, seqkit.
  • Caveat: assembly stats depend on the exact SPAdes version pulled by shovill 0.9 and thread count → expect small (within-tol) differences in contig count/N50, not bit-identical output. AMR-gene counts depend on ResFinder DB version (drift likely).
Figures / tables: Table
contigs_CFSAN059650
Reported
356
Reproduced
356
exact
contigs_CFSAN059651
Reported
359
Reproduced
356
within tolerance
contigs_CFSAN059652
Reported
384
Reproduced
382
within tolerance
length_CFSAN059650
Reported
4629951
Reproduced
4628510
within tolerance
length_CFSAN059651
Reported
4620515
Reproduced
4619421
within tolerance
length_CFSAN059652
Reported
4528736
Reproduced
4528916
within tolerance
n50_CFSAN059650
Reported
33749
Reproduced
33749
exact
n50_CFSAN059651
Reported
31916
Reproduced
31916
exact
n50_CFSAN059652
Reported
25093
Reproduced
25093
exact
gc_CFSAN059650
Reported
50.55
Reproduced
50.55
exact
gc_CFSAN059651
Reported
50.45
Reproduced
50.45
exact
gc_CFSAN059652
Reported
50.85
Reproduced
50.85
exact
coverage_CFSAN059650
Reported
52
Reproduced
52.0
exact
coverage_CFSAN059651
Reported
69
Reproduced
69.9
within tolerance
coverage_CFSAN059652
Reported
71
Reproduced
72.8
within tolerance
readcount_CFSAN059650
Reported
1021504
Reproduced
1021518
within tolerance
readcount_CFSAN059651
Reported
1373334
Reproduced
1373348
within tolerance
readcount_CFSAN059652
Reported
1388868
Reproduced
1388882
within tolerance
mlst_CFSAN059650
Reported
ST245
Reproduced
ST245
exact
mlst_CFSAN059651
Reported
ST245
Reproduced
ST245
exact
mlst_CFSAN059652
Reported
ST152
Reproduced
ST152
exact
amrcount_CFSAN059650
Reported
10
Reproduced
10
exact
amrcount_CFSAN059651
Reported
8
Reproduced
9
partial
amrcount_CFSAN059652
Reported
7
Reproduced
7
exact
amr_shared
Reported
dfrA1, sul2, CTX-M, tet in all 3
Reproduced
dfrA1+sul2+blaCTX-M+tet present in all 3
exact

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 92/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟢3. Location of the main deviation
🟢4. Cause of the deviation
🟢5. Derivability / plausibility
🟢6. Severity of the deviation
🟢7. Core claim
🟢8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Every question reproduced
-1 pts
From: “every question reproduced”
Total score -7

This MRA genome announcement reproduces essentially 1:1 from the deposited SRA data: all three N50, GC% and MLST sequence types are bit-exact, total lengths agree within 0.03%, and the shared AMR-gene claim holds. The only deviations — contig counts off by ≤3 and one AMR count off by one — are attributable to technical tool/DB version drift (SPAdes 4.3.0 vs ~3.12; newer ResFinder DB), not to authors or data availability. No fabrication concern: the near-exact agreement is strong positive evidence the Table 1 values are genuine pipeline output.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

275.7 k
tokens (I/O) · 65.9 M incl. cache
148 min
runtime · 1.45 CPU-h
10.3 GB
peak RAM
2
HPC jobs
hummel
machine