Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Comprehensive Genomic and Phenotypic Characterization of Escherichia coli O78:H9 Strain HPVN24 Isolated from Diarrheic Poultry in Vietnam.

Microorganisms · 2025
L1 86/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +2
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
  • Reported values are derivable from the shared data
  • Any deviation was negligible
  • The central claim held under reproduction
What did not (or only partly)
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
86/100
Reproducibility score
0.7 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 70% of all assessed papers rank 334 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough and reproduces 1:1 on the deterministic typing/assembly-metric outputs. Ran QUAST/mlst/ECTyper/ABRicate on the paper's own data (deposited assembly GCA_052827245.1 + raw reads SRR34216250, SPAdes) via «our HPC» «job». EXACT: N50 50829, L50 29, contigs>200bp 495, contigs>=1000bp 222, MLST ST23, serotype O78:H9, AMR dfrA1 & blaEC-13. Within-tol: GC 50.58 vs 50.57, genome size (diff = the 217 sub-200bp contigs NCBI drops; 712->495). Partial: 3/4 plasmid replicons (pSE11 absent), virulence (major iron/adhesion/curli systems present; iuc/hlyE/iss/bcs/gfc not in VFDB hits), gyrA/parC point mutations (ABRicate is acquired-gene BLAST, no SNP caller). Internal consistency strong -> no fabrication signal. NOT attempted (80/20): BUSCO 432/440, Prokka/Bakta CDS counts, CARD-RGI point mutations, phylogenetics/pangenome/COG/KEGG figures, MIC phenotypes (wet-lab/manual/version-sensitive).

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 86
    assessed: 2026-06-15 ⛓ 3eb528b66c38
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-15
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-16
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

This study aimed to comprehensively characterize the phenotypic features, antibiotic susceptibility profile, and whole-genome sequence of avian pathogenic Escherichia coli strain HPVN24 isolated from diarrheic broiler chickens in Hai Phong, Vietnam, to address the gap in genomic characterization of highly virulent, multidrug-resistant O78 APEC strains circulating in Vietnam.

Core claims
  • HPVN24 is an avian pathogenic E. coli serotype O78:H9, sequence type ST23, with a 5.05 Mb genome and 50.57% GC content. finding
  • HPVN24 is a multidrug-resistant strain, exhibiting resistance to trimethoprim, ampicillin, and ciprofloxacin with predicted resistance to 18 antibiotic classes and particularly strong fluoroquinolone resistance. finding
  • HPVN24 is highly virulent, displaying β-hemolytic activity and harboring a broad repertoire of virulence genes for adhesion, iron acquisition, hemolysin production, and stress response. finding
  • HPVN24 clusters phylogenetically with O78:H9 strains from poultry in other regions, suggesting potential cross-population transmission. finding
  • The genome encodes multiple secretion systems (T2SS most complete, T4SS ~75%, T1SS ~65%, T6SS least complete) plus Sec-SRP and Tat pathways. mechanism
  • Whole-genome sequencing combined with phenotypic assays and antibiotic susceptibility profiling characterizes the APEC strain. method
  • HPVN24 encodes multiple iron acquisition systems including enterobactin, salmochelin, aerobactin, and yersiniabactin. mechanism
Experimental setups
Assay System Perturbation Readout Platform
Strain identification by MALDI-TOF mass spectrometry E. coli isolates from diarrheic broiler chickens, Hai Phong, Vietnam none species identification MALDI-TOF Biotyper (Bruker)
Hemolysis assay on blood agar E. coli HPVN24 vs γ-hemolysis reference strain none hemolytic activity (zone of hemolysis) 5% sheep blood agar plate
Antibiotic susceptibility / MIC testing E. coli HPVN24 antibiotic exposure (CI, TC, DC, TS, AM) inhibition zone and MIC breakpoints E-test strips (bioMérieux); MHA plates; CLSI M100 2023 / EUCAST 2023
Whole-genome sequencing E. coli HPVN24 genomic DNA none genome sequence reads Illumina HiSeq 3000; GeneJET Genomic DNA Purification Kit; NanoDrop Lite
De novo genome assembly and quality assessment HPVN24 sequencing reads none contigs, N50/L50, completeness (QUAST, BUSCO) SPAdes, Ragtag, QUAST v5.3.0, BUSCO v5.8.0 (Enterobacterales_odb10)
MLST and serotyping HPVN24 genome none sequence type and O/H antigen serotype SeroTypeFinder v2.0, ChTyper v1.0, ECTyper; Center for Genomic Epidemiology
Virulence and antimicrobial resistance gene detection HPVN24 genome none virulence and resistance gene hits (≥90% identity) VirulenceFinder v2.0, ABRicate v1.0.1, CARD-RGI, ResFinder, VFDB
Comparative genomics and phylogenetic analysis (ANI, core-genome ML tree) 22 E. coli strains incl. HPVN24 and outgroup CFT073 none ANI values, phylogenetic clustering FastANI/ANIclustermap v2.0.1, progressiveMauve v2.4.0, RAxML-ng, iTOL
Key results
  • HPVN24 showed the strongest hemolytic activity (β-hemolysis) among hemolysin-producing isolates
  • No inhibition zones around trimethoprim and ampicillin (0 mm); minimal around ciprofloxacin (1 mm) and tetracycline (1 ± 0.5 mm); largest with doxycycline (10 ± 0.5 mm) 0–10 mm
  • Draft genome assembled at 5,053,087 bp with 50.57% GC content 5.05 Mb
  • BUSCO identified 432/440 complete genes (431 single-copy, 1 duplicate), 7 fragmented, 1 missing 432/440
  • Strain identified as serotype O78:H9 (wzx allele 3, wzy 6, fliC 1) and sequence type ST23
  • Genome predicted resistance to 18 antibiotic classes with particularly strong fluoroquinolone resistance 18 classes
  • Most abundant COG categories were carbohydrate transport/metabolism (~400 genes) and amino acid transport/metabolism (~370 genes) ~400 and ~370 genes
  • HPVN24 encodes complete Type I fimbrial operon (fimH allele 35) and multiple iron acquisition systems (enterobactin, salmochelin, aerobactin, yersiniabactin)
Key statistics
  • count 5,053,087 bp genome length (draft assembly total length)
  • other 50.57% GC content (genome GC content)
  • count 432/440 complete BUSCO genes (core single-copy ortholog completeness)
  • other N50/L50 = 50829/29; 10X coverage (assembly contiguity)
  • count 712 contigs >=0 bp; 495 >200 bp; 222 >=1000 bp (assembly contig counts)
  • other MIC: CI 12 µg/mL, TC 96 µg/mL, DC 11 µg/mL (minimum inhibitory concentration breakpoints)
  • count Prokka 4728 features (4646 CDS, 77 tRNAs, 4 rRNAs); Bakta 5022 features (4636 CDS, 95 tRNAs) (genome annotation)
  • other 99.52% reads mapped; 95.43% K-mer compliance; duplication ratio 1.002 (assembly quality metrics)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This is a single-strain genomic characterization study with no formal inferential statistics; the approach combines phenotypic assays (hemolysis on blood agar, MIC determination via E-test) with whole-genome sequencing and a suite of bioinformatics pipelines for assembly, annotation, and comparative genomics. Antibiotic resistance was classified by CLSI M100 2023 breakpoints, and phylogenetic relationships were inferred by maximum likelihood (RAxML-ng) on a core-genome alignment across a curated 22-strain panel. All results are reported descriptively as gene counts, BUSCO completeness fractions, KEGG pathway completeness scores, and MIC values; no p-values, confidence intervals, or formal effect sizes are presented.

Replicationunclear Sample sizeSingle focal strain (HPVN24); comparative genomics panel of 22 strains selected from NCBI; no power calculation or sample size justification described GroupsSingle strain HPVN24 characterized phenotypically and genomically; comparative panel spans O78:H9 and related APEC serotypes across Asia, Southeast Asia, Europe, the Americas, and Australia Pairingna Randomization/blindingnot stated Dispersionunclear Multiplicity correctionnone stated
Statistical tests used
Test Applied to n Assumptions
MIC determination via E-test with CLSI M100 2023 breakpoint classification (EUCAST 2023 as secondary verification) Antibiotic susceptibility profiling of five agents: ciprofloxacin, tetracycline, doxycycline, trimethoprim, ampicillin 1 strain (HPVN24) not stated
BUSCO completeness assessment (Enterobacterales_odb10, Prodigal gene prediction) Genome assembly quality evaluation 1 genome assembly na
Average Nucleotide Identity (ANI) analysis via FastANI Comparative genomics across 22 E. coli strains 22 strains not stated
Maximum likelihood phylogenetic tree construction (RAxML-ng) on core-genome alignment (progressiveMauve) Phylogenetic analysis of HPVN24 relative to 21 comparator strains plus outgroup CFT073 22 strains (plus 1 outgroup) not stated
Gene identity threshold filtering (≥90% identity cutoff) Virulence and resistance gene detection via VirulenceFinder, ABRicate (VFDB and ResFinder), and CARD-RGI 1 genome not stated
Pan-genome analysis (BPGA v1.3 with Usearch 50% cutoff; Roary v3.13.0 with 80% BLAST+ cutoff) Core/soft-core/shell/cloud gene distribution and new gene accumulation across studied strains not stated explicitly for BPGA; same comparative panel for Roary not stated
Approaches that could also have been used
  • Short-read Illumina HiSeq 3000 sequencing was used, yielding a draft assembly of 712 contigs
    Could also: Long-read sequencing (Oxford Nanopore or PacBio) or a hybrid short-read + long-read assembly strategy could also be used — Long-read approaches can span repetitive elements and resolve complete plasmid sequences and genomic islands into a closed chromosome, which would allow more definitive characterization of mobile genetic elements carrying virulence and resistance determinants
  • Phylogenetic inference was performed with maximum likelihood (RAxML-ng) without reported branch support values or substitution model details in the main text
    Could also: Bootstrap resampling (e.g., 1000 replicates reported on branches) or Bayesian inference (e.g., MrBayes) could also be applied — Reporting branch support values allows readers to assess topological confidence directly from the figure; Bayesian approaches additionally provide posterior probabilities and naturally incorporate substitution model uncertainty
  • Pan-genome analysis was performed with two tools using different identity cutoffs (BPGA at 50%, Roary at 80%), without a stated rationale for the differing thresholds
    Could also: A single consistently justified cutoff, or a sensitivity analysis across identity thresholds, could also be reported — Different cutoffs can yield substantially different core vs. accessory genome boundaries; documenting the rationale for each threshold or showing threshold sensitivity helps readers evaluate how cutoff choice influences pan-genome size and gene-sharing estimates
  • Inhibition zone diameters were reported with ± values (e.g., '1 ± 0.5 mm') without specifying whether these represent SD, SEM, or range, and without stating the number of technical replicates
    Could also: Explicit statement of the dispersion measure type and replicate number (e.g., 'mean ± SD, n = 3 independent readings') could also be included — Specifying the dispersion measure and replicate count allows readers to interpret measurement variability; for E-test MIC determinations, even a single determination is common practice but stating it explicitly supports reproducibility assessment
  • The 22-strain comparative panel was assembled by opportunistic selection from NCBI with geographic diversity as the criterion
    Could also: A systematically sampled panel — or a recombination-aware phylogenetic method (e.g., ClonalFrameML, gubbins) — could also be applied — Recombination-aware approaches distinguish vertical descent from horizontal gene transfer, which is particularly relevant for interpreting whether observed clustering of MDR O78:H9 strains reflects clonal spread versus convergent acquisition of mobile resistance elements
Software: R / RStudio R v4.5, RStudio v2024.12.1+563 · Falco v1.2.4+galaxy0 · fastp v0.24.0+galaxy4 · SPAdes · Ragtag v2.1.0+galaxy1 · QUAST v5.3.0+galaxy0 · BUSCO v5.8.0+galaxy1 · Roary v3.13.0+galaxy3 · progressiveMauve v2.4.0 · RAxML-ng · FastANI / ANIclustermap ANIclustermap v2.0.1 · BPGA (Bacterial Pan Genome Analysis) v1.3 · Prokka · Bakta · COGclassifier v2.0.0 · EggNog-mapper · KEGGaNOG · VirulenceFinder v2.0 · ABRicate v1.0.1 · CARD-RGI · SeroTypeFinder v2.0 · ChTyper v1.0 · ECTyper · iTOL (Interactive Tree of Life)

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
0
Impact: low
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

GCA_000307205.1 GCA in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

n50
Reported
50829
Reproduced
50829
exact
l50
Reported
29
Reproduced
29
exact
contigs_gt200bp
Reported
495
Reproduced
495
exact
contigs_ge1000bp
Reported
222
Reproduced
222
exact
gc_percent
Reported
50.57
Reproduced
50.58
within tolerance
genome_size_bp
Reported
5053087
Reproduced
5025120 (>=200bp deposited); 5041736 (my SPAdes)
within tolerance
mlst_st
Reported
ST23
Reproduced
ST23 (identical 7-allele profile, both assemblies)
exact
serotype
Reported
O78:H9
Reproduced
O78:H9 (ECTyper, both assemblies)
exact
amr_dfrA1
Reported
dfrA1
Reproduced
dfrA1_10 (ResFinder 100%/99.79%)
exact
amr_blaEC13
Reported
blaEC-13
Reproduced
EC-13 (CARD 99.91%)
exact
plasmid_replicons
Reported
Col156,IncFIB,pSE11,IncI1-Ialpha
Reproduced
Col156+IncFIB+IncI1-Ialpha (3/4)
partial
virulence_genes
Reported
fim,ent,iro,ybt,iuc,hlyE,iss,csg,bcs,gfc
Reproduced
fim,ent,iro,ybt,csg present (VFDB)
partial
amr_point_mutations
Reported
gyrA S83L,D87N;parC S80L
Reproduced
not detected (ABRicate lacks SNP caller)
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 86/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟢5. Derivability / plausibility
🟢6. Severity of the deviation
🟢7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +2

All directly-comparable, deterministic outputs reproduced exactly or within rounding (N50, L50, contig counts, GC%, genome size after standard NCBI filtering, MLST ST23, serotype O78:H9, AMR dfrA1 and blaEC-13), and the central genomic characterization holds from both the deposited assembly and an independent reassembly of the raw reads. The remaining partial matches — 3/4 plasmid replicons, a subset of virulence genes, and undetected gyrA/parC point mutations — sit on our side (different databases, and ABRicate's lack of a SNP caller), not on the authors' or data-availability side. Internal consistency is strong with no fabrication signal; severity is negligible, so the only reason this is not flat green is the genuinely incomplete reproduction of the secondary claims.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

96.9 k
tokens (I/O) · 4.6 M incl. cache
18 min
runtime · 0.09 CPU-h
4.2 GB
peak RAM
1
HPC jobs
hummel
machine