Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Prediction of Antibiotic Susceptibility Profiles of Vibrio cholerae Isolates From Whole Genome Illumina and Nanopore Sequencing Data: CholerAegon.

Front Microbiol · 2022
L1 86/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +1
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
  • No authors-side cause for any deviation
  • Reported values are derivable from the shared data
  • Any deviation was negligible
  • The central claim held under reproduction
What did not (or only partly)
  • 🟡A deviation arose in the data or preprocessing
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
86/100
Reproducibility score
0.7 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 70% of all assessed papers rank 334 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

STRONG 1:1 reproduction of the core results (prior run, outputs preserved here). Ran the authors' CholerAegon pipeline (github.com/RaverJay/CholerAegon @ d4ff04e) on the paper's own ENA data (PRJEB51675, 82 V. cholerae dual-platform WGS isolates, N matches exactly) on «our HPC» via a version-matched conda decomposition. ALL 11 paper AMR genes present 82/82 (100%); WHO ciprofloxacin prediction exact 82/82; drug-class exact 82/82; hybrid assembly mean length within 0.018%, median contigs exact (2), median ANI exact. C5 genotype-phenotype concordance reproduces the paper Table exactly (574/574 prediction agreement). Only systematic delta: abricate's newer bundled CARD db adds almE+almF (full almEFG operon) -> a downstream colistin prediction; no paper gene/drug dropped. NOW continuing on the harder 20%: C6 read-downsampling time-to-detection (re-running, «infra» was reclaimed). C4 (Illumina/ONT-only assembly stats) is heavier and supplementary. All grades PROVISIONAL pending human audit.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 90
    assessed: 2026-06-22 ⛓ 4ed516c1629a
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-26
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-22
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

Whole genome sequencing (short-read Illumina and long-read Oxford Nanopore MinION) can be used to predict antimicrobial resistance (AMR) profiles of Vibrio cholerae isolates as a substitute for phenotypic antibiotic susceptibility testing.

Core claims
  • CholerAegon, a Nextflow-based pipeline, predicts AMR profiles of V. cholerae from assembled genomes using CARD ontology method
  • In silico AMR prediction can replace in vitro susceptibility testing for five of seven tested antibiotics finding
  • Nanopore (ONT MinION) sequencing produces more contiguous/complete V. cholerae assemblies than Illumina alone, while Illumina achieves higher sequence identity to the reference finding
  • Hybrid assembly (long-read-first Flye+Pilon approach) combines contiguity of long reads with low error rate of short reads and outperforms Unicycler (short-reads-first hybrid assembler) in runtime and ANI method
  • CholerAegon combines Abricate and RGI results to detect more AMR genes than other tools such as AMRFinderPlus or Resfinder method
  • MinION sequencing, due to low cost, real-time analysis capability, and portability, is well suited for pathogen genomic surveillance in low-resource settings finding
Experimental setups
Assay System Perturbation Readout Platform
Whole genome sequencing (short-read) Vibrio cholerae isolates (n=82) from Ghana cholera outbreaks 2011/2012/2014 none genome assembly, AMR gene presence Illumina NextSeq 500/550 (Nextera XT Library Prep Kit)
Whole genome sequencing (long-read) Vibrio cholerae isolates (n=82) none genome assembly, AMR gene presence Oxford Nanopore Technologies MinION, SpotON Flow Cell R9.4.1
Hybrid genome assembly and polishing Vibrio cholerae isolates (n=82) none assembly length, contig number, N50, ANI to reference Flye, Medaka, Pilon (bioinformatics pipeline)
In silico AMR gene detection Assembled V. cholerae genomes none presence of resistance genes/variants, predicted resistance profile RGI and Abricate against CARD database
Phenotypic antibiotic susceptibility testing (Kirby-Bauer disk diffusion) Vibrio cholerae isolates (n=80) exposure to ampicillin, chloramphenicol, gentamicin, nalidixic acid, sulfamethoxazole/trimethoprim, tetracycline, ciprofloxacin susceptible/intermediate/resistant/susceptible dose-dependent classification CLSI 2015 / EUCAST 2015 breakpoints
Average nucleotide identity comparison Assembled genomes vs V. cholerae O1 biovar El Tor str. N16961 (NC_002505.1) none ANI percentage FastANI v1.32
Key results
  • In silico prediction can replace in vitro susceptibility testing for 5 of 7 antibiotics tested
  • Illumina sequencing yielded ~7 million reads of 73 nt length per isolate (throughput 522 Mb, 111X coverage) 111X
  • ONT MinION sequencing yielded on average 94,000 reads of ~5.8 kb length (throughput 920 Mb, 209X coverage) 209X
  • Illumina-only average assembly length was ~4,042 Mb versus ~4,107 Mb for long-read assemblies
  • Nanopore-based assemblies achieved complete contiguity of both V. cholerae chromosomes (~3 Mb and ~1 Mb) for all isolates, with erroneous fusion in two strains
  • Illumina assemblies were highly fragmented but achieved higher genome identity to the reference sequence than long-read assemblies
  • Long-reads-first hybrid approach (Flye+Pilon) outperformed Unicycler (short-reads-first hybrid assembler) in runtime and ANI
Key statistics
  • count 82 (number of V. cholerae isolates whole-genome sequenced)
  • count 7,055,057 (average number of Illumina reads per isolate)
  • count 93,778 (average number of ONT reads per isolate)
  • other 111X (Illumina sequencing coverage depth)
  • other 209X (ONT MinION sequencing coverage depth)
  • mean ~4,042 Mb (average Illumina-only assembly length)
  • mean ~4,107 Mb (average long-read/hybrid assembly length)
  • other 80% identity, 80% coverage (CholerAegon cutoff thresholds for filtering false-positive resistance gene hits)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The study compares in silico–predicted antimicrobial resistance (AMR) profiles, derived from whole-genome assemblies of 82 Vibrio cholerae isolates (Illumina, ONT, and hybrid assemblies), against phenotypic Kirby-Bauer disk-diffusion antibiotic susceptibility testing (AST) results for the same isolates across seven antibiotics. Agreement was assessed by classifying each isolate/antibiotic combination as a correct or false prediction (via custom Python scripts) rather than through inferential hypothesis testing. Sequencing and assembly quality metrics (read counts, lengths, throughput, assembly length, N50, ANI) were reported as averages describing the two assembly strategies and their hybrid combination.

Replicationbiological Sample size82 V. cholerae isolates whole-genome sequenced; phenotypic AST performed for 80 of these; no formal sample-size or power calculation described Groupsin silico–predicted resistance profile vs. phenotypic AST result, per isolate per antibiotic (7 antibiotics); also hybrid vs. short-read-only vs. long-read-only assemblies, and hybrid approach vs. Unicycler assembler Pairingpaired Randomization/blindingnot stated Dispersionnone Exact p-valuesno Effect sizesno Confidence intervalsno Multiplicity correctionnone stated
Statistical tests used
Test Applied to n Assumptions
not stated — descriptive concordance classification (correct vs. false prediction) between predicted and phenotypic resistance status comparison of in silico AMR predictions vs. Kirby-Bauer AST results across 7 antibiotics 82 sequenced isolates; phenotypic AST available for 80 isolates na
not stated — runtime and average nucleotide identity (ANI) comparison comparison of the hybrid assembly approach (Flye+Pilon) vs. Unicycler assembler na
Approaches that could also have been used
  • Agreement between predicted and phenotypic resistance was reported as raw counts of correct vs. false predictions per antibiotic.
    Could also: Standard diagnostic-agreement metrics such as sensitivity, specificity, positive/negative predictive value, and Cohen's kappa — These metrics are widely used for genotype-phenotype AMR concordance studies and allow direct comparison with other prediction tools or studies using the same standardized measures.
  • Per-antibiotic prediction accuracy was reported without an accompanying measure of precision.
    Could also: A 95% confidence interval around each proportion of correct predictions — With a moderate and antibiotic-specific sample size (up to 80 isolates), a CI would convey how precisely the observed concordance rate reflects the likely true rate.
  • Sequencing and assembly statistics (read length, throughput, assembly length, N50, ANI) were summarized as single average values per sequencing method.
    Could also: Reporting SD, IQR, or range alongside the means — The text itself notes read length and throughput varied by about an order of magnitude across isolates, so a dispersion measure would help convey this isolate-to-isolate variability alongside the averages.
  • The hybrid assembly approach was compared to Unicycler using runtime and ANI values without a formal statistical test.
    Could also: A paired test such as a paired t-test or Wilcoxon signed-rank test across the matched isolates — Since both assemblers were run on the same set of isolates, a paired test could formally characterize whether differences in ANI or runtime were consistent across isolates rather than comparing only summary values.
  • Two AMR gene-detection tools (Abricate and RGI) were combined and compared descriptively against other tools like AMRFinderPlus and Resfinder in a supplementary table.
    Could also: A formal statistical comparison of gene-detection sensitivity across tools (e.g., McNemar's test for paired detection outcomes) — McNemar's test is suited to paired binary detection/no-detection outcomes from different tools applied to the same isolates, and could complement the descriptive tool comparison.
Software: CholerAegon (custom Nextflow/Python3 pipeline) · custom Python scripts (for comparing predicted vs. phenotypic AST) · FastANI 1.32 · fastqc · pycoQC 2.5.0.3

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-35814690 (CholerAegon)

Paper: Fuesslin et al. 2022, Front Microbiol 13:909692. "Prediction of Antibiotic Susceptibility Profiles of Vibrio cholerae Isolates From Whole Genome Illumina and Nanopore Sequencing Data: CholerAegon." DOI 10.3389/fmicb.2022.909692 · PMC9257098.

Code: https://github.com/RaverJay/CholerAegon (GPL-3.0, public, default branch main). A Nextflow (DSL2) pipeline. P16 note: authors' own repo (not third-party), but either would be equally valid. The repo also ships its published result files under paper/results/ and the exact figure/table-generating scripts under paper/scripts/ — these are the 1:1 comparison targets.

Data: ENA project PRJEB51675 — 82 V. cholerae isolates, each with paired-end Illumina (NextSeq 500, WGS) and Oxford Nanopore (MinION, WGS) reads = 164 runs, ~97 GB. Open access. Isolate aliases Iso02501..Iso02597 map 1:1 to the paper's result rows.

Pipeline (what produces each result)

CholerAegon.nf per isolate, default profile (do_all_assemblies off → AMR on hybrid only):

  1. Long-read assembly: flye 2.9 (--nano-raw --plasmids) → medaka 1.5.0 polish (model r941_min_sup_g507).
  2. Hybrid polish: bwa-mem2 2.2.1 maps Illumina reads (fastp-trimmed) to the medaka assembly → samtools 1.14 → pilon 1.24*_hybrid_assembly.fasta.
  3. AMR detection: RGI 5.2.1 (rgi main --input_type contig --local, bundled CARD localDB) + Abricate 1.0.1 (--db card) → combine_results.py merges & filters at ≥80 % coverage and ≥80 % identity → aggregate_combined_results.py.
  4. Resistance prediction: predict_drug_resistance.py walks the CARD ARO ontology (data/CARD/aro.obo) gene→drug/drug-class edges → drug_resistance_prediction.tsv (headline result; 82 hybrid rows).
  5. ANI / assembly stats: fastANI 1.32 vs bundled reference V_cholerae_O1_biovar_El_Tor_str_N16961.fa; read/assembly stats via paper/scripts/*.

Tool versions are pinned in nextflow.config containers; note the paper text cites spades 3.15.2 / medaka 1.4.4 while the containers pin spades 3.15.3 / medaka 1.5.0 — the shipped result CSVs were produced with the container versions, so those are the reproduction target (delta recorded).

IN SCOPE (pipeline-derived, attempted)

# Result Paper location Pipeline Gold-standard file in repo
C1 Per-isolate AMR gene presence + predicted drug resistances (10–11 genes; SxT/quinolone/phenicol/sulfonamide/etc.) Sec 3.3, Fig/Table; Abstract full pipeline → predict_drug_resistance.py paper/results/drug_resistance_prediction.tsv (82 rows)
C2 Set of 10 resistance genes identified across the panel Sec 3.3 RGI+Abricate+CARD drug_resistance_prediction.tsv columns
C3 Hybrid assembly stats: mean length 4,106,391 nt, median 2 contigs, mean ANI 99.97385 % Table 1 flye+medaka+pilon, fastANI paper/results/table_stats_overall.csv, general_stats.txt
C4 Illumina & ONT assembly stats (Table 1, needs --do_all_assemblies) Table 1 spades / flye+medaka table_stats_overall.csv
C5 Genotype↔phenotype concordance per antibiotic (e.g. SxT 75/80 correct, CN 80/80, TE 80/80; "replace AST for 5 of 7 antibiotics") Sec 3.4, Abstract genotype from C1 vs shipped phenotype table, paper/scripts/compare_with_madrid_combined.py paper/results/comparison_madrid_result.csv
C6 (hard 20%) Downsampling / time-to-detection: "4 % of reads (8.36X) sufficient to detect all 10 AMR genes" Sec 3.5 paper/scripts/downsampled_multi.py, make_timeline_samples.py genes_found_subsampled_multi*.pdf, segment_completeness*.tsv

OUT OF SCOPE (not pipeline-derived → not attempted)

  • Phenotypic AST (disk diffusion / MIC at the Madrid reference lab) — wet-lab; the values are an input to C5, not regenerable. C5 reproduces only the comparison.
  • DNA extraction & sequencing (wet-lab).
Figures / tables: Tabletable_stats_overall
C1
Reported
82 hybrid isolates; AMR panel {APH(3'')-Ib,APH(6)-Id,CRP,E.coli parE,V.cholerae varG,almG,catB9,dfrA1,floR,rsmA,sul2} (11 cols); NUM_FOUND dist 74x11/5x7/3x4
Reproduced
82 isolates; all 11 paper genes present in 82/82 (902/902 cells = 100%); +almE+almF in every isolate (abricate 2025-Jan CARD db resolves full almEFG operon); dist 74x13/5x9/3x6
within tolerance
C1b
Reported
Predicted WHO=ciprofloxacin (via parE); per-isolate drug-class resistances
Reproduced
WHO_SUGGESTED 82/82 EXACT; DRUG_CLASS 82/82 EXACT; DRUG_RESISTANCES adds only colistin A/B (downstream of complete alm operon), nothing dropped
within tolerance
C2
Reported
11 AMR gene columns
Reproduced
13 columns = 11 paper genes (all present) + almE + almF (CARD db version)
within tolerance
C3a
Reported
hybrid mean assembly length 4,106,391.52 nt
Reproduced
4,107,118.88 nt (diff 0.0177%)
within tolerance
C3b
Reported
hybrid median fragments = 2 contigs
Reproduced
2 contigs
exact
C3c
Reported
hybrid 'mean_ANI' vs N16961 = 99.97385% (label is actually MEDIAN of per-isolate ANI)
Reproduced
median ANI 99.9739% (diff 5e-5%); true mean 99.90321% also matches gold true mean 99.90207%
exact
C5
Reported
genotype<->phenotype concordance vs Madrid AST: SxT 75/80, TE 80/80, CN 20/20, AMP 20/20, NA 63/80, C 3/80, CIP 0/18 (replace AST for 5/7 antibiotics)
Reproduced
OUR pipeline predictions match gold/paper predictions 574/574 (100%); per-antibiotic concordance with phenotype identical to paper Table in every antibiotic (SxT 75/80, TE 80/80, CN 20/20, ...)
exact
C6
Reported
downsampling Iso02507: ~4% of reads (8.36X) sufficient to detect all 10 AMR genes; assembly completeness rises with read fraction
Reproduced
IN PROGRESS (re-running downsampling pipeline on «our HPC»)
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 86/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟡3. Location of the main deviation
🟢4. Cause of the deviation
🟢5. Derivability / plausibility
🟢6. Severity of the deviation
🟢7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +1

Strong 1:1 reproduction: the authors' CholerAegon pipeline run on the paper's own ENA data (PRJEB51675, 82/82 isolates) reproduces all 11 AMR genes in 100% of isolates, the ciprofloxacin/WHO prediction exactly in 82/82, and assembly mean length (0.018%), median contigs (2), and median ANI (99.974%) within tolerance. The only deviation is on the technical/expected side: abricate's newer 2025 CARD database resolves the full almEFG operon, adding almE+almF and a downstream colistin call — a strict superset that contradicts no reported value. Judged yellow overall only because of this explainable database-version drift; the central conclusion is fully confirmed and nothing is on the authors' side.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

463.1 k
tokens (I/O) · 53.9 M incl. cache
164 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.