Essential Genes of Vibrio anguillarum and Other Vibrio spp. Guide the Development of New Drugs and Vaccines.
The main results reproduced, with only marginal, non-material deviations.
- ✓Same input data as the authors
- ✓Reported values were directly comparable
- ✓No relevant deviation in data/preprocessing
- ✓No authors-side cause for any deviation
- ✓Any deviation was negligible
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
PARTIAL (described well enough; mostly 1:1, one method-sensitive divergence). Third-party-tool-on-own-data repro: ran the paper's own pipeline (github.com/pseudogene/vibrio-tnseq @ 43d9241) on the paper's own ENA data (PRJEB39186, 12 MiSeq runs, pooled) on «our HPC»/«infra» -> cutadapt 3.7 -> bowtie2 2.3.5.1 (bundled NB10 GCF_000786425.1) -> sam_to_map.pl -> contrib/el-artist.py. READ PROCESSING & MAPPING REPRODUCE ESSENTIALLY EXACTLY: total reads EXACT (5,802,645), QC-retained within 0.006% (4,727,881 vs 4,727,608), unique insertion sites within 0.21% (52,773 vs 52,662). The paper's 329 essential = 25 rRNA + 89 tRNA + 212 protein-coding + 3 other; we reproduce rRNA (25/25 exact), tRNA (88 vs 89), and other RNA near-exactly. The ONE material divergence is protein-coding essential calls from the self-described BETA el-artist HMM port: 272 vs 212 (confident-only 250 vs 212), and domain-essential 82 vs 91 -> total essential 388 vs 329, domain 86 vs 95. Deterministic across 3 reruns (systematic, not stochastic). C4 annotation count within ~0.3%. NOT ATTEMPTED: C8 core-genome pangenome (105 isolates), C9 cross-Vibrio set comparisons (external gene lists), C10 vaccine-candidate annotation (SignalP/SecretomeP/LipoP not in repo) -- all out of pipeline scope. No fabrication indicated: all values derivable from shipped data+code, and 329 reconstructs cleanly from its feature-type components.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 67assessed: 2026-06-21 ⛓ 3e80f82c67da
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-21
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetEssential genes of Vibrio anguillarum, identified via a Tn-seq transposon mutagenesis approach, can serve as targets for new antibiotic drugs and as candidates for subunit vaccine development against vibriosis.
- ★ Tn-seq using the TnSC189 mariner transposon identified 329 essential genes in V. anguillarum NB10Sm from a library of 52,662 insertion mutants. finding
- ★ 34.7% of the essential genes are found within the core genome of V. anguillarum, marking them as strong potential drug targets. finding
- ★ Seven essential gene products are predicted to be membrane-localised or extracellularly released, making them putative vaccine candidates. finding
- ★ Comparison with five other Vibrio essential-gene studies revealed 13 proteins conserved across studies and 25 genes specific to V. anguillarum. finding
- ★ The Tn-seq and analysis pipeline (EL-ARTIST HMM classification, functional/subcellular annotation, comparative genomics) is a generalisable methodology applicable to other pathogens. method
- Ribosome and Sulfur relay system KEGG pathways are significantly enriched for essential genes. mechanism
- Transposases constitute the largest functionally annotated group of essential protein-coding genes (40 of 212, 18.9%). finding
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Transposon-insertion sequencing (Tn-seq) | Vibrio anguillarum NB10Sm (fish pathogen strain) | TnSC189 mariner transposon random insertion mutagenesis | Genome-wide insertion site frequency to classify genes as essential/domain-essential/non-essential | Illumina MiSeq, bowtie2 mapping, EL-ARTIST HMM analysis |
| Conjugative transposon mutagenesis | E. coli SM10λpir (donor) x V. anguillarum NB10Sm (recipient) | Mating/conjugation to transfer pSC189 plasmid carrying TnSC189 | Colony counts on selective media (KAN/STR) to construct three independent mutant libraries | — |
| In silico functional/domain annotation | V. anguillarum NB10Sm essential gene set | none | Functional reclassification of essential/hypothetical/pseudo-genes | InterProScan v5.44-79, KofamKOALA v95.0 |
| Subcellular localisation prediction | V. anguillarum NB10Sm essential gene products | none | Predicted signal peptides, secretion, lipoprotein status | SignalP v5.0, SecretomeP v2.0a, LipoP v1.0 |
| KEGG pathway and protein-protein interaction enrichment analysis | V. anguillarum NB10Sm essential genes vs. whole-genome reference | none | Statistically enriched pathways/interaction networks among essential genes | bioconductor/DOSE v3.10, clusterProfiler v3.14.3, STRING v11.5 |
| Comparative genomics (BlastN orthology search) | V. anguillarum vs. V. cholerae and V. parahaemolyticus essential gene lists from prior studies | none | Orthologous essential genes conserved/unique across Vibrio species | BlastN, jVenn |
| Pangenome/core genome analysis | 105 V. anguillarum genomes | none | Proportion of essential genes present in the core genome (≥95% of genomes) | PIRATE v1.0.4, R v4.0.2 |
- – 329 of 3,774 annotated genes (8.7%) classified as essential; 95 (2.5%) domain-essential 8.7%
- – All 25 rRNA genes and 89 of 93 tRNA genes classified as essential 89/93
- – 52,662 unique transposon insertion locations mapped from 4,727,608 quality-filtered reads (81.5% of 5,802,645 total) 81.5%
- – 3,100,490 reads (65.6%) aligned exactly once to the reference genome; 24.8% unmapped; 9.6% multi-mapped 65.6%
- ▲ 67-kb pJM1-like virulence plasmid had approximately twice the transposon insertions per gene compared to chromosomes ~2-fold
- ▲ Ribosome KEGG pathway significantly enriched for essential genes adjusted P=10-23
- ▲ Sulfur relay system pathway enriched, with 16 essential genes involved in tRNA thiolation and related metabolism 16 genes
- – Transposases were the largest annotated functional group among essential protein-coding genes 40/212 (18.9%)
- count 329 essential genes (Total essential genes identified out of all annotated genes in V. anguillarum NB10Sm)
- count 95 domain-essential genes (Genes with insertions only at sequence ends)
- fold_change 34.7% (Proportion of essential genes within the core genome (105 genomes))
- count 13 conserved proteins (Essential genes conserved across V. anguillarum and 5 other Vibrio studies)
- count 25 genes (Essential genes specific to V. anguillarum, not essential in other Vibrio spp. studies)
- pvalue P<0.001 (adjusted P=10-23 for Ribosome pathway) (KEGG pathway enrichment significance thresholds)
- pvalue ANOVA P-value <10-15 (Difference in regression slope of gene length vs. insertions for 67-kb plasmid vs. chromosomes)
- correlation R2 = 0.61, 0.70, 0.60, 0.58, 0.25 (Correlation between gene length and transposon insertion number for Chromosome I, II, Plasmid 67k, 8k, 6k respectively)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
The study employed Tn-seq with a mariner-based transposon (TnSC189) in three independent biological replicate libraries to generate 52,662 insertion mutants in V. anguillarum NB10Sm, then classified 3,774 annotated genes as essential, domain-essential, or non-essential using a hidden Markov model (HMM) implemented in EL-ARTIST. Linear regression and ANOVA were used to compare the relationship between gene length and insertion density across chromosomal and plasmid elements. KEGG pathway and STRING protein-protein interaction enrichment analyses with adjusted P-value thresholds identified functionally enriched processes among essential genes, and orthologous essential genes across Vibrio spp. were identified by BlastN with fixed identity-and-coverage thresholds.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Hidden Markov Model (HMM) classification via EL-ARTIST (Python implementation) | Classification of each gene as essential, domain-essential, or non-essential based on transposon insertion density across the V. anguillarum genome and plasmids | 3,774 genes across two chromosomes and three plasmids; 52,662 unique insertion sites | not stated |
| Linear regression with ANOVA comparison of slopes | Relationship between gene sequence length and number of transposon insertions, compared across Chromosome I, Chromosome II, and 67-kb plasmid (smaller plasmids excluded due to low gene count) | — | not stated |
| KEGG pathway enrichment analysis (R/clusterProfiler v3.14.3 and bioconductor/DOSE v3.10) | Identification of significantly enriched KEGG pathways (P < 0.001) among 329 essential genes relative to all V. anguillarum genes with KEGG annotation | 329 essential genes; reference set: all annotated V. anguillarum genes with KEGG annotation | not stated |
| Protein-protein interaction functional enrichment analysis (STRING v11.5) | Functional enrichment of essential gene products relative to all V. anguillarum annotated genes | 329 essential genes; reference set: all V. anguillarum annotated genes | not stated |
| BlastN sequence alignment with fixed threshold-based orthology assignment (>80% identity across >80% of gene length) | Identification of orthologous essential genes across six Vibrio spp. studies (V. cholerae, V. parahaemolyticus, V. anguillarum MVM425) | — | na |
-
Essential genes were classified with a single HMM (EL-ARTIST) using a fixed sliding-window P-value threshold of 0.005, applied once to the merged library↳ Could also: Alternative Tn-seq analysis platforms such as TRANSIT (offering Gumbel, HMM, resampling, and ZINB models) could also be applied, including sensitivity analyses across models — Using multiple statistical models for essentiality calling and comparing their agreement quantifies how robust the essential-gene list is to modelling assumptions, which is useful when comparing results across studies that used different tools
-
Regression slopes relating gene length to insertion count were compared across genomic elements using ANOVA, with the two small plasmids excluded due to low gene counts↳ Could also: An analysis of covariance (ANCOVA) with genomic element as a categorical predictor and an interaction term (length × element) could also formally test slope heterogeneity across all elements simultaneously — ANCOVA integrates slope and intercept comparisons into a single model and allows explicit post-hoc pairwise contrasts with multiplicity correction, facilitating direct comparison of all elements rather than sequential pairwise tests
-
Orthologous essential genes across Vibrio spp. were identified by one-directional BlastN with fixed 80%/80% identity-and-coverage thresholds↳ Could also: Reciprocal best-hit BLAST (RBH) or graph-based tools such as OrthoFinder could also define orthologs by requiring bidirectional best matches across species — Reciprocal best-hit methods reduce false-positive ortholog assignments that can arise from paralogs or gene family members, particularly relevant in a genus with diverse genome architectures
-
Inter-library reproducibility was not quantitatively reported; the three biological replicate libraries were merged without reporting concordance in essential-gene calls across replicates↳ Could also: Reporting the overlap in essential-gene classifications (e.g., number of genes classified as essential in all three libraries independently) or a Cohen’s kappa across replicates could also characterise reproducibility — Quantifying inter-library concordance directly informs confidence in the merged essential-gene list and helps distinguish robustly essential genes from borderline classifications
-
KEGG enrichment results were summarised by stating that two pathways passed an adjusted P-value threshold of 0.001, with only those significant pathways described↳ Could also: Reporting all tested pathways with their adjusted P-values (e.g., as a supplementary table) and explicitly naming the multiple-testing correction procedure would also convey the full scope of the enrichment analysis — Disclosing the complete set of tested hypotheses and the correction method allows readers to evaluate the stringency of false-discovery control and to identify nominally enriched pathways that did not reach the chosen threshold
-
Core-genome membership was defined using a pre-existing pangenome analysis (Coyle et al., 2020) with a fixed ≥95% genome prevalence threshold across 105 V. anguillarum genomes↳ Could also: A rarefaction analysis showing how core-genome size estimates change as additional genomes are included could also assess whether 105 genomes are sufficient to stabilise the core-genome definition — Core-genome rarefaction curves provide a visual and quantitative basis for judging whether the genome sample is large enough to reliably identify genes conserved across the species, which directly affects confidence in drug and vaccine target prioritisation
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — PMID 34745063
Title: Essential Genes of Vibrio anguillarum and Other Vibrio spp. Guide the Development of New Drugs and Vaccines. Bekaert M, Goffin N, McMillan S, Desbois A. Front. Microbiol. 12:755801 (2021). DOI 10.3389/fmicb.2021.755801 · PMCID PMC8564382.
Code: https://github.com/pseudogene/vibrio-tnseq (pinned commit
43d924199ad71f2b2e9fea150140213978f0df72, archived 2024-10-31)
Data: ENA project PRJEB39186 (Tn-seq reads, Illumina MiSeq).
The pipeline (as shipped in the repo)
A self-contained Docker pipeline. Per read file:
- cutadapt — strip transposon/adapter:
cutadapt -g TCGTCGGCAGCGTCAGATGTGTATAAGAGACAGGTTCAGAGTTCTACAGTCCGACGATCACAC -a TAACAGGTTGGATGATAAGTCCCCGGTCTCTGTCTCTTATACACATCTCCGAGCCCACGAGAC -O 3 -m 10 -M 18 -e 0.15 --times 2 --trimmed-only(Dockerfile pinscutadapt==3.7.) - bowtie2 — map to bundled NB10 genome:
bowtie2 --no-1mm-upfront --end-to-end --very-fast -x /databases/vibrio -U <reads> -S <sam>(apt bowtie2 on Ubuntu 20.04 = v2.3.5.1.) - sam_to_map.pl — SAM → insertion-site map / GFF coverage track (+ optional CGView PNG/SVG).
- contrib/el-artist.py — essential-gene calling: a beta Python port of EL-ARTIST (ARTIST; Pritchard et al. 2014). Counts insertions at TA sites within CDS features, bins into windows, fits an HMM → essential / domain-essential / non-essential.
Reference genome bundled in repo (docker/vibrio.fa.gz, vibrio.gff.gz) =
GCF_000786425.1 (V. anguillarum NB10 serovar O1): chromosomes I & II + plasmids.
In scope (pipeline-derived → attempt to reproduce)
| # | Reported result | Pipeline producing it |
|---|---|---|
| C1 | Total reads generated = 5,802,645 | raw ENA deposit (countable from FASTQ) |
| C2 | Reads retained after QC ≈ 4,727,608 (81.5%) | cutadapt (+ paper's fastp QC) |
| C3 | Unique insertion sites = 52,662 | bowtie2 + sam_to_map.pl |
| C4 | Total genes in genome = 3,774 | GFF annotation (GCF_000786425.1) |
| C5 | Essential genes = 329 (8.7%) | el-artist.py HMM |
| C6 | Domain-essential genes = 95 (2.5%) | el-artist.py HMM |
| C7 | rRNA 25/25, tRNA 89/93 (+4 dom.), protein-coding 212 ess. + 91 dom. | el-artist.py + GFF biotype |
| C8 | Core-genome essentiality 114/212 (53.8%), dom 68/91 (74.7%) | downstream comparative (roary/panaroo-like) |
| C9 | Cross-Vibrio overlaps (13 shared / 25 unique / 51 vs MVM425) | downstream set comparison vs external studies |
Primary targets (low-hanging, fully specified): C1, C3, C4, C5, C6.
Out of scope (not attempted; not pipeline-derived from the shipped code/data)
- Wet-lab: transposon library construction, colony counts (~5,500/9,500/15,300), MIC/antibiotic, vaccine/antigen work, growth assays.
- Functional annotation enrichment (InterProScan, KofamKOALA, SignalP, SecretomeP, LipoP, clusterProfiler) — described in Methods but not in the shipped run_pipeline; not reproduced unless time permits (annotation-only, no new essentiality claim).
- C8/C9 comparative-genomics across 105 isolates / 6 external studies — depends on external genome sets + other papers' gene lists not shipped here. Reproduce only if inputs are recoverable; otherwise documented as not-attempted.
Known reproducibility gaps (flagged up front)
- cutadapt version: paper Methods say v2.10; repo Dockerfile pins 3.7. We run the repo's pinned 3.7 (the shipped pipeline) and note the divergence.
- fastp QC step: paper describes fastp filtering (Q<25, len≥50, entropy>15);
run_pipeline.plperforms no fastp — only cutadapt. C2 (81.5% retained) may not be reproducible from the shipped pipeline alone. We will report cutadapt-only retention and note the gap. - el-artist.py is self-described "beta"; exact thresholds (P>0.005, 50 bp window) live in its CLI args — to be confirmed at run time.
Compute plan
All heavy compute on «our HPC»/SLURM; data + clone on «infra» (`«path»
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.