Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Analysis of the Taxonomy, Synteny, and Virulence Factors for Soft Rot Pathogen Pectobacterium aroidearum in Amorphophallus konjac&lt

Front Microbiol · 2022
L1 100/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Every question reproduced
-1 pts
From: “every question reproduced”
Total score -7
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
  • No relevant deviation in data/preprocessing
  • No authors-side cause for any deviation
  • Reported values are derivable from the shared data
  • Any deviation was negligible
  • The central claim held under reproduction
  • Overall, the reproduction was clean
What did not (or only partly)
  • Every checked point held up.
How its reproducibility compares
100/100
Reproducibility score
1.5 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 95% of all assessed papers rank 1 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

1:1 reproduction of the paper's primary computational result. Ran the named tool pyani v0.2.11 (ANIm; MUMmer 3.23 + BLAST+ 2.17.0) on the paper's own 11 deposited RefSeq assemblies (PRJNA794971) plus the P. aroidearum reference and 4 near-neighbour Pectobacterium species, all on «our HPC» («job»). C1: the 11 konjac isolates + reference form one cluster at >=98.30% ANI (all >95%) => species P. aroidearum, reproducing the central taxonomic claim. C1b: every near-neighbour species is <=90.86% ANI (<95%), confirming the 95% boundary separates them. C2a (genome size 4,865,541-5,057,072 bp) and C2b (GC 51.6-51.9%) reproduced EXACTLY from the deposited FASTA. Paper is described well enough; named third-party tool directly applicable (P16). NOT attempted (out of scope / unreproducible-as-specified): annotation-derived counts C2c/d/e, pangenome C3 (tool/version unpinned), virulence C4 (DB unpinned). Dataset PRJNA794971 profiled: 11/11 complete single-contig genomes, open RefSeq, sizes match Table 1 to the byte; grade A. Solver gotchas handled: front1 login-node tmpfs is 50M/full (built conda env on a compute node's 750G node-local /tmp); conda otherwise pulls python 3.14 + matplotlib>=3.9 which breaks pyani 0.2.x at import (pinned python=3.8 + matplotlib-base<3.9).

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 100
    assessed: 2026-06-21 ⛓ ea80c4dec884
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-21
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-21
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The study aims to definitively determine the causal agent of konjac (Amorphophallus konjac) soft rot using genomic data and to characterize the taxonomy, synteny, pangenome plasticity, and virulence factor variation of this pathogen through comparative genomics.

Core claims
  • The causal agent of konjac soft rot in China is Pectobacterium aroidearum, confirmed via in vitro/in vivo pathogenicity tests, ANI, dDDH, and phylogenomic analysis. finding
  • 11 complete P. aroidearum genomes were sequenced and assembled via hybrid Nanopore/Illumina sequencing. method
  • Synteny analysis reveals considerable chromosomal inversions among P. aroidearum strains, one caused by homologous recombination of the ribosome operon. finding
  • The pangenome of P. aroidearum is open, with accessory genes enriched for replication, recombination, and repair functions. finding
  • Type IV and type VI secretion systems show strain-level variation while plant cell wall degrading enzymes (PCWDEs) remain conserved. finding
  • Sequence analysis provides evidence for the presence of a type V secretion system in Pectobacterium. finding
  • P. aroidearum QJ036 exhibits a broader host range than previously reported, causing rot in sweet potato, jicama, yacon, and taro. finding
  • Three strains named P. carotovorum subsp. carotovorum (PC1, PCCS1, PCC21) cluster within P. aroidearum's clade and are likely misnamed. finding
Experimental setups
Assay System Perturbation Readout Platform
in vitro pathogenicity assay (tuber slice inoculation) Amorphophallus konjac tubers bacterial inoculation (11 isolates) soft rot symptom formation
in vivo pathogenicity assay (pin-prick inoculation) 6-month-old konjac seedlings (stems) bacterial inoculation (strain QJ036) stem rot and wilting symptoms
pectinolytic activity assay bacterial isolates on crystal violet pectate (CVP) medium none pit/cavity formation
16S rRNA gene sequencing and BLAST analysis 11 bacterial isolates none sequence identity to P. aroidearum reference
whole-genome sequencing and hybrid de novo assembly 11 P. aroidearum isolates none complete genome sequences Nanopore PromethION and Illumina NovaSeq PE150
ANI and digital DNA-DNA hybridization (dDDH) 64 Pectobacterium genomes (11 new + 53 from GenBank) none pairwise sequence identity for species delineation pyani v0.2.11; Genome-to-Genome Distance Calculator 3.0
whole-genome synteny/alignment analysis P. aroidearum genomes none chromosomal rearrangements/inversions progressiveMauve; D-Genies; BRIG
pangenome and COG functional enrichment analysis 64 Pectobacterium genomes / 14 P. aroidearum genomes none core/accessory gene counts, functional category enrichment Roary; eggNOG-mapper v2.1.4; Fisher's exact test
Key results
  • All 11 isolates caused typical soft rot symptoms on konjac tubers in vitro and formed pits on CVP medium
  • 16S rRNA genes of all 11 isolates share high identity with P. aroidearum reference NR_159926 ≥99% identity
  • Strain QJ036 caused stem rot and wilting in seedlings within 2 days in vivo; controls showed no symptoms
  • Clade IV strains (including P. aroidearum L6, PC1, PCCS1) show ANI and dDDH above species delineation cutoffs ANI >95%; dDDH >70%
  • Pectobacterium pangenome (64 genomes) contains a small proportion of core genes relative to total pangenome, indicating an open pangenome 2,228 core genes / 19,698 total (11.31%)
  • P. aroidearum pangenome (14 genomes) has a larger core gene proportion but remains open with continued gene accumulation 6,630 total genes; 3,575 core (53.92%)
  • QJ036 caused rot symptoms on sweet potato, jicama, yacon, and taro tuber slices, expanding known host range
  • Genomic inversion between QJ036 and QJ002 borders coincide with ribosome operon locations, implicating homologous recombination
Key statistics
  • other ≥99% 16S rRNA identity (species identification of konjac soft rot isolates as P. aroidearum)
  • other ANI >95%, dDDH >70% (species delineation cutoff met among clade IV strains)
  • count 2,228 core genes / 19,698 pangenome genes (11.31%) (Pectobacterium genus pangenome analysis)
  • count 6,630 pangenome genes; 3,575 core genes (53.92%) (P. aroidearum pangenome analysis)
  • other genome size range 4,865,541–5,057,072 bp (P. aroidearum genome assemblies)
  • other GC content 51.6%–51.9% (P. aroidearum genome assemblies)
  • count CDS range 4,277 (QJ311) to 4,469 (AK042) (protein-coding gene counts across strains)
  • count accessory genes per strain range 701 (PC1) to 898 (QJ003) (P. aroidearum pangenome accessory gene variation)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This is a comparative/functional genomics study of Pectobacterium aroidearum strains isolated from konjac, combining in vitro/in vivo pathogenicity assays with whole-genome sequencing, phylogenomics (parsimony and maximum-likelihood trees), average nucleotide identity/digital DNA-DNA hybridization for species assignment, pangenome analysis (Roary), and COG functional enrichment. The only inferential statistical test described is Fisher's exact test for COG category enrichment, with p-values adjusted for multiple comparisons. Most other results (genome features, ANI/dDDH values, phylogenetic placement, pangenome composition) are reported descriptively or against established cutoff thresholds rather than through hypothesis tests.

Replicationmixed Sample sizePathogenicity assays (konjac slice and in vivo seedling inoculation) were described as repeated at least twice independently with similar results; no formal sample size or power calculation was described for genomic/enrichment analyses GroupsBacterial isolates vs. sterile water negative control (pathogenicity assays); core vs. accessory/pangenome genes across strains; COG functional categories Pairingunclear Randomization/blindingnot stated Dispersionnone Multiplicity correctionBenjamini-Hochberg (FDR)
Statistical tests used
Test Applied to n Assumptions
Fisher's exact test COG functional category enrichment analysis of protein-coding genes (e.g., accessory genes enriched in replication, recombination, and repair) not stated
Approaches that could also have been used
  • Functional (COG) enrichment was assessed with Fisher's exact test corrected by the Benjamini-Hochberg method.
    Could also: A hypergeometric test or a gene-set enrichment approach (e.g., GSEA-style permutation testing) could also be used — These are standard alternatives for categorical enrichment testing and can be preferred when category sizes are large or when a continuous ranking of genes (rather than a binary presence/absence split) is available, potentially increasing sensitivity.
  • Species assignment relied on fixed ANI (>95%) and dDDH (>70%) cutoff thresholds combined with tree topology.
    Could also: Bootstrap or jackknife resampling of ANI/dDDH estimates, or a probabilistic/model-based species delimitation method (e.g., using confidence intervals around ANI values), could also be reported — Adding a measure of uncertainty around the ANI/dDDH point estimates would let readers gauge how close values are to the delineation threshold, which is informative near borderline cases.
  • Phylogenetic relationships were inferred using a parsimony method (kSNP3) and a maximum-likelihood method (FastTree) without reported bootstrap or posterior support values in the described results.
    Could also: Bayesian phylogenetic inference (e.g., MrBayes or BEAST) with posterior probabilities, or standard bootstrap resampling (e.g., 100-1000 replicates) for the ML tree, could also be applied — Support values quantify confidence in individual branches/clades, which can strengthen claims about monophyletic grouping such as the konjac pathogen clade.
  • Pathogenicity and pectinolytic assays were repeated at least twice independently and reported as consistent/qualitative outcomes (symptom presence/absence).
    Could also: Quantitative scoring of lesion size or symptom severity across a larger number of biological replicates, summarized with a dispersion measure (SD, SEM, or CI) and compared statistically (e.g., t-test or ANOVA across isolates/timepoints), could also be used — Quantitative, replicate-based comparisons allow formal statistical inference about differences in virulence between isolates or conditions, beyond descriptive agreement across two repeats.
  • Pangenome 'openness' was inferred qualitatively from the shape of the core/pangenome accumulation curves as more genomes were added.
    Could also: Fitting Heaps' law (or a similar power-law/exponential decay model) to the pangenome accumulation data and reporting the fitted exponent with a confidence interval could also be done — A fitted openness parameter with an interval estimate gives a quantitative, comparable measure of pangenome openness rather than a qualitative visual trend.
  • Multiple-testing correction (Benjamini-Hochberg) was applied only to the COG enrichment analysis.
    Could also: If additional comparisons were made across gene families, secretion-system presence/absence, or PCWDE conservation across the 14 genomes, a similar FDR or family-wise correction (e.g., Bonferroni) could also be applied to that broader set of comparisons — Extending multiplicity control to any additional comparison families would keep the overall false-positive rate consistent across all statistical claims in the study.
Software: Unicycler v0.4.8 · Raven v1.5.1 · Pilon v1.24 · BUSCO v5.2.2 · Prokka v1.14.5 · pyani v0.2.11 · Genome-to-Genome Distance Calculator 3.0 · progressiveMauve · D-Genies · BRIG · kSNP3 v3.1.2 · FastTree 2.1.11 · iTOL v5 · MEGA11 · Roary · eggNOG-mapper v2.1.4 · R v4.0.0 · TXSScan/MacSyFinder · BLAST+ v2.6.0

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-35910650

Paper: Zhang Y, Chu H, Yu L, He F, Gao Y, Tang L. Analysis of the Taxonomy, Synteny, and Virulence Factors for Soft Rot Pathogen Pectobacterium aroidearum in Amorphophallus konjac. Front Microbiol. 2022;13:868709. DOI 10.3389/fmicb.2022.868709 · PMID 35910650 · PMCID PMC9326479.

Named code artifact (BRIEF): https://github.com/widdowquinn/pyani (pyani v0.2.11). This is a third-party tool (per P16, applying an existing tool to the paper's own data is an equally valid reproduction). pyani computes Average Nucleotide Identity.

Data: BioProject PRJNA794971 — 11 newly sequenced P. aroidearum genomes (Nanopore PromethION + Illumina NovaSeq, hybrid assembly, "Complete Genome"). Resolved to 11 GenBank/RefSeq assemblies (GCF_024498675.1 … GCF_024498875.1), 22 SRA runs (11 strains × 2 platforms). Plus 53 publicly downloaded Pectobacterium reference genomes (listed in Supplementary Table; 64 total in the genus analysis).


In-scope (pipeline-derived, attempted)

# Result Pipeline / tool named in paper Reproducible?
C1 ANI species delineation: the 11 konjac isolates are P. aroidearum; all clade-IV pairwise comparisons exceed the 95% ANI species threshold (Results "Taxonomy"; Table 1 / Suppl. Table 4) pyani v0.2.11, ANIm method PRIMARY TARGET — directly re-runnable with the named tool on the deposited assemblies. The strongest 1:1 result.
C2 Genome characteristics (Table 1): size 4,865,541–5,057,072 bp; GC 51.6–51.9%; CDS 4,277–4,469; tRNA 77; rRNA 22 assembly stats / annotation (NCBI PGAP for RefSeq) Partially checkable directly from the deposited assemblies (size, GC, contig count, RefSeq CDS/tRNA/rRNA counts). Annotation tool not pinned by authors → treat as provisional supporting evidence.
C3 Pangenome: P. aroidearum (14 strains) core 3,575 / pan 6,630; genus (64 strains) core 2,228 / pan 19,698 (Roary-class; tool/version not clearly pinned) Secondary — attempt only if reference set fully resolvable; high parameter sensitivity.

Out of scope (not pipeline / not attempted, or weakly specified)

  • Wet-lab: DNA extraction, sequencing, pathogenicity assays — not computational.
  • Synteny (progressiveMauve / D-Genies dot plots, chromosomal inversions) — visual/ qualitative; no single numeric claim to grade 1:1.
  • Phylogenetic trees (kSNP3 v3.1.2, FastTree 2.1.11) — topology/clade membership is reproducible but graded qualitatively, not as a number; secondary.
  • Virulence factors (TXSScan/MacSyFinder, SecReT6, InterProScan, IslandViewer4, BLAST+ for PCWDE = 21 genes/strain) — multiple unpinned external web servers; PCWDE count is a candidate secondary check (BLAST+ is reproducible) but DB versions undefined → provisional at best.
  • dDDH >70% (Genome-to-Genome Distance Calculator web server) — external web service, not locally reproducible deterministically.

Reproduction strategy for C1 (primary)

  1. Download the 11 deposited assemblies (GCF_024498675.1 … _875.1) + a reference set anchoring the species boundary: the P. aroidearum type/reference strain and a few near-neighbour Pectobacterium species (e.g. P. carotovorum, P. brasiliense, P. parmentieri) — all onto «infra» via front1.
  2. Run pyani anim (MUMmer/nucmer) on the genome FASTAs on a «our HPC» compute node.
  3. Read the percentage-identity matrix; verify: (a) all 11 konjac isolates are mutually >95% ANI (one species) and >95% to the P. aroidearum reference, and (b) near-neighbour species fall below 95%. This reproduces the taxonomic assignment that is the paper's central computational claim.

Grade basis: the paper states "> 95%" rather than itemising every pairwise value in the main text (values in Suppl. Table 4), so C1 is graded against the threshold claim and species-assignment conclusion, not against per-cell decimals.

Figures / tables: TableFig 2
C1
Reported
all 11 konjac isolates = P. aroidearum, clade-IV pairwise ANI >95% (pyani v0.2.11 ANIm)
Reproduced
konjac mutual ANI min 98.32%; konjac vs P. aroidearum reference min 98.30%; clade range 98.30-100% — all >95%
exact
C1b
Reported
near-neighbour Pectobacterium species fall below the 95% ANI cutoff vs the konjac isolates
Reproduced
max konjac-to-neighbour ANI 90.86% (P. brasiliense); P. carotovorum 90.44%, P. atrosepticum 89.25%, P. parmentieri 88.91% — all <95%
exact
C2a
Reported
genome size 4,865,541-5,057,072 bp (Table 1)
Reproduced
4,865,541-5,057,072 bp across the 11 deposited RefSeq assemblies (FASTA lengths)
exact
C2b
Reported
GC content 51.6-51.9% (Table 1)
Reproduced
51.63-51.88% per-strain GC across 11 assemblies
exact

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 100/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟢3. Location of the main deviation
🟢4. Cause of the deviation
🟢5. Derivability / plausibility
🟢6. Severity of the deviation
🟢7. Core claim
🟢8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Every question reproduced
-1 pts
From: “every question reproduced”
Total score -7

1:1 reproduction of the primary result. The central taxonomic claim (P. aroidearum species assignment via >95% ANI) was reproduced exactly using the authors' own named tool (pyani v0.2.11 ANIm) on their own deposited RefSeq assemblies: ≥98.30% ANI within clade vs ≤90.86% to near-neighbours, with genome size matching to the byte and GC rounding exactly to Table 1. No deviation sits on the authors' side for any attempted value. The only limitation is on method specification: annotation counts, the pangenome, and PCWDE virulence counts were not gradable because the authors did not pin the annotation tool, Roary-class version, or reference DB — these are secondary and non-deterministic, not discrepancies, so the overall judgement remains green.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

298.3 k
tokens (I/O) · 26.1 M incl. cache
55 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.