Comparative Genome Analysis of 16SrXII-A 'Candidatus Phytoplasma solani' POT Transmitted by Hyalesthes obsoletus.
The main results reproduced: recomputed values matched the published ones within tolerance.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Same input data as the authors
- ✓Reported values were directly comparable
- ✓No relevant deviation in data/preprocessing
- ✓No authors-side cause for any deviation
- ✓Reported values are derivable from the shared data
- ✓Any deviation was negligible
- ✓The central claim held under reproduction
- ✓Overall, the reproduction was clean
- Every checked point held up.
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
DESCRIBED WELL ENOUGH; reproduced 1:1. PacBio Revio phytoplasma genome paper. Six clearly-specified pipeline outputs reproduced from public data on «our HPC» with the named/standard tools: genome size 832,614 bp (exact, GenBank CM135798.1), GC 28.21% (exact), 60,609 deposited long reads (exact count), pbmm2 v1.17.0 mapping depth 1009.86x vs reported ~1010x (exact, the named code), and fastANI strain comparison reproducing both the >98% within-A-type relatedness (lower bound 98.48% exact, upper ~99.93%==c1-c4 99.985%) and the ~82-83% A-vs-P divergence (82.64%). Two reporting/auditing notes: (1) the paper's read 'N50 13.91 kb' actually equals the MEAN read length (true N50 = 13.46 kb) - minor mislabel, not fabrication; (2) DATA-DEPOSIT GAP - the full 1,929,813 raw reads and the 316,493 MEGAN-binned Phytoplasma reads are NOT in the public SRA (only the 60,609-read >10kb subset is), so total-read, taxonomic-binning, and Canu coverage (245.67x) claims cannot be checked against public data. NOT ATTEMPTED (hard 20%): de-novo Canu assembly (needs undeposited raw set; product checked directly instead), DIAMOND/MEGAN binning, BUSCO/CheckM completeness, cBUSCO core-gene identity. No completeness claim - the in-scope, publicly-checkable pipeline results reproduce essentially perfectly.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 95assessed: 2026-06-16 ⛓ d81c1720b94f
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-16
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-16no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusWhat are the phylogenetic position and pathogen–host interaction features of the Hyalesthes obsoletus-transmitted 16SrXII-A 'Candidatus Phytoplasma solani' strain POT, the first such complete genome reported from Germany, relative to other 16SrXII-group phytoplasmas?
- ★ The complete 832,614 bp circular chromosome of the H. obsoletus-transmissible 'Ca. P. solani' 16SrXII-A strain POT was assembled and functionally reconstructed. resource
- ★ POT shares highest average nucleotide identity with Italian bindweed-associated genomes and displays strong synteny with the c5 strain. finding
- ★ Phylogenomic analysis confirms POT belongs to the 16SrXII-A lineage of 'Ca. P. solani', while 16SrXII-P strains (GOE, PENLEP) form a distinct, separate species lineage. finding
- ★ The POT genome combines mobile-element-driven instability with a conserved core metabolism, with virulence factors including transposon-linked effectors but lacking pathogenicity island organisation. mechanism
- ★ POT differs from other 16SrXII-group phytoplasmas through unique collagen-like proteins that could contribute to virulence. finding
- H. obsoletus collected from a symptomatic potato field experimentally transmitted stolbur phytoplasma to Catharanthus roseus. method
- The genomic framework improves diagnostics, enables strain-level resolution, and supports assessment of breeding materials under stolbur phytoplasma pressure. resource
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Experimental insect transmission (no-choice setup) | Catharanthus roseus plants with field-collected Hyalesthes obsoletus from a Bingen potato field | vector inoculation access period (IAP, 10 days) | phytoplasma infection rate and symptom development in recipient plants | insect-proof acrylic cages; Fitotron SGR233 climatic chamber |
| Endpoint PCR detection / Sanger sequencing | DNA from C. roseus midrib tissue | none | phytoplasma infection confirmation (rRNA operon regions) | primers P1/P7 and R16F2n/R2; Qubit fluorometer; NucleoBond HMW DNA Kit |
| SMRT (PacBio HiFi) whole-genome shotgun sequencing | DNA from infected C. roseus (strain POT) | none | long-read sequence data for genome assembly | PacBio Revio, 25M ZMW SMRT cell, SPRQ chemistry; SMRTbell prep kit v3.0 |
| Genome assembly and quality assessment | PacBio reads of strain POT | none | circular chromosome size, GC content, coverage, completeness, contamination | Canu v2.2, CheckM2 v1.1.0, BUSCO v5.8.1, pbmm2 v1.17.0 |
| Taxonomic binning | SMRT long reads | none | reads assigned to 'Candidatus Phytoplasma' | DIAMOND v0.9.30.131 (BLASTX vs NCBI nr), MEGAN v7.1.0 |
| Genome annotation and functional/metabolic reconstruction | POT genome and comparative 16SrXII genomes | none | CDS, RNA features, pathways, membrane/secreted proteins | RAST v2.0, PGAP v6.10, BlastKOALA v3.1, InterProScan v106.0, Phobius v1.01 |
| Comparative genomics: ANI and 16S rRNA identity | POT vs 16SrXII-A strains c1/c4/c5/o3 and 16SrXII-P strains GOE/PENLEP | none | average nucleotide identity and pairwise 16S rRNA identity | FastANI v1.34, BLASTN v2.11.0, iPhyClassifier |
| Whole-genome alignment / synteny and phylogenomics | Complete 16SrXII stolbur phytoplasma genomes | none | collinearity/LCBs, orthogroups, maximum-likelihood phylogeny | Mauve v2.4.0, OrthoFinder v2.5.5, BUSCO/Prodigal, MAFFT v7.505, IQ-TREE v2.4.0, MEGA v12.0.11 |
- – Five of seven exposed C. roseus plants tested positive for phytoplasma; all negative controls remained negative 71.43% (5/7)
- – Of 48 recovered H. obsoletus individuals (from 56 used), 13 tested positive for phytoplasma 27.08% (13/48)
- – POT genome assembled as a single circular contig of 832,614 bp with 28.21% G+C content 832,614 bp; 28.21% GC
- – POT 16S rRNA gene (PSOL_02220) matched the 'Ca. P. solani' reference strain AF248959, placing POT in the taxon 99.67% identity
- – ANI between POT and 16SrXII-A strains c1, c4, c5, o3 exceeds 98% 98.48–99.93%
- ▼ 16SrXII-P strains GOE and PENLEP show much lower ANI to POT/'Ca. P. solani', delineating a separate species lineage ~82–83%
- – POT shows strongest synteny/collinearity with strain c5; strain o3 is the most structurally divergent 16SrXII-A genome
- – Assembly completeness estimated high with low contamination by CheckM and BUSCO CheckM 99.45% complete, 0.43% contamination; BUSCO 95.3% complete
- count 832,614 bp circular chromosome (POT genome size)
- correlation 99.67% 16S rRNA identity to AF248959 (taxonomic placement of POT)
- other ANI 98.48–99.93% (POT vs 16SrXII-A strains c1,c4,c5,o3)
- other ANI ~82–83% (POT vs 16SrXII-P strains GOE/PENLEP)
- count 5/7 (71.43%) plants infected (C. roseus transmission success)
- count 13/48 (27.08%) insects positive (recovered H. obsoletus tested positive)
- other 245.67-fold assembly coverage; ~1010× read-mapping depth (POT genome sequencing coverage)
- count 316,493 reads assigned to 'Candidatus Phytoplasma' of 1,929,813 total long reads (taxonomic binning)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This is primarily a bioinformatic comparative genomic study of a newly assembled phytoplasma genome (strain POT, 832,614 bp), accompanied by a small-scale insect transmission experiment. Genome quality was benchmarked with BUSCO and CheckM2; phylogenomic placement used maximum-likelihood inference (IQ-TREE, 1000 bootstrap replicates) on a BUSCO-derived core gene set and on near-full-length 16S rRNA sequences; whole-genome relatedness was quantified via Average Nucleotide Identity (FastANI) and BLASTN pairwise identity; and gene content was compared across seven complete genomes using OrthoFinder orthogroup analysis. Transmission experiment outcomes were reported as simple proportions with no formal inferential test.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Maximum likelihood phylogenetic reconstruction with 1000 bootstrap replicates | Phylogenomic tree of complete phytoplasma genomes (BUSCO core genes, IQ-TREE v2.4.0) and separate 16S rRNA maximum-likelihood tree | Complete phytoplasma genomes from NCBI as of 15 July 2025; core comparative set of 7 strains | not stated |
| Average Nucleotide Identity (ANI) pairwise comparison | Whole-genome relatedness among POT and six other 16SrXII-A/-P strains (FastANI v1.34) | 7 complete genomes | na |
| BLASTN pairwise sequence identity | 16S rRNA gene identity among strains; taxonomic assignment via iPhyClassifier | — | na |
| BUSCO completeness scoring | Assembly quality assessment and core-gene phylogenomics using 151 Mollicutes single-copy orthologs (BUSCO v5.8.1) | 1 newly assembled genome; 7 comparative genomes | na |
| CheckM2 marker-gene completeness and contamination estimation | Assembly quality validation of strain POT (CheckM2 v1.1.0) | 1 genome | na |
| Proportional summary (no formal inferential test) | Transmission success: 5/7 C. roseus plants positive (71.43%); 13/48 insects positive (27.08%) | 7 exposed plants, 2 negative controls; 48 recovered insects | na |
-
Transmission success in the insect experiment was reported as proportions (5/7 plants positive; 13/48 insects positive) without a formal inferential test or uncertainty estimate.↳ Could also: Exact binomial confidence intervals (e.g., Clopper-Pearson 95% CI) could also be reported for each proportion. — With small denominators (n = 7 plants; n = 48 insects), a binomial CI explicitly quantifies sampling uncertainty around the observed rates, making cross-study comparisons and meta-analyses more informative.
-
Phylogenetic reconstruction used maximum likelihood with bootstrap support (IQ-TREE, 1000 replicates).↳ Could also: Bayesian phylogenetic inference (e.g., MrBayes or BEAST) with posterior probability branch support is a standard complementary approach. — Bayesian posteriors provide a probabilistic interpretation of clade membership and can converge efficiently; comparing ML bootstrap and Bayesian topology is a common way to assess robustness of inferred relationships.
-
Whole-genome similarity was quantified using ANI (FastANI), providing a single pairwise distance metric compared against a 95% species boundary.↳ Could also: Digital DNA-DNA hybridization (dDDH) via the Genome-to-Genome Distance Calculator (GGDC) is another widely used genomic relatedness metric for prokaryotic species delineation. — dDDH mirrors the historical 70% wet-lab DDH species threshold and is directly comparable to a large body of existing prokaryotic taxonomy literature, offering an alternative frame for the same relatedness question.
-
Core-gene phylogenomics was based on BUSCO single-copy orthologs identified within the set of seven compared genomes.↳ Could also: Multi-locus sequence analysis (MLSA) using phytoplasma-established housekeeping markers (e.g., tuf, secY, groEL, map) could also be used for phylogenetic placement. — MLSA with markers widely used in phytoplasma systematics allows direct integration with the large existing classification literature built on those specific loci, facilitating cross-study comparability.
-
Orthologous gene families across the seven complete genomes were identified using OrthoFinder with default parameters.↳ Could also: Pan-genome analysis tools such as Roary or PIRATE could also partition genomes into core, accessory, and unique fractions and produce explicit Venn-style summaries. — Pan-genome frameworks provide a standardised vocabulary (core/accessory/unique) for comparing gene content across strains and can directly visualise the shared and lineage-specific portions of each genome.
-
Assembly quality was assessed with BUSCO (gene-content completeness) and CheckM2 (marker-gene completeness and contamination).↳ Could also: K-mer-based quality value (QV) assessment using Merqury is an additional metric that estimates base-level assembly accuracy independently of gene models. — K-mer QV scores complement gene-content metrics by directly assessing sequence fidelity at base resolution, which is informative for long-read assemblies where systematic error patterns (e.g., homopolymer compression) can occur.
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-41597744
Paper: Ilic et al. 2026, Microorganisms. Comparative Genome Analysis of 16SrXII-A 'Ca. Phytoplasma solani' POT transmitted by Hyalesthes obsoletus. DOI 10.3390/microorganisms14010226.
Pipeline-derived results (what the bioinformatics produced): PacBio Revio HiFi sequencing → SMRTlink trimming → DIAMOND/MEGAN taxonomic binning → Canu v2.2 assembly → pbmm2 v1.17.0 mapping for validation → BLASTN circularity → BUSCO/CheckM completeness → ANI / cBUSCO strain comparison.
Data actually deposited (the obtainable surface)
- SRA run SRR35343499 (BioProject PRJNA1327367, BioSample SAMN51278244): ENA reports 60,609 reads / 843,125,788 bases / 199 MB gz, Revio, WGS. → This is the >10 kb long-read subset, NOT the full 1.9M-read raw set.
- Assembled genome GenBank CM135798.1 (POT chromosome), esummary slen = 832,614 bp (already matches reported genome size).
- Comparison genomes (public, small): CP103788/87/86/85 (c1/c4/c5/o3), CP155828 (GOE). All resolve via NCBI esummary.
IN SCOPE (attempted — clear, small data, named/standard tools)
| id | claim | how reproduced | data | weight |
|---|---|---|---|---|
| C1 | genome size 832,614 bp | seqkit stats on CM135798.1 | ~0.8 MB | high |
| C2 | GC content 28.21% | seqkit fx2tab -g on CM135798.1 | ~0.8 MB | high |
| C3 | 60,609 long reads (>10 kb), N50 13.91 kb | seqkit stats -a on SRR35343499 | 199 MB | high |
| C4 | read-mapping depth ~1010× | pbmm2 (the named code) align reads→CM135798.1, samtools depth | 199 MB | high |
| C5 | ANI POT vs c1/c4/c5/o3 = 98.48–99.93% (>98%) | fastANI (third-party) | ~4 MB | high |
| C6 | ANI POT vs GOE (16SrXII-P) ≈ 82–83% | fastANI | ~1.5 MB | med |
OUT OF SCOPE (the hard ~20% — not attempted, with reasons)
- Canu de-novo assembly (832,614 bp, 245.67× coverage): the full 1.9M-read raw set is not deposited (only the 60,609 long reads are). De-novo assembly from the partial set won't faithfully reproduce coverage/size. The product (CM135798.1) is checked directly instead (C1/C2).
- Total 1,929,813 PacBio reads and 316,493 Ca. Phytoplasma reads (N50 7.26 kb): derived from the full raw set + DIAMOND/MEGAN binning, which are NOT in the deposited SRA run → not verifiable from public data (flag for human audit; possible the deposit is only the filtered subset).
- BUSCO 95.3% / CheckM 99.45% / contamination 0.43%: completeness QC; doable but lower priority, skipped for 80/20 (DB downloads + runtime).
- cBUSCO core-gene identity (≥99.76% within A; ~84–85% A vs P): needs BUSCO core extraction; skipped, ANI (C5/C6) covers the strain-relatedness claim.
- Wet-lab / vector transmission / 16S phylogeny figures: out of scope (manual).
Code note
The brief's code link is pbmm2 (PacBio's minimap2 frontend) — a third-party tool, which per P16 is equally valid. We run pbmm2 v1.17.0 (C4) on the paper's own deposited reads, plus standard tools (seqkit, fastANI) for the other claims.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
The six publicly-checkable, in-scope pipeline claims reproduce 1:1 from the authors' deposited data with the named tools: genome size 832,614 bp (exact), GC 28.21% (exact), 60,609-read count (exact), pbmm2 depth 1009.86x vs ~1010x, and fastANI confirming both >98% within-16SrXII-A relatedness (98.48% lower bound exact) and ~82.64% divergence from the P-type. The only authors'-side blemish is a minor mislabel — the reported 'N50 13.91 kb' is actually the mean read length (true N50 13.46 kb), not fabrication. A data-deposit gap (full raw read set and MEGAN-binned subset not in SRA) prevents independent checking of total-read, binning, and Canu-coverage numbers, but these are peripheral to the central comparative-genomics conclusion, which holds fully. Overall a high-fidelity reproduction.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.