Microbial diversity of plant pathogens and insect endosymbionts in Reptalus artemisiae.
The main results reproduced, with only marginal, non-material deviations.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Reported values were directly comparable
- ✓Reported values are derivable from the shared data
- ✓Any deviation was negligible
- 🟡Could not use the authors’ exact input data
- 🟡A deviation arose in the data or preprocessing
- 🟡A deviation was attributed to the published material
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Described well enough -> reproduced 1:1 (within tolerance). The paper's named code artifact is Filtlong (github.com/rrwick/Filtlong, P16: a third-party QC tool applied to the paper's own data). I ran Filtlong v0.2.1 with the exact described parameters (--min_mean_q 12 --keep_percent 90) on the only publicly deposited SRA run, SRR32132256 (sample 135/24, ONT MinION R10.4.1), on «our HPC» (SLURM «job», 12 min). The headline QC claim reproduces well: reads after filtering 690,990 vs reported 678,825 (+1.8%); all five sequencing/QC metrics agree within ~5%. The small base-count surplus (2.59 vs 2.46 Gb) is fully explained by two known factors: (1) the paper's 'after filtering' figure adds a Porechop adapter-trim step after Filtlong, which is not part of the Filtlong repo and was not applied; (2) the public SRA deposit holds ~2% fewer raw reads than the paper's stated raw count. No fabrication concern. NOT attempted (80/20, out of scope): DIAMOND/MEGAN taxonomic binning and per-taxon read counts (authors' custom protein DB is not shipped -> not deterministically reproducible), Flye de-novo assembly + genome sizes/GC/depth/CDS (non-deterministic + a manual Geneious contig-merge step), BUSCO/CheckM2 completeness, Prokka/BlastKOALA/PHASTER/IslandViewer annotation, and the other 4 specimens (their reads are not deposited in SRA at all).
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 71assessed: 2026-06-16 ⛓ 7f13ffcf1a1a
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-16
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-16no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusThe microbiome of the emerging sugar beet disease vector Reptalus artemisiae is uncharacterized; this study uses PCR-free metagenomic long-read shotgun sequencing to investigate its bacterial diversity (plant pathogens and insect endosymbionts) and to obtain genomic insight into the plant pathogens 'Ca. Phytoplasma solani' and 'Ca. Arsenophonus phytopathogenicus'.
- ★ R. artemisiae harbors six prokaryotic taxa: two plant pathogens ('Ca. P. solani', 'Ca. A. phytopathogenicus') and four insect endosymbionts ('Ca. Vidania', 'Ca. Purcelliella', 'Ca. Karelsulcia', and Wolbachia). finding
- ★ The four endosymbionts form a stable consortium present in all five evaluated R. artemisiae individuals, whereas plant pathogen presence is variable (each detected in three individuals). finding
- ★ Phylogenies of primary endosymbiont 16S rRNA genes are congruent with the host insect COI phylogeny, indicating long coevolution and vertical transmission. finding
- ★ A complete 774 kb circular chromosome was assembled for 'Ca. P. solani' showing streamlined metabolism with limited biosynthetic pathways but a full arsenal of host-pathogen interaction/pathogenicity genes. resource
- ★ A draft genome of 'Ca. A. phytopathogenicus' (18 scaffolds, 3.11 Mb, two plasmids) shows self-sufficient metabolism with missing metabolic modules, genomic islands, virulence factors, and a dynamic mobilome, indicating a bacterium in genomic transition. resource
- ★ A PCR-free metagenomic singleplex/multiplex long-read shotgun sequencing approach (ONT MinION) characterizes the R. artemisiae microbiome with minimized taxonomic representation bias. method
- This is the first in-depth characterization of the R. artemisiae microbiome. finding
- Two of the three pathogen-positive individuals carried mixed infections of both plant pathogens. finding
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| PCR-free metagenomic long-read shotgun sequencing (singleplex and multiplex) | individual R. artemisiae planthoppers (5 specimens; 2 males, 3 females) collected from sugar beet fields, South Banat, Serbia | none | read counts/bp assigned to bacterial taxa; genome assembly | Oxford Nanopore MinION Mk1B, FLO-MIN114 flow cell (R10.4.1 Kit 14), ONT Rapid Barcoding Kit V14 (SQK-RBK114), MinKNOW v23.11.4 |
| PCR amplification / gel electrophoresis for species identification (ITS2 amplicon length) | individual Reptalus insects | none | ITS2 amplicon size for R. artemisiae identification | PCR Master Mix (Thermo Scientific); primers ITS2fw/ITS2rv; 1% agarose gel, ethidium bromide, UV transilluminator |
| Total nucleic acid (DNA) extraction | individual R. artemisiae specimens | none | metagenomic DNA | CTAB protocol; Select-a-Size DNA Clean & Concentrator MagBead Kit (Zymo) for >600 bp enrichment (specimen 135/24) |
| Taxonomic profiling / metagenomic binning | long reads from each R. artemisiae specimen | none | taxonomic assignment of reads to six prokaryotic taxa | DIAMOND v2.1.9.163 ('more sensitive'), MEGAN v6.25.10, NCBI BLASTn |
| 16S rRNA gene phylogenetic / co-phylogenetic analysis | endosymbiont sequences from R. artemisiae 135/24 | none | maximum-likelihood phylogeny, 1000 bootstraps, host-endosymbiont congruence | ClustalX, MEGA X |
| COI gene phylogenetic analysis for host identification | R. artemisiae 135/24 and cixiid reference sequences | none | sequence similarity and ML phylogeny | MegaBLAST, MEGA X |
| Genome assembly, polishing, completeness assessment and functional annotation | 'Ca. P. solani' strain 135/24 and 'Ca. A. phytopathogenicus' strain 135/24 from R. artemisiae 135/24 | none | assembled chromosome/scaffolds, genome completeness, KO/functional annotation | Flye v2.9.5, Racon v1.54.0, medaka v2.0.1, Prokka v1.14.6, BUSCO v5.8.0, CheckM2, BlastKOALA, Geneious Prime 2025.2.2 |
| Mobilome / genomic island / prophage prediction | 'Ca. A. phytopathogenicus' and 'Ca. P. solani' genomes | none | prophage regions, genomic islands, alignment vs A. nasoniae FIN | PHASTER, IslandViewer4 (SIGI-HMM, IslandPath-DIMOB) |
- – All four insect endosymbionts (Vidania, Purcelliella, Karelsulcia, Wolbachia) detected in each of the five R. artemisiae specimens (no community variability for endosymbionts).
- – Plant pathogens varied among individuals: 'Ca. P. solani' and 'Ca. A. phytopathogenicus' each detected in three of five specimens, two with mixed infection. 3 of 5 each
- – Complete 774 kb circular chromosome assembled for 'Ca. P. solani' with streamlined/limited biosynthetic metabolism. 774 kb
- – Draft genome of 'Ca. A. phytopathogenicus' totalling 3.11 Mb across 18 scaffolds plus two plasmids, with self-sufficient but partially incomplete metabolism. 3.11 Mb, 18 scaffolds, 2 plasmids
- – COI sequence of R. artemisiae 135/24 showed highest similarity to R. artemisiae, corroborating species identification. 97.4% similarity
- – Purcelliella 16S rRNA gene of strain 135/24 most similar to 'Ca. P. pentastirinorum' from R. cuspidatus; phylogeny congruent with host. 97.4% similarity
- – Singleplex sequencing of specimen 135/24 generated 1,316,420 reads (2.93 Gbp); after trimming, average read length 3,626 bp. 1,316,420 reads; 2.93 Gbp; 3,626 bp avg
- – Endosymbiont read assignment ranged from as few as six reads (Karelsulcia, of 5,995) to 3,410 reads (Wolbachia, of 678,825). 6 to 3,410 reads
- correlation 97.4% (COI similarity to R. artemisiae) (COI species-level identification of specimen 135/24)
- correlation 97.4% (16S rRNA similarity) (Purcelliella 135/24 vs 'Ca. P. pentastirinorum' from R. cuspidatus)
- count 1,316,420 reads (2.93 Gbp) (singleplex run, specimen 135/24)
- count 3,626 bp (average read length after adapter/barcode trimming, 135/24)
- count 93,050 reads (0.16 Gb) of 383,100 (0.46 Gbp) (reads assigned to barcode of specimen 93/24, first multiplex run)
- count 774 kb (complete circular chromosome size of 'Ca. P. solani')
- count 3.11 Mb, 18 scaffolds, 2 plasmids (draft genome of 'Ca. A. phytopathogenicus')
- count Purcelliella reads 0.36% (37 of 11,015) to 0.44% (3,012 of 678,825) (proportion of reads assigned to Purcelliella across individuals)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This observational metagenomic study characterized the bacterial microbiome of five Reptalus artemisiae individuals using PCR-free Oxford Nanopore long-read shotgun sequencing, with taxonomic assignment via DIAMOND alignment against a custom database parsed through MEGAN. Phylogenetic relationships of endosymbiont 16S rRNA and insect COI sequences were inferred using maximum-likelihood methods in MEGA X with 1,000 bootstrap replicates to assess co-phylogenetic congruence. Results were reported descriptively as read counts, total base pairs per taxon per specimen, and genome assembly statistics; no formal inferential hypothesis tests were applied to community composition data, and genome quality was assessed with BUSCO and CheckM2.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Maximum-likelihood phylogenetic inference with 1,000 bootstrap replicates; best-fit substitution model selected automatically via Neighbor-Join/BioNJ heuristic | 16S rRNA gene sequences of four primary endosymbionts (Purcelliella, Karelsulcia, Vidania, Wolbachia) from R. artemisiae 135/24 vs. reference strains | One focal specimen (135/24) provided sequences; number of reference sequences not stated | not stated |
| Maximum-likelihood phylogenetic inference with 1,000 bootstrap replicates; best-fit substitution model selected automatically | COI gene sequences for host insect species identification and co-phylogenetic congruence analysis with endosymbionts | One focal specimen (135/24) plus publicly available reference sequences; exact reference n not stated | not stated |
| DIAMOND protein alignment ('more sensitive' mode) + MEGAN taxonomic binning with default parameters | Read-level taxonomic assignment of metagenomic reads from all five R. artemisiae specimens against a custom endosymbiont/pathogen protein database | Five specimens; read counts ranged from approximately 5,995 to 678,825 per individual | not stated |
| BUSCO v. 5.8.0 marker-gene completeness assessment | Genome quality evaluation of assembled Ca. P. solani and Ca. A. phytopathogenicus genomes from specimen 135/24 | One genome assembly per pathogen | na |
| CheckM2 completeness and contamination assessment | Genome quality evaluation of assembled Ca. P. solani and Ca. A. phytopathogenicus genomes from specimen 135/24 | One genome assembly per pathogen | na |
| SIGI-HMM (codon usage bias) and IslandPath-DIMOB (sequence composition + mobile element genes) genomic island prediction | Genomic island identification in Ca. A. phytopathogenicus 135/24 draft genome, aligned against A. nasoniae FIN as reference | One draft genome assembly | not stated |
-
Community composition across five individuals was described by raw read counts and base-pair totals per taxon, without formal diversity metrics↳ Could also: Alpha-diversity indices (e.g., Shannon entropy, observed richness) and beta-diversity distances (e.g., Bray-Curtis dissimilarity with PERMANOVA) could also summarize community structure across specimens — Standardized diversity metrics provide a quantitative framework that facilitates direct comparison with published microbiome data from other cixiid species and enables more formal assessment of community variation across individuals
-
Phylogenetic support was estimated using 1,000 ML bootstrap replicates in MEGA X↳ Could also: Bayesian inference (e.g., MrBayes or BEAST) could also be used, reporting posterior probability as the support metric and optionally incorporating divergence time estimation — Bayesian posterior probabilities and ML bootstrap values convey complementary information about clade confidence; BEAST additionally enables molecular-clock-calibrated divergence dating, which is informative for co-evolutionary analyses of long-associated endosymbionts
-
Genome assembly and 16S rRNA phylogenetic analysis were performed on a single specimen (135/24), which received singleplex sequencing yielding substantially more reads than the multiplexed specimens↳ Could also: Deeper or additional singleplex sequencing of further specimens (e.g., 93/24 or 92/24) could also yield assemblies for intra-population genomic comparisons — Multiple genome assemblies per pathogen taxon would allow within-vector-population strain diversity to be assessed and would strengthen inferences about pathogen genomic plasticity beyond a single representative
-
Taxonomic profiling used a custom subset database restricted to taxa previously associated with cixiids (Mollicutes, specific endosymbiont genera, Fulgoridae, and one plant host)↳ Could also: Classification against a broader reference database (e.g., full NCBI RefSeq or GTDB-Tk) or k-mer-based classifiers (e.g., Kraken2/Bracken) could also be applied for unbiased taxon discovery — Broader or k-mer-based approaches reduce the risk of missing divergent or novel organisms not represented in a curated subset; they also provide a cross-check on assignments made against the custom database
-
Pathogen and endosymbiont presence was described as binary (detected/not detected) across the five individuals↳ Could also: Normalized relative abundance estimates (e.g., reads mapped per Mbp of reference genome, or coverage-based copy-number ratios) could also quantify within-host microbial load — Quantitative abundance or coverage metrics complement presence/absence calls by capturing variation in pathogen titre, which is biologically relevant to vector competence and transmission efficiency
-
Genome completeness was evaluated with two marker-gene tools (BUSCO and CheckM2)↳ Could also: Average nucleotide identity (ANI) comparison against publicly available reference genomes could also contextualize strain-level relatedness alongside the completeness metrics — ANI is a widely adopted genomic species-boundary metric that complements completeness assessment by placing a newly assembled genome within a phylogenomic framework, particularly useful when describing a draft genome of a partially characterized pathogen
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-41826827
Paper: Duduk et al. 2026, Microbial diversity of plant pathogens and insect endosymbionts in Reptalus artemisiae, BMC Microbiol. DOI 10.1186/s12866-026-04915-x. Named code artifact: https://github.com/rrwick/Filtlong (third-party long-read QC tool — P16-valid: applying an existing tool to the paper's own data). Data: BioProject PRJNA1215974 → SRA run SRR32132256 (sample "135/24 BC16", Oxford Nanopore MinION R10.4.1). This is the ONLY publicly deposited run.
Pipeline (from Methods)
ONT MinION basecall (MinKNOW super-accurate, Q10) → Filtlong v0.2.1 (mean quality ≥12, remove worst 10%) → Porechop v0.2.4 (adapter trim) → DIAMOND v2.1.9 + MEGAN v6.25 (taxonomic binning vs custom protein DB) / NCBI BLASTn → Flye v2.9.5 assembly + Racon/Medaka polish + Geneious merge → BUSCO v5.8 / CheckM2 → Prokka / BlastKOALA / PHASTER / IslandViewer4.
IN SCOPE (reproduce 1:1)
- Filtlong QC of run SRR32132256 (135/24). The named repo, applied to the
paper's own deposited data, with the exact described parameters
(
--min_mean_q 12 --keep_percent 90). Deterministic, fully specified.- Reported raw 135/24: 1,316,420 reads / 2.93 Gb.
- Reported after filtering: 678,825 reads / 2.46 Gb (avg read length ~3,624 bp, which the paper reports as 3,626 bp).
- Sanity check: raw read/base counts of SRR32132256 vs paper's raw.
OUT OF SCOPE (not attempted — 80/20)
- DIAMOND/MEGAN taxonomic binning & per-taxon read counts/percentages — requires the authors' custom protein database (plant pathogens + endosymbionts), which is NOT shipped. Not deterministically reproducible.
- Flye assembly + polishing → genome sizes (P. solani 774,238 bp; Arsenophonus 3,110,796 bp), GC%, depth, CDS counts — long, non-deterministic de-novo assembly with a manual Geneious contig-merge step; large compute, under-specified.
- BUSCO/CheckM2 completeness, Prokka/KEGG/PHASTER/IslandViewer annotation — downstream of the assembly; out of scope.
- The 4 other specimens (93/24, multiplex 2 etc.) — their reads are NOT deposited in SRA (only 135/24 BC16 present), so unreproducible by definition.
- Phylogenetics / ANI / % similarity values — wet-lab + curated-reference, manual.
Reproducibility note
Only one of five specimens' reads are public. The one clearly-specified, deterministic, named-tool pipeline output (Filtlong filtering of 135/24) is the target. Everything downstream depends on an unshipped custom DB or non-deterministic assembly + a manual GUI step.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
The named code artifact (Filtlong) reproduces faithfully on the paper's own deposited run: filtered reads 690,990 vs reported 678,825 (+1.8%) and all five QC metrics within ~5%, with no fabrication concern — the values are derivable from the shipped data + tool. The only deviations are explainable and on our/data side: an omitted Porechop adapter-trim step (filt_bases +5.2%) and a public SRA deposit ~2% smaller than the stated raw. However, the paper's central conclusions (microbial diversity, endosymbiont/pathogen taxonomy, assemblies) were not reproducible — custom DB not shipped, non-deterministic assembly with manual Geneious merge, and 4 of 5 specimens not deposited — so the QC headline holds within tolerance while the biological core stays untested, yielding an overall yellow.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.