Discovery of Early-Branching Wolbachia Reveals Functional Enrichment on Horizontally Transferred Genes.
Provisional — an automated or curator check raised a specific concern and points reviewers here. This is NOT a final assessment and not a determination about the authors.
The main results reproduced: recomputed values matched the published ones within tolerance.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Same input data as the authors
- ✓Reported values were directly comparable
- ✓Reported values are derivable from the shared data
- ✓Any deviation was negligible
- 🟡A deviation arose in the data or preprocessing
- 🟡A deviation was attributed to the published material
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
REPRODUCED (1:1, described well enough). Table 1 comparative genome stats for all 6 Wolbachia strains reproduced from the publicly deposited FASTAs using the LINKED third-party tool barrnap (rRNA) + Prokka 1.14.6 (genes/tRNA) + seqkit (length/GC/contigs/N50/max): 32 exact, 6 within-tol, 6 mismatch of 44 graded values. Length & contig counts EXACT for all 6 (wTex off 1 bp). barrnap rRNA = 3 EXACT for all. Prokka CDS = reported Genes EXACT for 5/6 (wTex within 2). tRNA EXACT for all once Prokka's separately-reported tmRNA (=1) is added (paper's 'tRNAs' = tRNA+tmRNA). GC within-tol, with shortfalls fully explained by N-base handling (exact where deposit has 0 N's). wTex N50 (10082) and max-scaffold (57862) EXACT. The ONLY mismatches are the 6 pseudogene counts: plain Prokka does not call pseudogenes (the paper's values come from NCBI PGAP, a different annotation source) -- an honest pipeline limitation, not a fabrication signal. Stretch 16S divergence partial (wTex-wPpe 3.805% vs paper's max-of-trio 4.211%; wRad genome not provided). NOT attempted: wet-lab, manual Geneious curation, de-novo wTex re-assembly, HGT calls, GO/dN-dS enrichment, phylogeny supports, the unpublished ~3,400-amplicon SRA screen, ModelSEED, CheckM/coverage.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 90assessed: 2026-06-21 ⛓ d95094057560
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-21
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-21no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetThe study tests whether early-branching Wolbachia symbionts of plant-parasitic nematodes (PPNs) show genomic and functional signatures reflecting plant-specialization as a key driver of early Wolbachia evolution, positioning them functionally between mutualist and reproductive-parasite Wolbachia strains.
- ★ A novel early-branching Wolbachia strain (wTex), one of the earliest supergroup L strains, was discovered in a plant-parasitic nematode community from 1 of 16 sampled sites. finding
- ★ wTex lacks known genes for cytoplasmic incompatibility, riboflavin, and biotin biosynthesis, placing it functionally between mutualist C+D strains and reproductive parasite A+B strains. finding
- ★ Functional terms enriched in supergroup L Wolbachia include protoporphyrinogen IX, thiamine, lysine, fatty acid, and cellular amino acid biosynthesis. finding
- ★ dN/dS analysis indicates the strongest purifying selection on arginine and lysine metabolism, and on vitamin B6, heme, and zinc ion binding pathways. finding
- ★ PPN Wolbachia contains putative horizontally transferred genes, including a Midichloria-like lysine biosynthesis operon, a spirochete-like thiamine operon shared with wCfeT, a Rickettsia-like ATP/ADP carrier, and a eukaryote-like gene implicated in the lysine-to-pipecolic acid systemic acquired resistance pathway. mechanism
- ★ Group L-like Wolbachia variants were identified in global rhizosphere sequence databases, suggesting additional undiscovered PPN Wolbachia diversity. finding
- An iterative subtractive mapping and reassembly approach (bwa mapping + metaSPAdes) was developed to recover Wolbachia metagenome-assembled genomes from mixed nematode community DNA. method
- ★ Higher dN/dS pathways distinguish group L Wolbachia from wPni (aphids), wFol (springtails), and wCfeT (cat fleas), indicating distinct functional shifts across early host transitions. finding
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Metagenomic shotgun sequencing / genome skimming | rhizosphere nematode community DNA (soil/root samples) | none | de novo assembled Wolbachia-like contigs/genome (MAG) | Illumina HiSeq, 150 PE cycles |
| PCR pre-screening (16S rRNA) | total nematode community DNA | none | presence/absence of Wolbachia 16S rRNA sequence | custom primers (Wol-mee-F/R) thermal cycler |
| Comparative genomics / gene presence-absence analysis | wTex genome vs. other Wolbachia and Anaplasmataceae outgroups | none | gene repertoire and pathway presence/absence | Roary v3.13 |
| Metabolic pathway modeling | wTex genome | none | metabolic pathway completeness (KEGG/MetaCyc) | ModelSEED v2.6.1 |
| Phylogenomic/phylogenetic analysis (16S rRNA, COI, concatenated genes) | Wolbachia strains and candidate nematode hosts | none | evolutionary relationships/tree topology, bootstrap/posterior support | RAxML v4 (ML), MrBayes v2.2.4 (Bayesian) |
| dN/dS selection analysis | Wolbachia coding genes across strains/pathways | none | signatures of purifying/positive selection per pathway | — |
| SRA metadata/read screening | public NCBI SRA sequencing projects (rhizosphere, soil, nematode) | none | presence of Wolbachia-like 16S rRNA reads | NCBI E-utilities, blastn |
| Community abundance correlation analysis | 21 sampled rhizosphere nematode communities (COI reads) | none | Spearman rank correlation between nematode taxon abundance and Wolbachia 16S rRNA coverage | R package Hmisc (rcorr), corrplot |
- – 1 of 16 sampled sites harbored nematode communities hosting a PPN Wolbachia strain (wTex) 1/16 sites
- ▼ wTex lacks genes for cytoplasmic incompatibility, riboflavin, and biotin biosynthesis
- ▲ Functional enrichment in group L for protoporphyrinogen IX, thiamine, lysine, fatty acid, and amino acid biosynthesis
- ▼ Strongest purifying selection (dN/dS) found on arginine and lysine metabolism, vitamin B6, heme, and zinc ion binding pathways
- – Identification of putative HGT genes: Midichloria-like lysine operon, spirochete-like thiamine operon (shared with wCfeT), Rickettsia-like ATP/ADP carrier, eukaryote-like pipecolic acid pathway gene
- – Group L-like Wolbachia variants detected in global rhizosphere sequence databases
- – Higher dN/dS pathways differentiate group L from wPni, wFol, and wCfeT, indicating distinct functional changes across early host transitions
- count 1 out of 16 (sampled sites with PPN Wolbachia-positive nematode communities)
- count 800–1,400 (nematodes processed per sample for DNA extraction)
- other 0.6–1 μg (DNA input for genomic library preparation)
- other 150 PE cycles (Illumina HiSeq sequencing read length)
- correlation Spearman rho with BH-adjusted p-values (correlation between nematode COI abundance and Wolbachia 16S rRNA coverage across communities)
- other up to 25% global crop yield loss / ~$100 billion annually (economic/agricultural impact of plant-parasitic nematodes)
- count 3 gene regions (sequence data available for comparison strain wRad)
- other chain length 1,100,000; 4 heated chains; burn-in 100,000 (MrBayes MCMC settings for COI phylogenetic analysis)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This study used metagenomic assembly and comparative/phylogenomic bioinformatics rather than classical wet-lab experimental statistics: nematode community composition (from COI barcode abundance) was related to Wolbachia abundance across 21 sampled communities using Spearman rank correlation with Benjamini-Hochberg-adjusted p-values, phylogenetic relationships were reconstructed with maximum-likelihood (RAxML, bootstrap support) and Bayesian (MrBayes, MCMC convergence diagnostics) methods, and gene/pathway-level selection was assessed via dN/dS analysis referenced in the abstract. Results are reported primarily as correlation coefficients (rho) with adjusted p-values, phylogenetic trees with bootstrap/posterior support values, and descriptive gene presence-absence and functional enrichment comparisons across strains.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Spearman rank correlation (rho) | Correlation between nematode taxon (COI-based) relative abundance and Wolbachia 16S rRNA relative abundance across sampled nematode communities | 21 sampled nematode communities | not stated |
| dN/dS (nonsynonymous/synonymous substitution rate) analysis | Comparison of selection pressure across metabolic pathways/genes between Wolbachia group L and related strains (wPni, wFol, wCfeT) | — | not stated |
| Maximum likelihood phylogenetic reconstruction with bootstrap support (RAxML, GTR+Gamma) | Phylogenies of candidate PPN hosts (COI) and of Wolbachia strains (16S rRNA and concatenated gene regions) | 500 bootstrap replicates | na |
| Bayesian phylogenetic inference (MrBayes, MCMC) | Same phylogenetic datasets as above (COI and Wolbachia strain relationships) | chain length 1,100,000; burn-in 100,000; convergence checked with minESS > 200 | na |
-
The relationship between nematode taxon abundance and Wolbachia abundance across communities was assessed using Spearman rank correlation on relative-abundance data, with BH FDR correction applied across the resulting p-value matrix.↳ Could also: Compositional-data-aware methods such as centered-log-ratio (CLR) transformation followed by correlation, or tools like SparCC/SCC designed for relative-abundance microbiome data — Relative abundance data are compositional (constrained to sum to a whole), and compositional-aware correlation methods can also help distinguish genuine co-occurrence from artifacts of the closure constraint, complementing the Spearman approach already used.
-
Node support for phylogenetic relationships was evaluated with maximum-likelihood bootstrapping (500 replicates) and separately with Bayesian posterior probabilities.↳ Could also: Additional support metrics such as SH-aLRT or ultrafast bootstrap approximations (e.g., in IQ-TREE), or running multiple independent MCMC chains with PSRF/ESS diagnostics — These approaches can also provide complementary or faster-converging measures of node reliability and additional convergence diagnostics alongside the bootstrap and minESS metrics already reported.
-
Selection pressure across pathways/genes was summarized using gene-wide or pathway-wide dN/dS ratios to identify signatures of purifying selection.↳ Could also: Site-specific or branch-site selection models such as PAML's branch-site test or HyPhy's FEL/MEME — These methods could also resolve selection signals at the level of individual codons or specific lineages, offering finer-grained detail than an aggregate dN/dS ratio for a gene or pathway.
-
Functional term enrichment in supergroup L genomes is described narratively (e.g., pathways 'enriched') based on gene presence-absence comparisons via Roary and ModelSEED.↳ Could also: A formal statistical enrichment test, such as a hypergeometric/Fisher's exact test or permutation-based gene-set enrichment analysis — A formal enrichment test could also quantify the statistical significance of pathway overrepresentation, complementing the descriptive presence-absence comparison already performed.
-
Correlation results are reported as Spearman rho and adjusted p-values without an explicit mention of confidence intervals.↳ Could also: Bootstrap-derived confidence intervals around the Spearman rho estimates — Reporting a confidence interval alongside rho and the p-value could also convey the precision of the correlation estimate, which is particularly informative given the modest number (21) of sampled communities.
-
Wolbachia detection was positive in only 1 of 16 sampled sites, with samples from that site subsequently combined/pooled to improve genome assembly coverage.↳ Could also: A formal detection-sensitivity or rarefaction/accumulation-curve analysis across sampling sites — Such an analysis could also characterize how sampling effort relates to the likelihood of detecting a low-prevalence symbiont, complementing the qualitative positive/negative site tally already presented.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-35547116
Paper: Lefoulon et al. (2022) Discovery of Early-Branching Wolbachia Reveals Functional Enrichment on Horizontally Transferred Genes. Front. Microbiol. 13:867392. DOI 10.3389/fmicb.2022.867392 · PMCID PMC9084900.
Linked code: https://github.com/tseemann/barrnap — a third-party rRNA-prediction tool (used by the authors inside Prokka). Per BRIEF rule P16, applying this existing third-party tool to the paper's own data is an equally valid reproduction. barrnap detects 5S/16S/23S rRNA in a genome FASTA → GFF3.
Linked data: SRA BioProject PRJNA687334 (raw Illumina HiSeq metagenome reads of nematode pools); plus a set of public NCBI genome assemblies used for comparative analysis (the wTex assembly the authors deposited + 5 reference Wolbachia genomes).
In scope (pipeline-derived, public data, reproducible)
A. Table 1 — comparative genome statistics for 6 Wolbachia strains ← PRIMARY 1:1 TARGET
The authors deposited/used 6 assembled genomes. All are public on NCBI. The Table 1 columns are deterministic functions of those FASTAs + the Prokka/barrnap annotation pipeline the Methods describe ("Prokka v1.14.6 using Prodigal, HMMER3, BLAST+, Barrnap for rRNAs, Aragorn for tRNAs"):
| column | how reproduced | determinism |
|---|---|---|
| Total length (bp) | sum of contig lengths from FASTA | exact (FASTA only) |
| %GC | base composition from FASTA | exact (FASTA only) |
| # contigs/scaffolds | count of records in FASTA | exact (FASTA only) |
| rRNAs | barrnap (the linked tool) | exact/within-tol |
| tRNAs | Prokka → Aragorn | within-tol (version-sensitive) |
| Genes | Prokka → Prodigal | within-tol (version-sensitive) |
| Pseudogenes | Prokka | within-tol |
Strains + accessions (Table 1 / Data availability):
- wTex — JAIXMJ000000000 (GenBank WGS; this study)
- wPpe — NZ_MJMG01000001.1 (RefSeq WGS)
- wPni — JACVWV010000040.1 (GenBank WGS)
- wFol — NZ_CP015510.2 (RefSeq complete)
- wCfeT — NZ_CP051156.1 (RefSeq complete)
- wChem — NZ_CP061738.1 (RefSeq complete)
Pipeline: barrnap 0.9+ (rRNA) + Prokka 1.14.6 (genes/tRNA/pseudogenes) + plain FASTA stats (length/GC/contigs/N50).
B. 16S rRNA sequence divergence (HARDER, stretch)
Reported: max 16S difference among wTex/wPpe/wRad = 4.211%; PPN-clade vs wPni avg 3.934%. Reproduce by extracting 16S with barrnap from each genome, aligning (MAFFT), computing pairwise %identity. Caveat: wRad accession not given in the paper's data table → may only partially reproduce (wTex vs wPpe).
C. wTex de-novo assembly from PRJNA687334 (HARDEST, stretch — likely partial)
Reported wTex stats (1,013,022 bp, 33.49% GC, 192 scaffolds, N50 10,082, 989 genes). Full iterative subtractive metaSPAdes pipeline. Exact reproduction is infeasible because the published Methods include manual curation in Geneious Prime (proprietary, no license) + a hand-tuned 3-round subtractive read-extraction loop. We can attempt a single metaSPAdes pass for an order-of-magnitude sanity check only, not a 1:1 match.
Out of scope (wet-lab / manual / external / non-reproducible)
- Field sampling, nematode isolation (Baermann/sucrose), DNA extraction, PCR pre-screen — wet-lab.
- Manual genome curation in Geneious Prime v2020.0.4 (ProgressiveMauve/LASTZ) — proprietary GUI, no license, manual.
- HGT identification (asd2/lysC, sdh, thiE/thiM/thiD, tlcA, BON/deoC, PCBD1) — manual BLAST + synteny interpretation, no scripted entrypoint.
- GO enrichment (topGO weight01.fisher) + dN/dS (KaKs Calculator) — depend on Roary ortholog sets and specific PRANK alignments not shipped; reproducible in principle but no released ortholog tables → not a clean 1:1; deferred.
- Phylogenetic bootstrap/posterior support (RAxML/MrBayes) — depend on specific alignments + curation not deposited; not a clean numeric comparison.
- SRA screening of "~3,400 amplicon experiments" → 81 runs / 61 unique seqs — the exact query list / 3,400-experiment set
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
Table 1 comparative genome statistics reproduce essentially 1:1 from the publicly deposited FASTAs — total length and contig counts exact for all 6 strains (wTex within 1 bp), barrnap rRNA=3 exact, Prokka CDS exact for 5/6, tRNA exact for all, GC differences fully explained by N-base handling. The only real deviation is the pseudogene column (reproduced 0 vs reported 2–55), which is on our methodology side — plain Prokka has no pseudogene caller while the paper used NCBI PGAP — not an authors' defect or fabrication. The reported values are derivable from the shared data (q5 green), but the paper's actual thesis (HGT functional enrichment, early-branching phylogeny) was out of scope, so the core claim is only partially confirmed (q7 yellow). Overall a solid reproduction with explainable deviations.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.