Discovery of a novel filamentous prophage in the genome of the Mimosa pudica microsymbiont Cupriavidus taiwanensis STM 6018.
Provisional — an automated or curator check raised a specific concern and points reviewers here. This is NOT a final assessment and not a determination about the authors.
The main results reproduced, with only marginal, non-material deviations.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Same input data as the authors
- ✓Reported values were directly comparable
- ✓Reported values are derivable from the shared data
- ✓The central claim held under reproduction
- 🟡A deviation arose in the data or preprocessing
- 🟡A deviation was attributed to the published material
- 🟡The deviation was non-trivial in magnitude
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
STRONG PARTIAL reproduction of a C. taiwanensis STM 6018 genome-announcement + comparative-genomics paper. ACCESSION NOTE: the brief's 'PRJNA61615' is NOT garbled - it is the reference type strain C. taiwanensis LMG19424T BioProject (the ANI/16S comparison target per Methods); the SUBJECT genome is PRJNA165307 / GCF_000472465.1. EXACT matches (FASTA/metadata): genome length 6,553,639 nt (C1), GC 66.90% (C2), 80 contigs (C3), 14,977,300 reads (C11). The paper's headline novel filamentous prophage (C6) reproduced cleanly: scaffold_0.1=NZ_AXAK01000001.1, 7,540 bp, 61.25% GC (vs 61.1%), exactly 11 genes, INCLUDING a zonular-occludens-toxin (Zot, pfam05707) - the exact Inoviridae marker the authors used - independently confirming the finding. Mu-like prophage (C7) confirmed on scaffold 19.20 (gpT major head gene), span ~31 kb vs reported 36,733 bp. ANI (C8) reproduced: ANIb 98.84% (paper 98.82%, ~exact), 92.5% aligned (paper 92.6%); ANIm 99.03% (paper 98.90%). 16S (C9) ~100% identity (paper 99.6%) over 1,075 bp (draft-assembly 16S fragmentation). DIVERGENCES, both explained by annotation-pipeline substitution (paper used JGI/IMG-ER + an unspecified pan-genome tool, both hosted/version-locked; we use the deposited RefSeq/PGAP annotation + prokka/roary): gene-function fraction C5 (RefSeq 90.6% vs IMG-ER 80.69%) and 3-strain pan-genome C10 (roary 8,571 vs 5,205). Gene totals C4 within ~0.8%. No fabrication concerns: every divergence traces to a known pipeline difference, not a non-derivable number. OUT OF SCOPE (not attempted): de-novo assembly (Velvet+Allpaths-LG, version-locked), JGI/IMG-ER annotation (hosted); wgsim (the brief's code repo) built cleanly at pinned commit a12da33 and was demonstrated - it produced intermediate simulated reads to aid assembly, no reported number to match.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-30
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-30no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetThe paper investigates the genome of the Mimosa pudica microsymbiont Cupriavidus taiwanensis STM 6018, focusing on whether and how integrated prophages (particularly a novel filamentous phage) are present in this and other rhizobial genomes and what genomic features support phage-host interactions.
- ★ The STM 6018 genome contains two prophages: a complete Mu-like capsular phage and a filamentous phage that integrates into a putative dif site. finding
- ★ This is the first characterization of a filamentous phage found within the genome of a rhizobial strain. finding
- ★ Filamentous prophage sequences were identified in several Beta-rhizobial strains but not in any Alphaproteobacterial rhizobia. finding
- ★ STM 6018 belongs to the same species as C. taiwanensis LMG19424T based on ANI and 16S rRNA analyses. finding
- ★ STM 6018 shares >99.97% bp identity in nod/nif/noeM gene clusters with C. taiwanensis LMG19424T and "Cupriavidus neocaledonicus" STM 6070, supporting horizontal gene transfer origin of symbiotic Cupriavidus populations. finding
- The STM 6018 draft genome (6,553,639 bp, 80 scaffolds, 5,864 protein-coding genes, 61 RNA genes) was sequenced as part of the GEBA-RNB project at JGI. resource
- STM 6018 possesses conserved type I, II, III, IV and VI secretion systems and Type IV pilus systems relevant to both symbiosis and phage interactions. finding
- STM 6018 out-competes Paraburkholderia phymatum STM815T for nodulation of M. pudica var. unijuga. finding
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| whole genome sequencing/draft assembly and annotation | Cupriavidus taiwanensis STM 6018 | none | genome size, GC content, gene counts, scaffold number | — |
| 16S rRNA gene phylogenetic/sequence identity analysis | STM 6018 vs. Cupriavidus/Ralstonia type strains | none | percent sequence identity, phylogenetic placement | — |
| Average Nucleotide Identity (ANIb, ANIm, ANIg) analysis | STM 6018 genome vs. other Cupriavidus and Ralstonia genomes | none | percent identical DNA / ANI values | MUMmer, BLASTN (jSpecies), nSimScan |
| pangenome analysis | STM 6018, C. taiwanensis LMG19424T, "C. neocaledonicus" STM 6070 | none | core, variable, and strain-unique gene counts | progressiveMauve |
| plant nodulation/nitrogen fixation phenotyping | STM 6018 inoculated onto multiple Mimosa species (M. pudica, M. pigra, M. caesalpiniaefolia, M. acustipulata, M. scabrella) | bacterial inoculation of host plants | nodulation (Nod+/-) and nitrogen fixation (Fix+/-) status | — |
| nodule occupancy competition assay | STM 6018 vs. gfp-marked Paraburkholderia phymatum STM815T on M. pudica varieties | co-inoculation competition | percent nodule occupation | — |
| comparative gene cluster/synteny analysis of secretion and pilus systems | STM 6018 vs. LMG19424T genomes | none | identification and synteny of T1SS, T2SS, T3SS, T4SS/T6SS and Type IV pilus gene clusters | — |
| gene expression analysis (referenced prior study) | C. taiwanensis LMG19424T cultures exposed to M. pudica root exudates | root exudate exposure | up-regulation of pil genes (pilVWXYE, pilQPONM) | — |
- – STM 6018 draft genome comprises 6,553,639 bp, 66.90% GC content, 80 scaffolds, 5,864 protein-coding genes and 61 RNA genes
- ▲ ANIb and ANIm values >98% (>92% conserved DNA) and ANIg >99% between STM 6018 and C. taiwanensis LMG19424T >98%/>99%
- ▲ 16S rRNA gene of STM 6018 shares 99.6% identity with C. taiwanensis LMG19424T over 1,426 bp 99.6%
- ▲ nod/nif/noeM gene clusters shared nearly 100% identity among STM 6018, LMG19424T and STM 6070 >99.97%
- – 244 genes unique to STM 6018 included two intact prophage regions (Mu-like and filamentous phage)
- – STM 6018 nodulates and fixes N2 with M. pudica and M. pigra, nodulates but does not fix with M. caesalpiniaefolia, variable nodulation without fixation on M. acustipulata, and does not nodulate M. scabrella
- ▲ STM 6018 out-competed gfp-marked P. phymatum STM815T for nodulation of M. pudica var. unijuga 80% nodule occupation
- ▼ Translated NifV identity was 100% among STM 6018, LMG19424T and STM 6070 but only 77.89% and 48.13% in AMP6 and UYPR2.512 respectively 77.89%/48.13% vs 100%
- count 6,553,639 bp genome, 80 scaffolds, 5,864 protein-coding genes, 61 RNA genes (STM 6018 draft genome assembly)
- other 449x sequence coverage (genome assembly coverage)
- correlation ANIb 95.16%, ANIm 98.90%, ANIg 99.03% vs LMG19424T (species assignment of STM 6018)
- correlation ANIb 82.59%, ANIm 93.99%, ANIg 94.68% vs STM 6070 (comparison to next closest rhizobial relative)
- other 16S rRNA identity 99.6% (LMG19424T), 99.37% (X1T), 99.14% (STM6070), 99.02% (ASC-732T), 98.5% (ATCC43291T) (16S rRNA phylogenetic comparison)
- count pangenome 5,205 genes; variable genome 438 genes; 244 genes unique to STM 6018 (pangenome analysis with LMG19424T and STM 6070)
- other nod/nif/noeM identity >99.97% bp over 100% coverage (symbiotic gene cluster conservation)
- other nodule occupation 80% (var. unijuga), 30% (var. tetrandra), 5% (var. hispida) (competition assay vs P. phymatum STM815T)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This paper is a descriptive/comparative genomics study rather than a hypothesis-testing study: it characterizes the draft genome of Cupriavidus taiwanensis STM 6018 and its prophages using sequence-based similarity metrics (average nucleotide identity, 16S rRNA % identity, pangenome gene counts, synteny) and reports plant symbiosis phenotypes (nodulation/nitrogen fixation) as categorical outcomes from small numbers of replicate plants. No inferential statistical hypothesis tests, p-values, or formal uncertainty estimates are reported in the provided text.
-
Average nucleotide identity (ANIm, ANIb, ANIg) and 16S rRNA % identity values are reported as single-point percentages for species/strain assignment.↳ Could also: Reporting these alongside a measure of alignment confidence (e.g., % of genome aligned already given, but also bootstrap or replicate-based variability) or supplementing with digital DNA-DNA hybridization (dDDH) — Would let readers gauge how much the species-boundary call depends on algorithm choice or alignment coverage, and cross-validate the identity-based classification with an independent metric.
-
Nodulation (Nod) and nitrogen-fixation (Fix) phenotypes for several Mimosa species were scored from triplicate plant trials and reported as categorical +/− outcomes (including a mixed '+/-' result for M. acustipulata) without a formal statistical comparison.↳ Could also: A categorical association test such as Fisher's exact test, or reporting occupancy/positivity as a proportion with an exact (e.g., Clopper–Pearson) confidence interval — Would quantify the certainty behind phenotype calls made from a small number of replicate plants, which is particularly informative for the ambiguous M. acustipulata result.
-
Nodule-occupancy percentages from a prior competition study (STM 6018 vs. P. phymatum STM815, 80%/30%/5% across three Mimosa varieties) are cited descriptively without a stated test of association between strain and host variety.↳ Could also: A chi-square or Fisher's exact test of occupancy counts across host varieties, or a generalized linear mixed model accounting for plant/pot-level clustering — Would formalize whether the apparent difference in competitiveness across host varieties exceeds what could arise from sampling variation alone.
-
Gene/protein sequence conservation across strains (e.g., NifV, NoeM percent identities) is used to support relatedness and horizontal-transfer claims without phylogenetic support values.↳ Could also: Reporting bootstrap or posterior-probability support values from a phylogenetic tree alongside pairwise % identity — Bootstrap/support values give a standard, widely used way to communicate confidence in the topology underlying relatedness and horizontal-gene-transfer inferences, complementing raw % identity.
-
Pangenome composition (core, variable, unique gene counts) is presented as fixed counts from a single pangenome analysis run.↳ Could also: Reporting pangenome accumulation/rarefaction curves or sensitivity of core/variable gene counts to clustering thresholds — Would show how stable the reported gene counts are to parameter choices, which is a common complementary check in pangenome studies.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — PMID 36925474
Paper: Discovery of a novel filamentous prophage in the genome of the Mimosa pudica microsymbiont Cupriavidus taiwanensis STM 6018. Front. Microbiol. 2023;14:1082107. DOI 10.3389/fmicb.2023.1082107 · PMCID PMC10011098.
This is a genome announcement + comparative-genomics paper. The wet-lab DNA
sequencing produced reads; the rest is a bioinformatic pipeline:
Illumina HiSeq2000 → DUK filter → Velvet assembly → wgsim (1–3 kb simulated PE reads) → Allpaths-LG final assembly → JGI/IMG-ER annotation (Prodigal, tRNAScan-SE, SILVA) → comparative genomics (ANI, 16S, pan-genome) + prophage discovery.
Accession reconciliation (IMPORTANT)
- Brief says
sra:PRJNA61615— this accession does NOT exist (NCBI SRA/BioProject return "phrase not found" / a spurious match to the generic RefSeq-annotation umbrella PRJNA224116). It is a garbled accession. - Correct genome project:
PRJNA165307("Cupriavidus taiwanensis STM 6018 Genome sequencing and assembly", JGI). It has 3 SRA runs (SRR3943816/17/18, SRP079279, BioSample SAMN02440787). - Deposited assembly:
GCA_000472465.1/GCF_000472465.1(ASM47246v1), WGS prefix AXAK01, submitter DOE-JGI, status "Contig". This is the assembled genome the paper describes; it is the ground-truth artifact for the genome-derived numbers.
In scope — pipeline-derived results we attempt
| id | reported result | paper loc | pipeline / how reproduced |
|---|---|---|---|
| C1 | Genome size 6,553,639 nt | Genome props | recompute total length from GCA_000472465.1 FASTA (seqkit) |
| C2 | GC content 66.90% | Genome props | recompute GC% from FASTA |
| C3 | 80 scaffolds / 80 contigs | Genome props | count records in FASTA |
| C4 | 5,925 genes (5,864 CDS + 61 RNA) | Genome props | count features in GFF (GCF annotation) |
| C5 | 80.69% genes w/ predicted function | Genome props | fraction of CDS w/ non-hypothetical product (approx) |
| C6 | Filamentous prophage: 7,540 bp, 61.1% GC, 11 genes, scaffold 0.1 pos 19,342–26,881 | Results | extract region from matching contig; measure len+GC; count CDS in GFF over region |
| C7 | Mu-like prophage ~36,733 bp, scaffold 19.20 | Results | locate region; measure length |
| C8 | ANIm 98.90% (95.16% aligned) / ANIb 98.82% (92.6%) vs LMG19424T | Results | pyani ANIm (MUMmer) + ANIb (BLAST) STM6018 vs C. taiwanensis LMG19424 |
| C9 | 16S rRNA identity 99.6% over 1,426 bp vs LMG19424T | Results | extract 16S, blastn/needle vs LMG19424 16S |
| C10 | Pan-genome 5,205 / variable 438 / 244 unique (3-strain) | Results | roary/panaroo over STM6018+LMG19424+STM6070 (stretch) |
| C11 | 14,977,300 reads totaling 2,245 Mb | Genome props | SRA metadata: SRR3943817 = 7,488,650 spots ×2 = 14,977,300; 2,246,595,000 bp ✓ |
Out of scope — not attempted (and why)
- De-novo assembly (Velvet + Allpaths-LG) producing the 6.55 Mb genome from raw reads: multi-tool, version-sensitive (Velvet 1.1.04, Allpaths-LG r39750), JGI-internal parameters; not 1:1 reproducible and not the point. We take the deposited assembly as the artifact and verify the numbers derived from it.
- JGI/IMG-ER annotation pipeline (functional annotation, COG/KEGG/Pfam): IMG-ER is a hosted JGI service, not runnable here. We use the deposited RefSeq/GenBank annotation (GFF) for gene counts.
- wgsim's role was to generate intermediate 1–3 kb simulated PE reads to aid the assembler — it produces no reported numeric result. We will BUILD wgsim (the named repo) and demonstrate the documented command on the assembly as an executability check, but there is no paper number to match against it.
- Wet-lab (DNA extraction, sequencing, Mimosa nodulation phenotype, microscopy): out of scope.
- 449× coverage: derived stat; recorded but note it is inconsistent with 2,245 Mb / 6.55 Mb ≈ 343× (paper's own arithmetic) — flag, not a repro target.
Reproduction substrate
- Heavy
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
This genome-announcement / comparative-genomics paper reproduces strongly: the deposited genome's headline properties (length 6,553,639 nt, GC 66.90%, 80 contigs, 14,977,300 reads) match exactly, the novel filamentous prophage (7,540 bp, 11 genes, exact coordinates, Zot marker) and ANI conspecificity (ANIb 98.84% vs 98.82%) reproduce cleanly, and the central conclusion holds fully. The two genuine divergences — C5 functional fraction (80.69% vs 90.63%) and C10 pan-genome (5,205 vs 8,571) — sit entirely on our methodology side, caused by substituting RefSeq/PGAP + prokka/roary for the paper's hosted IMG-ER + an unspecified pan-genome tool; both are pipeline/version effects, not non-derivable numbers, so there is no fabrication concern. One minor authors'-side anomaly: the reported 449x coverage is internally inconsistent with the paper's own 2,245 Mb / 6.55 Mb (=343x), flagged for human review.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.