Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Discovery of a novel filamentous prophage in the genome of the Mimosa pudica microsymbiont Cupriavidus taiwanensis STM 6018.

Front Microbiol · 2023
L1 74/100 3/4
⚑ Flagged for review — a reproduced result did not match the reported value

Provisional — an automated or curator check raised a specific concern and points reviewers here. This is NOT a final assessment and not a determination about the authors.

Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +3
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
  • Reported values are derivable from the shared data
  • The central claim held under reproduction
What did not (or only partly)
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡The deviation was non-trivial in magnitude
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
74/100
Reproducibility score
at the mean
vs. all fields · 1173 studies
🎯 Scores higher than 43% of all assessed papers rank 644 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

STRONG PARTIAL reproduction of a C. taiwanensis STM 6018 genome-announcement + comparative-genomics paper. ACCESSION NOTE: the brief's 'PRJNA61615' is NOT garbled - it is the reference type strain C. taiwanensis LMG19424T BioProject (the ANI/16S comparison target per Methods); the SUBJECT genome is PRJNA165307 / GCF_000472465.1. EXACT matches (FASTA/metadata): genome length 6,553,639 nt (C1), GC 66.90% (C2), 80 contigs (C3), 14,977,300 reads (C11). The paper's headline novel filamentous prophage (C6) reproduced cleanly: scaffold_0.1=NZ_AXAK01000001.1, 7,540 bp, 61.25% GC (vs 61.1%), exactly 11 genes, INCLUDING a zonular-occludens-toxin (Zot, pfam05707) - the exact Inoviridae marker the authors used - independently confirming the finding. Mu-like prophage (C7) confirmed on scaffold 19.20 (gpT major head gene), span ~31 kb vs reported 36,733 bp. ANI (C8) reproduced: ANIb 98.84% (paper 98.82%, ~exact), 92.5% aligned (paper 92.6%); ANIm 99.03% (paper 98.90%). 16S (C9) ~100% identity (paper 99.6%) over 1,075 bp (draft-assembly 16S fragmentation). DIVERGENCES, both explained by annotation-pipeline substitution (paper used JGI/IMG-ER + an unspecified pan-genome tool, both hosted/version-locked; we use the deposited RefSeq/PGAP annotation + prokka/roary): gene-function fraction C5 (RefSeq 90.6% vs IMG-ER 80.69%) and 3-strain pan-genome C10 (roary 8,571 vs 5,205). Gene totals C4 within ~0.8%. No fabrication concerns: every divergence traces to a known pipeline difference, not a non-derivable number. OUT OF SCOPE (not attempted): de-novo assembly (Velvet+Allpaths-LG, version-locked), JGI/IMG-ER annotation (hosted); wgsim (the brief's code repo) built cleanly at pinned commit a12da33 and was demonstrated - it produced intermediate simulated reads to aid assembly, no reported number to match.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-30
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-30
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The paper investigates the genome of the Mimosa pudica microsymbiont Cupriavidus taiwanensis STM 6018, focusing on whether and how integrated prophages (particularly a novel filamentous phage) are present in this and other rhizobial genomes and what genomic features support phage-host interactions.

Core claims
  • The STM 6018 genome contains two prophages: a complete Mu-like capsular phage and a filamentous phage that integrates into a putative dif site. finding
  • This is the first characterization of a filamentous phage found within the genome of a rhizobial strain. finding
  • Filamentous prophage sequences were identified in several Beta-rhizobial strains but not in any Alphaproteobacterial rhizobia. finding
  • STM 6018 belongs to the same species as C. taiwanensis LMG19424T based on ANI and 16S rRNA analyses. finding
  • STM 6018 shares >99.97% bp identity in nod/nif/noeM gene clusters with C. taiwanensis LMG19424T and "Cupriavidus neocaledonicus" STM 6070, supporting horizontal gene transfer origin of symbiotic Cupriavidus populations. finding
  • The STM 6018 draft genome (6,553,639 bp, 80 scaffolds, 5,864 protein-coding genes, 61 RNA genes) was sequenced as part of the GEBA-RNB project at JGI. resource
  • STM 6018 possesses conserved type I, II, III, IV and VI secretion systems and Type IV pilus systems relevant to both symbiosis and phage interactions. finding
  • STM 6018 out-competes Paraburkholderia phymatum STM815T for nodulation of M. pudica var. unijuga. finding
Experimental setups
Assay System Perturbation Readout Platform
whole genome sequencing/draft assembly and annotation Cupriavidus taiwanensis STM 6018 none genome size, GC content, gene counts, scaffold number
16S rRNA gene phylogenetic/sequence identity analysis STM 6018 vs. Cupriavidus/Ralstonia type strains none percent sequence identity, phylogenetic placement
Average Nucleotide Identity (ANIb, ANIm, ANIg) analysis STM 6018 genome vs. other Cupriavidus and Ralstonia genomes none percent identical DNA / ANI values MUMmer, BLASTN (jSpecies), nSimScan
pangenome analysis STM 6018, C. taiwanensis LMG19424T, "C. neocaledonicus" STM 6070 none core, variable, and strain-unique gene counts progressiveMauve
plant nodulation/nitrogen fixation phenotyping STM 6018 inoculated onto multiple Mimosa species (M. pudica, M. pigra, M. caesalpiniaefolia, M. acustipulata, M. scabrella) bacterial inoculation of host plants nodulation (Nod+/-) and nitrogen fixation (Fix+/-) status
nodule occupancy competition assay STM 6018 vs. gfp-marked Paraburkholderia phymatum STM815T on M. pudica varieties co-inoculation competition percent nodule occupation
comparative gene cluster/synteny analysis of secretion and pilus systems STM 6018 vs. LMG19424T genomes none identification and synteny of T1SS, T2SS, T3SS, T4SS/T6SS and Type IV pilus gene clusters
gene expression analysis (referenced prior study) C. taiwanensis LMG19424T cultures exposed to M. pudica root exudates root exudate exposure up-regulation of pil genes (pilVWXYE, pilQPONM)
Key results
  • STM 6018 draft genome comprises 6,553,639 bp, 66.90% GC content, 80 scaffolds, 5,864 protein-coding genes and 61 RNA genes
  • ANIb and ANIm values >98% (>92% conserved DNA) and ANIg >99% between STM 6018 and C. taiwanensis LMG19424T >98%/>99%
  • 16S rRNA gene of STM 6018 shares 99.6% identity with C. taiwanensis LMG19424T over 1,426 bp 99.6%
  • nod/nif/noeM gene clusters shared nearly 100% identity among STM 6018, LMG19424T and STM 6070 >99.97%
  • 244 genes unique to STM 6018 included two intact prophage regions (Mu-like and filamentous phage)
  • STM 6018 nodulates and fixes N2 with M. pudica and M. pigra, nodulates but does not fix with M. caesalpiniaefolia, variable nodulation without fixation on M. acustipulata, and does not nodulate M. scabrella
  • STM 6018 out-competed gfp-marked P. phymatum STM815T for nodulation of M. pudica var. unijuga 80% nodule occupation
  • Translated NifV identity was 100% among STM 6018, LMG19424T and STM 6070 but only 77.89% and 48.13% in AMP6 and UYPR2.512 respectively 77.89%/48.13% vs 100%
Key statistics
  • count 6,553,639 bp genome, 80 scaffolds, 5,864 protein-coding genes, 61 RNA genes (STM 6018 draft genome assembly)
  • other 449x sequence coverage (genome assembly coverage)
  • correlation ANIb 95.16%, ANIm 98.90%, ANIg 99.03% vs LMG19424T (species assignment of STM 6018)
  • correlation ANIb 82.59%, ANIm 93.99%, ANIg 94.68% vs STM 6070 (comparison to next closest rhizobial relative)
  • other 16S rRNA identity 99.6% (LMG19424T), 99.37% (X1T), 99.14% (STM6070), 99.02% (ASC-732T), 98.5% (ATCC43291T) (16S rRNA phylogenetic comparison)
  • count pangenome 5,205 genes; variable genome 438 genes; 244 genes unique to STM 6018 (pangenome analysis with LMG19424T and STM 6070)
  • other nod/nif/noeM identity >99.97% bp over 100% coverage (symbiotic gene cluster conservation)
  • other nodule occupation 80% (var. unijuga), 30% (var. tetrandra), 5% (var. hispida) (competition assay vs P. phymatum STM815T)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This paper is a descriptive/comparative genomics study rather than a hypothesis-testing study: it characterizes the draft genome of Cupriavidus taiwanensis STM 6018 and its prophages using sequence-based similarity metrics (average nucleotide identity, 16S rRNA % identity, pangenome gene counts, synteny) and reports plant symbiosis phenotypes (nodulation/nitrogen fixation) as categorical outcomes from small numbers of replicate plants. No inferential statistical hypothesis tests, p-values, or formal uncertainty estimates are reported in the provided text.

Replicationmixed Sample sizeSymbiotaxonomy (Nod/Fix) assays on Mimosa species other than M. pudica var. unijuga/tetrandra were performed 'using triplicates' per Table 2; sample sizes for genome-based comparisons (ANI, % identity) are not expressed as replicate n. GroupsGenomes/strains of Cupriavidus and related taxa compared by sequence identity; Mimosa host species compared for nodulation (Nod) and nitrogen-fixation (Fix) outcomes Pairingna Randomization/blindingnot stated Dispersionnone Exact p-valuesno Effect sizesno Confidence intervalsno Multiplicity correctionno
Approaches that could also have been used
  • Average nucleotide identity (ANIm, ANIb, ANIg) and 16S rRNA % identity values are reported as single-point percentages for species/strain assignment.
    Could also: Reporting these alongside a measure of alignment confidence (e.g., % of genome aligned already given, but also bootstrap or replicate-based variability) or supplementing with digital DNA-DNA hybridization (dDDH) — Would let readers gauge how much the species-boundary call depends on algorithm choice or alignment coverage, and cross-validate the identity-based classification with an independent metric.
  • Nodulation (Nod) and nitrogen-fixation (Fix) phenotypes for several Mimosa species were scored from triplicate plant trials and reported as categorical +/− outcomes (including a mixed '+/-' result for M. acustipulata) without a formal statistical comparison.
    Could also: A categorical association test such as Fisher's exact test, or reporting occupancy/positivity as a proportion with an exact (e.g., Clopper–Pearson) confidence interval — Would quantify the certainty behind phenotype calls made from a small number of replicate plants, which is particularly informative for the ambiguous M. acustipulata result.
  • Nodule-occupancy percentages from a prior competition study (STM 6018 vs. P. phymatum STM815, 80%/30%/5% across three Mimosa varieties) are cited descriptively without a stated test of association between strain and host variety.
    Could also: A chi-square or Fisher's exact test of occupancy counts across host varieties, or a generalized linear mixed model accounting for plant/pot-level clustering — Would formalize whether the apparent difference in competitiveness across host varieties exceeds what could arise from sampling variation alone.
  • Gene/protein sequence conservation across strains (e.g., NifV, NoeM percent identities) is used to support relatedness and horizontal-transfer claims without phylogenetic support values.
    Could also: Reporting bootstrap or posterior-probability support values from a phylogenetic tree alongside pairwise % identity — Bootstrap/support values give a standard, widely used way to communicate confidence in the topology underlying relatedness and horizontal-gene-transfer inferences, complementing raw % identity.
  • Pangenome composition (core, variable, unique gene counts) is presented as fixed counts from a single pangenome analysis run.
    Could also: Reporting pangenome accumulation/rarefaction curves or sensitivity of core/variable gene counts to clustering thresholds — Would show how stable the reported gene counts are to parameter choices, which is a common complementary check in pangenome studies.
Software: MUMmer (for ANIm) · BLASTN via jSpecies (for ANIb) · nSimScan (pairwise bidirectional best hits, for ANIg)

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — PMID 36925474

Paper: Discovery of a novel filamentous prophage in the genome of the Mimosa pudica microsymbiont Cupriavidus taiwanensis STM 6018. Front. Microbiol. 2023;14:1082107. DOI 10.3389/fmicb.2023.1082107 · PMCID PMC10011098.

This is a genome announcement + comparative-genomics paper. The wet-lab DNA sequencing produced reads; the rest is a bioinformatic pipeline: Illumina HiSeq2000 → DUK filter → Velvet assembly → wgsim (1–3 kb simulated PE reads) → Allpaths-LG final assembly → JGI/IMG-ER annotation (Prodigal, tRNAScan-SE, SILVA) → comparative genomics (ANI, 16S, pan-genome) + prophage discovery.

Accession reconciliation (IMPORTANT)

  • Brief says sra:PRJNA61615 — this accession does NOT exist (NCBI SRA/BioProject return "phrase not found" / a spurious match to the generic RefSeq-annotation umbrella PRJNA224116). It is a garbled accession.
  • Correct genome project: PRJNA165307 ("Cupriavidus taiwanensis STM 6018 Genome sequencing and assembly", JGI). It has 3 SRA runs (SRR3943816/17/18, SRP079279, BioSample SAMN02440787).
  • Deposited assembly: GCA_000472465.1 / GCF_000472465.1 (ASM47246v1), WGS prefix AXAK01, submitter DOE-JGI, status "Contig". This is the assembled genome the paper describes; it is the ground-truth artifact for the genome-derived numbers.

In scope — pipeline-derived results we attempt

id reported result paper loc pipeline / how reproduced
C1 Genome size 6,553,639 nt Genome props recompute total length from GCA_000472465.1 FASTA (seqkit)
C2 GC content 66.90% Genome props recompute GC% from FASTA
C3 80 scaffolds / 80 contigs Genome props count records in FASTA
C4 5,925 genes (5,864 CDS + 61 RNA) Genome props count features in GFF (GCF annotation)
C5 80.69% genes w/ predicted function Genome props fraction of CDS w/ non-hypothetical product (approx)
C6 Filamentous prophage: 7,540 bp, 61.1% GC, 11 genes, scaffold 0.1 pos 19,342–26,881 Results extract region from matching contig; measure len+GC; count CDS in GFF over region
C7 Mu-like prophage ~36,733 bp, scaffold 19.20 Results locate region; measure length
C8 ANIm 98.90% (95.16% aligned) / ANIb 98.82% (92.6%) vs LMG19424T Results pyani ANIm (MUMmer) + ANIb (BLAST) STM6018 vs C. taiwanensis LMG19424
C9 16S rRNA identity 99.6% over 1,426 bp vs LMG19424T Results extract 16S, blastn/needle vs LMG19424 16S
C10 Pan-genome 5,205 / variable 438 / 244 unique (3-strain) Results roary/panaroo over STM6018+LMG19424+STM6070 (stretch)
C11 14,977,300 reads totaling 2,245 Mb Genome props SRA metadata: SRR3943817 = 7,488,650 spots ×2 = 14,977,300; 2,246,595,000 bp ✓

Out of scope — not attempted (and why)

  • De-novo assembly (Velvet + Allpaths-LG) producing the 6.55 Mb genome from raw reads: multi-tool, version-sensitive (Velvet 1.1.04, Allpaths-LG r39750), JGI-internal parameters; not 1:1 reproducible and not the point. We take the deposited assembly as the artifact and verify the numbers derived from it.
  • JGI/IMG-ER annotation pipeline (functional annotation, COG/KEGG/Pfam): IMG-ER is a hosted JGI service, not runnable here. We use the deposited RefSeq/GenBank annotation (GFF) for gene counts.
  • wgsim's role was to generate intermediate 1–3 kb simulated PE reads to aid the assembler — it produces no reported numeric result. We will BUILD wgsim (the named repo) and demonstrate the documented command on the assembly as an executability check, but there is no paper number to match against it.
  • Wet-lab (DNA extraction, sequencing, Mimosa nodulation phenotype, microscopy): out of scope.
  • 449× coverage: derived stat; recorded but note it is inconsistent with 2,245 Mb / 6.55 Mb ≈ 343× (paper's own arithmetic) — flag, not a repro target.

Reproduction substrate

  • Heavy
C1
Reported
6,553,639 nt genome
Reproduced
6,553,639 nt (seqkit/biopython on deposited FASTA)
exact
C2
Reported
66.90% GC
Reproduced
66.90% (seqkit)
exact
C3
Reported
80 scaffolds/contigs
Reproduced
80 contigs (FASTA record count)
exact
C4
Reported
5,925 genes (5,864 CDS + 61 RNA)
Reproduced
5,881 genes (5,819 CDS + 61 RNA + 100 pseudogenes)
within tolerance
C5
Reported
80.69% genes with predicted function
Reproduced
90.63% (RefSeq non-hypothetical) - PGAP vs IMG-ER annotation
did not match
C6
Reported
filamentous prophage 7,540 bp, 61.1% GC, 11 genes, scaffold 0.1 19,342-26,881
Reproduced
7,540 bp, 61.25% GC, 11 genes at NZ_AXAK01000001.1 (=scaffold_0.1) 19,342-26,881; Zot/pfam05707 marker present
within tolerance
C7
Reported
Mu-like prophage 36,733 bp, scaffold 19.20
Reproduced
scaffold 19.20 = NZ_AXAK01000020.1; Mu-like prophage confirmed (gpT major head gene); phage-gene span 31,201 bp
partial
C8
Reported
ANIm 98.90% (95.16% aln); ANIb 98.82% (92.6% aln); ANIg 99.03%
Reproduced
ANIm 99.03% (92-93% aln); ANIb 98.84% (91.6-92.5% aln) via pyani
within tolerance
C9
Reported
16S 99.6% over 1,426 bp
Reproduced
100% identity over 1,075 bp (barrnap+blastn); shorter aln = draft 16S fragmentation
within tolerance
C10
Reported
pan 5,205 / variable 438 / unique 244 (3 strains)
Reproduced
pan 8,571 / core 3,602 / unique 413 (prokka+roary defaults)
did not match
C11
Reported
14,977,300 reads totaling 2,245 Mb
Reproduced
SRR3943817 = 7,488,650 spots x2 = 14,977,300 reads; 2,246,595,000 bp
exact
C12
Reported
449x coverage
Reproduced
n/a; 2,245 Mb/6.55 Mb = 343x (internally inconsistent)
m.public.grade.flag

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 74/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟢5. Derivability / plausibility
🟡6. Severity of the deviation
🟢7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +3

This genome-announcement / comparative-genomics paper reproduces strongly: the deposited genome's headline properties (length 6,553,639 nt, GC 66.90%, 80 contigs, 14,977,300 reads) match exactly, the novel filamentous prophage (7,540 bp, 11 genes, exact coordinates, Zot marker) and ANI conspecificity (ANIb 98.84% vs 98.82%) reproduce cleanly, and the central conclusion holds fully. The two genuine divergences — C5 functional fraction (80.69% vs 90.63%) and C10 pan-genome (5,205 vs 8,571) — sit entirely on our methodology side, caused by substituting RefSeq/PGAP + prokka/roary for the paper's hosted IMG-ER + an unspecified pan-genome tool; both are pipeline/version effects, not non-derivable numbers, so there is no fabrication concern. One minor authors'-side anomaly: the reported 449x coverage is internally inconsistent with the paper's own 2,245 Mb / 6.55 Mb (=343x), flagged for human review.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.