Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

The first single-stranded DNA virus targeting Pectobacterium belongs to the family Microviridae and demonstrates a broad host range to Pectobacterium brasiliens

Arch Virol · 2026
L1 90/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Every question reproduced
-1 pts
From: “every question reproduced”
Total score -7
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
  • No relevant deviation in data/preprocessing
  • No authors-side cause for any deviation
  • Reported values are derivable from the shared data
  • Any deviation was negligible
  • The central claim held under reproduction
  • Overall, the reproduction was clean
What did not (or only partly)
  • Every checked point held up.
How its reproducibility compares
90/100
Reproducibility score
0.9 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 79% of all assessed papers rank 211 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough -> 1:1 EXACT on the primary, deterministic pipeline. The artifact (juliestenp/mimer_supplementary_data) is a supplementary-data repo (P16 third-party-tool case) shipping inputs + outputs of two phylogenetics pipelines on 4 conserved Microviridae proteins (gpA/gpB/gpF/gpH) over 5 phages (Mimer + phiX174/G4/alpha3 + phiMH2K outgroup). I re-ran the named third-party tool DessimozLab fold_tree (foldtree v1.1.0) on the shipped .pdb structures on «our HPC»; its foldtree-metric output (foldtree_struct_tree.nwk) reproduces ALL FOUR shipped structural trees EXACTLY -- identical unrooted topology (Robinson-Foulds 0/4) AND identical branch lengths to the printed 5-decimal precision. This also correctly distinguishes the foldtree-metric tree from the LDDT/alnTMscore variants (RF 0-4), confirming which output the authors shipped. NOT attempted/left as the optional 20%: the Phylogeny.fr One-Click sequence trees -- it is a web service (MUSCLE+Gblocks+PhyML+TreeDyn) not exactly CLI-reproducible; I ran the CLI tool chain (MUSCLE+Gblocks succeeded) but PhyML did not emit trees cleanly in batch, and the paper itself notes these aa-sequence trees have low statistical support, so per 80/20 I stopped. Also out of scope: upstream AlphaFold2 PDB prediction (shipped as inputs) and all wet-lab/manual numbers (genome 5879nt, 12 ORFs, virion 28.1nm, latent 65min, burst ~79, adsorption 17%, VIRIDIC similarities). No fabrication flags: the shipped structural trees are exactly regenerable from the shipped PDBs by the named tool. All grades provisional; human reviewer assigns final ground truth (AUDIT.md).

💻 Code ↗ 🗄 Data: 10.5281/zenodo.8128917

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 90
    assessed: 2026-06-16 ⛓ a85ca2cf4637
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-16
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-16
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

Can a single-stranded DNA bacteriophage targeting Pectobacterium be isolated and characterized, and does it represent a novel taxon within the family Microviridae with potential as a biocontrol agent against Pectobacterium brasiliense soft rot pathogens?

Core claims
  • Phage Mimer is the first described single-stranded DNA phage targeting Pectobacterium and belongs to the family Microviridae, subfamily Bullavirinae. finding
  • Phage Mimer represents a proposed new genus, 'Mimervirus', within the subfamily Bullavirinae, with intergenomic similarity <70 to all genus representatives. finding
  • Phage Mimer shares gene synteny and conserved protein structure (e.g., gpG, gpH) with Bullavirinae members like phiX174 despite negligible nucleotide/aa sequence similarity. finding
  • Phage Mimer infects a broad range of Pectobacterium brasiliense (Pbr) isolates. finding
  • Phage Mimer is a promising biocontrol agent with therapeutic potential owing to its small genome size and well-characterized phiX174-like genome architecture. finding
  • Manual gene calling/curation (GeneMark, Glimmer, DNA Master) plus HHpred and structural prediction (AlphaFold2, Foldseek) is needed to annotate Microviridae genomes with overlapping genes and low similarity. method
  • Phage Mimer has a 5879 nt genome encoding twelve predicted gene products, seven of which could be assigned a function. resource
Experimental setups
Assay System Perturbation Readout Platform
Phage isolation and purification (double agar overlay) Pectobacterium brasiliense (Pbr) J47; organic waste sample none plaque formation / lytic activity
Whole-genome sequencing (Illumina) Phage Mimer DNA from lysate none genome sequence/assembly NextSeq500, Illumina Mid Output Kit v2 (300 cycles); NEBNext Ultra II DNA Library Prep Kit
Genome assembly and annotation Phage Mimer reads none ORFs/gene functions FastQC, Trimmomatic, SPAdes, CLC Genomics Workbench, GeneMark, Glimmer, DNA Master, HHpred, blastp
Comparative genomics and phylogenetic analysis Phage Mimer vs phiX174, alpha3, G4, phiMH2K none intergenomic similarity, gene synteny, protein structure phylogeny VIRIDIC, clinker, AlphaFold2, Foldseek, Phylogeny.fr, Foldtree, ITOL
Transmission electron microscopy (TEM) CsCl-purified Phage Mimer particles none morphology and particle size (n=21) Talos L120C TEM (ThermoFisher), Ceta 4k camera, 80 kV; Velox v3.9.0
Phage growth kinetics (single burst-size, adsorption) Pbr J47, MOI 0.75 phage infection adsorption rate, latent period, burst size plaque assay (double agar overlay); RStudio/ggplot2
Host range analysis (spot test and efficiency of plaquing) Multiple Pectobacterium species/isolates phage infection (with Pbr J47 filtrate prophage control) clearance/lytic activity, EOP
Key results
  • Phage Mimer genome is 5879 nt with twelve predicted gene products; seven assigned a function. 5879 nt; 7/12 genes
  • Intergenomic similarity score of Mimer to phiX174 was 18.9 and ~14 to alpha3 and G4, below the 70 genus threshold, supporting a new genus. 18.9 (phiX174); ~14 (alpha3, G4)
  • GpG of Mimer matches GpG of phiX174 in structure with strong significance despite no aa similarity. e-value 3.59 × 10^-12
  • GpH of Mimer shows somewhat shared structure with GpH of phiX174. e-value 2.17 × 10^-1
  • Phage Mimer showed poor adsorption, with only 17% of phage particles adsorbed within 10 min on the isolation host. 17% in 10 min
  • Phage Mimer has a latent period of 65 min. 65 min
  • Phage Mimer has an average burst size of approximately 79 virions per cell. ~79 virions/cell
  • megablast of Mimer against nr/nt yielded only one hit (phiX174, 1% query cover, 38 bp match with 1 gap), indicating little nucleotide similarity. 1% query cover; 38 bp, 1 gap
Key statistics
  • count 5879 nt genome size (Phage Mimer genome length)
  • count twelve predicted gene products (7 with assigned function) (Mimer gene content)
  • other intergenomic similarity 18.9 (phiX174), ~14 (alpha3, G4) (VIRIDIC, genus threshold <70)
  • pvalue e-value 3.59 × 10^-12 (Foldseek GpG Mimer vs phiX174)
  • pvalue e-value 2.17 × 10^-1 (Foldseek GpH Mimer vs phiX174)
  • other 17% adsorbed within 10 min (adsorption rate on Pbr J47)
  • other latent period 65 min; burst size ~79 virions/cell (single burst-size experiment, MOI 0.75)
  • count mean/SD of 21 phage particles (TEM particle size measurement)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This is a descriptive virology characterization study of a novel ssDNA phage (Mimer) with no formal inferential hypothesis testing. Quantitative analyses were limited to: (1) Poisson distribution applied to the actual MOI to estimate the probability of multiple infections per bacterium during the single-step growth experiment; (2) mean ± SD of phage capsid diameter measured from 21 TEM-imaged particles; and (3) triplicate plaque assays to estimate adsorption rate, latent period, and burst size. Phylogenetic and genomic relationships were assessed using bioinformatic similarity metrics (VIRIDIC intergenomic similarity scores, Foldseek e-values, maximum-likelihood trees) rather than statistical tests.

Replicationmixed Sample sizeSingle burst-size experiment (one experimental run); plaque assays performed in triplicates at each time point; TEM size measured on n=21 particles; EOP dilutions spotted in triplicates on each host GroupsPhage Mimer growth kinetics on single host (Pbr J47); host range across Pectobacterium brasiliense isolates via EOP Pairingna Randomization/blindingnot stated DispersionSD Exact p-valuesno Effect sizesno Confidence intervalsno Multiplicity correctionnone stated
Statistical tests used
Test Applied to n Assumptions
Poisson distribution (probability mass function) Single-step growth experiment — estimating probability of multiple phage adsorptions per bacterial cell given actual MOI MOI actual derived from adsorption assay triplicates; approximately 6.67 × 10^7 CFU/ml not stated
Mean and standard deviation TEM morphological analysis — phage capsid diameter n = 21 phage particles not stated
VIRIDIC intergenomic similarity score (nucleotide similarity + genome length comparison) Comparative genomics — genus-level classification of phage Mimer against Bullavirinae representatives 4 phage genomes compared na
Foldseek structural alignment e-value Protein structure similarity — GpG and GpH in Mimer vs. phiX174 Pairwise protein comparisons na
Maximum likelihood phylogenetic tree (PhyML via Phylogeny.fr; Foldtree for structure-based trees) Phylogenetic analysis of conserved Microviridae proteins across 5 phages; structure-based phylogeny via Foldtree/Foldseek 5 phages, 4 conserved protein groups na
Approaches that could also have been used
  • The burst-size experiment was performed as a single experimental run, with triplicates taken at each time point within that run
    Could also: Multiple independent biological replicates of the one-step growth experiment (typically n ≥ 3 independent runs) could also have been conducted, with mean burst size and latent period reported across runs with a measure of variability (SD or range) — Independent biological replicates would allow estimation of run-to-run variability in burst size and latent period, providing a measure of uncertainty around the reported point estimates that triplicates within a single run cannot supply
  • Phage capsid diameter was summarised as mean ± SD from 21 TEM-imaged particles
    Could also: A 95% confidence interval around the mean could also have been reported alongside or instead of SD — For small n (here n=21), a CI directly communicates uncertainty about the true mean diameter, whereas SD describes particle-to-particle spread; both convey complementary information and CI is increasingly preferred in morphological characterisation studies
  • Adsorption rate was expressed as a single percentage (17% adsorbed within 10 min) derived from triplicate plaque assays
    Could also: The three replicate plaque counts could also be used to report mean ± SD (or a 95% CI) of the adsorption percentage, or a one-sample test against a reference adsorption rate could be performed — Reporting variability across triplicates would convey the measurement precision of the adsorption estimate; many phage characterisation studies report the range or SD of triplicate EOP/adsorption values
  • Host range was assessed qualitatively via spot tests (presence/absence of clearance) followed by EOP estimation from triplicate dilution spots
    Could also: EOP values could also be analysed with a one-way ANOVA or Kruskal-Wallis test across host strains, or EOP ratios could be reported with confidence intervals relative to the isolation host — Formal comparison of EOP across strains would allow statements about whether differences in plaquing efficiency between Pbr isolates are larger than assay noise, supporting quantitative host-range claims beyond binary lytic/non-lytic classification
  • Phylogenetic trees were built using the maximum likelihood method (PhyML) without reporting bootstrap support values in the methods text
    Could also: Bootstrap resampling (e.g., 100–1000 replicates) or Bayesian posterior probabilities (e.g., MrBayes) could also be used to quantify branch support — Branch support values are standard in phylogenetic publications and allow readers to assess confidence in inferred relationships, particularly important when the dataset is small (5 phages, 4 protein groups) and divergence is high
  • Intergenomic similarity scores from VIRIDIC were used to assign genus-level taxonomy by comparing against a fixed threshold (< 70 = new genus)
    Could also: Average nucleotide identity (ANI) via tools such as FastANI, or protein-based clustering (e.g., vConTACT2) could also be applied as orthogonal genus demarcation approaches — Multiple demarcation metrics applied in parallel increase confidence in novel genus assignment, particularly for phages with very low nucleotide similarity to known taxa, as is the case here
Software: RStudio 2023.6.1.524 · R/ggplot2 · Velox (TEM analysis, size measurement) 3.9.0 · VIRIDIC (intergenomic similarity) · Phylogeny.fr (MUSCLE + Gblocks + PhyML + TreeDyn) · Foldseek / Foldtree (structural alignment and phylogeny) · AlphaFold2 (protein structure prediction) · clinker (genome synteny comparison) 0.0.28 · SPAdes (assembly) 3.14.1 · FastQC (QC) 0.11.9 · Trimmomatic (adapter/quality trimming) 0.39 · CLC Genomics Workbench 22.0

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
1
Impact: low
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-41739249

Paper: Pedersen et al. 2026, Arch Virol 171(3):97. "The first single-stranded DNA virus targeting Pectobacterium belongs to the family Microviridae..." (phage Mimer). PMCID PMC12935840 · DOI 10.1007/s00705-026-06543-2.

Code/data artifact: github.com/juliestenp/mimer_supplementary_data (commit 0b4fca08) — a supplementary-data repo, not authors' analysis code. Ships the inputs and the output trees of two third-party-tool pipelines. Per brief rule P16, applying the named third-party tools to the paper's own shipped data is an equally valid reproduction.

In scope (pipeline-derived, reproducible)

# Result Pipeline (tool) Shipped input Shipped output (compare target)
S-gpA structural phylogeny of gpA Foldseek + fold_tree (DessimozLab) Foldseek/gpA/*.pdb (5 structures) Foldseek/gpA/foldtree_tree_gpA.nwk
S-gpB structural phylogeny of gpB Foldseek + fold_tree Foldseek/gpB/*.pdb Foldseek/gpB/foldtree_struct_tree_gpB.nwk
S-gpF structural phylogeny of gpF Foldseek + fold_tree Foldseek/gpF/*.pdb Foldseek/gpF/foldtree_struct_tree_gpF.nwk
S-gpH structural phylogeny of gpH Foldseek + fold_tree Foldseek/gpH/*.pdb Foldseek/gpH/foldtree_tree_gpH.nwk
Q-gpA sequence phylogeny VP4/gpA Phylogeny.fr One-Click (MUSCLE+Gblocks+PhyML+TreeDyn) Phylogeny_fr/gpA/gpA_VP4_aa.fasta Phylogeny_fr/gpA/phylo_fr_gpA_VP4.nwk
Q-gpB sequence phylogeny VP3/gpB Phylogeny.fr One-Click Phylogeny_fr/gpB/gpB_VP3_aa.fasta Phylogeny_fr/gpB/gpB.nwk
Q-gpF sequence phylogeny VP1/gpF Phylogeny.fr One-Click Phylogeny_fr/gpF/gpF_VP1_aa.fasta Phylogeny_fr/gpF/phylo_fr_gpF_VP1.nwk
Q-gpH sequence phylogeny VP2/gpH Phylogeny.fr One-Click Phylogeny_fr/gpH/gpH_VP2_aa.fasta Phylogeny_fr/gpH/phylo_fr_gpH_VP2.nwk

All 8 trees: 5 taxa = Mimer, phiX174, G4, alpha3, phiMH2K (outgroup). Comparison metric: unrooted Robinson-Foulds (RF) topology distance after relabelling leaves to the 5 canonical taxa (max RF for 5 unrooted taxa = 4). Primary biological claim to check qualitatively: Mimer falls inside Bullavirinae (with phiX174/G4/alpha3), phiMH2K basal/outgroup; paper notes per-protein topology varies & support is low.

Out of scope (not attempted, with reason)

  • AlphaFold2/ColabFold PDB prediction — the .pdb are shipped as inputs (filenames carry _rank_1/plddts); regenerating them is stochastic + GPU-heavy and is an upstream step. We treat the structures as given inputs (honest: we reproduce the tree-building, not structure prediction).
  • Wet-lab / manual results — genome size (5,879 nt), 12 gene products, virion diameter 28.1 nm, latent period 65 min, burst size ~79, adsorption 17%, intergenomic similarity scores (computed in VIRIDIC, not in this repo) — not pipeline-derived from the shipped artifact → out of scope.
  • Phylogeny.fr exact web run — the One-Click web service is the exact path; we reproduce its documented tool chain (MUSCLE+Gblocks+PhyML) on CLI. Differences in default versions/params are expected → graded leniently (this is the optional ~20%).

Primary / strongest data point

The Foldseek/fold_tree structural trees (S-gpA..S-gpH): deterministic tool, all inputs shipped, exact same pipeline. This is the clear 1:1 comparison.

S-gpA
Reported
shipped foldtree_tree_gpA.nwk (structural phylogeny)
Reproduced
RF=0/4, branch lengths identical to 5dp
exact
S-gpB
Reported
shipped foldtree_struct_tree_gpB.nwk
Reproduced
RF=0/4, byte-identical
exact
S-gpF
Reported
shipped foldtree_struct_tree_gpF.nwk
Reproduced
RF=0/4, branch lengths identical to 5dp
exact
S-gpH
Reported
shipped foldtree_tree_gpH.nwk
Reproduced
RF=0/4, branch lengths identical to 5dp
exact
Q-gpA..gpH
Reported
Phylogeny.fr One-Click aa-sequence trees (4 proteins)
Reproduced
not completed (PhyML batch tooling quirk; MUSCLE+Gblocks ran)
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 90/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟢3. Location of the main deviation
🟢4. Cause of the deviation
🟢5. Derivability / plausibility
🟢6. Severity of the deviation
🟢7. Core claim
🟢8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Every question reproduced
-1 pts
From: “every question reproduced”
Total score -7

This is among the cleanest cases: re-running the named, deterministic DessimozLab fold_tree tool on the shipped, identical PDB inputs regenerated all four published structural phylogenies (gpA/gpB/gpF/gpH) exactly — unrooted RF=0/4 and branch lengths identical to 5 dp (gpB byte-identical) — strong evidence the trees are genuine pipeline output, no fabrication concern. The single gap is the optional Phylogeny.fr One-Click sequence trees, which did not complete due to an our-side PhyML batch tooling quirk; the paper itself flags these aa trees as low-support, so this falls in the 80/20 optional tail and does not touch the central claim. Net: a 1:1 reproduction on the primary deterministic pipeline.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

255.1 k
tokens (I/O) · 19 M incl. cache
24 min
runtime · 0.3 CPU-h
31.2 GB
peak RAM
5 (4 failed)
HPC jobs
hummel
machine