A whole genome duplication drives the genome evolution of Phytophthora betacei, a closely related species to Phytophthora infestans.
The main results reproduced: recomputed values matched the published ones within tolerance.
- ✓Same input data as the authors
- ✓Reported values were directly comparable
- ✓No relevant deviation in data/preprocessing
- ✓No authors-side cause for any deviation
- ✓Reported values are derivable from the shared data
- ✓Any deviation was negligible
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
REPRODUCED (clean 1:1). All 12 Table-1 assembly statistics for BOTH deposited genomes (P. betacei GCA_011320135.1, P. infestans RC1-10 GCA_011316315.1) recompute EXACTLY via two independent tools (seqkit + QUAST): total length, #contigs, N50, L50, %>=50kb, %N. PacBio yield (17.57 Gb) within-tol of '~17.5 Gb'. Linked third-party tool pbmm2 exercised on the paper's own PacBio subreads: 95.18% primary mapping, 61.18x depth (reads<->assembly consistent). BUSCO (stramenopiles_odb10) reproduces the Fig-1 signal underpinning the central whole-genome-duplication claim: P. betacei 71% duplicated single-copy orthologs vs P. infestans 15% (both 100% complete, 0% missing). No examined value was non-derivable from the shipped data; no fabrication concern. NOT attempted (out of scope, heavier multi-tool/wet-lab): full Canu re-assembly+polishing, Table-2 re-annotation, Table-3 TE catalog, Ks/synteny/phylogenomics. All grades PROVISIONAL pending human audit.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 97assessed: 2026-06-21 ⛓ ff3ab63af646
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-25
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetGiven that P. betacei P8084 has an estimated nuclear DNA content almost twice that of P. infestans T30-4, the study asks whether a transposon-driven genomic expansion or a whole genome duplication (WGD) independently occurred in the P. betacei lineage to explain this size difference.
- ★ P. betacei P8084 has the largest sequenced genome in the Phytophthora genus (270 Mb) finding
- ★ A whole genome duplication plus moderate transposable element invasion explains P. betacei's genome size expansion relative to P. infestans mechanism
- ★ P. infestans RC1-10 genome expansion is instead driven by transposable element activity finding
- ★ P. betacei is genomically supported as a standalone species and sister group to P. infestans finding
- ★ P. betacei P8084 does not follow the classic two-speed (gene-dense/gene-sparse) genome architecture reported for P. infestans and other filamentous plant pathogens finding
- ★ Long-read PacBio SMRT sequencing was used to assemble and annotate the genomes of P. betacei P8084 and P. infestans RC1-10 method
- P. betacei P8084 has nearly twice the number of RxLR effector genes compared to P. infestans strains finding
- P. betacei proteins show disproportionately higher representation in the shared core ortholog genome relative to other Phytophthora species finding
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Whole genome sequencing and de novo assembly | P. betacei P8084 and P. infestans RC1-10 (oomycete plant pathogens) | none | genome assembly size, contig number, N50, contiguity | PacBio Sequel SMRT long-read sequencing plus Illumina paired-end sequencing |
| K-mer coverage frequency analysis | P. betacei P8084 genomic DNA | none | genome heterozygosity rate and ploidy estimate | Illumina paired-end reads |
| BUSCO gene-completeness assessment | P. betacei P8084, P. infestans RC1-10, and RefSeq Phytophthora genomes | none | counts of complete, duplicated, fragmented, and missing single-copy orthologs | BUSCO, Stramenopila-Alveolata dataset |
| Gene prediction and functional annotation | P. betacei P8084, P. infestans RC1-10, and P. infestans T30-4 genome assemblies | none | predicted gene models with functional annotation (InterPro, GO, Pfam, SUPERFAMILY, BLAST/UniRef90) | MAKER2 pipeline; InterProScan; BLAST vs UniRef90 |
| CAZyme annotation | predicted proteomes of P. betacei P8084 and P. infestans (RC1-10, T30-4) | none | counts of carbohydrate-active enzyme family proteins | — |
| Trophic lifestyle classification | P. betacei P8084 and P. infestans (RC1-10, T30-4) CAZyme profiles | none | relative centroid distance (RCD) score and trophic category assignment | CATAStrophy |
| Orthologous gene cluster analysis | predicted proteomes of P. betacei P8084, P. infestans RC1-10, P. infestans T30-4, P. nicotianae, P. sojae, P. ramorum | none | number of shared and exclusive ortholog clusters, core genome composition | — |
- ▲ P. betacei P8084 genome assembled at 270.89 Mb, the largest in the Phytophthora genus, 42 Mb larger than P. infestans T30-4 and 70 Mb larger than P. infestans RC1-10 270.89 Mb (+42/+70 Mb)
- – Illumina k-mer analysis of P. betacei indicates a highly heterozygous diploid genome 3.51% heterozygosity
- ▼ P. infestans RC1-10 assembly (203.29 Mb) is smaller and more fragmented than the T30-4 reference (228.54 Mb) 203.29 Mb vs 228.54 Mb
- ▲ P. betacei predicted more genes than either P. infestans assembly 23,457 vs 15,893 (RC1-10) vs 17,475 (T30-4)
- ▲ P. betacei has nearly double the RxLR effector count of both P. infestans strains 201 vs 107
- ▲ P. betacei has more CAZyme-annotated proteins than both P. infestans assemblies 760 vs 478 (RC1-10) vs 492 (T30-4)
- ▲ P. betacei proteins are disproportionately represented in the shared core ortholog genome across six Phytophthora species 21.99% vs ~15% each for other species
- – BUSCO analysis shows P. betacei has the fewest missing single-copy orthologs but the highest number of duplicated genes among compared genomes
- count 270.89 Mb (total assembly size of P. betacei P8084)
- other 3.51% (heterozygosity rate of P. betacei P8084 from Illumina k-mer analysis)
- other N50 = 737.97 Kb (P. betacei P8084 assembly contiguity)
- count 23,457 genes (total predicted genes in P. betacei P8084 (MAKER2))
- fold_change 201 vs 107 (RxLR effector gene counts, P. betacei vs P. infestans (RC1-10/T30-4))
- count 760 proteins (total CAZyme-annotated proteins in P. betacei P8084)
- count 5,881 clusters; 42,198 proteins (core genome ortholog clusters shared among all six compared Phytophthora genomes)
- other 21.99% (P. betacei relative proportion of proteins in core genome ortholog clusters)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This is a comparative genomics study reporting de novo long-read (PacBio SMRT) assembly and annotation of two Phytophthora species. Analyses are predominantly descriptive: assembly quality metrics (N50, BUSCO completeness), gene counts, and ortholog cluster distributions are compared numerically across six genomes. No formal inferential statistical tests are applied; conclusions are drawn from bioinformatics pipeline outputs and direct numerical comparisons of genomic features.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| BUSCO completeness assessment (proportion of single-copy orthologs classified as complete/duplicated/fragmented/missing against a Stramenopila-Alveolata dataset) | Evaluation of assembly semantic completeness across all six Phytophthora genomes (Fig. 1) | — | na |
| k-mer frequency analysis for heterozygosity and ploidy estimation | Characterization of P. betacei P8084 genome from Illumina data prior to assembly (Supplementary Fig. 1); heterozygosity reported as 3.51% | — | not stated |
| Ortholog clustering with proportional overlap summarized via UpSet diagram | Comparative ortholog analysis across six Phytophthora proteomes (Fig. 2A–C) | 122,388 proteins grouped into 20,738 clusters | na |
| Relative centroid distance (RCD) scoring via CATAStrophy for trophic classification | CAZyme-based trophic phenotype assignment for all three assemblies (Supplementary Table 1) | — | na |
| Descriptive count comparison of virulence gene family sizes (RxLR effectors, CRNs, and other categories) | Comparison across P. betacei P8084, P. infestans RC1-10, and P. infestans T30-4 (Supplementary Table 2) | — | na |
-
Virulence gene family sizes (e.g., 201 RxLR effectors in P. betacei vs. ~107 in the two P. infestans assemblies) were compared by direct counts without normalizing to genome size or applying a statistical test↳ Could also: Poisson or negative-binomial regression normalizing gene counts to genome size (or predicted gene count) could also be used, as could a chi-squared goodness-of-fit test comparing observed proportions — Because P. betacei has a genome ~33% larger than P. infestans T30-4, count-based normalization would help distinguish whether higher effector copy numbers reflect genome-size expansion or true gene-family enrichment
-
BUSCO completeness proportions were compared visually across genomes (bar chart, Fig. 1) without a formal test of whether differences exceed chance↳ Could also: A chi-squared or Fisher's exact test on the counts of complete/duplicated/fragmented/missing BUSCO genes could also be applied for each pairwise assembly comparison — Formal testing would quantify whether observed differences in completeness proportions are larger than expected under a null model, complementing the descriptive bar-chart summary
-
Ploidy and heterozygosity were inferred from a k-mer frequency histogram interpreted visually (diploid model assumed, heterozygosity = 3.51%)↳ Could also: Model-based fitting tools such as GenomeScope2 or Smudgeplots could also be applied to estimate ploidy with explicit uncertainty bounds — These approaches provide standard errors on the heterozygosity estimate and test alternative ploidy models formally, making the diploid assumption more explicitly supported by the data
-
Ortholog cluster sharing was summarized as raw counts and proportions in an UpSet diagram without assessing whether any pairwise overlap is larger than expected↳ Could also: Hypergeometric tests or permutation-based methods (e.g., shuffling gene assignments across genomes) could also assess whether the observed cluster overlap between any pair of species exceeds chance expectation — Such tests would indicate whether the large exclusive-to-long-read partition (2,137 clusters shared only by P. betacei and P. infestans RC1-10) reflects a genuine biological or technological signal beyond what random sampling would produce
-
The trophic classification via CATAStrophy assigned a single best-matching centroid (monomertroph) to each assembly based on RCD scores, with no confidence estimate reported↳ Could also: Bootstrap resampling of the CAZyme profile or a probabilistic classifier outputting posterior probabilities per trophic class could also provide uncertainty estimates around the assigned class — Confidence intervals or posterior probabilities would convey how decisively each assembly falls into the monomertroph category relative to other trophic classes, particularly useful when assemblies of different completeness are compared
-
The study is based on single representative isolates per novel species, precluding within-species variance estimation↳ Could also: Population-level sampling with multiple isolates per species followed by pan-genome analyses and permutation or bootstrap tests on gene-count differences could also be applied — Multiple isolates per species would allow estimation of within-species genomic variation, providing statistical power to distinguish species-level from isolate-level genome size and gene content differences
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-34740326
Paper: Ayala-Usma et al. 2021, A whole genome duplication drives the genome evolution of Phytophthora betacei… BMC Genomics. DOI 10.1186/s12864-021-08079-y.
Linked code (P16, third-party tool): https://github.com/PacificBiosciences/pbmm2 — a PacBio minimap2 wrapper used in the paper's assembly-polishing step. Applying it to the paper's own deposited reads is a valid reproduction.
Deposited artifacts located
- P. betacei P8084 assembly: GenBank GCA_011320135.1 (UAnd_PBet_P8084), BioProject PRJNA608953. 270,893,651 bp, 802 contigs.
- P. infestans RC1-10 assembly: GenBank GCA_011316315.1 (UAnd_PInf_RC1-10.1), BioProject PRJNA517953. 203,291,745 bp, 2902 contigs.
- Raw reads (PRJNA608953, P. betacei): 6 PacBio Sequel II WGS runs (SRR14352910-915, ~17.57 Gb) + 3 Illumina HiSeq runs (SRR14352916/17/18).
IN SCOPE (pipeline-derived, attempted)
- Table 1 assembly statistics for both genomes — total length, # contigs, N50, L50, % in contigs ≥50kb, % N's — recomputed directly from the deposited FASTAs (seqkit). [PRIMARY, cleanest 1:1]
- pbmm2 read-mapping of the deposited PacBio subreads back onto GCA_011320135.1 — exercises the linked tool on the paper's data; report mapping rate + mean depth.
- BUSCO completeness (Fig 1) on both assemblies — stramenopiles/eukaryota odb10.
- Table 2 gene counts — IF a structural annotation (GFF) is deposited with the GenBank assembly; otherwise re-annotation is heavier (Augustus/BRAKER) and noted.
- Table 3 TE coverage (RepeatMasker/EDTA) — heavier, attempted after the above.
OUT OF SCOPE (not attempted / why)
- Full de-novo Canu assembly from 17.5 Gb subreads (400 Mb genome estimate, 30 cores /128 GB, multi-day) — we reproduce the DEPOSITED assembly's properties, not re-run the months-long assembly+polish+purge_haplotigs pipeline end-to-end.
- Wet-lab: DNA extraction, library prep, sequencing.
- Ks/synteny/MCScanX WGD analysis (Fig 5/6) and phylogenomics (Fig 3) — pursued only if annotations are available and after the floor results; large multi-tool pipelines.
Heavy compute → «our HPC»/SLURM; downloads → front1→«infra». («our HPC» tunnel currently down.)
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.