Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Population differentiation and epidemic tracking of Bursaphelenchus xylophilus in China based on chromosome-level assembly and whole-genome sequencing data.

Pest Manag Sci · 2021
L1 44/100 3/4
⚑ Flagged for review — a reproduced result did not match the reported value

Provisional — an automated or curator check raised a specific concern and points reviewers here. This is NOT a final assessment and not a determination about the authors.

Why this verdict

The main result did not reproduce in this reproduction attempt. Where our recomputation produced values that differ from the published ones, those discrepancies are listed below. This is a single automated attempt — not peer review and not a finding of error or misconduct — and differences can also arise from data access, undocumented parameters or the computing environment. The verdict can be contested via “report an error”.

Reproduced on the brainbox compute brainarbeit.com
✓ What held up
  • Nothing in this column.
What did not (or only partly)
  • 🔴Could not use the authors’ exact input data
  • 🔴Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🔴A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
44/100
Reproducibility score
1.7 SD below mean
vs. all fields · 1173 studies
🎯 Scores higher than 5% of all assessed papers rank 1109 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

PARTIAL reproduction with a strong, honest data-completeness finding. Paper (Ding et al. 2021, Pest Manag Sci) reports population genomics of 181 resequenced B. xylophilus strains, but BioProject PRJNA524063 publicly deposits only 15 WGS runs (8.3% of the cohort), submitted ~2 years post-publication — so the headline cohort numbers (7,395,877 biallelic SNPs, median 288,883 SNPs/strain, PCA->4 populations, Treemix 3 spread routes, XGBoost 90%-within-727km) CANNOT be reproduced 1:1. What WAS done faithfully on «our HPC» (SLURM, BWA-MEM v0.7.17 + SAMtools v1.6 + freebayes v1.1.0 with the paper's EXACT flags '-u -C 5 -e 50 --standard-filters --min-coverage 10', then the named repo Dsuite built from source): aligned + jointly called all 15 public strains -> 36,985 biallelic SNPs, median 15,068 non-ref genotypes/strain, mean depth 115.6X; ran Dsuite Dtrios (10 trios, no significant introgression among the 15 China strains, as expected for a clonal invasion). Two checkable claims verified: reference genome size 77.27 Mb ~ reported 77.1 Mb (within-tol). Two data-gap mismatches recorded: (C2) the deposited assembly GCA_022702815.1 is CONTIG-level (129 contigs), NOT the claimed 6-chromosome scaffolded assembly; (C4) only 15/181 strains public. No fabrication concern: the smaller SNP catalogue is fully explained by 12x fewer samples + no divergent USA outgroup + superlinear catalogue growth. NOT attempted (out of scope): BUSCO, PCA, Treemix, XGBoost (full cohort / non-repo models). All 30 FASTQs md5-verified. Honest verdict: pipeline faithfully reproducible on public data; cohort-level results not reproducible because the data the paper relied on was largely never deposited.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 44
    assessed: 2026-06-22 ⛓ 9957f20e736e
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-22
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The study aims to characterize the population structure, genetic variation, and geographic spread of the invasive pinewood nematode Bursaphelenchus xylophilus in China using a chromosome-level genome assembly and whole-genome resequencing, to understand adaptation and enable epidemic source tracking.

Core claims
  • Generated the first chromosome-level genome assembly (AH1) of B. xylophilus using PacBio, Illumina, BioNano, and Hi-C data resource
  • Sequenced 181 B. xylophilus strains and identified ~7.8 million SNPs finding
  • The B. xylophilus population in China divides into geographically bounded subpopulations with severe cross-infection and potential migrations finding
  • Geographically associated SNPs are mainly located in adaptation-related GPCR gene families and distribution is dominated by temperature zones, suggesting temperature adaptation mechanism
  • Established a machine-learning (XGBoost) based epidemic tracking method to predict geographical origins, applicable to other species method
  • The AH1 assembly shows high continuity and completeness and contains novel genes compared with previous assemblies finding
Experimental setups
Assay System Perturbation Readout Platform
SMRT long-read DNA sequencing B. xylophilus inbred strain AMA3 (AH1) none genome assembly PacBio RS II
Short-read DNA-Seq B. xylophilus inbred strain AMA3 none genome polishing Illumina (100 bp paired-end)
BioNano optical mapping B. xylophilus inbred strain AMA3 none physical map for hybrid scaffolding BioNano IrysView, NT.BssSI nicking enzyme
Hi-C sequencing B. xylophilus inbred strain AMA3 none chromosome-level scaffolding Illumina Novaseq/MGI-2000
RNA-Seq B. xylophilus inbred strain AMA3 none transcript assembly, gene expression, exon-skipping events TopHat/Cufflinks/Cuffdiff/GESS
Whole-genome resequencing (DNA-Seq) 181 B. xylophilus strains from China (16 provinces) and USA none SNP genotyping, population structure, PCA, Treemix, GWAS Illumina Hiseq 4000 (150 bp paired-end)
Pyrosequencing validation 181 B. xylophilus samples none validation of two SNP sites PyroMark Q96 ID
Machine learning geographic prediction SNP genotype data from 181 B. xylophilus strains none predicted latitude/longitude of sample origin XGBoost via scikit-learn
Key results
  • AH1 chromosome-level assembly comprises 6 chromosomes totaling 77.1 Mb with chromosome N50 of 12 Mb 77.1 Mb, chromosome N50 12 Mb
  • Identified ~7.8 million SNPs across 181 resequenced strains ~7.8 million SNPs
  • Population structure analysis revealed geographically bounded subpopulations with cross-infection and migration among provinces
  • Geographically associated SNPs enriched in GPCR gene families
  • Nematode distribution pattern is dominated by temperature zones
  • PacBio subreads for AH1 had mean length 10.1 kb and N50 13.8 kb at 123X coverage 10.1 kb mean, 13.8 kb N50, 123X coverage
  • Short-read polishing corrected sequencing errors in the assembly 5395 SNVs and 102 719 indels corrected
Key statistics
  • count ~7.8 million SNPs (SNPs identified across resequenced B. xylophilus strains)
  • other 123X coverage (PacBio sequencing depth for AH1 inbred strain genome)
  • mean 10.1 kb mean subread length, 13.8 kb N50 (PacBio subread length statistics after quality filtering)
  • other 46X coverage (Illumina short-read sequencing depth used for genome polishing)
  • other 77.1 Mb assembly, 6 chromosomes, chromosome N50 12 Mb (AH1 chromosome-level assembly statistics)
  • count 5395 SNVs and 102 719 indels corrected (Errors corrected during short-read polishing of the assembly)
  • pvalue 1.6e-9 (Manhattan plot significance threshold for geographically associated SNPs (0.05/(SNP numbers*4)))
  • count 370 million paired-end reads (Hi-C library sequencing output used for chromosome scaffolding)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The study re-sequenced 181 Bursaphelenchus xylophilus strains (180 from 16 Chinese provinces, 1 from the USA) and called ~7.8 million SNPs against a new chromosome-level reference assembly. Population structure was characterised by PCA and hierarchical clustering (R/SNPRelate), population splitting and migration by Treemix with bootstrapping, and gene flow by ABBA-BABA D-statistics (Duite). Geographically associated SNPs were identified by a GWAS-style scan (Plink) with a Bonferroni-style genome-wide threshold, environmental correlations by Pearson correlation and Duncan's test in SPSS, and geographic origin was predicted with XGBoost using 5-fold cross-validation and bootstrapped haversine-distance errors.

Replicationbiological Sample size180 strains collected from 16 Chinese provinces plus 1 USA strain (n=181 total for re-sequencing); Treemix analysis restricted to provinces with ≥2 strains; no formal sample-size or power calculation stated GroupsGeographic provinces/subpopulations across China and one USA strain; temperature zones for environmental association Pairingunpaired Randomization/blindingnot stated Dispersionnone Effect sizesno Confidence intervalsyes Multiplicity correctionBonferroni-style threshold: 0.05 / (number of SNPs × 4), yielding 1.6e-9; basis for the ×4 factor not explained in text
Statistical tests used
Test Applied to n Assumptions
Principal component analysis (PCA) Population structure of 181 strains from genome-wide SNPs after MAF, missing-rate, and LD filtering (R/SNPRelate) 181 strains not stated
Hierarchical clustering Population grouping alongside PCA (R/SNPRelate) 181 strains not stated
Treemix maximum-likelihood network with bootstrap (-noss -bootstrap) and migration-edge inference (-m) Population splitting events among provinces; migration edges added incrementally Provinces with ≥2 strains (exact per-province n not stated) not stated
ABBA-BABA D-statistics (Duite) Detection of gene flow and admixture among subpopulations 181 strains not stated
Pearson correlation Association between B. xylophilus distribution and environmental/climatic factors (temperature zones, WorldClim variables) in SPSS not stated
Duncan's multiple range test Post-hoc group comparisons of environmental/distribution data in SPSS not stated
GWAS-style association scan (Plink linear/logistic regression, Manhattan plot) with genome-wide threshold 1.6e-9 = 0.05 / (SNP count × 4) Identification of geographically associated SNPs 181 strains not stated
XGBoost gradient boosting with randomised hyperparameter search (500 rounds, ranked by MSE) and 5-fold cross-validation Geographic origin prediction (latitude and longitude separately) from SNP allele counts 181 strains not stated
Bootstrapping (haversine distance between true and predicted locations) Estimation of error distribution and confidence intervals for geographic predictions 181 strains not stated
Approaches that could also have been used
  • Population structure was inferred using PCA and hierarchical clustering of genome-wide SNPs
    Could also: Model-based clustering methods such as ADMIXTURE or STRUCTURE could also be applied to the same SNP data — Model-based approaches explicitly estimate individual ancestry proportions and the number of genetic clusters K, providing probabilistic population assignments that complement the geometric separation shown by PCA and are widely used alongside it in population-genomics studies
  • Migration and population splitting were inferred with Treemix, which fits a graph of discrete migration edges onto a population tree
    Could also: f3/f4 statistics (ADMIXTOOLS) or demographic modelling with fastsimcoal2 could also be used — f-statistics can quantify admixture proportions and test specific mixture hypotheses with formal statistical tests and uncertainty estimates, while coalescent-based demographic models allow estimation of migration rates and divergence times with confidence intervals
  • Environmental associations were assessed with Pearson correlation
    Could also: Spearman rank correlation could also be applied to the same data — Spearman correlation is non-parametric and does not assume linearity or normality of variables, which may be preferable for ecological or climatic data that are often skewed or non-normally distributed
  • Post-hoc group comparisons used Duncan's multiple range test
    Could also: Tukey HSD or Games-Howell test could also serve as post-hoc procedures following ANOVA — Duncan's test does not maintain a constant family-wise error rate across all pairwise comparisons; Tukey HSD controls the family-wise error rate more strictly and is more widely reported in the ecological and genomics literature, making results easier to compare across studies
  • Geographic origin prediction used 5-fold cross-validation with folds assigned by random partitioning
    Could also: Spatially stratified (block) cross-validation, where folds are defined by geographic region rather than random assignment, could also be used — When samples exhibit spatial autocorrelation, random CV folds can place geographically neighbouring samples in both training and test sets, yielding optimistic error estimates; block CV holds out entire geographic regions and produces error estimates that better reflect performance on truly novel locations
  • The genome-wide association threshold was defined as 0.05 / (SNP count × 4)
    Could also: A standard Bonferroni correction (0.05 / number of independent SNPs after LD pruning) or a Benjamini-Hochberg FDR procedure could also be applied — FDR control can increase power to detect true associations compared to Bonferroni when many SNPs are tested; the basis for the additional ×4 factor in the paper's threshold is not explained, and using the conventional formulation would aid reproducibility and comparison with other GWAS studies
  • Prediction accuracy was evaluated using mean squared error of haversine distances between true and predicted locations
    Could also: Root mean squared error (RMSE) or median absolute error of haversine distance could also be reported alongside bootstrapped distributions — Median absolute error is more robust to the influence of outlier predictions and is easier to interpret as a typical prediction error in km; reporting both mean and median errors gives a fuller picture of model performance across samples
Software: R/SNPRelate 1.14.0 · Treemix 1.13 · Duite · Plink 1.90b6.1 · R/qqman 0.1.4 · SPSS · scikit-learn 0.19.1 · XGBoost 0.7 · FreeBayes 1.1.0 · VCFtools 0.1.15 · ANNOVAR 2018Apr16 · BWA-MEM 0.7.15-r1140 · SAMtools 1.6 · BUSCO 3.0.2 · TopHat 2.1.1 · Cufflinks/Cuffdiff 2.2.2.1

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-34839581

Paper: Ding et al. 2021, Pest Manag Sci — "Population differentiation and epidemic tracking of Bursaphelenchus xylophilus in China based on chromosome-level assembly and whole-genome sequencing data." DOI 10.1002/ps.6738.

Named code link: https://github.com/millanek/Dsuite (third-party tool — D-statistics / ABBA-BABA introgression). Per brief P16, applying this existing tool to the paper's data is a fully valid reproduction.

Data accession: SRA/BioProject PRJNA524063.


Critical data-availability finding (drives scope)

The paper analyzes 181 strains (180 China + 1 USA) of resequenced B. xylophilus. The public BioProject PRJNA524063 contains only 15 Illumina WGS resequencing runs (strains: HEN09, HB10, AH32, AH03, LN17, ZJ29, JS08, LN15, LN12, LN18, LN16, JS20, HB06, ZJ27, HEN17 — provinces Henan, Hubei, Anhui, Liaoning, Zhejiang, Jiangsu). The runs (SRR24954462–476) were submitted ~2023, two years after the 2021 paper.

  • ~8.3% (15/181) of the resequencing cohort is public.
  • The genome-assembly raw data (PacBio SMRT, BioNano Irys, Hi-C) is NOT in this BioProject — only Illumina resequencing reads are deposited.
  • The chromosome-level reference assembly IS public: GenBank GCA_022702815 (ASM2270281v1), 77,269,747 bp ≈ 77.1 Mb (matches the paper's stated 77.1 Mb genome size).

Consequence: the paper's headline cohort-level numbers (the 7,395,877 biallelic SNP catalogue, the PCA→4 populations, the Treemix routes, the Dsuite introgression events, the XGBoost geolocation) are all computed over 181 strains and cannot be reproduced 1:1 from the public deposit. What CAN be reproduced is the per-sample SNP-calling pipeline on the 15 public strains, and a demonstration of the named Dsuite tool on the resulting 15-strain VCF.


In scope (pipeline-derived, attempted)

# Result Pipeline Reproducibility
R1 Reference genome size 77.1 Mb / 6 chromosomes (assembly, downloaded not re-run) Verify GCA_022702815 metadata 1:1
R2 Per-strain SNP calling (BWA-MEM v0.7.15 + SAMtools v1.6 + freebayes v1.1.0, params -u -C 5 -e 50 --standard-filters --min-coverage 10) align→call Reproduce on 15 public strains; compare per-strain SNP counts vs Table S1/S9 if those strains are listed; report median
R3 Dsuite D-statistics (ABBA-BABA) Dsuite Dtrios on VCF Demonstrate named tool on 15-strain VCF (cannot match 181-strain groupings)

Out of scope (not attempted — why)

Result Reason
Genome assembly (minimap/miniasm/Racon/LACHESIS, N50, BUSCO 78.4%) Raw long-read/BioNano/Hi-C data not deposited; assembly is downloaded as reference, not re-run
7,395,877 biallelic SNP catalogue (181 strains) 166/181 strains not public
PCA → 4 populations (3137 SNPs) requires full 181-strain genotype matrix
Treemix spread routes; environmental/temperature association requires 181-strain populations
XGBoost geolocation (90% within 727 km) requires full cohort + lat/lon labels; wet-lab/ML model not in named repo

Notes

  • Reference assembly accession confirms a clean, checkable 1:1 metadata claim (R1).
  • All heavy compute (align + freebayes + Dsuite) runs on «our HPC»/SLURM; data on «infra».
C1
Reported
77.1 Mb reference, 6 chromosomes, chr N50 12 Mb
Reproduced
77,269,747 bp total (GCA_022702815.1, strain AMA3/AH1)
within tolerance
C2
Reported
six chromosomes (chromosome-level assembly)
Reproduced
129 contigs (assembly_level=Contig); 6 contigs >5 Mb sum to only 46.9 Mb; no chromosome scaffolds deposited
did not match
C4
Reported
181 strains resequenced (180 China + 1 USA)
Reproduced
15 public WGS runs in PRJNA524063 (8.3%); submitted ~Jul 2023, 2y post-paper
did not match
C5
Reported
average 102X coverage (181 strains)
Reproduced
mean 115.6X over 15 public strains (range 88.6-157.9X)
partial
C7
Reported
7,395,877 biallelic SNPs (181 strains)
Reproduced
36,985 biallelic SNPs from 15 public strains (freebayes, paper's exact flags)
partial
C8
Reported
median 288,883 SNPs/strain (181-strain catalogue)
Reproduced
median 15,068 non-ref genotypes/strain (15-strain catalogue; range 9,710-25,646)
partial
C10
Reported
Dsuite ABBA-BABA D-statistics introgression
Reproduced
Dsuite Dtrios ran cleanly on 15-strain VCF: 10 trios, all |D|<0.075, none significant after correction (max Z=2.89)
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 44/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🔴1. Data identity
🔴2. Endpoint comparability
🟡3. Location of the main deviation
🔴4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

459.1 k
tokens (I/O) · 46.2 M incl. cache
264 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.