A Reference DNA Barcode Library for UK Fungi associated with Bark and Ambrosia Beetles.
The main result did not reproduce in this reproduction attempt. Where our recomputation produced values that differ from the published ones, those discrepancies are listed below. This is a single automated attempt — not peer review and not a finding of error or misconduct — and differences can also arise from data access, undocumented parameters or the computing environment. The verdict can be contested via “report an error”.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Same input data as the authors
- ✓Reported values were directly comparable
- 🟡A deviation arose in the data or preprocessing
- 🟡A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The deviation was non-trivial in magnitude
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Described well enough to reproduce: YES for the VSEARCH read-processing core (authors' own repo theo-llewellyn/UK_Survey gives exact params); PARTIALLY for taxonomy (reference DB version/source not pinned by paper); NO for OTU clustering (relies on external AusMycobiome helper scripts not in the repo). RESULT = same pipeline, same-order-of-magnitude, structurally faithful PARTIAL reproduction. Headline final-ASV count reproduced 10,056 vs reported 11,606 (87%). The whole count cascade runs ~13-20% low, traced to read YIELD at merge+filter (we retained 74% of 13.1M raw reads vs the paper's 93%); critically, the INTERNAL proportions match almost exactly (chimera fraction 3.00% vs 3.21%; fungal fraction 67.4% vs 68.0%), confirming the pipeline logic is reproduced and the gap is an upstream merge/primer-trim parameter difference, NOT fabrication. Tier-2 taxonomic rank counts (C8) were NOT captured this run due to a fixable wiring miss: the UNITE2024 ITS2 reference FASTA+classification downloaded correctly from Zenodo 13336328 (378MB+229MB, on «infra») but the dnabarcoder cutoffs JSON path data/UNITE_2024_cutoffs/ was absent in the shallow git clone, so the classify guard skipped it; a targeted re-run would supply it. NOT ATTEMPTED: dynamic OTU clustering (C9-C11, hard 20%), wet-lab steps, beetle morphology, RDP/T-BAS comparison, vegan rarefaction. No completeness claim. Fabrication concern: none.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 44assessed: 2026-06-15 ⛓ 06b0f7df44d2
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-15
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-16no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusCan passive insect trapping combined with whole-beetle ITS2 metabarcoding generate a comprehensive reference DNA barcode library of fungal communities associated with native UK bark and ambrosia beetles, providing a baseline for biodiversity and forest-health monitoring?
- ★ A reference DNA barcode library of fungi associated with over 1000 UK bark and ambrosia beetles (25 native Scolytinae species) was generated, yielding 5274 identified fungal OTUs. resource
- ★ Whole-beetle metabarcoding using passive insect traps is a high-throughput method to assess fungal communities that are otherwise difficult to sample. method
- ★ Dynamic OTU clustering using dnabarcoder taxon-specific sequence-similarity thresholds best approximates fungal species diversity. method
- ★ The most abundant fungal OTUs are known bark/ambrosia beetle associates rather than random environmental taxa, validating the dataset. finding
- Current best-practice identification (dnabarcoder) is more conservative at species level than RDP and T-BAS, and the three tools show low consistency in assignments. finding
- The dataset provides a baseline to detect invasive fungi, track forest disease outbreaks, and inform biosecurity monitoring. resource
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| ITS2 fungal metabarcoding (whole-beetle) | 1064 bark and ambrosia beetle specimens (25 Scolytinae weevil species) from 20 UK forest sites | none | fungal ITS2 amplicon sequences / OTU richness and composition | Illumina HiSeq 2500 v2 chemistry (Rapid mode), 2×250 bp paired-end; QIAGEN DNeasy Blood & Tissue kit; primers ITS86F/ITS4; TaKaRa Taq |
| Passive insect trapping / specimen collection | 20 UK sites across four forest composition types (monospecific conifers, mixed conifers, mixed broadleaves, conifer+broadleaves) | ethanol + a-pinene bait | beetle specimens collected and morphologically identified | Lindgren multiple-funnel traps (Phero Tech) |
| Sequence denoising / ASV (zOTU) generation | 96 sequencing libraries | none | merged, quality-filtered, chimera-removed ASVs/zOTUs | VSEARCH v2.30.0 (UNOISE3, UCHIME3); Cutadapt |
| Taxonomic identification | fungal ASVs | none | taxonomic assignments per ASV/OTU | dnabarcoder pipeline, ITSx, BLASTn against UNITE+INSD 2024 ITS2 database |
| Comparative identification benchmarking | fungal OTUs | none | proportion and consistency of taxon assignments per rank | RDP Classifier (UNITE-trained); T-BAS v2.3 (Fungi v3, RAxML EPA) |
| Alpha-diversity estimation | per beetle sample | none | raw and rarefied OTU richness | vegan v2.6-4 rarefy function (subsampled to 10 reads) |
- – Final dataset of 5274 fungal OTUs after removing OTUs with >1 read in negative controls 5274 OTUs
- – 7896 of 11606 final ASVs identified as Fungi 7896 ASVs
- – OTU taxonomic resolution: 100% phylum, 94% class, 85% order, 71% family, 63% genus, 17% species 17% to species
- – Ascomycota is the most abundant phylum 83% of all OTU reads
- – Phialophoropsis and Ambrosiella were the most abundant genera 17% and 8% of OTU reads
- – 12 of 32 Microascales/Ophiostomatales species-level OTUs previously reported with same beetle species/genus 12 of 32
- – Average OTUs per beetle 65 ± 38 (range 0–209; median 59)
- – dnabarcoder more conservative at species level than RDP and T-BAS; T-BAS and RDP assignments differ drastically
- count 12,210,693 total paired reads (mean 12,158 per sample) (after merging and quality filtering)
- count 5274 fungal OTUs (final dataset after negative-control removal)
- count 1064 beetle specimens, 25 Scolytinae species (specimens collected across 20 sites)
- mean 65 ± 38 OTUs per beetle (range 0–209, median 59, IQR 38–87) (OTU richness per beetle)
- mean prevalence 13 ± 36 beetles per OTU (range 1–730) (average beetles a fungal OTU occurs in)
- count 6939 ASVs to phylum, 6463 class, 5893 order, 4978 family, 4459 genus, 1214 species (ASV identification depth)
- count 29 OTUs removed for >1 read in negative controls (contamination filtering)
- other Sordariomycetes 34%, Saccharomycetes 16%, Dothideomycetes 16%, Leotiomycetes 9%, Eurotiomycetes 4% (most abundant classes by OTU reads)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This is a data descriptor paper presenting a national biodiversity survey of fungi associated with UK bark and ambrosia beetles via ITS2 metabarcoding; it relies primarily on descriptive statistics and bioinformatic pipelines rather than inferential hypothesis testing. Sequence processing used VSEARCH for denoising and chimera removal, dynamic OTU clustering with dnabarcoder taxon-specific similarity thresholds, and BLAST against the UNITE ITS2 database for taxonomic assignment. Results are reported as counts, proportions, means ± SD, and medians with IQR/range; OTU richness was also computed in rarefied form using vegan. Comparisons among three identification tools (dnabarcoder, RDP, T-BAS) were presented visually without formal agreement statistics.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Descriptive statistics: mean ± SD, median, IQR, range | OTU richness per beetle sample (mean 65 ± 38, median 59, IQR 38–87) and OTU prevalence across samples (mean 13 ± 36, range 1–730) | 1064 beetle specimens; 5274 fungal OTUs | na |
| Rarefaction (vegan::rarefy) to 10 reads | Rarefied OTU richness per sample, shown alongside raw richness in Fig. 2 right column | Subsampling depth of 10 reads chosen based on a visual natural cut-off in the read abundance distribution across all samples | not stated |
-
OTU richness per sample was rarefied to a depth of 10 reads, chosen by visual inspection of a natural cut-off in the read abundance distribution↳ Could also: Rarefaction to a higher standard depth (e.g., the minimum observed library size after quality filtering), or use of sample-accumulation/rarefaction curves (e.g., vegan::rarecurve) to evaluate whether chosen depths approach an asymptote; alternatively, model-based normalization such as DESeq2 variance-stabilizing transformation or cumulative-sum scaling (metagenomeSeq) could be applied — Rarefying to very low depth retains more samples but discards substantial sequencing information per sample; rarefaction curves and model-based approaches allow readers to assess sampling completeness and are increasingly recommended alternatives for metabarcoding richness estimation
-
Differences in fungal OTU richness across metadata groups (country, beetle genus, site, forest composition) were summarized with boxplots (Fig. 2) without formal statistical tests↳ Could also: Permutation-based multivariate tests such as PERMANOVA (vegan::adonis2) for beta-diversity, and non-parametric Kruskal-Wallis tests or linear mixed-effects models for alpha-diversity comparisons accounting for site or beetle-species as random effects — Formal tests would quantify effect sizes and sampling uncertainty, enabling readers to assess whether observed differences among groups exceed what would be expected from random variation, which is particularly useful given uneven group sizes
-
Agreement among the three identification pipelines (dnabarcoder, RDP, T-BAS) was presented as proportion of OTUs assigned at each taxonomic rank (Fig. 1b,c), assessed visually↳ Could also: Pairwise quantitative agreement metrics such as percent exact agreement, Cohen's kappa, or Krippendorff's alpha computed at each taxonomic rank for each tool pair — Numeric agreement statistics would provide a concise, reproducible summary of concordance across tools and allow comparison with benchmarks from other metabarcoding studies
-
Alpha diversity was measured solely as OTU richness (a species-count metric treating all OTUs equally)↳ Could also: Complementary abundance-weighted indices such as Shannon entropy (H') or the inverse Simpson index, which also capture evenness of OTU read distributions across samples — Richness does not distinguish communities dominated by one or two highly prevalent OTUs from those with even representation; abundance-weighted indices capture additional community structure dimensions and are routinely reported alongside richness in metabarcoding studies
-
Sample sizes per species per site were determined by a pragmatic cap (8–12 for common species, 1 for rare species) rather than a formal sampling-sufficiency assessment↳ Could also: Species or OTU accumulation curves (vegan::specaccum) computed per beetle species or per site to assess whether sampling effort at each level approached an asymptote in detected fungal OTU richness — Accumulation curves would allow readers to gauge whether additional beetle specimens would substantially increase detected fungal diversity, providing context for interpreting richness comparisons across beetle species or forest types
-
Dispersion of OTU richness per beetle was reported using both SD (paired with the mean) and IQR (paired with the median) for the same distribution↳ Could also: Consistently reporting a single pairing—either mean ± SD or median with IQR—supplemented by a bootstrap or percentile-based 95% confidence interval for the central tendency — A confidence interval around the central tendency distinguishes uncertainty in estimating the population mean or median from the natural variability among individual beetles, which aids readers in comparing this dataset to future surveys
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
scope.md — pmid-41022899
Paper: Ceballos-Escalera, Llewellyn, Richards, Inward, Vogler (2025). A Reference DNA Barcode Library for UK Fungi associated with Bark and Ambrosia Beetles. Sci Data. DOI 10.1038/s41597-025-05845-5. PMCID PMC12480477.
Type: Scientific Data Descriptor. The "result" is a curated ITS2 reference barcode library of fungi associated with UK Scolytinae (bark/ambrosia beetles), built from Illumina metabarcoding of 1,064 beetle specimens.
Data & code
- Raw data: SRA PRJNA1291454 (SRP600814) = 1,064 Illumina ITS2 amplicon PE runs, ~3.27 GB, 13,108,648 raw reads (ENA verified). Primers ITS86F/ITS4.
- Authors' pipeline: https://github.com/theo-llewellyn/UK_Survey
(
vsearch.sh,dnabarcoder.sh,dynamic_clustering.R) — branchmain. - Classification tool: https://github.com/vuthuyduong/dnabarcoder (P16: third-party tool applied to the paper's data — equally valid to reproduce).
- Final library / figures: Figshare 10.6084/m9.figshare.29505365 (OTU table, taxonomy, representative FASTAs, KRONA). NOT in GenBank.
Pipeline (verbatim from the repo)
Per sample: cutadapt -g ITS86F (R1) / -g ITS4 (R2) → vsearch --fastq_mergepairs --fastq_allowmergestagger → vsearch --fastq_filter --fastq_maxee 2.0 --fastq_maxns 0.
Pooled: cat → all.fq → vsearch --fastx_uniques (derep) → vsearch --cluster_unoise --minsize 4 (zOTUs) → vsearch --uchime3_denovo (final ASVs) → vsearch --usearch_global --id 1.0 (ASV table). Then ITSx --saveregions ITS2 → fungal
(awk col3=="F") + seqkit seq -m 100 -M 500 → dnabarcoder.py search -r unite2024ITS2.fasta -ml 50 → dnabarcoder.py classify -cutoffs unite2024ITS2.unique.cutoffs.best.json.
Then dynamic_clustering.R (blastclust) → OTUs.
IN SCOPE (pipeline-derived, attempted)
Tier 1 — deterministic VSEARCH read processing (HIGH confidence): C1 total paired reads (12,210,693), C2 unique seqs (1,591,669), C3 post-singleton (119,015), C4 zOTUs (11,991), C5 chimeras removed (385), C6 final ASVs (11,606, HEADLINE). Fully scripted, exact params, no external DB needed.
Tier 2 — ITSx + dnabarcoder classification (MEDIUM confidence):
C7 fungal ASVs (7,896); C8 ASV rank counts (6,939 phylum / 6,463 class / 5,893
order / 4,978 family / 4,459 genus / 1,214 species). Needs the UNITE2024 ITS2
reference DB (unite2024ITS2.fasta + .classification) and cutoffs JSON
(unite2024ITS2.unique.cutoffs.best.json, in dnabarcoder repo). The reference
FASTA download URL is not pinned by the paper — sourced from the dnabarcoder/UNITE
2024 distribution (Zenodo 13336328 / dnabarcoder data dir). Exact-version risk.
OUT OF SCOPE (not attempted, with reason)
Tier 3 — dynamic OTU clustering (C9 5,303 OTUs, C10 5,274 final, C11 OTU rank %):
dynamic_clustering.R sources ANALYSIS_SCRIPTS/get_cutoffs.R and taxa_match.R
which are NOT in the UK_Survey repo (must come from AusMycobiome
github.com/LukeLikesDirt/AusMycobiome), plus several manual reformatting steps
(.dynamic.format, json→csv cutoff conversion) that are not scripted, plus
blastclust (legacy BLAST). This is the hard ~20%; skipped per the 80/20 rule.
Wet-lab / external (never in scope): DNA extraction, PCR, Illumina sequencing, beetle morphological ID (1,064 specimens → 25 Scolytinae species), Table 2 literature cross-check, vegan rarefaction figures, RDP/T-BAS qualitative comparison.
Compute plan
Single «our HPC» SLURM job (partition=std, 1 node, ~32 cpus, ≤12h). Build conda env
in-job (cutadapt, vsearch=2.30.0, itsx=1.1.3, seqkit, blast, python3+biopython/
scikit-learn/scipy/matplotlib, clone dnabarcoder). Download SRA reads to «infra»
in-job. Emit counts.json after Tier 1 (checkpoint) then Tier 2. Compare to claims.tsv.
Blocker
«our HPC» VPN tunnel currently DOWN — needs operator 2FA to bring up.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
The VSEARCH/ITSx/dnabarcoder pipeline reproduces structurally tightly on the public SRA data (PRJNA1291454): the headline 11,606 final ASVs came out 10,056 (87%), and chimera (3.00% vs 3.21%) and fungal (67.4% vs 68.0%) fractions are nearly identical, so the gap is real but uniform. The deviation sits upstream at merge+filter read yield (74% vs 93% retained), most plausibly a primer re-trim/merge parameter difference that the paper did not fully pin — an explainable methodological deviation on our/processing side, not authors' fabrication. Severity is moderate (magnitude and direction hold, ~13-20% low), and the core deliverable holds in substance though exact counts differ and the taxonomy rank table (C8, a fixable wiring miss) and OTU clustering (C9-C11, out of scope) were not captured.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.