Genome of the Asian longhorned beetle (Anoplophora glabripennis), a globally significant invasive species, reveals key functional and evolutionary innovations a
The main results reproduced: recomputed values matched the published ones within tolerance.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Same input data as the authors
- ✓Reported values were directly comparable
- ✓No relevant deviation in data/preprocessing
- ✓No authors-side cause for any deviation
- ✓Reported values are derivable from the shared data
- ✓Any deviation was negligible
- ✓The central claim held under reproduction
- ✓Overall, the reproduction was clean
- Every checked point held up.
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Described well enough; 1:1 reproduction of the clearly-specified pipeline-derived summary statistics from the paper's OWN deposited data using standard third-party tools (P16). Assembly GCA_000390285.1 (Agla_1.0) recomputed with seqkit v2.13.0, assembly-stats v1.0.1, and an independent python script (all agree): total 707,712,193 bp (paper '710 Mb', within-tol 0.32%), scaffold N50 658,851 bp ('659 kb', exact), contig N50 16,544 bp ('16.5 kb', exact); cross-checked against NCBI's own assembly_stats. Official Gene Set v1.2 (i5k NAL) recounted two independent ways: 22,253 protein-coding genes (paper 22,253, EXACT) and 66 pseudogenes (paper 66, EXACT). Four claims exact, two within-tol, zero mismatch, no fabrication flags. NOT attempted (hard-20% / non-deposited): MAKER pre-curation count 22,035 (intermediate not deposited), BUSCO completeness (version-locked benchmark set, cannot match 1:1), and all wet-lab + de-novo-assembly + orthology/phylogenomics analyses (out of scope, see scope.md). The de-novo assembly and MAKER annotation themselves were not re-run; their deposited outputs were measured instead.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 84assessed: 2026-06-16 ⛓ c34fbdb1c7f0
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-16
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-16no human curator yet
- Last updated
- 2026-07-31
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusThe paper investigates the genomic basis and evolution of specialized phytophagy/wood-feeding (xylophagy) in beetles, asking what genomic features enable the Asian longhorned beetle (Anoplophora glabripennis) to digest woody plant tissues, detoxify plant allelochemicals, and succeed as a globally significant invasive polyphage.
- ★ The A. glabripennis genome encodes a uniquely diverse arsenal of enzymes that degrade the main plant cell wall polysaccharide networks (cellulose, hemicellulose, pectin) and detoxify plant allelochemicals. finding
- ★ Amplification and functional divergence of feeding-associated genes, including PCWDEs originally acquired via horizontal gene transfer from fungi and bacteria, expanded the beetle's metabolic repertoire. mechanism
- ★ Large expansions of chemosensory genes for reception of pheromones and plant kairomones reflect the complex chemical cues used to find host plants and mates. finding
- ★ The genome provides metabolic plasticity enabling feeding on diverse woody plant species, contributing to the beetle's highly invasive nature. finding
- ★ A draft reference genome and official gene set (OGS v1.2) were generated and annotated for A. glabripennis, plus first genomes of emerald ash borer and bull-headed dung beetle. resource
- GH-family HGTs are ancient insertions that evolved into functional genes, whereas eight bacterial HGT candidates are recent/degrading and not significantly expressed. mechanism
- A. glabripennis possesses an incomplete DNA methylation machinery, retaining DNMT1 but lacking de novo methyltransferase DNMT3. finding
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Whole-genome shotgun sequencing and assembly | Anoplophora glabripennis (single female larva) | none | genome assembly size, contig/scaffold N50, coverage | — |
| Genome annotation (MAKER pipeline) | A. glabripennis genome | none | protein-coding gene models and pseudogenes | customized MAKER pipeline |
| Genome size estimation (flow cytometry) | A. glabripennis male and female | none | genome size (Mb) | — |
| Genome/gene set completeness assessment (BUSCO) | A. glabripennis plus 14 other insect genomes | none | percent missing/complete universal single-copy orthologs | 2675 arthropod BUSCOs |
| Orthology delineation (OrthoDB) and phylogenomic analysis | 15 insect genomes | none | orthologous groups, gene duplications, ML phylogenetic tree | OrthoDB; ML tree from 523 orthologs |
| Horizontal gene transfer detection (DNA-based HGT pipeline) | A. glabripennis genome | none | HGT candidates and bacterial source, sequence similarity | DNA-based HGT pipeline |
| RNA-seq gene expression | A. glabripennis adult males, females, and larvae (whole organism) | none | transcript expression of HGT candidates and genes | — |
- – Draft reference assembly of 710 Mb generated from 134× coverage with contig/scaffold N50 of 16.5 kb and 659 kb 710 Mb; 134× coverage
- ▲ Official gene set (OGS v1.2) contains 22,253 protein-coding gene models plus 66 pseudogenes, more than other published beetle genomes (13,526–19,222) 22,253 genes
- ▼ A. glabripennis gene set had slightly fewer missing BUSCOs (~3.3%) than most other genomes studied ~3.3% missing
- ▲ A. glabripennis has the most Coleoptera-specific genes (5229) of the five beetle genomes studied, suggesting high adaptive novelty 5229 genes
- – Eight bacterial HGT candidates found; two show 95% similarity to Wolbachia (recent), two show 70–71% similarity with indels (older/degrading); none significantly expressed 95%; 70–71% similarity
- – Conserved core of 5029 orthologs maintained across all 15 species; 6880 widespread orthologs of which 3346 single-copy and 3534 duplicated 5029; 3534 duplicated
- – A. glabripennis placed sister to Dendroctonus ponderosae with 100% ML bootstrap support 100% bootstrap
- – A. glabripennis lacks de novo methyltransferase DNMT3 in both assembly and raw reads but retains DNMT1
- other 710 Mb assembly; contig N50 16.5 kb; scaffold N50 659 kb (draft genome assembly metrics)
- count 134× sequence coverage (genome sequencing depth from a single female larva)
- mean female 981.42 ± 3.52 Mb; male 970.64 ± 3.69 Mb (estimated A. glabripennis genome size)
- count 22,253 protein-coding gene models; 66 pseudogenes (official gene set OGS v1.2)
- count 22,035 gene models annotated; 1144 manually curated (MAKER automated annotation and manual curation)
- other ~3.3% missing BUSCOs (completeness vs 2675 arthropod BUSCOs)
- count 5229 Coleoptera-specific genes; 1210 with orthologs in other beetles; 1003 unique with no homology (lineage-restricted gene counts)
- other $889 billion (conservatively estimated potential US economic impact, inflation-adjusted May 2016)
Statistical methods review
Model: opusA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This is primarily a genome sequencing, annotation, and comparative genomics study rather than a hypothesis-testing experimental paper. The reported approach centers on genome assembly metrics (coverage, contig/scaffold N50), gene-model annotation (MAKER pipeline plus manual curation), completeness assessment with BUSCO, orthology delineation (OrthoDB), maximum-likelihood phylogenomics with bootstrap support, and HGT detection; genome-size estimates are reported as a mean with a dispersion value. No conventional inferential statistical tests (e.g., t-tests, ANOVA) for group comparisons are described in the provided text.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Maximum likelihood phylogenetic inference with bootstrap support | Fig. 2a ML tree from amino acid sequences of 523 orthologs (all nodes 100 % ML bootstrap support) | 523 orthologs | not stated |
-
Genome-size estimates are reported as a mean followed by a single ± dispersion value (e.g., 981.42 ± 3.52 Mb) without specifying whether it is SD or SEM.↳ Could also: Explicitly labeling the dispersion as SD, SEM, or reporting a 95 % confidence interval, and stating the number of measurements. — Naming the dispersion statistic and n makes the spread unambiguous and lets readers gauge measurement precision; SD or a CI is often preferred for conveying variability.
-
Node support on the maximum-likelihood phylogeny is summarized with bootstrap percentages.↳ Could also: Complementary support assessment such as Bayesian posterior probabilities, approximate likelihood-ratio tests (aLRT/SH-aLRT), or ultrafast bootstrap. — Multiple, methodologically distinct support measures can corroborate one another and provide additional perspective on branch reliability.
-
Gene-family and ortholog counts are compared descriptively across the 15 genomes.↳ Could also: Model-based gene-family expansion/contraction analyses (e.g., CAFE) or phylogenetically informed comparative methods. — Such approaches place count differences in an explicit evolutionary and statistical framework, accounting for shared ancestry when interpreting lineage-specific expansions.
-
Genome completeness is assessed using BUSCO ortholog presence/absence proportions.↳ Could also: Reporting these proportions with binomial confidence intervals, alongside complementary metrics such as read-mapping rates or k-mer-based completeness. — Adding interval estimates and orthogonal metrics conveys the uncertainty around completeness percentages and cross-validates the assessment.
Result convergence & founder nodes
Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.
-
The gene set has fewer missing universal single-copy orthologs (~3.3%) than most other insect genomes, indicating high completeness.WGS insect anoplophora glabripennis down 2016×1papers★ This paper is the founder (earliest)
-
Anoplophora glabripennis has the most Coleoptera-specific genes (5229) among five beetle genomes, indicating high adaptive novelty.WGS insect anoplophora glabripennis up 2016×1papers★ This paper is the founder (earliest)
-
Anoplophora glabripennis lacks the de novo methyltransferase DNMT3 while retaining DNMT1.WGS insect anoplophora glabripennis down 2016×1papers★ This paper is the founder (earliest)
-
Anoplophora glabripennis is placed sister to Dendroctonus ponderosae with 100% ML bootstrap support.WGS insect anoplophora glabripennis none 2016×1papers★ This paper is the founder (earliest)
-
The official gene set contains 22,253 protein-coding genes, more than other published beetle genomes.WGS insect anoplophora glabripennis up 2016×1papers★ This paper is the founder (earliest)
-
Eight bacterial HGT candidates detected in the genome, including recent Wolbachia-derived transfers (~95% similarity) and older degrading ones; none significantly expressed.WGS insect anoplophora glabripennis mixed 2016×1papers★ This paper is the founder (earliest)
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
Assessed papers, coloured by verdict. Click a node to open it.
- No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-27832824 (Asian longhorned beetle genome, McKenna et al., Genome Biol 2016)
Paper: doi:10.1186/s13059-016-1088-8 · PMCID PMC5105290 Code (supp scripts): https://github.com/NAL-i5K/AGLA_GB_supp-scripts @ f93a3dc (master, 2018-03-20) Data:
- Genome assembly: GenBank GCA_000390285.1 (Agla_1.0) — exact accession cited in paper.
- OGS v1.2 (official gene set): i5k NAL
BCM-After-Atlas/.../OGS_v1_2/agla_OGS_v1_2.gff3.gz(the paper's own gene set; NAL "Current" has moved to v1.3.1 — we deliberately use v1.2). - Expression: GEO GSE68149 (RNA-seq, used for annotation) — not reproduced (see out-of-scope).
In scope (pipeline-derived, clearly specified, deterministic)
These are summary statistics computed by standard pipelines over deposited data. We reproduce them by running standard third-party tools on the paper's own deposited data (brief rule P16: third-party tool on the paper's data is equally valid).
| id | reported (paper) | location | how reproduced |
|---|---|---|---|
| asm_total | draft assembly "710 Mb" | Results, ¶ "draft genome reference assembly of 710 Mb" | total bp of GCA_000390285.1 FASTA (seqkit/assembly-stats + own python) |
| asm_scaf_n50 | scaffold N50 = 659 kb | same ¶ (Add. file1 Table S3) | scaffold N50 over FASTA records |
| asm_contig_n50 | contig N50 = 16.5 kb | same ¶ | contig N50 (scaffolds split on N-gaps) |
| ogs_genes | OGS v1.2 = 22,253 protein-coding gene models | Results ¶ "official gene set (OGS v1.2)" | count gene features in agla_OGS_v1_2.gff3 |
| ogs_pseudo | 66 pseudogenes | same ¶ | count pseudogene features in OGS v1.2 GFF3 |
| maker_models | 22,035 gene models (MAKER, pre-curation) | Results ¶ "Using a customized MAKER pipeline, 22,035 gene models" | NOT directly reproduced — MAKER intermediate not deposited; cross-check only |
Cross-check (not a paper-tool reproduction, but confirms the deposited data is the paper's): NCBI's own assembly_stats.txt for GCA_000390285.1 reports total-length 707,712,193; scaffold-N50 658,851; contig-N50 16,551; gc 32.5% — i.e. "710 Mb", "659 kb", "16.5 kb" rounded. We independently recompute from the FASTA to confirm.
Optional / hard last-20% (attempt lightly or skip with reason)
- BUSCO completeness (paper: arthropod set of 2675 BUSCOs; A. glabripennis gene set ~3.3% missing, Fig.2). BUSCO is version-sensitive (paper used BUSCO v1 / early arthropoda set, 2675 BUSCOs; modern BUSCO uses OrthoDB v10 with a different, larger arthropoda_odb10 set of ~1013/5235 BUSCOs). A modern rerun cannot match "2675" or "3.3%" 1:1 — it is a different benchmark set. Recorded as a known version-drift limitation; not run as a 1:1 claim.
Out of scope (not pipeline-reproducible from deposited data)
- Wet-lab: flow-cytometry genome size (981 Mb female / 970 Mb male), in-vitro enzyme assays, manual curation of 1144 gene models.
- De-novo genome assembly itself (ALLPATHS-LG + Atlas-Link/Atlas-gapfill over raw SRA reads): the assembly output is deposited and we reproduce stats over it, but re-running ALLPATHS-LG at 222× coverage is heavy, non-deterministic, and out of the 80/20 budget. Not attempted.
- MAKER re-annotation (would require the full repeat library, training, RNA-seq evidence; the OGS output is deposited and we count it instead).
- Gene-family expansions / OrthoDB orthology / HGT / phylogenomics: depend on the 87-species OrthoDB v8 build + the supp perl scripts with hardcoded inputs; the authors' numbers are not regenerable from the shipped scripts alone. Not attempted.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
This is a clean, high-quality reproduction: every in-scope summary statistic was recomputed from the authors' own exactly-cited deposited data (GCA_000390285.1 assembly + OGS v1.2 GFF3) using multiple independent tools that all agree, with four exact and two within-rounding matches and zero mismatches. The only deviation — 707.7 Mb vs the reported 710 Mb (0.32%) — is the paper rounding to whole Mb, confirmed against NCBI's own assembly_stats. The unreproduced items (MAKER 22,035, BUSCO ~3.3%) are honest 80/20 exclusions due to a non-deposited intermediate and BUSCO version drift, not authors-side or fabrication concerns. Central claims fully hold; no derivability or core-claim issues.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.