Genome of the endangered Guatemalan Beaded Lizard, Heloderma charlesbogerti, reveals evolutionary relationships of squamates and declines in effective populatio
Provisional — an automated or curator check raised a specific concern and points reviewers here. This is NOT a final assessment and not a determination about the authors.
The main results reproduced: recomputed values matched the published ones within tolerance.
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
STRONG REPRODUCTION of a single-individual PacBio genome+annotation paper from its figshare deposit (doi:10.25387/g3.20092538), all compute on «our HPC». Of 22 in-scope attempted claims: 18 EXACT, 6 within-tol, 2 partial, 2 mismatch. Assembly stats (size/contigs/N50 1358783/L50 517/N90/L90/longest/GC 45.05) all EXACT by independent pure-python recompute from the deposited FASTA. Annotation: 31411 genes / 32205 mRNA EXACT; exon 230.9bp/intron 1963bp ~exact; InterPro 28923 (89.81%) EXACT. BUSCO (5.4.7 vertebrata_odb10) C95.6 S93.3 D2.3 F1.6 M2.8 vs reported 93.0/2.4/1.8/2.8 -> within-tol (M exact). Repeat classes match authors' RepeatMasker .tbl exactly (tautological); independent softmask fraction 57.70% within tol of 57.54%. Venom jackhmmer vs shipped DB -> 311 unique GBL proteins vs reported 312 (off by 1, used deposited proteome). THREE AUDITABLE FLAGS for a human: (A7) reported '83%' of contigs >=50kb is internally inconsistent with its own count 2801 (=78.9%) -- count exact, percentage is a reporting error; (G7/G8) reported 337 tRNA genes + 208 pseudogenes are NOT REPRODUCIBLE -- they match neither the deposited GFF (305/0), nor raw tRNAscan-SE 2.0.9 (20241/16834), nor the EukHighConfidenceFilter (1/16834); the genome is tRNA-SINE-rich and the reported counts cannot be located in the deposit nor regenerated by the cited tool (likely an undocumented custom filter; possibly a numeric error -- human review recommended, not a certain fabrication). NOT attempted (deliberately, heavy/out of cheap scope): E1 dN/dS (13-genome PAML), P1-P3 PSMC/heterozygosity/Ne (needs 232Gb raw reads). Overall: a well-described, deposit-backed genome paper whose headline assembly/annotation/completeness numbers reproduce essentially 1:1, with a notable irreproducible tRNA result worth a closer human look.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-29
- Rubric version
- not recorded
- Assessed by
- —
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetThe study characterizes the draft genome of the highly endangered Guatemalan Beaded Lizard (Heloderma charlesbogerti) to resolve its phylogenetic relationships among squamates, identify genes under differential evolutionary constraint, and reconstruct its historical effective population size for conservation purposes.
- ★ The assembled draft genome of H. charlesbogerti totals 2.31 Gb, similar in size to related species finding
- ★ Phylogenomic analysis places H. charlesbogerti in a clade with the Asian Glass Lizard (Anguidae), closely associated with the Komodo Dragon (Varanidae) and Chinese Crocodile Lizard (Shinisauridae) finding
- ★ 31,411 protein-coding genes were identified in the genome finding
- ★ 504 genes show a differential evolutionary constraint (dN/dS) specifically on the branch leading to H. charlesbogerti finding
- ★ Effective population size declined approximately 400,000 years ago, stabilized, then began declining again approximately 60,000 years ago finding
- Genome assembled from PacBio long-read/HiFi sequencing using Flye, annotated with BRAKER2 using OrthoDB protein hints method
- A custom venom protein database (Heloderma + viper venom sequences) was built to identify putative venom genes via jackhmmer/Exonerate homology search resource
- PSMC-based demographic reconstruction used ROH-masked, heterozygosity-filtered consensus sequences scaled with a 12.3-year generation time method
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Whole-genome long-read sequencing | Blood DNA from wild-caught adult male H. charlesbogerti | none | Raw long reads / HiFi reads for assembly | PacBio Sequel II SMRTCell |
| De novo genome assembly and quality assessment | H. charlesbogerti draft genome contigs | none | Assembly contiguity/completeness (N50, L50, BUSCO score) | Flye v2.8.1-b1676; QUAST v5.1.0rc1; BUSCO v5.2.2 |
| Repeat annotation and gene prediction | H. charlesbogerti draft genome | none | Softmasked repeats, predicted protein-coding genes and tRNAs | RepeatModeler v2.0.3; RepeatMasker v4.1.2; TRF v4.09.1; BRAKER2 v2.1.6; InterProScan v55.5-88.0; tRNAscan-SE v2.0.9 |
| Comparative phylogenomics (ortholog identification and tree building) | 13 reptile species genomes including H. charlesbogerti | none | Maximum likelihood phylogenetic tree topology | OrthoFinder v2.3.8; MUSCLE v5.1; PRANK v170427; IQ-TREE v2.2.0; MEGA-X v10.2.6 |
| Molecular evolution / dN-dS branch analysis | H. charlesbogerti, Dopasia gracilis, Varanus komodoensis, Anolis carolinensis (outgroup) | none | Branch-specific dN/dS ratios and likelihood ratio test for differential selection | Clustal Omega v1.2.4; IQ-TREE v2.2.0-beta; PAML v4.9 (codeml) |
| Venom gene homology search | Annotated H. charlesbogerti proteome vs custom venom protein database (Heloderma + viper venoms) | none | Putative venom protein hits (e-value thresholds) | jackhmmer v3.1b2; Exonerate v2.4.0 |
| Historical effective population size inference (PSMC) | Whole-genome long reads aligned to H. charlesbogerti assembly | none | Ne trajectory over time, heterozygosity | pbmm2 v1.30.0; longshot v0.4.1; bcftools v1.14-36; PSMC v0.6.5; bedtools v2.30.0 |
- – Draft genome assembled totaling 2,308,465,658 bp with 86x coverage 2.31 Gb
- – 31,411 protein-coding genes identified in the annotated genome 31,411 genes
- – H. charlesbogerti clades with Asian Glass Lizard and is closely associated with Komodo Dragon and Chinese Crocodile Lizard
- – 504 genes identified with differential branch-specific dN/dS constraint on the H. charlesbogerti lineage 504 genes
- ▼ Effective population size declined starting approximately 400,000 years ago ~400,000 years ago
- ▼ After stabilization, effective population size began declining again approximately 60,000 years ago ~60,000 years ago
- ▲ Genome-wide heterozygosity estimated before and after ROH masking 1.454×10^-4 to 1.465×10^-4
- – Assembly contiguity statistics reported (contig N50, L50, GC content) N50=1,358,783 bp; GC=45.05%
- count 2,308,465,658 bp (Total assembled genome size)
- other 86x (Sequencing coverage of the assembly)
- other N50 = 1,358,783 bp; L50 = 517 (Contig-level assembly contiguity statistics)
- count 31,411 (Number of predicted protein-coding genes)
- count 504 (Genes with significant differential dN/dS on the H. charlesbogerti branch (Bonferroni-corrected LRT, α=0.01/N))
- other 1.454 × 10^-4 (Heterozygosity after variant filtering, before ROH masking)
- other 1.465 × 10^-4 (Heterozygosity after ROH masking)
- other 45.05% (GC content of the assembled genome)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This is a genome report describing assembly, annotation, and comparative/phylogenomic analysis of a single wild-caught Guatemalan Beaded Lizard specimen. Statistical inference is used in a few specific places: a likelihood ratio test to detect branch-specific dN/dS shifts, Bonferroni-corrected significance/e-value thresholds for gene evolution and venom-protein homology searches, bootstrap support for a maximum-likelihood phylogeny, and bootstrap resampling within a PSMC analysis of historical effective population size. Results are reported primarily as point estimates, thresholds, and phylogenetic/demographic trajectories rather than as group-mean comparisons.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Likelihood ratio test (LRT), comparing a one-ratio PAML/codeml model (model=0) to a free-ratio branch model (model=1) | Detecting genes with differential dN/dS on the branch leading to the Guatemalan Beaded Lizard | Based on single-copy orthologs identified among the Guatemalan Beaded Lizard, Asian Glass Lizard, Komodo Dragon, and Green Anole (outgroup); N = number of tested orthologs, used in the Bonferroni-corrected significance level | not stated |
| Bonferroni-corrected significance threshold (α = 0.01/N) applied to the LRT statistic vs χ² critical value, k=1 df | Genome-wide screen for genes with differential evolutionary constraint | N = number of tested single-copy orthologs | not stated |
| Bonferroni-corrected e-value thresholds (0.0001/n for full query, 0.01/n for best-scoring domain) applied to jackhmmer similarity search | Identification of putative venom proteins by similarity to a custom Heloderma/viper venom protein database | n = number of queries | not stated |
| KEGG pathway analysis (DAVID) | Genes in the top/bottom 10% of branch-specific dN/dS values on the Guatemalan Beaded Lizard branch | List of all tested single-copy orthologs used as background | not stated |
| Maximum-likelihood phylogenetic inference with bootstrap support (500 replicates, JTT model) | Species tree of 13 reptile species built from 3,324 BUSCO single-copy orthologs | 3,324 orthologous genes across 13 species | not stated |
| PSMC (Pairwise Sequentially Markovian Coalescent) with bootstrap resampling | Inference of historical effective population size (Ne) trajectory over time | 100 bootstrap replicates from a single individual's consensus sequence | not stated |
-
Genes with branch-specific dN/dS shifts were identified using a likelihood ratio test with a Bonferroni-corrected significance threshold (α = 0.01/N) across all tested orthologs.↳ Could also: A false discovery rate (FDR/Benjamini-Hochberg) correction — For large numbers of simultaneous per-gene tests, an FDR-based correction is also commonly used and can offer more power to detect true signals while still controlling for multiple comparisons, as an alternative to the more conservative family-wise error control of Bonferroni.
-
Putative venom proteins were identified via jackhmmer similarity search using Bonferroni-corrected e-value thresholds.↳ Could also: An FDR-based threshold (e.g. on e-values or bit scores) for the homology search — FDR control is another standard way to manage the multiple comparisons inherent in large-scale sequence similarity searches and could be reported alongside Bonferroni-based thresholds for comparison.
-
Pathway enrichment for fast/slow-evolving genes was performed using DAVID/KEGG on genes in the top/bottom 10% of branch-specific dN/dS values.↳ Could also: Gene Set Enrichment Analysis (GSEA) using the full ranked dN/dS distribution rather than a top/bottom percentile cutoff — Rank-based enrichment methods use the continuous ranking of all genes rather than a hard threshold, which can capture enrichment signals that are distributed across the full range of values.
-
Phylogenetic relationships were inferred with maximum likelihood (MEGA-X, JTT model) and node support assessed via 500 standard bootstrap replicates.↳ Could also: Bayesian phylogenetic inference (e.g. MrBayes, BEAST) or ultrafast bootstrap approximation (as implemented in IQ-TREE, which was already used elsewhere in the pipeline) — Bayesian methods yield posterior probabilities as an alternative measure of node support, and ultrafast bootstrap can offer comparable support estimates with reduced computation, providing a cross-check against standard bootstrap values.
-
Historical effective population size was inferred from a single individual's genome using PSMC with 100 bootstrap replicates.↳ Could also: Multi-sample sequentially Markovian coalescent methods (e.g. MSMC2) if additional individuals' genomes become available — Methods that incorporate multiple individuals can improve resolution of effective population size estimates in more recent time periods, complementing the single-genome PSMC approach used here.
-
Genome size was estimated from k-mer frequency distributions computed in R using a single formula-based approximation.↳ Could also: A dedicated k-mer-based genome size estimation tool (e.g. GenomeScope) fitting a mixture model to the k-mer histogram — Mixture-model based tools can also provide estimates of heterozygosity and repeat content alongside genome size, along with model-fit diagnostics, as a complement to a direct formula-based calculation.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-36226801
Paper: Dyson et al. 2022, Genome of the endangered Guatemalan Beaded Lizard, Heloderma charlesbogerti, reveals evolutionary relationships of squamates and declines in effective population size. G3 (Bethesda) 12(12):jkac276. DOI: 10.1093/g3journal/jkac276 · PMID 36226801 · PMCID PMC9713440
Code: https://github.com/LachanceLab/GBL (HEAD commit
d7676db2851f1a946a5b85851970a363dd728102, 2023-01-31) — authors' own Snakemake
workflows (genome_annotation, dNdS, effective_population_size, venom).
Data:
- Raw reads: NCBI SRA / BioProject PRJNA834834 (PacBio Sequel II, 1 male).
- Assembly + annotation: figshare doi:10.25387/g3.20092538 (12 files incl. assembly FASTA, gene GFF, protein FAA, CDS, softmasked FASTA, RepeatMasker .tbl/.out/.gff, TRF gff, InterProScan gff/tsv).
Study design (one individual)
A single wild-caught male (voucher UTA R-15002) was PacBio-sequenced (Sequel II, HiFi + CLR). This is a single-genome de novo assembly + annotation + single- sample PSMC demography paper — NOT a population/cohort study. N(individuals)=1.
In-scope (pipeline-derived computational results)
| group | result | reported | pipeline | cheap? |
|---|---|---|---|---|
| Assembly | size, #contigs, N50, L50, longest, N90, L90 | 2,308,465,658 bp; 3,551; 1,358,783; 517; 7,420,054; 389,083; 1,704 | Flye v2.8.1 → QUAST v5.1.0 | yes (recompute from deposited FASTA) |
| Assembly | GC content | 45.05% | QUAST / FASTA | yes |
| Completeness | BUSCO complete/dup/frag/missing | 93.0% / 2.4% / 1.8% / 2.8% | BUSCO v5.2.2 Vertebrata_odb10 | moderate (rerun on «our HPC») |
| Repeats | total repetitive DNA + classes | 57.54%; LINE 20.81; SINE 1.41; LTR 1.61; DNA 2.59; retro 23.83 | RepeatModeler+RepeatMasker v4.1.2 | yes (deposited .tbl/.out; recompute softmask frac) |
| Annotation | protein-coding genes / mRNAs | 31,411 / 32,205 | BRAKER2 v2.1.6 | yes (parse deposited GFF) |
| Annotation | exon/intron stats | 4 exons/gene; exon 231 bp; intron 1,967 bp | GFF | yes |
| Annotation | proteins with InterPro domains | 28,923 (89.81%) | InterProScan v55 | yes (parse deposited tsv + faa) |
| Annotation | tRNA genes / pseudogenes | 337 / 208 | tRNAscan-SE v2.0.9 | moderate (rerun) / parse |
| Selection | genes with differential dN/dS constraint | 504 (p<5.46e-6) | OrthoFinder+codeml(PAML) over 13 spp | hard (multi-species) |
| Venom | venom-like protein sequences | 312; 5/6 toxin classes | jackhmmer v3.1 vs shipped DB | moderate (DB shipped in repo) |
| Demography | current heterozygosity | 1.454e-4 (1.465e-4 post-ROH) | pbmm2+longshot+bcftools | hard (needs raw reads/BAM) |
| Demography | current Ne | ~3,839 (θ=4Neμ) | derived from het | follows from het |
| Demography | PSMC trajectory (declines ~400k/60k/10k ya) | qualitative curve | PSMC v0.6.5 | hard (full alignment) |
| Phylogeny | squamate tree topology | Anguidae sister; near Varanidae/Shinisauridae | OrthoFinder+IQ-TREE 3,324 BUSCO genes | hard (13 genomes) |
Out of scope (not pipeline-reproducible here)
- Wet-lab: DNA extraction, library prep, PacBio sequencing.
- Pathway enrichment via DAVID web tool (non-significant after FDR anyway).
- Qualitative biological interpretation.
Reproduction tiers
- CORE (quick ~80% floor): assembly stats, GC, gene/mRNA/tRNA counts, InterPro count, repeat content — all recomputable from the deposited figshare files with light compute. These are the clear, low-hanging 1:1 data points.
- BUSCO: rerun on assembly (moderate, few hours on «our HPC»).
- HARDER (push beyond floor): venom screen (DB shipped → feasible), then PSMC heterozygosity/Ne and dN/dS/phylogeny (need raw reads / many genomes; heavy).
Honesty notes
- The RepeatMasker
.tblis the authors' own summary file — comparing the paper's 57.54% to that file is near-tautological; the independent check is recomputing the soft-masked fraction from*_genomic_softmasked.fasta.gz. Both recorded. - Assembl
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.