Whole-Genome Sequence of Cervid atadenovirus A from the Initial Cases of an Adenovirus Hemorrhagic Disease Epizootic of Black-Tailed Deer in Canada.
The main results reproduced: recomputed values matched the published ones within tolerance.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Reported values were directly comparable
- ✓The central claim held under reproduction
- 🟡Could not use the authors’ exact input data
- 🟡A deviation arose in the data or preprocessing
- 🟡A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The deviation was non-trivial in magnitude
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
GOOD partial reproduction of an MRA genome announcement (Cervid atadenovirus A, Lung et al. 2022) from PRJNA803320 via a faithful nf-villumina-class core (fastp -> minimap2 -> samtools consensus / metaSPAdes -> MUMmer) on «our HPC». EXACT: genome length 30,616 nt (both runs), GC 33.2% (33.23%), 99.96% identity to KY748210.1, and both deposited read counts (md5-verified). WITHIN-TOL: coverage 4,038.8x->3,980x (-1.5%) for SRR17880609. PARTIAL: coverage 4,590.5x->3,605x (-21.5%) for SRR21177721 (real gap; same method matched the other sample within 1.5% -> flagged, likely mapper/trimming/dup differences in the authors' pipeline). UNREPRODUCIBLE: the 3rd described sample (109.4x / 6,536,978 reads) has no raw data in PRJNA803320. De novo metaSPAdes fragmented the genome at the adenovirus ITRs under ultra-high (~4000x) coverage, so the reference-guided consensus is the faithful route to the single 30,616 nt genome. Out of scope: GenBank deposit, MAFFT/Geneious/IQ-TREE phylogeny.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 81assessed: 2026-06-22 ⛓ 3db301d31346
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-22
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-22no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetThe study aims to determine the complete whole-genome sequence of Cervid atadenovirus A from the initial 2020 black-tailed deer adenovirus hemorrhagic disease (AHD) cases in British Columbia, Canada, and compare it to previously sequenced U.S. isolates.
- ★ A complete 30,616-nucleotide genome of Cervid atadenovirus A was determined from lung tissue of black-tailed deer that died of AHD in British Columbia in 2020 finding
- ★ The Canadian genome contains unique nonsynonymous SNPs in the E1B, IVa2, and E4.3 coding regions not observed in moose and red deer isolates finding
- ★ The genome contains deletions totaling 74 nucleotides in the AT-rich noncoding region compared to moose and red deer isolates finding
- ★ The genome shares 99.96% pairwise nucleotide identity with the closest reference OdAdV-1 (KY748210.1), including 100% identity in the ITR finding
- Viral sequence enrichment with a custom capture-probe set combined with the nf-villumina Nextflow workflow was used for genome recovery and analysis method
- ★ The Canadian isolate phylogenetically clusters with black-tailed deer and mule deer isolates finding
- ★ This is the first reported genome sequence of Cervid atadenovirus A from a Canadian AHD case resource
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| High-throughput sequencing (viral capture-probe enrichment sequencing) | lung tissue from black-tailed deer (Odocoileus hemionus columbianus) | none (natural infection/mortality) | complete viral genome sequence | Illumina MiSeq (v2 flow cell, 500-cycle kit) |
| Metagenomic bioinformatic analysis (read classification, de novo assembly, BLAST homology search) | sequencing reads from black-tailed deer lung tissue | none | assembled contigs with homology to cervid atadenoviruses | nf-villumina v2.0.0 (Nextflow) |
| Reference mapping and coverage depth analysis | assembled adenovirus genome vs. sequencing reads | none | mean coverage depth, sequence comparison to reference | Geneious v9.1.8 |
| Multiple sequence alignment and phylogenetic analysis | 10 Deer atadenovirus A whole genomes | none | maximum likelihood phylogenetic tree, pairwise nucleotide identity | MAFFT and IQ-TREE with ModelFinder (HKY-F model), 1,000 ultrafast bootstraps |
- – Assembled genome length 30,616 bp
- – GC content of assembled genome 33.2%
- – Pairwise nucleotide identity to closest reference OdAdV-1 (KY748210.1) 99.96%
- – Mean coverage depth across three samples 109.4x, 4038.8x, 4590.5x
- – SNP at genome position 1,409 in E1B causing potential truncation 12-amino-acid truncation
- – Nonsynonymous substitutions found in IVa2 (T5486C, Lys to Glu) and E4.3 (C25817T, Met to Ile)
- ▼ Four deletions in AT-rich variable region relative to moose/red deer isolates 74 nucleotides total
- – Reads used for reference mapping across three samples 6,536,978; 7,576,032; 9,124,356 reads
- other 99.96% pairwise nucleotide identity (comparison to closest reference OdAdV-1 (KY748210.1))
- other 100% identity (ITR region comparison to reference)
- count 30,616 nucleotides (complete assembled genome length)
- other 33.2% GC content (assembled genome composition)
- other coverage depths of 109.4x, 4038.8x, and 4590.5x (mean coverage depth per sample after reference mapping)
- count 74 nucleotides total (four deletions) (deletions in AT-rich noncoding region vs. moose/red deer isolates)
- count 6,536,978; 7,576,032; 9,124,356 (total reads per sample used for reference mapping)
- other 1,000 ultrafast bootstraps (phylogenetic tree support values via IQ-TREE)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This is a genome sequence announcement rather than a hypothesis-testing study. It reports whole-genome sequencing and assembly of Cervid atadenovirus A from three deer tissue samples, with descriptive comparison (pairwise nucleotide identity, coverage depth) to reference genomes and a maximum-likelihood phylogenetic analysis to place the new genome relative to other Deer atadenovirus A sequences. No inferential statistical hypothesis tests (e.g., t-tests, ANOVA) were performed; results are reported as sequence metrics, coverage depths, and phylogenetic support values.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Maximum likelihood phylogenetic inference (IQ-TREE, ModelFinder for model selection, model HKY-F chosen) | Figure 1c, whole-genome phylogenetic tree of Deer atadenovirus A isolates | 10 publicly available complete genomes plus the genome from this study | stated (best-fit substitution model selected via ModelFinder) |
| Ultrafast bootstrap approximation (1,000 replicates) | Figure 1c, node support for the maximum likelihood tree | 1,000 bootstrap replicates | stated |
| Multiple sequence alignment (MAFFT, default parameters) | Figure 1b/1c and comparison of the three assembled contigs | na | not stated |
| Pairwise nucleotide identity / BLAST homology comparison | Comparison of the assembled genome to reference OdAdV-1 (KY748210.1) | na | na |
-
Phylogenetic relationships were inferred using a single maximum-likelihood method (IQ-TREE with ModelFinder-selected model and ultrafast bootstrap support).↳ Could also: A Bayesian phylogenetic approach (e.g., MrBayes or BEAST) could also be used, or a complementary distance-based method (e.g., neighbor-joining). — Bayesian methods provide posterior probability support values as an alternative to bootstrap proportions, and comparing tree topologies across methods can illustrate the robustness of the inferred relationships, which some readers find informative alongside a single ML tree.
-
Node support was assessed using ultrafast bootstrap approximation (1,000 replicates).↳ Could also: SH-aLRT (Shimodaira-Hasegawa-like approximate likelihood ratio test) support, often reported alongside ultrafast bootstrap in IQ-TREE, could also be presented. — Reporting both support metrics together is a common practice that can offer a fuller picture of branch support, since the two methods can occasionally disagree for individual nodes.
-
Coverage depth was summarized as mean values per sample (e.g., 109.4x, 4,038.8x, 4,590.5x) without a measure of spread.↳ Could also: Reporting the distribution of per-base coverage (e.g., via a coverage plot with standard deviation, interquartile range, or minimum coverage) could also be included. — A dispersion measure or coverage-uniformity plot can help readers assess assembly confidence across the genome, particularly in low-coverage regions, complementing the mean value already shown in Figure 1a.
-
Sequence identity between the new genome and the closest reference was reported as a single pairwise percentage (99.96%).↳ Could also: A sliding-window identity plot along the genome could also be used to visualize where divergence is concentrated. — This would let readers see whether the SNPs and deletions already tabulated in Table 1 are clustered in particular genomic regions, adding spatial context to the single overall identity figure.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-36129291
Paper: Lung et al. 2022, Microbiol Resour Announc. "Whole-Genome Sequence of Cervid atadenovirus A from the Initial Cases of an Adenovirus Hemorrhagic Disease Epizootic of Black-Tailed Deer in Canada." DOI 10.1128/mra.00662-22.
This is a Genome Announcement (MRA) — a short paper whose entire content is a small set of pipeline-derived numbers about a viral whole-genome assembly. There are no figures/tables of statistical results; the "results" are the numbers in the running text.
Pipeline
- nf-villumina v2.0.0 (https://github.com/CFIA-NCFAD/nf-villumina) — a Nextflow viral-Illumina pipeline: read QC/trim, host/taxonomic filtering, de novo assembly (SPAdes/Unicycler-class) + reference-guided mapping. This is a third-party/own CFIA tool applied to the paper's own data — valid per P16.
- Downstream manual tools: MAFFT (align), Geneious v9.1.8 (curate/annotate), IQ-TREE (phylogeny). These are interactive/manual → out of scope.
Data
- PRJNA803320 (SRA/ENA), METAGENOMIC WGS, Illumina MiSeq, paired-end.
- Observed runs (ENA filereport): 2 runs only
SRR17880609(SRX14039924, SAMN25644053) — 4,562,178 spots — earlier sample → OM470968SRR21177721(SRX17188812, SAMN30469495) — 3,788,016 spots — later sample → OP289523
- Paper describes 3 samples (two earlier 2020-09-11, one later 2020-09-16). The third sample's raw reads (the 109.4x / 6,536,978-read one) are NOT deposited.
In scope (pipeline-derived → attempt to reproduce)
| id | result | how |
|---|---|---|
| C3a/C3b | total read counts of the 2 deposited samples (9,124,356 / 7,576,032) | already reproduced from ENA metadata: paper value = 2 x ENA spots (paired) — EXACT, no compute |
| C1 | genome length 30,616 nt | run nf-villumina (or reference-guided assembly to KY748210.1) on each deposited run |
| C5b/C5c | mean coverage 4,038.8x / 4,590.5x | map each deposited run to reference, compute mean depth |
| C6 | 99.96% identity to KY748210.1 | align reconstructed assembly to KY748210.1 |
| C7 | GC content 33.2% | compute from reconstructed 30,616 nt assembly |
Out of scope / not attempted
- C2 GenBank accessions (OM470968/OP289523): manual deposit, not a pipeline output.
- C5a/C3c (109.4x, 6,536,978 reads): data not deposited — cannot reproduce.
- Phylogeny / MAFFT / Geneious annotation: manual/interactive.
Compute plan (HARD RULE: heavy compute on «our HPC» only)
- (front1) clone nf-villumina on «infra», pin commit; prefetch SRR17880609 + SRR21177721 fastqs + reference KY748210.1 into «infra» work dir.
- (SLURM) run pipeline / reference-guided mapping per run; pull back small outputs (assembly FASTA, samtools depth summary, length, GC, identity).
- Compare to C1/C5b/C5c/C6/C7; fill claims.tsv + agreement.json + AUDIT.md.
Status note
ENA metadata-based claims (C3a/C3b) reproduce EXACTLY today. Compute-based claims are PENDING a «our HPC» run; the VPN tunnel was down at first attempt (waiting for the central fix, retrying — not touching the VPN per rule 1d).
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
Strong, faithful partial reproduction of an MRA genome announcement: genome length (30,616 nt), GC (33.23%→33.2%), identity (99.96%), and both deposited read counts reproduce exactly and md5-verified, with de novo contigs 100% identical to the deposited genomes (OM470968/OP289523) — a robust anti-fabrication signal, so the central conclusion holds. Two explainable deviations keep this at yellow: C5c coverage is −21.5% low (4,590.5x→3,604.55x) for SRR21177721 while the same method matched the other sample within 1.5%, pointing to a mapper/trimming/duplicate-handling difference in the authors' nf-villumina pipeline. Separately, the 3rd sample the paper analyzes (109.4x / 6,536,978 reads) was never deposited in PRJNA803320 — an authors'-side completeness gap, not a fabrication concern. Overall: core claim confirmed; deviations sit on input/preprocessing and data-availability, not on the genome itself.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.