Genomic Correlates of Virulence Attenuation in the Deadly Amphibian Chytrid Fungus, Batrachochytrium dendrobatidis.
Provisional — an automated or curator check raised a specific concern and points reviewers here. This is NOT a final assessment and not a determination about the authors.
The main results reproduced, with only marginal, non-material deviations.
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
DESCRIBED WELL ENOUGH TO RUN; partial reproduction on the paper's own data (SRP049423: SRR1639248=P9, SRR1639277=P39) vs JEL423 GCA_000149865.1, full named pipeline on «our HPC»/SLURM «job» (n093, 1h02m, self-contained env+download+compute), confirming an earlier independent run (2211889) with bit-identical key numbers. COVERAGE (paper: 'depth of coverage for aligned reads of 33X and 20X'): tested 7 coverage definitions. P39=20X reproduces EXACTLY under every definition (20.5-21.6X) -> within-tol. P9=33X reproduces under NONE: standard aligned depth 20.24X, max 28.36X (soft-clip-inclusive), and 33X equals only P9's RAW theoretical depth (34.39X). Since P39's 20X is genuine aligned depth, the paper's two coverage figures appear to use DIFFERENT definitions (P9 theoretical, P39 aligned) -> P9's 33X is MISMATCH, flagged for human review as a probable theoretical-vs-aligned mislabel (not data fabrication). VARIANTS (2231 SNP / 470 indel between P9 and P39): reproduced in magnitude -- bcftools isec P9-private = 2698 SNP (within ~21%) and 406 indel (within ~14%), both bit-identical across the two independent runs -- but NOT 1:1: GATK1.4 unobtainable (bcftools+gatk4 proxies) and P39's low alignment (53%) injects asymmetric reference-divergence noise. Deposit integrity: read-pair counts match SRA spot counts exactly; no fabrication signal in the data itself. NOT ATTEMPTED: C3b (149 exonic indels; needs snpEff+annotation) and out-of-scope wet-lab/mutation-rate/CNV/expression items. Third-party tool (seqyclean, P16) on paper data. All grades provisional, pending human audit.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 49assessed: 2026-06-21 ⛓ 2819a928a057
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-24
- Rubric version
- not recorded
- Assessed by
- —
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetThe paper tests whether genomic processes (chromosome copy number variation, loss of heterozygosity, and elevated mutation rates in virulence-associated genes) are correlated with the observed attenuation of virulence in a Batrachochytrium dendrobatidis (Bd) isolate maintained over 39 vs. 9 laboratory passages.
- ★ Virulence attenuation in the longer-passaged Bd isolate (JEL427-P39) is associated with loss of chromosome copy number relative to the shorter-passaged, more virulent isolate (JEL427-P9) finding
- ★ Loss of heterozygosity (LOH) is not associated with virulence attenuation between the two isolates finding
- ★ Genes previously found to be up-regulated on frog skin show elevated rates of nonsynonymous mutations and indels in the attenuated isolate, consistent with relaxed selection in culture finding
- ★ Genomic processes (copy number lability, heterozygosity change) proposed as mechanisms for rapid evolution in Bd occur over extremely short timescales (~6 years/30 generations) and correlate with virulence shifts within a single lineage mechanism
- ★ Whole-genome resequencing and comparison of two samples of the same isolate differing only in passage history isolates genomic correlates of virulence attenuation without confounds of ancestry, host, or environment method
- Indel-containing genes are overrepresented for GO terms related to proteolysis and metabolic process, including an M36 metallopeptidase gene implicated in Bd virulence finding
- Both JEL427 isolates cluster together with high bootstrap support within the Global Panzootic Lineage (GPL) finding
- New Bd isolates should be immediately cryo-archived and fresh isolates used for infection experiments rather than long-term cultured samples resource
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Whole-genome resequencing (Illumina MiSeq, 2x250bp paired-end) | Bd strain JEL427, isolates P9 (9 passages) and P39 (39 passages) | laboratory passage history (duration in culture) | SNPs and indels genome-wide | Illumina MiSeq |
| SNP calling / variant calling (Bowtie2 alignment, GATK UnifiedGenotyper/VariantFiltration, snpEff annotation) | Bd JEL427-P9 vs JEL427-P39 aligned to JEL423 reference genome | passage history | number and genomic location/type of SNPs (synonymous/nonsynonymous/intronic) | GATK v.1.4, Bowtie2 v.2.1.0, snpEff v.2.0.5 |
| Indel detection (bedtools) | Bd JEL427-P9 vs JEL427-P39 | passage history | number of indels and coding exonic indels | bedtools |
| Phylogenetic analysis (maximum parsimony, SNP character states) | Bd JEL427-P9, JEL427-P39, plus 29 previously sequenced isolates | none (comparative) | tree topology and bootstrap node support | — |
| Loss of heterozygosity (LOH) analysis (hidden Markov model) | Bd JEL427-P9 and JEL427-P39, 15 largest supercontigs | passage history | genomic regions of long homozygosity stretches | — |
| Chromosome copy number estimation (SNP allele frequency distribution, Gaussian kernel density) | Bd JEL427-P9 and JEL427-P39, 17 supercontigs | passage history | estimated ploidy/copy number per supercontig (disomic to pentasomic) | R package KernSmooth |
| Gene expression correlation analysis (negative binomial GLM regression) using prior microarray data | Bd genes, referencing Rosenblum et al. 2012 frog-skin vs. tryptone growth microarray data | none/comparative (frog skin vs. standard growth medium expression) | correlation between differential expression category and number of nonsynonymous SNPs or indels per gene | — |
| GO term overrepresentation (hypergeometric test) | Bd genes with nonsynonymous mutations or indels that were up-regulated on frog skin | none | enriched GO biological process/molecular function terms | GOstats R package |
- ▼ JEL427-P9 (more virulent) showed chromosome copy number up to pentasomic, while JEL427-P39 (less virulent) showed up to tetrasomic, with 9 of 17 supercontigs having decreased copy number in P39
- – Synonymous change rate was nearly identical in LOH vs non-LOH regions, indicating no major effect of LOH on genotype change 0.00102 vs 0.00105
- ▲ Moderately up-regulated (frog skin) genes had significantly more nonsynonymous mutations than genes with no differential expression P=0.00648
- ▲ Increased gene expression on frog skin was weakly but significantly associated with greater indel occurrence P=0.048
- – 2,231 SNP changes identified between the two isolates, with ~9% nonsynonymous, ~20% synonymous, ~18% intronic, and roughly half in up/downstream regions 2231 SNPs
- – 470 indels identified between isolates, 149 in coding exonic regions 470 indels
- ▲ GO terms proteolysis and macromolecule/metabolic process were overrepresented among genes containing exonic indels P=0.0010 (proteolysis)
- – Both JEL427 isolates clustered together with high bootstrap support, nested within the Global Panzootic Lineage
- count 2,231 SNPs (total SNP changes between JEL427-P9 and JEL427-P39)
- count 470 indels (149 coding exonic) (total indels between isolates)
- pvalue P=0.00648 (negative binomial regression, nonsynonymous mutations vs. moderate up-regulation on frog skin)
- pvalue P=0.048 (regression of indel occurrence on gene expression level)
- pvalue P=0.0010 (GO term proteolysis enrichment among indel-containing genes (count 9 of 282 term size))
- other 0.00102 vs 0.00105 (synonymous change rate in LOH vs non-LOH regions)
- other 33X and 20X (depth of coverage for JEL427-P9 and JEL427-P39 respectively)
- count 84 genes (genes up-regulated on frog skin with ≥1 nonsynonymous change)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
The study compared whole-genome resequencing data from two samples of the same Bd isolate differing only in laboratory passage history (n = 2 samples, one per condition) to identify genomic correlates of virulence attenuation. Chromosome copy number per supercontig was inferred from SNP allele-frequency distributions smoothed with Gaussian kernel density; loss-of-heterozygosity regions were mapped with a hidden Markov model. The relationship between prior microarray-derived gene-expression categories and per-gene mutation counts was tested with negative binomial generalized linear models (gene length as covariate), and GO-term overrepresentation was assessed with hypergeometric tests. Results were reported as exact p-values; model parameter 95% confidence intervals were stated to have been estimated.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Negative binomial generalized linear model (GLM) | Association between differential-expression category (5 levels) and number of nonsynonymous, synonymous, or combined nucleotide changes per gene (Figure 3) | Full gene-set size not stated; 84 genes met up-regulated + ≥1 nonsynonymous-change criteria | not stated |
| Negative binomial generalized linear model (GLM) | Association between differential-expression category (3 levels) and indel count per gene (Figure 4); repeated with all indel-containing genes | null | not stated |
| Hypergeometric test (GOstats R package) | GO-term overrepresentation in (a) up-regulated genes with nonsynonymous changes (Table S2) and (b) all genes with exonic indels (Table 1) | null | not stated |
| Gaussian kernel density estimation (R package KernSmooth) | Smoothing of per-supercontig SNP allele-frequency distributions to assign chromosome copy number (Figure 2) | null | na |
| Hidden Markov Model (HMM) | Detection of LOH regions (long homozygous stretches) across the 15 largest supercontigs | null | na |
| Maximum parsimony phylogenetic analysis with 200 bootstrap replicates | Phylogenetic placement of both JEL427 isolates among 31 total Bd isolates (Figure 1) | 31 isolates (29 from prior study + 2 focal isolates) | na |
-
The genomic comparison rested on n = 2 sequenced samples (one per passage condition) with no biological replicates of the sequencing↳ Could also: Sequencing multiple independently cultured replicates of each passage stage would constitute biological replication — With a single observation per condition, observed genomic differences cannot be distinguished from normal culture-to-culture variation; replication would enable formal statistical inference rather than descriptive comparison of two data points
-
Multiple GO-term overrepresentation tests were performed (Table 1 reports 7 significant terms; Table S2 covers a second gene set) without a stated multiple-testing correction↳ Could also: Applying Benjamini-Hochberg FDR correction across all tested GO terms is standard practice in enrichment analyses — Testing many GO terms simultaneously inflates the expected number of false positives; FDR control quantifies that inflation and is the dominant convention in GO enrichment literature
-
Chromosome copy number was inferred from the modal peaks of SNP allele-frequency distributions smoothed by Gaussian kernel density↳ Could also: Read-depth-based copy number tools (e.g., CNVkit, Control-FREEC, or BIC-seq2) estimate copy number from normalized per-window read depth independently of local SNP density — Read-depth and allele-frequency methods capture different signals; using both provides cross-validation and is less sensitive to regions with few heterozygous SNPs
-
Phylogenetic relationships were inferred with maximum parsimony and bootstrap support↳ Could also: Maximum likelihood (e.g., IQ-TREE, RAxML) or Bayesian (e.g., MrBayes) methods could also be applied to the SNP matrix — Likelihood and Bayesian approaches model substitution processes explicitly, typically yield better-calibrated branch-length estimates for genome-wide SNP data, and generate posterior probabilities or likelihood-ratio-based support
-
The negative binomial GLM results were summarized by a p-value; regression coefficients (the magnitude of association between expression category and mutation count) were not reported numerically in the text↳ Could also: Reporting the estimated regression coefficients and their 95% CIs would also convey the effect magnitude directly — Effect-size estimates allow readers to evaluate biological relevance independently of sample size; the paper states CIs were computed, so presenting them explicitly would make the results more interpretable
-
The fit of the negative binomial distribution was assumed but not formally evaluated or compared against alternatives↳ Could also: A Poisson GLM or zero-inflated negative binomial model could also be fit, with model selection via AIC or a likelihood-ratio test — Reporting a dispersion parameter or comparing model fits would let readers assess whether overdispersion was substantial enough to require the negative binomial over the simpler Poisson, or whether zero-inflation was present given the many genes with zero mutations
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-26333840
Paper: Refsnider, Poorten, Langhammer, Burrowes, Rosenblum (2015). Genomic Correlates of Virulence Attenuation in the Deadly Amphibian Chytrid Fungus, Batrachochytrium dendrobatidis. G3 5(11):2291–2298. DOI 10.1534/g3.115.021808 · PMID 26333840 · PMCID PMC4632049.
The experiment
A single Bd isolate (JEL427) was serially passaged in culture; an early passage (P9, virulent) and a late passage (P39, attenuated) were whole-genome resequenced on Illumina MiSeq (2×250 bp PE). The paper looks for genomic differences (SNPs, indels, CNV, expression) that correlate with the in-vivo loss of virulence.
Code & data availability
- Code link (BRIEF): https://github.com/ibest/seqyclean — this is a generic third-party read-cleaning tool, NOT the authors' analysis code. Per P16 of the brief, applying this shipped tool to the paper's own data is an equally valid reproduction. The rest of the pipeline (Bowtie2, GATK, snpEff) is named in Methods by tool+version but no author script/repo is shipped.
- Data: SRA study SRP049423 (BioProject). Run accessions resolved via ENA:
SRR1639248= JEL427-P9 (virulent) — 1,586,395 spots, 821,752,610 bpSRR1639277= JEL427-P39 (attenuated) — 2,043,538 spots, 1,058,552,684 bp- (study also has SRR1636746 "Section Line", SRR1636747 "Carter Meadow", SRR1639274 "JEL427-P17" — not used for the P9-vs-P39 comparison.)
- Reference genome: Bd JEL423, Broad Institute v.17-Jan-2007 = GenBank assembly GCA_000149865.1 (BD_JEL423).
Reported pipeline (Methods)
- SeqyClean v1.8.10 — clean reads, remove PCR duplicates / contaminants / adaptors, quality-trim.
- Bowtie2 v2.1.0 — align cleaned reads to JEL423.
- GATK v1.4 best-practices — Picard MarkDuplicates → RealignerTargetCreator
- IndelRealigner → UnifiedGenotyper + VariantFiltration.
- snpEff v2.0.5 + bedtools — variant annotation.
In scope (pipeline-derived, attempted)
| # | Reported result | Pipeline | Feasibility |
|---|---|---|---|
| C1 | Mean genome coverage 33× (P9) and 20× (P39) | seqyclean→bowtie2→dedup→depth | HIGH — direct, low-hanging. Primary target. |
| C2 | 2,231 SNP changes between P9 and P39 | + GATK variant call & compare | MEDIUM — needs full variant pipeline; GATK 1.4 (2012) unobtainable, use modern caller w/ noted divergence. Stretch. |
| C3 | 470 indels (149 in coding exons) | + GATK indel call | MEDIUM/LOW — same caveat as C2. Stretch. |
Out of scope (not pipeline-derived from the deposited reads / not attempted)
- Mutation-rate extrapolation (0.096 changes/kb → 1.6×10⁻⁵ /site/yr) — derived arithmetic on top of C2, plus passage-time assumptions (wet-lab metadata).
- In-vivo virulence / infection-load phenotyping — wet-lab.
- Any CNV / expression / functional annotation interpretation beyond raw counts.
Reproduction strategy (80/20)
Primary: reproduce C1 (coverage) end-to-end with the named tools — this is the cleanest 1:1 and exercises the actual shipped tool (seqyclean). Stretch: attempt C2/C3 with a modern variant caller and report honestly as tool-version-divergent (we will NOT chase exact GATK-1.4 byte-parity — that is the hard last 20%). All heavy compute on «our HPC»/SLURM; «infra» work dir.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.