A Deeper Insight into Evolutionary Patterns and Phylogenetic History of ASFV Epidemics in Sardinia (Italy) through Extensive Genomic Sequencing.
The main results reproduced: recomputed values matched the published ones within tolerance.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- Nothing in this column.
- 🟡Could not use the authors’ exact input data
- 🟡Reported values were only indirectly comparable
- 🟡A deviation arose in the data or preprocessing
- 🟡A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The deviation was non-trivial in magnitude
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
PARTIAL (honest, data-availability-limited). «our HPC» reachable; SLURM «job» ran the full per-sample pipeline on n094. Pipeline: ENA fastq -> Trim Galore --paired -> bwa-mem vs ASFV ref KX354450 (strain 47/Ss/2008, Stintino/Sardinia, genotype I, 184,638 bp -- CORRECTED: scope.md had mislabeled it Georgia 2007/1) -> samtools sort -> Picard MarkDuplicates -> FreeBayes(--ploidy 1 -X -u -m 20 -q 20 -F 0.2) -> bcftools consensus (QUAL>=20, zero-cov masked). Single SRA run SRR13976567 (isolate NU1981_2, 68,097 read pairs, 99.98% mapped). RESULT: 3/4 per-sample metrics land inside the paper's reported cohort ranges -- length 181,640 bp non-N (reported 181,684-181,925), GC 38.60% (reported 38.56-38.60), median depth 80-82x (reported 6-250x). C4 (95 'new' mutations) is a cohort aggregate over all 71 genomes and is NOT reproducible from one genome; per-sample freebayes with the paper's exact params yields 42 confident SNPs vs KX354450. Described well enough to reproduce the per-sample pipeline 1:1. NOT attempted: cohort phylogenetics/dating and the 58-isolate cohort assembly (only 1 of 71 genomes provided). DATA NOTE: NCBI sra-tools fasterq-dump de-paired this run (R1 != R2); ENA direct fastq gave clean pairs.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 50assessed: 2026-06-18 ⛓ 3db117afccb7
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-25
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetThe study tests whether genetic variability within a single, spatially-localized ASFV outbreak can be adequately captured by a small number of whole genomes, or whether a much larger sequencing effort improves phylogenetic resolution and evolutionary/epidemiological inference of the Sardinian ASFV epidemic.
- ★ 58 new whole genomes of Sardinian ASFV isolates were sequenced, the largest ASFV whole-genome sequencing effort to date resource
- ★ The new sequences capture more than twice the genomic and phylogenetic diversity of all previously published Sardinian ASFV sequences finding
- ★ The increased diversity enables, for the first time, resolution of genetic substructure within the outbreak, revealing multiple ASFV subclusters, some coexisting in space and time finding
- ★ This is the first phylodynamic inference of ASFV based on whole genomes applied to a closed epidemic system method
- Sardinian ASFV isolates after 1990 (modern strains) show deletions in both the B602L and EP402R genes compared to historical strains isolated before 1990 finding
- ★ Ninety-five new point mutations were detected, located within MGF360, MGF110 and MGF505 gene families, along with seven URFs and indels finding
- Genome annotation identified 231 ORFs (165 protein-coding genes and 66 uncharacterized reading frames) finding
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Whole-genome sequencing (NGS) | ASFV isolates from porcine monocyte/macrophage cultures (domestic pig, wild boar, illegal free-ranging pig samples, Sardinia 1978-2018) | none | full viral genome sequence, nucleotide mutations, indels | Illumina HiSeq 2500 and NovaSeq 6000, Nextera XT/DNA Flex library prep |
| Sanger sequencing | B602L and EP402R gene regions of ASFV isolates | none | confirmation of gene sequence/deletions | Sanger sequencing |
| Haemoadsorption test (Malmquist test) | porcine monocyte/macrophage monolayers | none | presence of infectious ASFV | — |
| Bayesian phylogenetic reconstruction | 71 Sardinian ASFV WGSs plus 3 outgroup genomes (L60, E75, Benin97) | none | tree topology, clustering/subgroup structure | MrBayes 3.2.7 |
| Molecular dating / phylodynamic analysis | same 71-genome WGS dataset | none | time-scaled MCC tree, evolutionary substitution rate, lineage-through-time dynamics | BEAST 1.8.2, Tracer 1.7 |
| Principal coordinate analysis (PCoA) | ASFV genomes from 2002-2008 cluster (n=46) | none | genetic dissimilarity/subgroup structure | GenAlEx 6.5 |
| Likelihood-mapping analysis | genome alignment of ASFV dataset | none | phylogenetic signal reliability (noise/saturation/recombination check) | TreePuzzle |
| Pearson's Chi-squared test | phylogenetic clusters vs. geographic origin of isolates | none | statistical association between genetic structuring and geography | R stats package |
- – 58 new ASFV whole genomes sequenced, representing the largest ASFV WGS effort to date
- ▲ New sequences capture substantially more genomic and phylogenetic diversity than all previously published Sardinian sequences combined more than twice
- – Multiple ASFV subclusters identified within the Sardinian epidemic phylogeny, some coexisting in space and time
- – Complete genome sequences obtained for a subset of the 58 isolates analyzed 40 of 58
- – New point mutations detected within MGF360, MGF110 and MGF505 gene families 95
- – Genome annotation yielded protein-coding genes and uncharacterized reading frames 231 ORFs (165 protein-coding, 66 URFs)
- – Sequenced genome lengths fell within a narrow range 181,684–181,925 bp
- – GC content of sequenced genomes was consistent across isolates 38.56–38.60%
- count 58 (new ASFV whole genomes sequenced in this study)
- count 40/58 (isolates with complete genome sequences obtained)
- other 181,684–181,925 bp (range of assembled genome lengths)
- other 38.56–38.60% (range of GC content across sequenced genomes)
- count 231 ORFs (165 protein-coding, 66 URFs) (GATU genome annotation results)
- count 95 (new point mutations detected across MGF360, MGF110, MGF505)
- count 2253 (total ASFV outbreaks reported in Sardinia, 1978-2018)
- count n=46 (genomes from 2002-2008 cluster analyzed by PCoA)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This phylogenomic study characterized the evolutionary history of ASFV in Sardinia (1978–2018) using 74 whole-genome sequences (58 newly generated, 13 published, plus 3 outgroup genomes). Phylogenetic structure was inferred via Bayesian MCMC (MrBayes) and molecular dating via Bayesian phylodynamics (BEAST), with clock and demographic model selection by Bayes Factor comparison. Genetic diversity was quantified using Faith's phylogenetic diversity index and pairwise nucleotide statistics, and a Pearson's chi-squared test with Monte Carlo permutation was used to test geographic–phylogenetic association.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Bayesian MCMC phylogenetic inference (GTR+I+G model, two independent runs of 4 Metropolis-Coupled chains, 5 million generations, 25% burn-in; convergence via ASDSF and PSRF) | Main phylogeny of 74 ASFV whole genomes (71 Sardinian + 3 outgroups) | 74 | not stated |
| Bayesian molecular dating with strict and uncorrelated log-normal relaxed clock models under constant, exponential, expansion, and Bayesian Skyline demographic models; clock/demographic model selected by Bayes Factor (2lnBF via Tracer); final MCMC 100 million generations, ESS >200 required | Estimation of evolutionary rate and time-calibrated phylogeny of Sardinian ASFV; also applied to subsets to investigate temporal changes in genetic diversity | 71 | not stated |
| Likelihood mapping of 10,000 random quartets (TreePuzzle) | Verification of phylogenetic signal in the whole-genome alignment before phylodynamic inference | 74 | not stated |
| Faith's phylogenetic diversity index on neighbor-joining trees (Hamming distance per base); quantiles from 100 random subsamples of each size | Quantification of phylogenetic information contributed by additional sequences across different subsample sizes | Up to 71 Sardinian sequences (subsamples of varying size) | not stated |
| Principal coordinate analysis (PCoA, GenAlEx 6.5) based on genetic dissimilarity among genomes | Identification of virus subgroups within the 2002–2008 genetic cluster | 46 | not stated |
| Pearson's chi-squared test with Monte Carlo simulation (2000 replicates; R stats package) | Test of association between phylogenetic cluster membership and geographic origin of isolates | — | not stated |
-
Clock and demographic model selection was performed using the Bayes Factor (2lnBF) estimated from MCMC log-likelihoods in Tracer, which corresponds to the harmonic mean estimator↳ Could also: Stepping-stone or path-sampling marginal likelihood estimation in BEAST could also be used for model comparison — Stepping-stone and path-sampling are considered more numerically stable estimators of the marginal likelihood than the harmonic mean; they are increasingly recommended in phylodynamic workflows when computing 2lnBF for clock and demographic model selection
-
The main phylogeny was inferred using Bayesian MCMC with MrBayes under GTR+I+G↳ Could also: Maximum likelihood phylogeny with ultrafast bootstrap support (e.g., IQ-TREE 2) could also have been constructed as a complementary or independent tree — ML inference is computationally faster and produces bootstrap support values that are straightforward to interpret; it is widely used as a cross-check on Bayesian topologies, especially for large genome alignments
-
Phylogeographic association between cluster membership and geographic origin was assessed with a Pearson's chi-squared test (with Monte Carlo permutation)↳ Could also: Bayesian Tip Significance testing (BaTS) or a Mantel test on patristic genetic distances versus geographic distances could also be applied — BaTS explicitly accounts for phylogenetic uncertainty across the posterior tree distribution when testing trait–phylogeny associations, and a Mantel test evaluates continuous isolation-by-distance; either approach would complement a chi-squared test that treats cluster assignment as discrete and fixed
-
Subgroup structure in the 2002–2008 cluster was visualized using PCoA in GenAlEx↳ Could also: A neighbor-net phylogenetic network (e.g., SplitsTree4) or PCA directly on the SNP matrix could also depict genetic structure — Network methods represent reticulate relationships and homoplasy more explicitly than distance-based PCoA, which can be particularly informative for dense, low-divergence viral genomes where tree-like signal may be incomplete
-
Clock signal was assessed indirectly through BEAST model comparison and ESS diagnostics rather than a dedicated pre-analysis check↳ Could also: A root-to-tip regression against sampling date (e.g., in TempEst) prior to BEAST analysis could also be used to evaluate clock signal and flag temporal outliers — Root-to-tip regression explicitly tests whether collection dates correlate with genetic divergence — a prerequisite for reliable Bayesian molecular dating — and can identify sequences that are inconsistent with a molecular clock before investing substantial computational resources
-
The sequencing effort (n = 58 new isolates, ~15% of archived strains, at least 5% per year) was determined by pragmatic sampling criteria without a formal power analysis↳ Could also: Simulation-based power analyses for Bayesian phylodynamic inference could also be used to estimate how many sequences are needed to recover reliable evolutionary rate and demographic history estimates — Because the paper explicitly poses the question of how many sequences suffice for reliable phylodynamic inference, a simulation framework (e.g., using BEAST with synthetic data of varying n) would provide a quantitative answer alongside the empirical rarefaction analysis already performed
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-34696424
Paper: Fiori et al. 2021, Viruses 13:1994. "A Deeper Insight into Evolutionary Patterns and Phylogenetic History of ASFV Epidemics in Sardinia (Italy) through Extensive Genomic Sequencing." DOI 10.3390/v13101994 · PMCID PMC8539718.
Code link in record: https://github.com/FelixKrueger/TrimGalore (third-party read-trimming tool — valid per brief rule P16: applying an existing third-party tool to the paper's own data is a legitimate reproduction).
Data handed to this room: sra:SRR13976567 — a SINGLE run = sample NU1981_2
(BioSample SAMN18312889), ASFV WGS, Illumina NovaSeq 6000, paired-end 2×150,
68,156 read pairs / 20.06 Mbp. This is ONE of the 57/58 newly sequenced isolates.
Pipeline described (Methods 2.3, 2.5)
- Trim Galore — quality-trim + adapter removal of raw reads.
- bwa-mem — align reads; reads mapping uniquely to the ASFV genome retained. (Reference accession KX354450 = ASFV strain 47/Ss/2008 (Stintino, Sardinia, genotype I, 184,638 bp) — CORRECTED 2026-06-25: an earlier draft of this scope mislabeled KX354450 as "Georgia 2007/1"; that is wrong, KX354450 is the Sardinian genotype-I reference, which is why the new isolates' lengths (~181.6 kb) sit just below it. Verified against PMC8539718 Methods + GenBank.) Note: Methods text mentions Sus scrofa 10.2 to subtract host; the ASFV consensus is built against KX354450.
- samtools sort+index; Picard MarkDuplicates (dedup).
- GEM re-alignment (host/uniqueness filtering).
- Freebayes variant calling:
--ploidy 1 -X -u -m 20 -q 20 -F 0.2. - Consensus genome per sample → MSA → phylogenetics (MrBayes 3.2.7: NST=6, rates=invgamma, ngammacat=4; BEAST 1.8.2 molecular dating).
IN SCOPE (pipeline-derived, attempted on SRR13976567 = NU1981_2)
| # | Reported result | Pipeline | Note |
|---|---|---|---|
| C1 | ASFV genome length per isolate: 181,684–181,925 bp (range over all) | TrimGalore→bwa-mem→consensus vs KX354450 | check NU1981_2 length within/near range |
| C2 | GC content 38.56–38.60% | consensus | check NU1981_2 GC in range |
| C3 | Median coverage 6–250× (Table S2) | bwa-mem depth | check NU1981_2 depth plausible |
| C4 | Point mutations vs reference (95 new across dataset; Table S3) | Freebayes | count high-confidence variants in NU1981_2 vs KX354450 |
OUT OF SCOPE (not attempted)
- Wet-lab: DNA extraction, library prep, the actual Illumina sequencing.
- Aggregate / multi-genome results: the 95 new point mutations, MrBayes tree topology, BEAST tMRCA / molecular dating, recombination, selection analyses — these require ALL 71 WGS genomes (only 1 run was provided to this room) and external reference panels. Out of scope: single-accession room.
- N reported = 58 new isolates / 71 total: only 1 SRA run available here; we profile that one run, not the cohort.
Honesty note
Single-run reproduction: we can faithfully run the per-sample read→consensus→variant pipeline on NU1981_2 and check the derived metrics fall in the paper's reported ranges. We CANNOT reproduce cohort-level phylogenetics from one accession. This is a genuine partial reproduction by data availability, not a failure.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
This is an incomplete, in-progress reproduction: the pipeline (TrimGalore→bwa-mem vs KX354450→samtools/Picard→Freebayes) is scoped and scripted, reported cohort values (genome 181,684–181,925 bp; GC 38.56–38.60%; coverage 6–250x; 95 new mutations) are recorded, but no reproduced values exist because «our HPC» compute is pending and compare.py has not run. The only deviation so far is on our/data-availability side: 1 of 58 SRA runs was handed to the room, so cohort-level and phylogenetic claims are out of scope rather than refuted. There is no fabrication signal — the underlying SRA data is public and the reported per-isolate ranges are plausibly derivable once the single run completes. Graded all-yellow: a solid setup with explainable scoping/coverage deviations and nothing yet computed to mark green or red.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.