Towards reliable whole genome sequencing for outbreak preparedness and response.
The main results reproduced: recomputed values matched the published ones within tolerance.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Same input data as the authors
- ✓Reported values were directly comparable
- ✓Any deviation was negligible
- ✓The central claim held under reproduction
- 🟡A deviation arose in the data or preprocessing
- 🟡A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Described well enough to reproduce. Ran the authors' own pipeline (github dnieuw/platform-comparison-arbovirus @7bfb2d0, env pinned to their environment.yml versions) verbatim on all 60 PRJEB47177 read sets on «our HPC». Strong 1:1: the primary result (Fig 3 per-sample %reliable-genome-coverage) and 10/11 attempted claims reproduce exact/within-tol -- notably Meta-Illumina Ct33 range 3.2-46.1% vs reported 3-46% (exact), Meta-Nanopore only USUV+ZIKV complete (exact), amplicon Nanopore exception read counts 682k/917k (exact), dedup 87.4% vs 88%, capture 35.2x vs 'up to 35x', and all 5 reported false-positive variant positions recovered in high-Ct Illumina samples (exact). The single mismatch (R2c: meta-Nanopore '5.2M reads' vs 1.49M deposited) is a deposit-completeness nuance (the 5.2M is a run-level/pre-demultiplex count not in PRJEB47177), flagged as possible overstatement, not a wrong recomputed value. NOT attempted: Table 1 turn-around-time/cost (operational) and all wet-lab steps (out of scope). Only reconstruction needed was the Ct/approach->run mapping, recovered from ENA library_selection + run_alias order and triple-validated.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 84assessed: 2026-06-22 ⛓ 6c6040694756
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-22
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-22no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetThe paper tests how five whole-genome sequencing approaches (amplicon, metagenomic, and capture-based, on Nanopore and Illumina platforms) compare in performance for detecting and generating complete genomes of four flaviviruses at varying viral loads, to determine the optimal method for outbreak preparedness and public health decision making.
- ★ Amplicon-based Nanopore sequencing can rapidly generate whole genome sequences in samples with viral load up to Ct 33. finding
- ★ Capture-based Illumina sequencing (VirCapSeq-VERT) is the most sensitive method for initial virus determination. finding
- ★ The optimal sequencing approach and platform depends on the purpose (rapid outbreak response vs. detailed phylogenetic analysis). finding
- ★ A custom script trimming primer sequences from BAM files using primer coordinates reduces false-positive variant calls from poorly trimmed primers. method
- ★ Sambamba 'markdup' dereplication resolves PCR amplification errors in metagenomic/capture consensus sequences but greatly reduces mapped read counts. method
- ★ Nanopore sequencing has particular difficulty accurately calling insertions/deletions and low-frequency substitutions compared to Illumina. finding
- ★ Amplicon-based Nanopore sequencing is the preferred approach for rapid contact tracing when the target virus is known a priori. finding
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Amplicon-based whole genome sequencing | Cell culture supernatants of USUV, WNV, YFV, ZIKV diluted to Ct 25/29/33 | dilution (viral load titration) | genome coverage %, read counts, coverage depth | Illumina and Oxford Nanopore |
| Metagenomic (shotgun) sequencing | Cell culture supernatants of USUV, WNV, YFV, ZIKV diluted to Ct 25/29/33 | dilution (viral load titration) | % viral reads, genome coverage, read counts | Illumina and Oxford Nanopore |
| Capture-based sequencing (VirCapSeq-VERT probes) | Cell culture supernatants of USUV, WNV, YFV, ZIKV diluted to Ct 25/29/33, pooled before capture | dilution (viral load titration) | % viral reads, genome coverage | Illumina only |
| Variant calling / consensus sequence quality analysis | Aligned reads (BAM files) from all sequencing approaches for the four viruses | none (cross-method comparison) | major variant positions/frequencies, error classification | PySam; Sambamba (markdup) |
- – Illumina amplicon sequencing yielded equal or greater genome coverage than Nanopore in 10 of 12 sample comparisons; Nanopore outperformed for WNV Ct29 and Ct33. 5-9% more coverage for Nanopore in exceptions
- – Coverage varied greatly across amplicons, with extreme differences between well- and poorly-performing amplicons. >51,000-fold; up to 800,000x vs 0x in WNV Ct33 Illumina
- ▲ Metagenomic Nanopore sequencing produced a higher percentage of virus-specific reads than metagenomic Illumina on average. 0.8 percentage points more (avg 1.8% viral reads)
- ▼ Metagenomic Illumina gave reliable complete genomes for Ct25/29 samples (except WNV); Ct33 samples had much lower coverage. 3-46% genome coverage at Ct33
- ▼ Metagenomic Nanopore only achieved reliable complete genomes for the highest viral load USUV and ZIKV samples. na
- – Capture-based Illumina had a much higher percentage of viral reads than metagenomic Illumina, comparable to amplicon approaches, but performed worse at Ct33. na
- ▼ Sambamba markdup dereplication removed most consensus sequence errors from PCR amplification but drastically reduced mapped reads. 88% reduction in mapped reads on average
- – Illumina had only five erroneous false-positive substitutions across all data, versus more frequent indel/substitution errors in Nanopore data. 5 false positives (WNV 4509, 5034, 7088; YFV 822, 3711)
- count ~3 million reads per sample (Illumina amplicon), with one 5.4M exception (amplicon sequencing read yield, Illumina)
- count ~400k reads per sample (Nanopore amplicon), with 682k and 917k exceptions (amplicon sequencing read yield, Nanopore)
- fold_change >51,000-fold coverage difference between performant and failing amplicons (amplicon coverage depth variability)
- other 800,000x coverage in one genome region vs 0x elsewhere (Illumina amplicon WNV Ct33)
- other 88% reduction in mapped reads on average after Sambamba markdup (PCR duplicate removal effect on read depth)
- count 50-61 million reads (Illumina metagenomic/capture) vs 5.2 million reads (Nanopore metagenomic) (total sequencing data volume comparison)
- fold_change 15x more reads but only 5x more nucleotides for Illumina vs Nanopore at Ct25 metagenomic (read count vs nucleotide yield comparison due to Nanopore's longer reads)
- fold_change 2-fold variance (Nanopore metagenomic) vs 35-fold variance (capture-based Illumina) between highest and lowest viral load samples (pooling bias across viral loads)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
The study compared five virus whole-genome sequencing approaches (amplicon-Illumina, amplicon-Nanopore, metagenomic-Illumina, metagenomic-Nanopore, capture-Illumina) across four flaviviruses diluted to three Ct values (25, 29, 33). Results were reported descriptively using read counts, percentages of viral reads, percentage genome coverage, fold-differences between approaches/platforms, and coverage-depth profiles across the genome, along with a variant/error comparison across approaches using PySam and Sambamba. No formal inferential statistical tests (e.g., t-tests, ANOVA) or p-values were reported in the text.
-
Differences in read counts, percentage coverage, and fold-changes between sequencing approaches/platforms were summarized descriptively (e.g., 'Nanopore covered 5 and 9% more of the genome').↳ Could also: A formal statistical model (e.g., ANOVA or a generalized linear/mixed-effects model with approach, platform, virus, and Ct as factors, or a nonparametric Kruskal-Wallis/Wilcoxon test) could also be applied to these metrics — This would let readers see a quantified estimate (with a p-value or confidence interval) of how much of the observed differences might reflect systematic effects of approach/platform versus run-to-run variability, complementing the descriptive comparison already presented.
-
Each virus/Ct/approach combination appears to represent a single sequencing run, without reported replicate variability.↳ Could also: Incorporating biological or technical replicates and reporting a dispersion measure (SD, SEM, or a 95% CI) for metrics like percentage coverage or read counts could also be used — Replicate-based variance estimates would allow the magnitude of run-to-run variation to be distinguished from differences attributable to the sequencing approach itself.
-
Detection/coverage sensitivity was expressed as Ct-value thresholds (e.g., 'up to Ct 33') for each approach.↳ Could also: A logistic/probit regression modeling detection probability (or reliable-coverage probability) as a function of Ct value could also be used to estimate a formal limit of detection (e.g., LOD50/LOD95) with confidence intervals — This would provide a continuous, statistically estimated sensitivity threshold rather than a discrete Ct cutoff, which can be useful when comparing methods with only a few dilution points.
-
True versus erroneous variants were distinguished using a rule-based approach (agreement across approaches, primer-trim correction, PCR-duplicate removal, softclip filtering).↳ Could also: A probabilistic variant-calling framework (e.g., a binomial/beta-binomial model of allele frequency, or tools such as LoFreq/iVar that incorporate base-quality-aware statistical thresholds) could also be used to flag true variants versus sequencing errors — A model-based approach can provide an estimated false-positive/false-negative rate or confidence score for each variant call, complementing the heuristic filtering steps already used.
-
Pooling-related discrepancies in read yield across Ct values (e.g., up to 35-fold difference in capture-based Illumina sequencing) were described qualitatively.↳ Could also: A regression analysis relating input viral load (Ct) to resulting read yield or genome coverage could also be fitted — This would let the pooling-bias trend be expressed as a quantitative relationship (slope/coefficient) rather than as fold-difference examples, which can help predict behavior at Ct values not directly tested.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
Running the authors' own Snakemake pipeline verbatim on the exact deposited data (PRJEB47177, all 60 read sets), 10/11 attempted in-scope claims reproduce exact or within-tolerance, including the primary Fig 3 coverage result and all five reported false-positive variant positions. The sole mismatch (R2c) — reported ~5.2M meta-Nanopore reads vs 1.49M deposited — is a deposit-completeness nuance (the 5.2M is a run-level/pre-demux count not in the SRA deposit), i.e. one value not derivable from shared data, not a wrong recomputed value. The only methodological assumption was reconstructing the Ct→run mapping (ENA anonymised), independently validated three ways. Net: a strong, solid reproduction; deviations are negligible and the central conclusion holds, with criticality kept to yellow only by the non-derivable read-count figure.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.