Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Towards reliable whole genome sequencing for outbreak preparedness and response.

BMC Genomics · 2022
L1 84/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +4
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
  • Any deviation was negligible
  • The central claim held under reproduction
What did not (or only partly)
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
84/100
Reproducibility score
0.6 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 63% of all assessed papers rank 392 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough to reproduce. Ran the authors' own pipeline (github dnieuw/platform-comparison-arbovirus @7bfb2d0, env pinned to their environment.yml versions) verbatim on all 60 PRJEB47177 read sets on «our HPC». Strong 1:1: the primary result (Fig 3 per-sample %reliable-genome-coverage) and 10/11 attempted claims reproduce exact/within-tol -- notably Meta-Illumina Ct33 range 3.2-46.1% vs reported 3-46% (exact), Meta-Nanopore only USUV+ZIKV complete (exact), amplicon Nanopore exception read counts 682k/917k (exact), dedup 87.4% vs 88%, capture 35.2x vs 'up to 35x', and all 5 reported false-positive variant positions recovered in high-Ct Illumina samples (exact). The single mismatch (R2c: meta-Nanopore '5.2M reads' vs 1.49M deposited) is a deposit-completeness nuance (the 5.2M is a run-level/pre-demultiplex count not in PRJEB47177), flagged as possible overstatement, not a wrong recomputed value. NOT attempted: Table 1 turn-around-time/cost (operational) and all wet-lab steps (out of scope). Only reconstruction needed was the Ct/approach->run mapping, recovered from ENA library_selection + run_alias order and triple-validated.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 84
    assessed: 2026-06-22 ⛓ 6c6040694756
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-22
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-22
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The paper tests how five whole-genome sequencing approaches (amplicon, metagenomic, and capture-based, on Nanopore and Illumina platforms) compare in performance for detecting and generating complete genomes of four flaviviruses at varying viral loads, to determine the optimal method for outbreak preparedness and public health decision making.

Core claims
  • Amplicon-based Nanopore sequencing can rapidly generate whole genome sequences in samples with viral load up to Ct 33. finding
  • Capture-based Illumina sequencing (VirCapSeq-VERT) is the most sensitive method for initial virus determination. finding
  • The optimal sequencing approach and platform depends on the purpose (rapid outbreak response vs. detailed phylogenetic analysis). finding
  • A custom script trimming primer sequences from BAM files using primer coordinates reduces false-positive variant calls from poorly trimmed primers. method
  • Sambamba 'markdup' dereplication resolves PCR amplification errors in metagenomic/capture consensus sequences but greatly reduces mapped read counts. method
  • Nanopore sequencing has particular difficulty accurately calling insertions/deletions and low-frequency substitutions compared to Illumina. finding
  • Amplicon-based Nanopore sequencing is the preferred approach for rapid contact tracing when the target virus is known a priori. finding
Experimental setups
Assay System Perturbation Readout Platform
Amplicon-based whole genome sequencing Cell culture supernatants of USUV, WNV, YFV, ZIKV diluted to Ct 25/29/33 dilution (viral load titration) genome coverage %, read counts, coverage depth Illumina and Oxford Nanopore
Metagenomic (shotgun) sequencing Cell culture supernatants of USUV, WNV, YFV, ZIKV diluted to Ct 25/29/33 dilution (viral load titration) % viral reads, genome coverage, read counts Illumina and Oxford Nanopore
Capture-based sequencing (VirCapSeq-VERT probes) Cell culture supernatants of USUV, WNV, YFV, ZIKV diluted to Ct 25/29/33, pooled before capture dilution (viral load titration) % viral reads, genome coverage Illumina only
Variant calling / consensus sequence quality analysis Aligned reads (BAM files) from all sequencing approaches for the four viruses none (cross-method comparison) major variant positions/frequencies, error classification PySam; Sambamba (markdup)
Key results
  • Illumina amplicon sequencing yielded equal or greater genome coverage than Nanopore in 10 of 12 sample comparisons; Nanopore outperformed for WNV Ct29 and Ct33. 5-9% more coverage for Nanopore in exceptions
  • Coverage varied greatly across amplicons, with extreme differences between well- and poorly-performing amplicons. >51,000-fold; up to 800,000x vs 0x in WNV Ct33 Illumina
  • Metagenomic Nanopore sequencing produced a higher percentage of virus-specific reads than metagenomic Illumina on average. 0.8 percentage points more (avg 1.8% viral reads)
  • Metagenomic Illumina gave reliable complete genomes for Ct25/29 samples (except WNV); Ct33 samples had much lower coverage. 3-46% genome coverage at Ct33
  • Metagenomic Nanopore only achieved reliable complete genomes for the highest viral load USUV and ZIKV samples. na
  • Capture-based Illumina had a much higher percentage of viral reads than metagenomic Illumina, comparable to amplicon approaches, but performed worse at Ct33. na
  • Sambamba markdup dereplication removed most consensus sequence errors from PCR amplification but drastically reduced mapped reads. 88% reduction in mapped reads on average
  • Illumina had only five erroneous false-positive substitutions across all data, versus more frequent indel/substitution errors in Nanopore data. 5 false positives (WNV 4509, 5034, 7088; YFV 822, 3711)
Key statistics
  • count ~3 million reads per sample (Illumina amplicon), with one 5.4M exception (amplicon sequencing read yield, Illumina)
  • count ~400k reads per sample (Nanopore amplicon), with 682k and 917k exceptions (amplicon sequencing read yield, Nanopore)
  • fold_change >51,000-fold coverage difference between performant and failing amplicons (amplicon coverage depth variability)
  • other 800,000x coverage in one genome region vs 0x elsewhere (Illumina amplicon WNV Ct33)
  • other 88% reduction in mapped reads on average after Sambamba markdup (PCR duplicate removal effect on read depth)
  • count 50-61 million reads (Illumina metagenomic/capture) vs 5.2 million reads (Nanopore metagenomic) (total sequencing data volume comparison)
  • fold_change 15x more reads but only 5x more nucleotides for Illumina vs Nanopore at Ct25 metagenomic (read count vs nucleotide yield comparison due to Nanopore's longer reads)
  • fold_change 2-fold variance (Nanopore metagenomic) vs 35-fold variance (capture-based Illumina) between highest and lowest viral load samples (pooling bias across viral loads)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The study compared five virus whole-genome sequencing approaches (amplicon-Illumina, amplicon-Nanopore, metagenomic-Illumina, metagenomic-Nanopore, capture-Illumina) across four flaviviruses diluted to three Ct values (25, 29, 33). Results were reported descriptively using read counts, percentages of viral reads, percentage genome coverage, fold-differences between approaches/platforms, and coverage-depth profiles across the genome, along with a variant/error comparison across approaches using PySam and Sambamba. No formal inferential statistical tests (e.g., t-tests, ANOVA) or p-values were reported in the text.

Replicationunclear Groups5 sequencing approach/platform combinations (amplicon-Illumina, amplicon-Nanopore, metagenomic-Illumina, metagenomic-Nanopore, capture-Illumina) across 4 viruses (USUV, WNV, YFV, ZIKV) at 3 Ct dilutions (25, 29, 33) Pairingna Randomization/blindingnot stated Dispersionnone Exact p-valuesno Effect sizesno Confidence intervalsno
Approaches that could also have been used
  • Differences in read counts, percentage coverage, and fold-changes between sequencing approaches/platforms were summarized descriptively (e.g., 'Nanopore covered 5 and 9% more of the genome').
    Could also: A formal statistical model (e.g., ANOVA or a generalized linear/mixed-effects model with approach, platform, virus, and Ct as factors, or a nonparametric Kruskal-Wallis/Wilcoxon test) could also be applied to these metrics — This would let readers see a quantified estimate (with a p-value or confidence interval) of how much of the observed differences might reflect systematic effects of approach/platform versus run-to-run variability, complementing the descriptive comparison already presented.
  • Each virus/Ct/approach combination appears to represent a single sequencing run, without reported replicate variability.
    Could also: Incorporating biological or technical replicates and reporting a dispersion measure (SD, SEM, or a 95% CI) for metrics like percentage coverage or read counts could also be used — Replicate-based variance estimates would allow the magnitude of run-to-run variation to be distinguished from differences attributable to the sequencing approach itself.
  • Detection/coverage sensitivity was expressed as Ct-value thresholds (e.g., 'up to Ct 33') for each approach.
    Could also: A logistic/probit regression modeling detection probability (or reliable-coverage probability) as a function of Ct value could also be used to estimate a formal limit of detection (e.g., LOD50/LOD95) with confidence intervals — This would provide a continuous, statistically estimated sensitivity threshold rather than a discrete Ct cutoff, which can be useful when comparing methods with only a few dilution points.
  • True versus erroneous variants were distinguished using a rule-based approach (agreement across approaches, primer-trim correction, PCR-duplicate removal, softclip filtering).
    Could also: A probabilistic variant-calling framework (e.g., a binomial/beta-binomial model of allele frequency, or tools such as LoFreq/iVar that incorporate base-quality-aware statistical thresholds) could also be used to flag true variants versus sequencing errors — A model-based approach can provide an estimated false-positive/false-negative rate or confidence score for each variant call, complementing the heuristic filtering steps already used.
  • Pooling-related discrepancies in read yield across Ct values (e.g., up to 35-fold difference in capture-based Illumina sequencing) were described qualitatively.
    Could also: A regression analysis relating input viral load (Ct) to resulting read yield or genome coverage could also be fitted — This would let the pooling-bias trend be expressed as a quantitative relationship (slope/coefficient) rather than as fold-difference examples, which can help predict behavior at Ct values not directly tested.
Software: PySam · Sambamba (markdup)

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Figures / tables: Fig S1Table
R1b
Reported
Meta-Illumina Ct33: 3-46% of genome >=5x
Reproduced
3.2-46.1% (USUV 32.0/WNV 3.2/YFV 12.4/ZIKV 46.1)
exact
R1c
Reported
Meta-Nanopore: only USUV+ZIKV give complete (100x) genomes
Reproduced
Ct25 USUV 99.5 & ZIKV 98.8 complete; WNV 59.7/YFV 70.3 not; all Ct29/33=0
exact
R1a
Reported
Meta-Illumina Ct25/29 complete except WNV
Reproduced
USUV/YFV/ZIKV 99.5-99.9%; WNV exception (Ct29 64.3)
within tolerance
R1d
Reported
>51,000x amplicon depth difference; ~800,000x WNV Ct33
Reproduced
WNV Ct33 Illumina amplicon max 839,833x vs min 0
within tolerance
R1e
Reported
Capture comparable to amplicon Illumina, worse at Ct33
Reproduced
Ct25/29 comparable; Ct33 capture 43/14/24/59 vs amplicon 94/25/95/99.5
within tolerance
R2a
Reported
Amplicon Illumina ~3M reads/sample
Reproduced
mean 3.05M (2.24-5.40M)
within tolerance
R2b
Reported
Amplicon Nanopore ~400k (exceptions 682k,917k)
Reproduced
mean 434k; exceptions 681,556 & 917,426 present
exact
R2c
Reported
Meta-Nanopore total ~5.2M reads
Reproduced
deposited per-barcode sum = 1.49M (5.2M likely run-level/pre-demux, absent from deposit)
did not match
R3
Reported
Dedup removes 88% of mapped reads (meta/capture)
Reproduced
read-weighted 87.4%
within tolerance
R4
Reported
Capture pooling up to 35-fold difference
Reproduced
YFV 35.1x; overall 35.2x
within tolerance
R5
Reported
5 Illumina false-positive substitutions WNV 4509/5034/7088, YFV 822/3711
Reproduced
all 5 recovered as MAJOR variants only in high-Ct Illumina samples
exact

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 84/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟡5. Derivability / plausibility
🟢6. Severity of the deviation
🟢7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +4

Running the authors' own Snakemake pipeline verbatim on the exact deposited data (PRJEB47177, all 60 read sets), 10/11 attempted in-scope claims reproduce exact or within-tolerance, including the primary Fig 3 coverage result and all five reported false-positive variant positions. The sole mismatch (R2c) — reported ~5.2M meta-Nanopore reads vs 1.49M deposited — is a deposit-completeness nuance (the 5.2M is a run-level/pre-demux count not in the SRA deposit), i.e. one value not derivable from shared data, not a wrong recomputed value. The only methodological assumption was reconstructing the Ct→run mapping (ENA anonymised), independently validated three ways. Net: a strong, solid reproduction; deviations are negligible and the central conclusion holds, with criticality kept to yellow only by the non-derivable read-count figure.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

370.6 k
tokens (I/O) · 49.9 M incl. cache
153 min
runtime · 6.53 CPU-h
7.7 GB
peak RAM
2
HPC jobs
hummel
machine