Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Performance of methods for SARS-CoV-2 variant detection and abundance estimation within mixed population samples.

PeerJ · 2023
L1 50/100 3/4
Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
✓ What held up
  • Nothing in this column.
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🟡Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
50/100
Reproducibility score
1.4 SD below mean
vs. all fields · 1173 studies
🎯 Scores higher than 8% of all assessed papers rank 1026 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

REPRODUCED (partial, headline claim). Ran Freyja (third-party tool as embedded in authors' C-WAP pipeline; P16-valid) on the paper's OWN 39 ARTICv4 empirical wastewater runs (PRJNA765612, exact Table S1 list) on «our HPC»: bowtie2->NC_045512.2 -> ivar trim -e -b ARTICv4 -> freyja variants -> freyja demix --confirmedonly -> count non-zero WHO-level 'summarized' groups, mean over 39. Reproduced mean = 1.87 +/- 1.24 (>0) / 1.18 +/- 0.39 (>0.01) vs reported 2.2 +/- 0.45 (Table 7). Central tendency matches (~2 WHO variants/sample; samples are mono-Omicron Feb-2022 wastewater) and the qualitative ranking holds (Freyja most conservative, far below Kallisto's 10.1). Point estimate ~15% low and SD inflated by 3 low-coverage samples (genome cov 1-8% -> spurious 4/5/7 counts) -- all consistent with the documented Freyja UShER-barcode drift (ran 06_29_2026 barcodes vs paper's early-2022 DB; visible as an 'Other' residual bin). No fabrication concern: 2.2 is plausibly derivable from the public data with the period-appropriate DB. NOT attempted: Kallisto/LCS/LINDEC/ALLCB arms (stretch) and the simulated RRMSE/CCC results (simulated reads + random true proportions not deposited -> not 1:1 reproducible).

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 50
    assessed: 2026-06-19 ⛓ dc4285e28e79
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-29
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The paper tests which of several SARS-CoV-2 variant composition estimation methods (VCEs) most accurately identifies variants and estimates their relative abundance within mixed population samples such as wastewater, comparing a newly introduced linear regression method (LINDEC) against Kallisto, Kraken2, LCS, and Freyja.

Core claims
  • Kallisto was the most accurate VCE on simulated data, having the lowest RRMSE, followed by Freyja finding
  • Kallisto and Freyja produced the most similar predictions to each other, reflected by the highest concordance correlation coefficient (CCC) between estimators finding
  • VCE accuracy is dependent on sequencing platform and amplicon panel used finding
  • Freyja's accuracy was significantly higher with Illumina data compared to Oxford Nanopore (ONT) data finding
  • Kallisto's performance was best when used with the ARTICv4 amplicon panel finding
  • On empirical wastewater data there was poor agreement among VCE methods and large variation in the number of variants each method detected finding
  • LINDEC, a linear deconvolution by least-squares VCE, was developed and implemented within the custom C-WAP bioinformatics pipeline method
  • A custom NextFlow-based pipeline (C-WAP, CFSAN Wastewater Analysis Pipeline) was developed for bioinformatic analysis of SC2 wastewater sequence data across Illumina, ONT, and PacBio platforms resource
Experimental setups
Assay System Perturbation Readout Platform
Simulated Illumina sequencing (ART v2.5.8) in silico SARS-CoV-2 mixed-variant sequence data (Alpha, Beta, Delta, Epsilon, Omicron) simulated variant abundance mixtures across amplicon panels (ARTICv4, QIAseq DIRECT, NEB VSS v1a) and read length/fragment parameters VCE-estimated variant relative abundance vs known simulated abundance (RRMSE) Illumina MiSeq (simulated via ART v2.5.8)
Simulated Oxford Nanopore sequencing (DeepSimulator v1.5) in silico SARS-CoV-2 mixed-variant sequence data (same five variants) simulated variant abundance mixtures with NEB VSS v1a amplicon panel VCE-estimated variant relative abundance vs known simulated abundance (RRMSE) Oxford Nanopore (simulated via DeepSimulator v1.5)
Amplicon-based targeted sequencing of empirical wastewater samples publicly available wastewater samples (GenomeTrakr/FDA CFSAN surveillance, NCBI BioProjects PRJNA765612, PRJNA767800, PRJNA757447) none (natural community-level SC2 variant mixtures) VCE-predicted variant composition, number of variants detected, agreement (CCC) between VCE methods Illumina (ARTICv4, NEB VSS v1a, QIAseq DIRECT amplicon panels)
Variant composition estimation (bioinformatic method comparison) simulated and empirical SC2 sequence reads processed via C-WAP pipeline method/tool variation: LINDEC, Kallisto, Kraken2+Bracken (ALLCB), LCS, Freyja relative root mean squared error (RRMSE) and concordance correlation coefficient (CCC) between estimated and known/other-method abundances custom NextFlow pipeline C-WAP with Bowtie2 (Illumina) or Minimap2 (ONT) alignment, iVar trimming
Key results
  • Kallisto had the lowest RRMSE among all VCEs on simulated data
  • Freyja had the second-lowest RRMSE after Kallisto
  • Kallisto and Freyja showed the highest CCC agreement between each other's predictions
  • Freyja accuracy was significantly higher on Illumina data than ONT data
  • Kallisto's best performance occurred with the ARTICv4 amplicon panel
  • On empirical ARTICv4 data, Freyja detected far fewer variants on average than Kallisto Freyja mean 2.2 variants vs Kallisto mean 10.1 variants
  • Simulated mutation frequencies (observed) agreed strongly with expected frequencies from known lineage compositions, validating the simulation approach CCC range 0.8397–0.9487, mean 0.8857 ± 0.0249
Key statistics
  • correlation Lin's CCC range 0.8397–0.9487, mean 0.8857 ± 0.0249 (agreement between observed simulated mutation frequencies and expected frequencies, validating simulations)
  • mean 2.2 variants detected (Freyja ARTICv4 empirical data mean number of variants detected)
  • mean 10.1 variants detected (Kallisto ARTICv4 empirical data mean number of variants detected)
  • count 39 SRRs (empirical ARTICv4 amplicon panel samples analyzed)
  • count 15 SRRs (empirical NEB VSS v1a amplicon panel samples analyzed)
  • count 69 SRRs (empirical QIAseq DIRECT amplicon panel samples analyzed)
  • count 100 simulated datasets (total simulated sequencing datasets generated across platforms/panels for VCE evaluation)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The paper is a computational/simulation-based benchmarking study rather than a classical hypothesis-testing experiment: five SARS-CoV-2 variant composition estimators (VCEs) were compared using 100 in-silico simulated sequencing datasets (Illumina and Oxford Nanopore) of known variant abundance, plus a set of empirical wastewater samples of unknown composition. Performance/accuracy was quantified with relative root mean squared error (RRMSE) against the known simulated abundances, and agreement between estimator pairs (and between simulated vs. expected mutation frequencies) was quantified with Lin's concordance correlation coefficient (CCC). Results are reported primarily as descriptive summary metrics (RRMSE, CCC, mean ± SD, means/counts of detected variants) rather than through classical inferential test statistics with p-values.

Replicationunclear Sample size100 in-silico simulated datasets were generated with randomly assigned variant abundances (not derived from a stated power calculation); empirical data comprised variable numbers of SRA runs per amplicon panel (39 for ARTICv4, 15 for NEB VSS v1a, 69 for QIAseq DIRECT) Groups5 variant composition estimators (LINDEC, Kraken2/Bracken, Kallisto, Freyja, LCS) across 2 sequencing platforms and up to 3 amplicon panels Pairingunclear Randomization/blindingna DispersionSD Confidence intervalsno
Statistical tests used
Test Applied to n Assumptions
Relative root mean squared error (RRMSE) Accuracy of each VCE's abundance estimates vs. known simulated variant abundances, across platforms (Illumina/ONT) and amplicon panels (ARTICv4, NEB VSS v1a, QIAseq DIRECT) 100 simulated datasets (5 platform/panel/parameter combinations x replicates per Table 2) not stated
Lin's concordance correlation coefficient (CCC) (1) Validation of simulations by comparing observed vs. expected mutation frequencies; (2) agreement between pairs of VCE abundance estimates 100 simulations for validation; not explicitly stated for pairwise VCE agreement not stated
Approaches that could also have been used
  • Estimator accuracy against known simulated abundances was quantified using RRMSE alone.
    Could also: Reporting additional error metrics such as mean absolute error (MAE) or an R²/correlation between observed and expected abundances — Different error metrics emphasize different aspects of estimator performance (e.g., MAE is less sensitive to large outlier deviations than RMSE-based metrics), so presenting more than one can give a fuller picture of accuracy.
  • Agreement between pairs of VCEs was summarized with a single Lin's CCC value per comparison.
    Could also: A Bland-Altman style analysis plotting difference vs. mean abundance for each estimator pair — Bland-Altman plots can reveal whether agreement (bias, limits of agreement) varies across the range of estimated abundance values, which a single overall CCC does not show.
  • The abstract states that Freyja's accuracy was 'significantly higher' with Illumina than ONT data, without a formal statistical test being detailed in the available methods text.
    Could also: A paired test on RRMSE across matched simulation replicates (e.g., paired t-test or Wilcoxon signed-rank test) with a reported p-value and effect size — Pairing simulations by their known variant-composition scenario and applying a formal paired test would let readers see both the magnitude and the statistical uncertainty behind the platform comparison.
  • 100 simulated datasets were generated with randomly assigned abundances, and the number of simulations was not linked to a stated power calculation.
    Could also: A pre-specified power/precision analysis, or bootstrap resampling of the simulation results to obtain confidence intervals for RRMSE — This would quantify how precisely the chosen number of simulations estimates the true RRMSE/CCC for each estimator, complementing the point estimates already reported.
  • Comparisons were made across 5 estimators, 2 platforms, and up to 3 amplicon panels without a described multiplicity adjustment.
    Could also: A false-discovery-rate procedure such as Benjamini-Hochberg, applied across the set of estimator/platform/panel comparisons — When many comparisons are examined together, an FDR or family-wise error correction helps readers calibrate confidence in any individual comparison highlighted as notable.
  • Dispersion of Lin's CCC across the 100 simulations was expressed as mean ± SD.
    Could also: A 95% confidence interval for the mean CCC — A CI communicates the precision of the estimated mean agreement directly, which some readers find more directly interpretable than an SD alongside the mean.
Software: Python scikit-learn (linear regression, for the LINDEC estimator) · C-WAP / NextFlow bioinformatics pipeline (Bowtie2, Minimap2, iVar) · Kraken2 / Bracken · Kallisto · Freyja · LCS

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-36721781

Paper: Kayikcioglu et al. 2023, PeerJ 11:e14596. "Performance of methods for SARS-CoV-2 variant detection and abundance estimation within mixed population samples." PMCID PMC9884472 · DOI 10.7717/peerj.14596 · License CC0.

What the paper does

Benchmarks 5 variant-composition estimators (VCEs) for estimating the relative abundance of SARS-CoV-2 (SC2) lineages in mixed samples:

  • Freyja (UShER-tree SNPs + weighted least-absolute-deviation)
  • Kallisto (pseudo-alignment, RNA-seq tool repurposed)
  • ALLCB = Kraken2 + Bracken
  • LINDEC = authors' linear deconvolution over a mutation×lineage logical matrix (mutation list pre-compiled from the constellations repo — the brief's "code" link)
  • LCS = max-likelihood mixture model

Two data regimes:

  1. Simulated (300 datasets: 200 Illumina + 100 ONT) — 5 variants (Alpha/Beta/Delta/ Epsilon/Omicron) mixed at random abundances; in-silico PCR + ART (Illumina) / DeepSimulator (ONT); 3 amplicon panels (ARTICv4, QIAseq DIRECT, NEB VSS v1a) × read/fragment lengths. Metric: RRMSE vs known abundance (Tables 5,6), CCC between estimators & truth (Table S2 / Fig 2).
  2. Empirical (123 wastewater samples, Illumina): ARTICv4=39, NEB VSS v1a=15, QIAseq DIRECT=69. No ground truth → metrics are number of variants detected per method (Table 7) and pairwise CCC between methods (Fig 3).

Pipeline = C-WAP (CFSAN Wastewater Analysis Pipeline, NextFlow). Lineage outputs collapsed by a custom Python script to WHO-defined level (e.g. B.1.351.1+B.1.351.3 → B.1.351). "All detected mutations included … without explicit depth/frequency cutoff, as long as iVar indicates statistical significance."

IN SCOPE (pipeline-derived, attempted)

P16-valid: running the third-party tool Freyja (exactly as embedded in the authors' own C-WAP pipeline) on the paper's own empirical data.

  • R1 (primary): average number of variants detected by Freyja on ARTICv4 empirical wastewater = 2.2 ± 0.45 (Table 7; abstract: "Freyja ARTICv4 had a mean of 2.2 variants"). Data = 39 ARTICv4 runs (PRJNA765612), exact run list = Table S1 (supp-3). Pipeline: bowtie2→NC_045512.2 → samtools sort → ivar trim -e -b ARTICv4freyja variantsfreyja demix --confirmedonly → count non-zero WHO-level variants per sample → mean over 39.
  • R2 (stretch): average number of variants detected by Kallisto on ARTICv4 = 10.1 ± 2.43 (Table 7) — needs C-WAP's custom kallisto lineage index.
  • R3 (stretch): empirical pairwise CCC Freyja↔Kallisto behaviour (Fig 3) if R1+R2 succeed.

OUT OF SCOPE / not deterministically reproducible

  • Simulated RRMSE (Tables 5,6) & simulated CCC (Table S2 / Fig 2): the simulated reads and the random true mixing proportions were NOT deposited (only summary PDFs + the CCC table shipped). ww_simulations can regenerate new random mixtures but not the paper's exact ones → cannot be a 1:1 reproduction. Documented, not attempted as 1:1.
  • 30 GISAID reference genomes (EPI_ISL_*) used to build the simulations — GISAID is registration-gated; not needed for the empirical reproduction.
  • LINDEC/ALLCB/LCS exact configs — secondary; Freyja+Kallisto are the headline methods.

Reproducibility surface summary

Empirical SRA data = fully public (123/123 runs resolve, exact list shipped). Pipeline = public (C-WAP) wrapping public tools. Main fidelity risk: Freyja barcode/UShER DB drift since early-2022 (today's DB has thousands of post-2022 lineages) → "number of variants detected" may differ; reported honestly as a caveat.

Figures / tables: TableFig 2
R1
Reported
Freyja ARTICv4 mean WHO-level variants detected = 2.2 +/- 0.45 (Table 7, n=39)
Reproduced
1.87 +/- 1.24 (count>0, incl 'Other' bin); 1.18 +/- 0.39 (count>0.01); 39/39 samples ran
partial
R2
Reported
Kallisto ARTICv4 = 10.1 +/- 2.43
Reproduced
not attempted (stretch; needs C-WAP custom kallisto index)
partial
C1/C2
Reported
simulated RRMSE/CCC (Tables 5,6,S2)
Reproduced
not 1:1 reproducible (simulated reads + random true proportions not deposited)
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 50/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟡2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

142.7 k
tokens (I/O) · 8.6 M incl. cache
16 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.