Wochenende - modular and flexible alignment-based shotgun metagenome analysis.
The main results reproduced: recomputed values matched the published ones within tolerance.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Same input data as the authors
- ✓Reported values are derivable from the shared data
- ✓The central claim held under reproduction
- 🟡Reported values were only indirectly comparable
- 🟡A deviation arose in the data or preprocessing
- 🟡A deviation was attributed to the published material
- 🟡The deviation was non-trivial in magnitude
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Reproduced the Wochenende short-read mock-community result (Fig 2) by running the authors' own pipeline (nf_wochenende @372bbe2, run_Wochenende.py bwamem+mq30+PE+mismatches5 defaults + basic_reporting.py) on the paper's own public data (SRR11207337, 12,527,638 PE pairs, md5-verified) and the authors' combined 2021_12 human+bact+arch+fungi+vir reference (md5-verified, bwa-indexed from scratch) on «our HPC». P16-favourable case. 3 of 4 in-scope claims reproduce CLEARLY: C1 7/8 bacteria detected (exact); C3 B. subtilis -> B. intestinalis cross-assignment (exact: 855 vs 2,438,682 reads); C4 fungi underrepresented vs 2% with C. neoformans hardly detected (exact: 1.33% / 0.63%). C2 (count within 12%+/-1.8%) is PARTIAL: we get 2/8 (read-fraction) to 3/8 (length-normalized) vs the reported 4/8 - the exact count brackets the paper but is sensitive to the unspecified normalization metric; the qualitative abundance pattern (incl. strong Pseudomonas over-representation) reproduces. No fabrication indicators - all reported findings are derivable from the shipped data+code. Memory 23.8 GB ~ paper's 23.4 GB. NOT attempted: C5 competitor tools (stretch), C6 runtime (hardware-bound), clinical Fig 3/4 (patient data not deposited = out of scope).
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-30
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-30no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetThe paper asks whether a modular, transparent alignment-based shotgun metagenomics pipeline (Wochenende), combined with a novel host-genome-based normalization, can reduce false-positive taxonomic assignments and enable absolute (rather than only relative) microbial abundance quantification compared to existing k-mer/classifier-based tools.
- ★ Wochenende is a modular, transparent alignment-based pipeline for shotgun metagenome analysis supporting short and long reads across all kingdoms of life method
- ★ Stringent filtering of mismatches and mapping quality sharply reduces the number of false positive taxonomic assignments finding
- ★ A novel normalization (Bphc, bacterial cells per human cell) enables calculation of absolute abundance profiles by comparing microbial reads to human host reads method
- ★ Blacklister masks contaminant sequences (e.g., Illumina/Nanopore adapters) present in reference genomes, improving alignment accuracy resource
- ★ Genomic coverage plots and the integrated tool raspir help distinguish true rare species detections from false positives due to mismapping mechanism
- Bacterial growth rates can be inferred from peak-to-trough ratios of mapped read coverage, implemented in Python3 within the pipeline method
- ★ Pipeline transparency allowed identification of two mislabeled E. coli genomes (incorrectly annotated as Pseudomonas aeruginosa) that masked true P. aeruginosa reads finding
- Wochenende was benchmarked against KrakenUniq, Kaiju, MetaPhlAn3, and Centrifuge on a Zymo mock community dataset method
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| shotgun metagenomic sequencing (short-read) | Zymo mock community (8 bacteria, 2 yeasts) | none | taxonomic classification / abundance profile | — |
| shotgun metagenomic sequencing (long-read) | Zymo mock community (same composition) | none | taxonomic classification / abundance profile | Oxford Nanopore GridION |
| comparative metagenomic classifier benchmarking | Zymo mock community single-end reads | tool comparison (Wochenende vs KrakenUniq, Kaiju, MetaPhlAn3, Centrifuge) | taxonomic assignment accuracy, processing time, RAM usage | — |
| reference genome contaminant masking (Blacklister) | bacterial, fungal, viral, and human reference genomes | none | masked (N-substituted) contaminant/adapter bases in reference genomes | Bowtie2, Bedtools |
| simulated read recall testing | reference genomes (in silico reads) | none | read recall/uniqueness of alignment per genome | InsilicoSeq (Illumina NovaSeq 2x150bp profile, trimmed to 75bp single-end) |
| near-duplicate/mislabeled genome detection | bacterial reference genomes (e.g., E. coli, P. aeruginosa) | none | genome-to-genome similarity used to flag masking of reads between species | fastANI |
- ▼ Stringent mismatch and mapping quality (MQ30) filtering sharply reduced false positive taxonomic assignments
- ▼ Two mislabeled E. coli genomes (annotated as P. aeruginosa) completely masked true P. aeruginosa genomes, causing no reads to be attributed to P. aeruginosa in mock communities
- ▲ Blacklister masking of contaminant adapter regions in reference genomes markedly improved quality of Wochenende results (data not shown)
- ▲ KrakenUniq required the most RAM of the compared tools due to its large index
- – Kaiju could only assign reads at genus level, unlike other tools which achieved species-level resolution, so it was excluded from the main species-level comparison
- – Processing times of all compared metagenomic tools were short less than one hour
- count 8 bacteria at 12% each, 2 yeasts at 2% each (theoretical genomic DNA composition of the Zymo mock community)
- other <1 hour (processing time of all compared metagenomic tools on the mock community dataset)
- other 75 bp (simulated single-end read length used for reference genome recall testing)
- count 2 (number of mislabeled E. coli genomes found masking P. aeruginosa reads)
- other 6139 Mb (diploid human genome size (i) used in the Bphc absolute abundance formula)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This is a bioinformatics pipeline/tool description paper (Wochenende, an alignment-based shotgun metagenome pipeline) rather than a hypothesis-testing study. Evaluation is based on benchmarking the pipeline against other established metagenome classifiers (KrakenUniq, Kaiju, MetaPhlAn3, Centrifuge) using mock community datasets (Zymo mock community, short- and long-read) with a known theoretical taxonomic composition. Results are reported primarily through descriptive comparison (e.g., figures/tables of detected taxa, processing time, RAM use) rather than formal inferential statistical tests, based on the text provided.
-
Wochenende's performance was benchmarked against other classifiers using a single mock community dataset (one short-read, one long-read run) with results presented descriptively.↳ Could also: Using multiple replicate sequencing runs of the mock community and summarizing performance metrics (e.g., recall/precision) with a measure of variability such as SD or a bootstrap confidence interval — This would let readers distinguish consistent tool-level differences from run-to-run variability inherent to a single sequencing experiment.
-
The internal module raspir computes correlation coefficients and p-values per genome to help distinguish rare-but-present taxa from false positives.↳ Could also: Applying a multiple-testing correction (e.g., Benjamini-Hochberg FDR) across the set of p-values generated for all candidate genomes in a sample — When many taxa are tested simultaneously, an FDR or similar correction controls the expected proportion of false discoveries across the full set of comparisons.
-
Absolute abundance normalization (Bphc, RPMM) is presented as a point-estimate formula applied per sample without an accompanying uncertainty estimate.↳ Could also: Reporting confidence intervals or variance estimates for the normalized abundance values, e.g., via replicate library preparations or resampling of reads — This would convey how much sampling or library-prep variability contributes to a given Bphc or RPMM value, in addition to the central estimate.
-
Tool comparisons (Wochenende vs. KrakenUniq, MetaPhlAn3, Centrifuge) rely on default parameters and a fixed reference set without a formal significance test for differences in detection accuracy.↳ Could also: A paired statistical comparison (e.g., McNemar's test or a paired permutation test on per-taxon detection/abundance across samples) between tools — A paired test would formally quantify whether one tool's detection accuracy differs from another's beyond what might be expected by chance on this dataset.
-
Genome masking/quality control decisions (e.g., excluding mislabeled genomes via FastANI comparisons) are described procedurally rather than with a quantitative similarity threshold and associated statistic.↳ Could also: Reporting a defined sequence-identity/coverage threshold with, e.g., a distribution or histogram of pairwise ANI values used to justify the cutoff — This would make the genome-exclusion criterion more quantitatively reproducible for other users building their own reference databases.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-36368923 (Wochenende metagenome pipeline)
Paper: Rosenboom et al. 2022, Wochenende — modular and flexible alignment-based shotgun metagenome analysis. BMC Genomics 23:760. DOI 10.1186/s12864-022-08985-9. PMID 36368923 · PMCID PMC9650795.
Tool paper. Wochenende is an alignment-based shotgun-metagenome pipeline (bwa-mem / minimap2 alignment vs a combined human+microbial reference, duplicate removal, mismatch filtering, MQ filtering, RPMM/bphc normalization, reporting). This is a P16 case in the favourable direction: the code IS the authors' own pipeline, and the headline benchmark is the tool applied to a public mock-community dataset. Reproducing = run the tool on the paper's data with the described params.
Pipeline-derived results (IN SCOPE)
| id | result | paper loc | pipeline | feasibility |
|---|---|---|---|---|
| C1 | Wochenende detects 7/8 mock bacterial species in SRR11207337 (Zymo mock) | Fig 2 / Results "Mock community" | Wochenende bwamem+mq30 on SRR11207337 vs 2021_12 ref | HIGH — ref + data + pinned env all public |
| C2 | 4/8 bacteria within manufacturer tolerance (12% ±1.8% genomic abundance) | Results / Fig 2 | same | HIGH |
| C3 | B. subtilis underrepresented and reads assigned to B. intestinalis (all tools) | Results / Fig 2 | same — needs full ref containing B. intestinalis CP011051 | MEDIUM-HIGH |
| C4 | Fungi (S. cerevisiae, C. neoformans) strongly underrepresented vs the 2% input | Results / Fig 2 | same | MEDIUM |
| C5 | Benchmark vs MetaPhlAn3 (2/8), Centrifuge (1/8), KrakenUniq (0/8) accurate | Results / Fig 2 | each competitor tool on same data | LOWER — separate heavy installs; stretch goal |
| C6 | Runtime 15+4 min, memory 23.4 GB (Wochenende) | Results | wall-clock/RSS of run | NOT reproducible 1:1 — hardware-dependent; report qualitatively only |
Primary minimum target: C1–C4 (the Wochenende mock-community abundance profile, Figure 2, short-read panel). C5 is a stretch goal. C6 is hardware-bound → not a 1:1 claim.
OUT OF SCOPE
- Fig 3 (cystic-fibrosis vs healthy infant respiratory) and Fig 4 (ICU respiratory, 26 h turnaround): clinical patient data, not deposited as a public accession in the paper → cannot obtain → out of scope (data_restricted/not-deposited).
- Long-read mock (ENA ERR3152364, Suppl Fig S2): secondary; attempt only if short-read reproduction succeeds and time allows.
- Growth-rate, raspir, heat-tree/heatmap visual outputs: downstream illustrative outputs, not quantitative claims with reported numbers → not graded.
Key resources (pinned)
- Code: github.com/MHH-RCUG/nf_wochenende (HEAD main 372bbe2, 2026-03-04) + run_Wochenende.py
- Env: env.wochenende.minimal.yml (python3.7, bwa0.7.17, samtools1.11, bamtools2.5.1, nextflow22.04)
- Reference: 2021_12_human_bact_arch_fungi_vir.fa.gz (3.58 GB, GDrive folder 1q1btJCxtU15XXqfA-iCyNwgKgQq0SrG4, md5sum file provided) — must build bwa index
- Data: SRR11207337 (ENA, PE, 12,527,638 read pairs, synthetic metagenome WGS)
- Default params (nextflow.config): aligner=bwamem, mq30, PE, mismatches filter, no_prinseq, no_fastqc, abra=false
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
Running the authors' own Wochenende pipeline (nf_wochenende @372bbe2) on the paper's own public data (SRR11207337, md5-verified) and reference reproduces 3 of 4 in-scope claims exactly: 7/8 bacteria detected (C1), the signature B. subtilis→B. intestinalis cross-assignment (855 vs 2,438,682 reads, C3), and fungal under-representation vs 2% input (C4), with memory 23.8 GB ≈ the reported 23.4 GB. The only deviation is C2 — our within-tolerance count is 2–3/8 vs the reported 4/8 — and it brackets the paper, flipping purely on the unspecified normalization metric (read-fraction vs length-normalized), so the gap is on the methodology/paper-underspecification side, not a data or fabrication defect. The central conclusion holds fully; overall a solid reproduction with one explainable, normalization-sensitive deviation.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.