Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Wochenende - modular and flexible alignment-based shotgun metagenome analysis.

BMC Genomics · 2022
L1 88/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +4
✓ What held up
  • Same input data as the authors
  • Reported values are derivable from the shared data
  • The central claim held under reproduction
What did not (or only partly)
  • 🟡Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡The deviation was non-trivial in magnitude
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
88/100
Reproducibility score
0.8 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 74% of all assessed papers rank 276 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Reproduced the Wochenende short-read mock-community result (Fig 2) by running the authors' own pipeline (nf_wochenende @372bbe2, run_Wochenende.py bwamem+mq30+PE+mismatches5 defaults + basic_reporting.py) on the paper's own public data (SRR11207337, 12,527,638 PE pairs, md5-verified) and the authors' combined 2021_12 human+bact+arch+fungi+vir reference (md5-verified, bwa-indexed from scratch) on «our HPC». P16-favourable case. 3 of 4 in-scope claims reproduce CLEARLY: C1 7/8 bacteria detected (exact); C3 B. subtilis -> B. intestinalis cross-assignment (exact: 855 vs 2,438,682 reads); C4 fungi underrepresented vs 2% with C. neoformans hardly detected (exact: 1.33% / 0.63%). C2 (count within 12%+/-1.8%) is PARTIAL: we get 2/8 (read-fraction) to 3/8 (length-normalized) vs the reported 4/8 - the exact count brackets the paper but is sensitive to the unspecified normalization metric; the qualitative abundance pattern (incl. strong Pseudomonas over-representation) reproduces. No fabrication indicators - all reported findings are derivable from the shipped data+code. Memory 23.8 GB ~ paper's 23.4 GB. NOT attempted: C5 competitor tools (stretch), C6 runtime (hardware-bound), clinical Fig 3/4 (patient data not deposited = out of scope).

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-30
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-30
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The paper asks whether a modular, transparent alignment-based shotgun metagenomics pipeline (Wochenende), combined with a novel host-genome-based normalization, can reduce false-positive taxonomic assignments and enable absolute (rather than only relative) microbial abundance quantification compared to existing k-mer/classifier-based tools.

Core claims
  • Wochenende is a modular, transparent alignment-based pipeline for shotgun metagenome analysis supporting short and long reads across all kingdoms of life method
  • Stringent filtering of mismatches and mapping quality sharply reduces the number of false positive taxonomic assignments finding
  • A novel normalization (Bphc, bacterial cells per human cell) enables calculation of absolute abundance profiles by comparing microbial reads to human host reads method
  • Blacklister masks contaminant sequences (e.g., Illumina/Nanopore adapters) present in reference genomes, improving alignment accuracy resource
  • Genomic coverage plots and the integrated tool raspir help distinguish true rare species detections from false positives due to mismapping mechanism
  • Bacterial growth rates can be inferred from peak-to-trough ratios of mapped read coverage, implemented in Python3 within the pipeline method
  • Pipeline transparency allowed identification of two mislabeled E. coli genomes (incorrectly annotated as Pseudomonas aeruginosa) that masked true P. aeruginosa reads finding
  • Wochenende was benchmarked against KrakenUniq, Kaiju, MetaPhlAn3, and Centrifuge on a Zymo mock community dataset method
Experimental setups
Assay System Perturbation Readout Platform
shotgun metagenomic sequencing (short-read) Zymo mock community (8 bacteria, 2 yeasts) none taxonomic classification / abundance profile
shotgun metagenomic sequencing (long-read) Zymo mock community (same composition) none taxonomic classification / abundance profile Oxford Nanopore GridION
comparative metagenomic classifier benchmarking Zymo mock community single-end reads tool comparison (Wochenende vs KrakenUniq, Kaiju, MetaPhlAn3, Centrifuge) taxonomic assignment accuracy, processing time, RAM usage
reference genome contaminant masking (Blacklister) bacterial, fungal, viral, and human reference genomes none masked (N-substituted) contaminant/adapter bases in reference genomes Bowtie2, Bedtools
simulated read recall testing reference genomes (in silico reads) none read recall/uniqueness of alignment per genome InsilicoSeq (Illumina NovaSeq 2x150bp profile, trimmed to 75bp single-end)
near-duplicate/mislabeled genome detection bacterial reference genomes (e.g., E. coli, P. aeruginosa) none genome-to-genome similarity used to flag masking of reads between species fastANI
Key results
  • Stringent mismatch and mapping quality (MQ30) filtering sharply reduced false positive taxonomic assignments
  • Two mislabeled E. coli genomes (annotated as P. aeruginosa) completely masked true P. aeruginosa genomes, causing no reads to be attributed to P. aeruginosa in mock communities
  • Blacklister masking of contaminant adapter regions in reference genomes markedly improved quality of Wochenende results (data not shown)
  • KrakenUniq required the most RAM of the compared tools due to its large index
  • Kaiju could only assign reads at genus level, unlike other tools which achieved species-level resolution, so it was excluded from the main species-level comparison
  • Processing times of all compared metagenomic tools were short less than one hour
Key statistics
  • count 8 bacteria at 12% each, 2 yeasts at 2% each (theoretical genomic DNA composition of the Zymo mock community)
  • other <1 hour (processing time of all compared metagenomic tools on the mock community dataset)
  • other 75 bp (simulated single-end read length used for reference genome recall testing)
  • count 2 (number of mislabeled E. coli genomes found masking P. aeruginosa reads)
  • other 6139 Mb (diploid human genome size (i) used in the Bphc absolute abundance formula)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This is a bioinformatics pipeline/tool description paper (Wochenende, an alignment-based shotgun metagenome pipeline) rather than a hypothesis-testing study. Evaluation is based on benchmarking the pipeline against other established metagenome classifiers (KrakenUniq, Kaiju, MetaPhlAn3, Centrifuge) using mock community datasets (Zymo mock community, short- and long-read) with a known theoretical taxonomic composition. Results are reported primarily through descriptive comparison (e.g., figures/tables of detected taxa, processing time, RAM use) rather than formal inferential statistical tests, based on the text provided.

Replicationunclear GroupsWochenende pipeline vs. other metagenome classification tools (KrakenUniq, Kaiju, MetaPhlAn3, Centrifuge) on mock community sequencing datasets Pairingna Randomization/blindingna Dispersionnone
Approaches that could also have been used
  • Wochenende's performance was benchmarked against other classifiers using a single mock community dataset (one short-read, one long-read run) with results presented descriptively.
    Could also: Using multiple replicate sequencing runs of the mock community and summarizing performance metrics (e.g., recall/precision) with a measure of variability such as SD or a bootstrap confidence interval — This would let readers distinguish consistent tool-level differences from run-to-run variability inherent to a single sequencing experiment.
  • The internal module raspir computes correlation coefficients and p-values per genome to help distinguish rare-but-present taxa from false positives.
    Could also: Applying a multiple-testing correction (e.g., Benjamini-Hochberg FDR) across the set of p-values generated for all candidate genomes in a sample — When many taxa are tested simultaneously, an FDR or similar correction controls the expected proportion of false discoveries across the full set of comparisons.
  • Absolute abundance normalization (Bphc, RPMM) is presented as a point-estimate formula applied per sample without an accompanying uncertainty estimate.
    Could also: Reporting confidence intervals or variance estimates for the normalized abundance values, e.g., via replicate library preparations or resampling of reads — This would convey how much sampling or library-prep variability contributes to a given Bphc or RPMM value, in addition to the central estimate.
  • Tool comparisons (Wochenende vs. KrakenUniq, MetaPhlAn3, Centrifuge) rely on default parameters and a fixed reference set without a formal significance test for differences in detection accuracy.
    Could also: A paired statistical comparison (e.g., McNemar's test or a paired permutation test on per-taxon detection/abundance across samples) between tools — A paired test would formally quantify whether one tool's detection accuracy differs from another's beyond what might be expected by chance on this dataset.
  • Genome masking/quality control decisions (e.g., excluding mislabeled genomes via FastANI comparisons) are described procedurally rather than with a quantitative similarity threshold and associated statistic.
    Could also: Reporting a defined sequence-identity/coverage threshold with, e.g., a distribution or histogram of pairwise ANI values used to justify the cutoff — This would make the genome-exclusion criterion more quantitatively reproducible for other users building their own reference databases.
Software: Python3 · Python Matplotlib 3.2.1 · Python pandas 1.1.1 · R (base heatmap function / heatmaply) · R package metacoder · InsilicoSeq (read simulation)

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-36368923 (Wochenende metagenome pipeline)

Paper: Rosenboom et al. 2022, Wochenende — modular and flexible alignment-based shotgun metagenome analysis. BMC Genomics 23:760. DOI 10.1186/s12864-022-08985-9. PMID 36368923 · PMCID PMC9650795.

Tool paper. Wochenende is an alignment-based shotgun-metagenome pipeline (bwa-mem / minimap2 alignment vs a combined human+microbial reference, duplicate removal, mismatch filtering, MQ filtering, RPMM/bphc normalization, reporting). This is a P16 case in the favourable direction: the code IS the authors' own pipeline, and the headline benchmark is the tool applied to a public mock-community dataset. Reproducing = run the tool on the paper's data with the described params.

Pipeline-derived results (IN SCOPE)

id result paper loc pipeline feasibility
C1 Wochenende detects 7/8 mock bacterial species in SRR11207337 (Zymo mock) Fig 2 / Results "Mock community" Wochenende bwamem+mq30 on SRR11207337 vs 2021_12 ref HIGH — ref + data + pinned env all public
C2 4/8 bacteria within manufacturer tolerance (12% ±1.8% genomic abundance) Results / Fig 2 same HIGH
C3 B. subtilis underrepresented and reads assigned to B. intestinalis (all tools) Results / Fig 2 same — needs full ref containing B. intestinalis CP011051 MEDIUM-HIGH
C4 Fungi (S. cerevisiae, C. neoformans) strongly underrepresented vs the 2% input Results / Fig 2 same MEDIUM
C5 Benchmark vs MetaPhlAn3 (2/8), Centrifuge (1/8), KrakenUniq (0/8) accurate Results / Fig 2 each competitor tool on same data LOWER — separate heavy installs; stretch goal
C6 Runtime 15+4 min, memory 23.4 GB (Wochenende) Results wall-clock/RSS of run NOT reproducible 1:1 — hardware-dependent; report qualitatively only

Primary minimum target: C1–C4 (the Wochenende mock-community abundance profile, Figure 2, short-read panel). C5 is a stretch goal. C6 is hardware-bound → not a 1:1 claim.

OUT OF SCOPE

  • Fig 3 (cystic-fibrosis vs healthy infant respiratory) and Fig 4 (ICU respiratory, 26 h turnaround): clinical patient data, not deposited as a public accession in the paper → cannot obtain → out of scope (data_restricted/not-deposited).
  • Long-read mock (ENA ERR3152364, Suppl Fig S2): secondary; attempt only if short-read reproduction succeeds and time allows.
  • Growth-rate, raspir, heat-tree/heatmap visual outputs: downstream illustrative outputs, not quantitative claims with reported numbers → not graded.

Key resources (pinned)

  • Code: github.com/MHH-RCUG/nf_wochenende (HEAD main 372bbe2, 2026-03-04) + run_Wochenende.py
  • Env: env.wochenende.minimal.yml (python3.7, bwa0.7.17, samtools1.11, bamtools2.5.1, nextflow22.04)
  • Reference: 2021_12_human_bact_arch_fungi_vir.fa.gz (3.58 GB, GDrive folder 1q1btJCxtU15XXqfA-iCyNwgKgQq0SrG4, md5sum file provided) — must build bwa index
  • Data: SRR11207337 (ENA, PE, 12,527,638 read pairs, synthetic metagenome WGS)
  • Default params (nextflow.config): aligner=bwamem, mq30, PE, mismatches filter, no_prinseq, no_fastqc, abra=false
Figures / tables: Fig 2
C1
Reported
7/8 mock bacteria detected by Wochenende (Fig 2)
Reproduced
7/8 detected; B. subtilis absent (855 reads) as its reads map to B. intestinalis
exact
C2
Reported
4/8 bacteria within 12% +/-1.8% tolerance (Fig 2)
Reproduced
2/8 (read-fraction) to 3/8 (length-normalized) within 10.2-13.8%; brackets but != reported 4/8, normalization-sensitive; Pseudomonas over-represented ~27%
partial
C3
Reported
B. subtilis reads assigned to B. intestinalis (Fig 2)
Reproduced
B. subtilis 855 reads vs B. intestinalis 2,438,682 reads - cross-assignment reproduced exactly
exact
C4
Reported
Fungi (S. cerevisiae, C. neoformans) strongly underrepresented vs 2% input (Fig 2)
Reproduced
S. cerevisiae 248,786 reads (1.33%), C. neoformans 118,145 reads (0.63%) - both <2%, C. neoformans hardly detected
exact
C5
Reported
Wochenende 4/8 vs MetaPhlAn3 2/8, Centrifuge 1/8, KrakenUniq 0/8 (Fig 2)
Reproduced
not attempted (stretch goal; competitor tools not installed)
m.public.grade.not-attempted
C6
Reported
~15+4 min, 23.4 GB memory
Reproduced
33:40 wall, MaxRSS 23.8 GB (16-cpu node); hardware-dependent, not 1:1 but memory strikingly close
m.public.grade.not-1to1

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 88/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟡2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟢5. Derivability / plausibility
🟡6. Severity of the deviation
🟢7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +4

Running the authors' own Wochenende pipeline (nf_wochenende @372bbe2) on the paper's own public data (SRR11207337, md5-verified) and reference reproduces 3 of 4 in-scope claims exactly: 7/8 bacteria detected (C1), the signature B. subtilis→B. intestinalis cross-assignment (855 vs 2,438,682 reads, C3), and fungal under-representation vs 2% input (C4), with memory 23.8 GB ≈ the reported 23.4 GB. The only deviation is C2 — our within-tolerance count is 2–3/8 vs the reported 4/8 — and it brackets the paper, flipping purely on the unspecified normalization metric (read-fraction vs length-normalized), so the gap is on the methodology/paper-underspecification side, not a data or fabrication defect. The central conclusion holds fully; overall a solid reproduction with one explainable, normalization-sensitive deviation.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

62.8 k
tokens (I/O) · 3.5 M incl. cache
8 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.