SnakeMAGs: a simple, efficient, flexible and scalable workflow to reconstruct prokaryotic genomes from metagenomes.
The main results reproduced, with only marginal, non-material deviations.
- Nothing in this column.
- 🟡Could not use the authors’ exact input data
- 🟡Reported values were only indirectly comparable
- 🟡A deviation arose in the data or preprocessing
- 🟡A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The deviation was non-trivial in magnitude
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
SnakeMAGs (authors' own Snakemake workflow) reproduced faithfully via verbatim per-rule commands+params on «our HPC». R-test (functional floor): pipeline runs end-to-end on bundled mock -> 1 high-quality MAG = EXACT. R-srr-nmags (headline per-sample): from SRR10402454 (85.7M read pairs) MEGAHIT(min1000,k21-119)->MetaBAT2(minContig2500,seed19860615)->CheckM recovered 11 MAGs at >=50%/<10% vs 7 deposited (Zenodo 7661004) = PARTIAL (over-recovery; same major lineages incl. an identical-stat Spirochaetota MAG 81.15/1.20; +4 explained by tool-version drift + authors' deposited set being post-GUNC). GTDB-Tk r207 taxonomy (R-srr-taxa) running. Aggregate 65/46/59 over all 10 samples NOT attempted (out of single-RU data scope). Cleared 5 infra blockers (corrupt conda cache, iu/py3.14, CheckM DATA_CONFIG, numpy2 ast, fasterq 64-thread deadlock).
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 50assessed: 2026-06-19 ⛓ 37ef6c55ea1c
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-30
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19no human curator yet
- Last updated
- 2026-07-04
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetCan a simplified, flexible Snakemake-based workflow (SnakeMAGs) that integrates a minimal set of state-of-the-art tools reconstruct metagenome-assembled genomes (MAGs) from Illumina metagenomic reads as effectively as (or better than) existing, more complex workflows like ATLAS?
- ★ SnakeMAGs is a simple, efficient, flexible and scalable Snakemake workflow that processes Illumina reads from raw data to MAG classification and relative abundance estimation method
- ★ Compared to ATLAS, SnakeMAGs recovers more MAGs encompassing more diverse bacterial phyla from termite gut metagenomes finding
- ★ SnakeMAGs is slower than ATLAS but has similar memory usage finding
- ★ MAGs recovered only by SnakeMAGs do not significantly differ from MAGs recovered by both workflows in completeness, contamination, genome size or relative abundance finding
- ★ The advantage of SnakeMAGs in MAG yield and phylum diversity over ATLAS is robust to different MAG quality criteria (CheckM alone, quality score threshold, or CheckM+GUNC) finding
- SnakeMAGs integrates illumina-utils, Trimmomatic, Bowtie2 (optional), MEGAHIT, MetaBAT2, CheckM, GUNC (optional), GTDB-Tk and CoverM into 17 sequential Snakemake rules method
- SnakeMAGs source code, test files and tutorial are freely available on GitHub and archived on Zenodo under CeCILL v2.1 license resource
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Workflow performance comparison (CPU time, memory usage, MAG count, phylum diversity) | 10 publicly available termite gut metagenomes (10 termite species) | workflow choice: SnakeMAGs v1.1.0 vs ATLAS v2.9.1 | CPU time, memory usage, number/quality/diversity of reconstructed MAGs | Slurm cluster, Intel Xeon CPU E7-8890 v4 (96 cores/192 threads), 512GB RAM |
| Genome bin quality assessment | Bins generated from termite gut metagenomes | none | completeness, contamination, chimerism | CheckM v1.1.3, GUNC v1.0.5 |
| MEGAHIT metagenomic assembly | Quality-filtered, host-depleted Illumina reads from termite gut metagenomes | none | contigs/scaffolds | MEGAHIT v1.2.9 |
| Metagenomic binning | Assembled contigs from termite gut metagenomes | none | genome bins | MetaBAT2 v2.15 |
| Taxonomic classification of MAGs | Quality-filtered MAGs | none | taxonomic assignment (phylum-level and below) | GTDB-Tk v2.1.0 |
| Relative abundance estimation | MAGs mapped against metagenomic reads | none | relative abundance of MAGs | CoverM v0.6.1 |
| Read quality control and adapter trimming | Raw Illumina reads from termite gut metagenomes | none | quality-filtered, adapter-trimmed reads | illumina-utils v2.12, Trimmomatic v0.39 |
| Host sequence removal (read mapping) | Termite gut metagenomic reads vs termite host genome | host genome as reference for subtraction | host-depleted reads | Bowtie2 v2.4.5 |
- ▲ SnakeMAGs produced 65 MAGs total from 10 metagenomes vs 37 MAGs for ATLAS ~1.76-fold
- ▲ SnakeMAGs recovered MAGs from 15 bacterial phyla vs 11 phyla for ATLAS; ATLAS uniquely recovered only Patescibacteria (1 MAG) while SnakeMAGs uniquely recovered Verrucomicrobiota, Planctomycetota, Synergistota, Elusimicrobiota and Acidobacteriota
- ▼ ATLAS was faster than SnakeMAGs at reconstructing MAGs Wilcoxon P=0.002
- – Memory usage was similar between the two workflows Wilcoxon P=0.393
- – No significant difference between workflows in MAG completeness, contamination or genome size P=0.15 (completeness), P=0.60 (contamination), P=0.64 (genome size)
- – MAGs recovered by SnakeMAGs only did not significantly differ from MAGs recovered by both workflows in completeness, contamination, relative abundance or genome size P=0.19 (completeness), P=0.43 (contamination), P=0.51 (relative abundance), P=0.19 (genome size)
- ▲ Using an estimated quality threshold ≥50, SnakeMAGs still recovered more MAGs and phyla than ATLAS 46 vs 31 MAGs; 13 vs 10 phyla
- ▲ Using GUNC combined with CheckM, SnakeMAGs still produced more and more diverse MAGs than ATLAS 59 MAGs/13 phyla vs 29 MAGs/9 phyla
- pvalue P=0.002 (Wilcoxon test, CPU time: ATLAS faster than SnakeMAGs)
- pvalue P=0.393 (Wilcoxon test, memory usage comparison between workflows)
- count 65 MAGs (SnakeMAGs) vs 37 MAGs (ATLAS) (Total MAGs reconstructed from 10 termite gut metagenomes)
- count 15 phyla (SnakeMAGs) vs 11 phyla (ATLAS) (Bacterial phylum diversity of recovered MAGs)
- pvalue P=0.15 completeness; P=0.60 contamination; P=0.64 genome size (Wilcoxon tests, MAG quality/size comparison between workflows)
- pvalue P=0.19 completeness; P=0.43 contamination; P=0.51 relative abundance; P=0.19 genome size (Wilcoxon tests comparing SnakeMAGs-only MAGs vs MAGs recovered by both workflows)
- count 46 MAGs/13 phyla (SnakeMAGs) vs 31 MAGs/10 phyla (ATLAS) (Using quality threshold ≥50 (completeness − 5×contamination))
- count 59 MAGs/13 phyla (SnakeMAGs) vs 29 MAGs/9 phyla (ATLAS) (Using GUNC combined with CheckM for quality assessment)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
The paper benchmarks two metagenome-assembly workflows (SnakeMAGs vs. ATLAS) applied to the same 10 publicly available termite gut metagenomes. Differences in CPU time, memory usage, and recovered MAG characteristics (completeness, contamination, genome size, relative abundance) between the two workflows, and between MAG subsets, were assessed using Wilcoxon tests, with exact p-values reported in the text. Results are additionally displayed as boxplots (Figure 2) with lines linking paired per-metagenome outcomes across workflows. No correction for multiple comparisons, effect sizes, or confidence intervals were reported.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Wilcoxon test | CPU time to reconstruct MAGs, SnakeMAGs vs ATLAS (Figure 2A) | 10 metagenomes | not stated |
| Wilcoxon test | Memory usage, SnakeMAGs vs ATLAS | 10 metagenomes | not stated |
| Wilcoxon test | MAG completeness, contamination, and genome size, SnakeMAGs vs ATLAS | — | not stated |
| Wilcoxon test | Completeness, contamination, relative abundance, and genome size of MAGs recovered by SnakeMAGs only vs MAGs recovered by both workflows | — | not stated |
-
Paired per-metagenome comparisons between the two workflows (CPU time, memory, MAG quality, genome size) were assessed with the Wilcoxon test.↳ Could also: A linear mixed-effects model with metagenome as a random effect — would model the paired, repeated-measures structure directly within a single unified framework and could provide effect-size estimates with confidence intervals alongside significance testing.
-
Several Wilcoxon tests were performed across related outcome metrics (CPU time, memory, completeness, contamination, genome size, relative abundance) without a stated multiplicity adjustment.↳ Could also: A false-discovery-rate procedure such as Benjamini-Hochberg, or a Bonferroni correction — would account for the increased chance of a nominally significant result arising when testing multiple related outcomes from the same dataset.
-
MAGs recovered from the same metagenome appear to be compared as individual data points in the quality/abundance comparisons.↳ Could also: A hierarchical or mixed model with metagenome as a random effect — would explicitly account for non-independence among MAGs drawn from the same sample, since genomes from one metagenome may be more similar to each other than to genomes from other metagenomes.
-
Exact p-values from Wilcoxon tests are reported without accompanying effect-size estimates.↳ Could also: Reporting an effect size such as the rank-biserial correlation or the Hodges-Lehmann estimator of the paired difference — would communicate the magnitude of the difference between workflows in addition to its statistical significance.
-
Data spread across the two workflows is conveyed only through boxplots (Figure 2), without a numeric dispersion measure stated in text.↳ Could also: Reporting median and IQR, or mean and 95% CI, directly in the text or a table — would make the reported variability easier to reference and compare across metrics without needing to read values off a figure.
-
A non-parametric Wilcoxon test was used throughout without discussion of the underlying data distribution.↳ Could also: Checking normality (e.g., Shapiro-Wilk) and using a paired t-test where assumptions are met — could increase statistical power if the differences are approximately normally distributed, though the non-parametric choice is a robust default when distribution shape is uncertain or sample size is small (n=10).
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-36875992 (SnakeMAGs)
Paper: Tadrent, Dedeine, Hervé (2022). SnakeMAGs: a simple, efficient, flexible and scalable workflow to reconstruct prokaryotic genomes from metagenomes. F1000Research 11:1522. DOI 10.12688/f1000research.128091.2 · PMID 36875992 · PMCID PMC9978240.
Code: https://github.com/Nachida08/SnakeMAGs (Snakemake workflow; the authors' OWN tool). Results archive: Zenodo 10.5281/zenodo.7661004 (deposited MAGs + taxonomy CSV = ground truth). Data accession (this RU): SRA SRR10402454 — one of TEN termite-gut metagenomes used in the paper.
What SnakeMAGs is
An 8-step Snakemake pipeline turning raw metagenomic reads into MAGs:
- Quality filtering — illumina-utils v2.12
- Adapter trimming — Trimmomatic v0.39 (params
2:40:15) - (optional) Host removal — Bowtie2 v2.4.5 + SAMtools + BEDtools
- Assembly — MEGAHIT v1.2.9 (min contig 1000 bp; k=21..119)
- Binning — BWA + MetaBAT2 v2.15 (min contig 2500 bp; seed 19860615)
- Bin QC — CheckM v1.1.3 (keep ≥50% completion, ≤10% contamination); optional GUNC v1.0.5
- Taxonomy — GTDB-Tk v2.1.0 (GTDB r207+, tested r214)
- Abundance — CoverM v0.6.1 Driver: Snakemake v7.0.0.
Reported results (candidate claims; full text + figures)
- R-test (functional): repo ships a 250K-read ZymoBiomics mock subset
(
insub732_2_R1/R2.fastq) + hostchr19.fa. README reports a full run in 1159.32 s on an Intel Xeon Silver 4210 (40 cores) / 96 GB. No MAG count given for test. - R-65 (headline, AGGREGATE over all 10 samples): SnakeMAGs recovered 65 MAGs (>50% completion, <10% contamination, CheckM) spanning 15 phyla. (Results / Fig 2.)
- R-46: stricter "≥50 estimated quality" (Parks score) → 46 MAGs / 13 phyla.
- R-59: after GUNC chimera filtering → 59 of 65 MAGs / 13 phyla.
- R-perSample: Fig 2B shows per-sample MAG counts (exact per-sample numbers only in the figure / Zenodo deposit, not in body text).
In scope (pipeline-derived, attempted here)
| id | result | pipeline | tractability |
|---|---|---|---|
| R-test | end-to-end run on bundled mock data completes + emits MAGs | full SnakeMAGs | QUICK — small data, FLOOR |
| R-srr | # MAGs (≥50%/≤10%) from SRR10402454 alone + their CheckM stats + GTDB taxonomy | full SnakeMAGs | MAIN per-sample 1:1 vs Zenodo deposit |
| R-65* | aggregate 65 MAGs / 15 phyla over all 10 samples | full SnakeMAGs ×10 | partial — only the SRR10402454 slice is in this RU's data scope |
Ground truth for R-srr: the authors' own MAGs_SnakeMAGs.zip + taxonomic_assignment_MAGs.csv
on Zenodo, filtered to bins originating from SRR10402454 (filename prefix). This lets a
genuine per-sample 1:1 comparison even though the body text only reports the aggregate.
Out of scope (not attempted / not a SnakeMAGs pipeline output)
- ATLAS-vs-SnakeMAGs benchmark + Wilcoxon p-values (runtime/memory/quality) — a comparison to a different tool and manual statistics, not a SnakeMAGs-derived value. Out of scope (non_pipeline).
- The full 10-sample aggregate (65/46/59) — would require running all ten large metagenomes; the data scope of THIS RU is SRR10402454 only. Reported as context; the per-sample slice is attempted.
Notes
- All heavy compute on «our HPC» («infra» work dir); downloads on front1; SLURM std partition, no --mem.
- GTDB-Tk needs the ~70+ GB GTDB DB; CheckM needs its DB. Env built with conda on front1.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.