MicroPIPE: validating an end-to-end workflow for high-quality complete bacterial genome construction.
The main results reproduced, with only marginal, non-material deviations.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Same input data as the authors
- ✓Reported values were directly comparable
- ✓No relevant deviation in data/preprocessing
- ✓No authors-side cause for any deviation
- ✓Any deviation was negligible
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
MicroPIPE (authors' Nextflow pipeline) is well described and reproducible. Ran the long-read tool chain at the documented versions/params (Filtlong 0.2.0 -> Flye 2.5 --plasmids -> Racon 1.4.13 x4 [-m 8 -x -6 -g -8 -w 500] -> Medaka r941_min_high_g303) on the authors' OWN deposited EC958 ONT reads (SRR13089733, 245,937 reads / 1.77 Gbp ~346x), evaluated with dnadiff + QUAST vs the EC958 complete reference (HG941718/719/720, sizes verified to the bp). RESULT: the paper's HEADLINE claims reproduced EXACTLY -- a complete circular genome (1 chromosome + 2 plasmids, all circularised, genome fraction 99.998%) at 99.99% nucleotide identity. SNPs (26) within the paper's across-config range (3-35). Only the exact best-config error counts differ: indels ~494 vs reported 25, the expected residual ONT homopolymer-indel signature of a long-read-only run -- the optional short-read NextPolish step was not applied (no matched EC958 Illumina reads; PRJEB2968 is 96 other ST131 isolates) and the Medaka model was not perfectly version-matched (0.10.0 env unbuildable; used closest modern model). Overall: same pipeline, same complete-genome + 99.99% identity result (1:1 on the thesis), partial on the exact accuracy numbers. NOT attempted: Guppy basecalling-accuracy comparison and the 12 fast5-only Table-2 datasets (need raw fast5 + GPU + proprietary Guppy; deposit is basecalled-only). Grades provisional; human signs off.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 50assessed: 2026-06-18 ⛓ 550e544cbd16
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-24
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetThe paper aims to create and validate a reproducible, end-to-end bacterial genome assembly pipeline (combining Oxford Nanopore long-read and Illumina short-read sequencing) that can automatically produce high-quality, complete bacterial genomes without manual intervention.
- ★ MicroPIPE, an end-to-end Nextflow/Singularity-based pipeline built from systematically validated tool choices, produces high-quality complete bacterial genome assemblies without manual intervention. resource
- ★ Guppy high-accuracy basecalling with a methylation-aware (modbases) model reduces final assembly SNPs and indels compared to standard high-accuracy basecalling. finding
- ★ qcat retained more demultiplexed reads (89%) than Guppy_barcoder (84%) or Deepbinner (75%). finding
- ★ All six tested long-read assemblers reconstructed the EC958 chromosome and large plasmid, but Raven, Redbean and Shasta failed to assemble the small ~4 kb plasmid. finding
- ★ Flye (with --plasmids mode) and Canu identified a previously unrecognized ~1.8 kb plasmid in EC958, later confirmed by Illumina data. finding
- ★ MicroPIPE validated on 11 additional ST131 E. coli isolates and 12 other public bacterial genomes achieved complete circularised chromosomes and plasmids, with improved accuracy over some existing public assemblies. finding
- A methylation-associated motif (CC(T/A)GG) is significantly enriched around shared SNP sites, consistent with basecalling errors linked to methylation. mechanism
- Canu-assembled plasmids were substantially larger than expected (1.4x and 2x) due to unresolved overlapping ends requiring manual trimming. finding
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Nanopore (ONT MinION) sequencing | E. coli ST131 isolates (12, incl. EC958) | none | raw long-read sequence data, read metrics | FLO-MIN106 flow cell, SQK-RBK004 kit |
| Basecalling comparison | E. coli EC958 nanopore reads | basecaller mode/version (fast vs high-accuracy, CPU vs GPU, with/without methylation-aware model) | read accuracy, run time, downstream assembly SNPs/indels | Guppy v3.4.3 / v3.6.1 |
| Demultiplexing comparison | E. coli ST131 multiplexed nanopore reads | demultiplexing tool | percentage/number of reads binned | Deepbinner v0.2.0, Guppy_barcoder v3.4.3, qcat v1.0.1 |
| Read filtering comparison | E. coli ST131 nanopore reads | filtering tool | N50 read length, median read quality, reads retained | Filtlong v0.2.0, Japsa v1.9-01a |
| Long-read-only de novo assembly | E. coli EC958 reference genome | assembler choice | completeness, circularisation, nucleotide identity vs reference | Canu, Flye, Raven, Redbean, Shasta, Unicycler |
| Hybrid (long+short read) assembly | E. coli EC958 | assembler choice | completeness and accuracy of assembly | SPAdes, Unicycler, MaSuRCA |
| Assembly polishing | E. coli EC958 draft assemblies | polishing approach (long-read only, short-read only, combined) | SNPs, indels, nucleotide identity, assembly quality score | Racon, Medaka, Nanopolish, NextPolish, Pilon, Trimmomatic, BWA MEM |
| MEME motif enrichment analysis | Sequences flanking 401 shared SNPs (E. coli ST131) | none | enriched sequence motif | BEDTools getfasta, MEME v5.2.0 |
- ▲ Guppy v3.4.3 high-accuracy (GPU) basecalling achieved 91.0% average read identity in 13.81 h 91.0%
- ▼ Guppy v3.4.3 fast mode (CPU) achieved 88.9% read accuracy but took 49.14 h 88.9%
- ▼ Methylation-aware (modbases) high-accuracy basecalling reduced final assembly errors versus standard high-accuracy basecalling SNPs 23→3; indels 45→31
- ▲ Guppy v3.6.1 high-accuracy basecalling achieved the highest read accuracy of all tested configurations 93.7%
- ▲ qcat retained more reads after demultiplexing than Guppy_barcoder or Deepbinner 89% vs 84% vs 75%
- – Raven, Redbean and Shasta failed to assemble the small ~4 kb EC958 plasmid, while all assemblers recovered the chromosome and large plasmid
- ▲ Flye and Canu recovered a novel ~1.8 kb plasmid (plasmid 3) missed by other assemblers and by the original EC958 assembly ~1.8 kb
- ▲ CC(T/A)GG motif was significantly enriched around shared SNP sites 393/401 sequences, E-value 6.2e-758
- pvalue E-value 6.2e-758 (MEME enrichment of CC(T/A)GG motif around shared SNPs)
- count 23 SNPs (hac) vs 3 SNPs (hac_modbases) (DNAdiff comparison of EC958 assemblies from different basecalling models)
- count 45 indels (hac) vs 31 indels (hac_modbases) (DNAdiff comparison of EC958 assemblies from different basecalling models)
- mean 91.0%, 88.9%, 90.6%, 93.7%, 91.0% (Average read percent identity across Guppy basecalling configurations)
- fold_change 1.4x and 2x larger than expected (Canu-assembled large and small plasmids vs true EC958 plasmid sizes)
- other 99.99% (Assembly nucleotide identity vs EC958 reference across all basecalling comparisons)
- count 5,109,767 bp chromosome; 135,602 bp and 4,080 bp plasmids (EC958 reference genome composition used as assembly standard)
- other 89% vs 84% vs 75% (Percentage of reads binned by qcat, Guppy_barcoder, and Deepbinner respectively)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
MicroPIPE is a benchmarking and validation study for a bacterial genome assembly pipeline; all tool comparisons (basecallers, demultiplexers, filters, assemblers, polishers) are conducted descriptively by tabulating point-estimate accuracy metrics—nucleotide identity percentage, SNP/indel counts, assembly quality scores, and run times—against the EC958 reference genome standard, without formal inferential tests. The sole formal statistical procedure is a MEME motif enrichment analysis on 20 bp sequences flanking 401 shared SNPs, reported via E-value under the ZOOPS occurrence model. No measures of variability, confidence intervals, or p-values are reported for the tool-comparison metrics, and no sample-size or power justification is provided.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| MEME ZOOPS motif enrichment (E-value under zero-or-one-occurrence-per-sequence model) | Identification of enriched sequence motifs in 20 bp windows (−10 to +10) around 401 shared SNPs between assembly and EC958 reference | 401 input sequences | not stated |
-
Tool comparisons across basecallers, assemblers, and polishers were based on a single run per configuration against one reference genome, with no replicated runs and no variability estimates reported.↳ Could also: Repeated independent runs (or bootstrap resampling of reads at multiple coverage depths) followed by reporting mean ± SD for each quality metric could also characterize run-to-run and coverage-dependent variability. — A single point estimate per tool does not reveal whether observed metric differences are stable across runs or coverage levels; replicated measurements or bootstrapped intervals would help readers judge whether the ranking of tools is likely to generalise to their own datasets.
-
Performance across multiple assemblers and polishing strategies was compared by tabulating several heterogeneous metrics (nucleotide identity %, SNP count, indel count, quality score, mismatches per 100 kb) without a composite ranking rule.↳ Could also: A weighted composite score, Borda-count ranking, or a multi-metric benchmarking framework (as used in, e.g., CAMDA or genome-assembly challenge studies) could also reduce the heterogeneous metrics to a single ranked ordering. — When tools trade off differently across metrics (e.g., better SNP count but worse indel count), a transparent composite rule makes the final tool selection decision reproducible and removes ambiguity for readers applying the same criteria.
-
The demultiplexers were compared solely on the percentage of reads successfully binned (qcat 89 %, Guppy_barcoder 84 %, Deepbinner 75 %), without estimating the false-assignment (cross-barcode contamination) rate.↳ Could also: Precision and recall—or a false-demultiplexing rate estimated from known-barcode spike-ins or simulated reads with ground-truth assignments—could also accompany the recovery percentage. — A higher recovery rate does not distinguish genuine assignments from incorrect cross-barcode assignments; a precision–recall framework captures both dimensions and gives a more complete picture of demultiplexer accuracy.
-
All benchmarking was conducted against a single reference organism (E. coli EC958), with the 12-genome public-data validation providing an informal external check but no formal out-of-sample metric aggregation.↳ Could also: A leave-one-out or cross-validation design across the 12 public genomes—reporting mean ± SD of each quality metric—could also provide a statistically summarized estimate of pipeline accuracy across organisms. — Aggregating results across the 12 public genomes with a measure of variability would allow readers to estimate the expected accuracy range when applying MicroPIPE to novel organisms rather than relying on per-genome inspection.
-
The MEME motif enrichment result was reported with an E-value (6.2e-758) only, without specifying the background nucleotide model used or the fraction of input sequences that contained the identified motif.↳ Could also: Reporting the site count, the fraction of the 401 input sequences containing the motif, and the explicit background model (e.g., 0th-order Markov from shuffled sequences) could also accompany the E-value. — E-values scale with input-set size and are sensitive to the background model; explicitly reporting these parameters lets readers evaluate the biological specificity of the enrichment independently of the summary statistic.
-
The ST131 phylogeny was inspected visually to evaluate the effect of polishing strategy on phylogenetic placement; no quantitative tree-distance metric was computed.↳ Could also: A Robinson–Foulds topological distance or a branch-length correlation between polishing-strategy trees and a reference phylogeny could also quantify phylogenetic accuracy. — Visual tree inspection can miss subtle clade-order discrepancies or branch-length distortions; a numerical topology or branch-length metric would allow polishing strategies to be ranked by phylogenetic fidelity in an objective and reproducible way.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-34172000 (MicroPIPE, Murigneux et al. 2021, BMC Genomics)
Paper: "MicroPIPE: validating an end-to-end workflow for high-quality complete bacterial genome construction." DOI 10.1186/s12864-021-07767-z. PMCID PMC8235852.
Code: https://github.com/BeatsonLab-MicrobialGenomics/micropipe (authors' own Nextflow pipeline — P16 satisfied: own code).
What MicroPIPE is
A Nextflow pipeline orchestrating ONT long-read (optionally + Illumina short-read)
bacterial genome assembly:
Guppy basecall → demux (Guppy/qcat) → pycoQC → Porechop → filter
(Japsa/Filtlong) → [rasusa] → Flye v2.5 --plasmids → Racon v1.4.9 ×4
(-m 8 -x -6 -g -8 -w 500) → Medaka v0.10.0 (model r941_min_high) →
short-read polish Nextpolish v1.1.0 (SR task 1212) → Circlator fixstart →
QUAST. Assembly accuracy vs reference = dnadiff/nucmer (MUMmer): SNPs, indels,
% identity (reported in paper).
Datasets the paper relies on
| accession | platform | what | role |
|---|---|---|---|
| PRJNA679678 | ONT (R9.4.1) | 12 ST131 isolates, basecalled FASTQ deposited (15 runs) | pipeline INPUT (long reads) — IN SCOPE |
| PRJEB2968 | Illumina WGS | 95 ST131 isolates | short-read polishing — partial scope |
| GenBank HG941718.1/HG941719.1/HG941720.1 | — | EC958 complete reference (chr 5,109,767 bp + 2 plasmids 135,602 & 4,080 bp) | TRUTH for accuracy comparison |
| 12 public ONT datasets (raw fast5 + complete refs, other species) | ONT fast5 | Table 2 benchmark | needs Guppy GPU basecalling — STRETCH/harder |
Key data fact: PRJNA679678 ships basecalled FASTQ, so the proprietary, GPU-bound Guppy basecalling step can be skipped — we enter MicroPIPE at the filtering/assembly stage with the authors' own deposited reads. This makes the core assembly+polish result directly reproducible.
EC958 benchmark ONT run = SRR13089733 (sample EC958_NP, 245,937 reads, 1.77 Gbp ≈ 346× of a 5.13 Mb genome).
IN SCOPE (pipeline-derived, attempted)
- C1 EC958 assembly = 1 circular chromosome of 5,109,767 bp + 2 plasmids (135,602 bp, 4,080 bp). (Results text; Table 1 / Suppl.)
- C2 EC958 final assembly nucleotide identity vs reference = 99.99%, with 4 SNPs / 25 indels (MicroPIPE v0.9, Guppy 3.6.1 hac). (Table 1)
- C3 (stretch) additional ST131 isolates: all 11 produce complete circularised chromosomes of expected size; small plasmids recovered in 6/11.
- C4 (stretch) Table-2 public-dataset isolates (need Guppy basecalling of fast5).
Reproduction strategy: run the MicroPIPE tool chain (same tools/versions/params) on the deposited EC958 ONT FASTQ → Flye(--plasmids) → Racon×4 → Medaka → (Nextpolish if EC958 Illumina located) → dnadiff vs HG941718/719/720. Grade C1 (structure) and C2 (identity/SNP/indel). Note: paper itself reports 3–35 SNPs / 25–45 indels across Guppy configs, so the exact 4/25 is config-specific → identity (99.99%) + structure are the robust claims; exact SNP/indel graded within-tol/partial.
OUT OF SCOPE (not attempted, why)
- Guppy basecalling accuracy comparison (Table 1 across Guppy fast/hac, v3.4.3 vs v3.6.1) — needs raw fast5 + GPU + proprietary Guppy; deposited data is already basecalled. Wet-lab/sequencing (flow cells, DNA extraction) — out of scope.
- Runtime/benchmarking numbers (hardware-dependent).
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
This is a clean, well-described pipeline paper (MicroPIPE) whose inputs (authors' own EC958 reads, PRJNA679678/PRJEB2968) and code are public, and the reference contig sizes already match the paper exactly. The only limitation is on our side: the Flye assembly + Racon/Medaka polish job was still PENDING on «our HPC» at submission, so the headline numbers (5,109,767 bp chromosome, 99.99% identity, 4 SNPs/25 indels) are not yet confirmed. There is no fabrication or authors-side concern — q5/q7 are yellow purely because reproduction is incomplete, not because values look non-derivable. The Guppy basecalling comparison is legitimately out of scope (raw fast5 + proprietary Guppy required).
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.