The COMBAT-TB Workbench: Making Powerful Mycobacterium tuberculosis Bioinformatics Accessible.
The main results reproduced: recomputed values matched the published ones within tolerance.
- ✓Same input data as the authors
- ✓Reported values were directly comparable
- ✓No authors-side cause for any deviation
- 🟡A deviation arose in the data or preprocessing
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The deviation was non-trivial in magnitude
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
REPRODUCED. The COMBAT-TB Workbench is a software/Galaxy-deployment paper; its only deterministic pipeline-derived numbers are the two read-mapping percentages and the 30->28 QC count from running PRJNA633244 (30 Indonesian M.tb WGS runs) through the 'Sample Report' mapping step (Trimmomatic -> snippy 4.4.5 [bwa mem -Y -M | samclip --max 10] -> samtools flagstat) against the Comas inferred-ancestral reference (Zenodo 3497110). All three reproduced on «our HPC» (SLURM «job» mapped all 30 runs; 2248035 ran real snippy 4.4.5 on the 2 flagged runs): C3 EXACT (30 uploaded, exactly SRR12416842+SRR12416824 are the low-mapping outliers -> 2 excluded -> 28 to phylogeny, identities match the paper); C1 within-tol (9.37% vs 9.63%, abs 0.26pp); C2 within-tol (1.35% vs 1.47%, abs 0.12pp). KEY METHOD POINT: plain bwa-mem over-counts (29.88%/10.17%) on these contamination-heavy samples; snippy's samclip filter (drops heavily soft-clipped spurious alignments) is what produces the paper's low values - reproducing it brought us to within 0.1-0.3pp. No fabrication indicators: every reported value is derivable from the shipped public data via the described pipeline. NOT attempted (out of scope): server wall-clock/performance numbers (hardware-dependent), the workbench deployment (irida-galaxy-deploy, infrastructure not a numeric result), the Xu/Spain 117-sample performance demo (different BioProject, no pinnable per-sample claim), and per-sample lineage/drug-resistance calls (the paper reports no specific values to grade against). Both datasets profiled: complete, deliver exactly what the paper describes, grade A.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 50assessed: 2026-06-19 ⛓ 20152de6a1a4
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-25
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetCan combining the IRIDA data-management platform with the Galaxy workflow platform into a single Docker-deployable application make Mycobacterium tuberculosis whole-genome sequencing bioinformatics (variant calling, drug resistance prediction, lineage typing, phylogenetics) accessible to public health laboratories in low- and middle-income countries that lack dedicated bioinformatics expertise?
- ★ The COMBAT-TB Workbench combines the IRIDA web platform and the Galaxy workflow platform into a single easy-to-install, Docker-based application resource
- ★ Two Galaxy workflows were implemented: M. tuberculosis sample analysis and M. tuberculosis phylogeny method
- ★ Existing Galaxy tools (Trimmomatic, snippy, snp-sites) were updated and new Galaxy tools (snp-dists, TB-Profiler, tb_variant_filter, TB Variant Report) were written to build the workflows method
- ★ irida-wf-ga2xml was updated for recent Galaxy versions and used to build IRIDA plugins for both workflows, including metadata update integration for the sample analysis plugin method
- ★ The whole Workbench (IRIDA, IRIDA plugins, MariaDB, Galaxy) can be deployed with a single docker-compose command, unlike comparator platforms Innuendo and IRIDA alone resource
- ★ Reanalysis of a 30-sample Indonesian MDR-TB dataset with the Workbench was broadly concordant with the original published results finding
- ★ Advanced phylogeny visualization showed no clear relationship between phylogenetic clustering (lineage) and hospital collection site among the Indonesian isolates finding
- Updating TB-Profiler from version 2.8.4 to 3.0.6 improved concordance of streptomycin resistance prediction with phenotypic MGIT results finding
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Whole-genome sequencing variant calling and drug resistance/lineage prediction (Galaxy workflow: snippy mapping/variant calling + TB-Profiler) | M. tuberculosis clinical isolates, 30 samples, Java, Indonesia (Tania et al. dataset) | none | variants, drug resistance phenotype prediction, lineage assignment, % reads mapped | snippy; TB-Profiler; COMBAT-TB Workbench/Galaxy |
| Maximum-likelihood phylogenetic analysis from SNVs | M. tuberculosis isolates, 28 samples (Indonesia) and 117 samples (Xu et al., Valencia, Spain) | none | Newick phylogenetic tree, lineage clustering, association with metadata (hospital site) | snippy, snp-sites, snp-dists via Galaxy/IRIDA phylogeny pipeline |
| Read quality control | M. tuberculosis WGS reads, uploaded samples | none | sequence quality metrics | FastQC |
| Taxonomic classification of sequencing reads | Two outlier samples (SRR12416824, SRR12416842) with poor M. tuberculosis mapping | none | species/taxon assignment of reads | kraken2 (standard database, 14 April 2020) |
| Phenotypic drug susceptibility testing (from source studies, reanalyzed for comparison) | M. tuberculosis cultured clinical isolates | none | MDR/sensitive/mono-resistant phenotype | MGIT |
- ▼ Two Indonesian samples had very low proportions of reads mapping to M. tuberculosis 9.63% (SRR12416824) and 1.47% (SRR12416842)
- – Majority of reads from SRR12416824 classified as Mycobacterium avium complex by kraken2 63.55%
- – Majority of reads from SRR12416842 classified as Mycolicibacterium fortuitum by kraken2 75.39%
- ▼ Upgrading TB-Profiler version reduced discordance between predicted and phenotypic streptomycin resistance reduced discordance by 2 samples
- – No clear relationship found between phylogenetic clustering and hospital collection site despite site distances up to 107 km 10 km to 107 km between sites
- – Phylogeny of Spanish (Xu et al.) samples showed clustering by M. tuberculosis lineage as expected
- – Sample processing runtimes scaled with dataset size 30 samples/25GB: upload 9 min, processing 3h23min, phylogeny 3h43min; 117 samples/42GB: upload 21 min, processing 6h, phylogeny 8h4min
- other 9.63% (reads mapped to M. tuberculosis genome for sample SRR12416824)
- other 1.47% (reads mapped to M. tuberculosis genome for sample SRR12416842)
- other 63.55% (SRR12416824 reads classified as Mycobacterium avium complex by kraken2)
- other 75.39% (SRR12416842 reads classified as Mycolicibacterium fortuitum by kraken2)
- count 30 samples, 25 GB (Indonesian MDR-TB dataset (Tania et al.) analyzed)
- count 117 samples, 42 GB (Spanish transmission dataset (Xu et al.) analyzed)
- count reduced discordance by two samples (effect of updating TB-Profiler from v2.8.4 to v3.0.6 on streptomycin resistance concordance with MGIT)
- other 10 km to 107 km (range of distances between hospital collection sites of Indonesian isolates)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This is a software resource paper describing the COMBAT-TB Workbench, a bioinformatics platform for M. tuberculosis WGS analysis. The paper contains no inferential statistics; analytical outputs consist of bioinformatics pipeline results (drug resistance prediction via TB-Profiler, maximum-likelihood phylogeny, and per-sample QC metrics). Results from two published datasets are reported descriptively, primarily as pipeline run times, percentage of reads mapped, and qualitative concordance with previously published findings.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Maximum-likelihood phylogeny (via SNV-based pipeline, tool not named in extracted text) | Phylogeny of 28 M. tuberculosis samples (Tania et al. dataset) and 117 samples (Xu et al. dataset) | 28 samples (after QC exclusion) and 117 samples respectively | not stated |
| Taxonomic read classification (Kraken2) | Quality-control investigation of two samples with low M. tuberculosis read mapping rates (SRR12416824 and SRR12416842) | 2 samples | na |
| Drug resistance phenotype prediction (TB-Profiler computational pipeline) | All 30 samples from Tania et al. dataset; results compared qualitatively to published MGIT phenotypic DST | 30 samples (28 after exclusions) | not stated |
-
Phylogenetic trees were inferred using a maximum-likelihood approach from SNV data↳ Could also: Bayesian phylogenetic inference (e.g., BEAST2, MrBayes) could also be applied to the same SNV data — Bayesian methods additionally produce posterior probability support values and can incorporate molecular clock models for dated phylogenies, which may be informative for transmission dynamics analyses; the trade-off is substantially higher computational cost
-
Concordance between pipeline drug resistance predictions and published phenotypic DST results was described qualitatively (e.g., 'broadly concordant', difference of two samples after version update)↳ Could also: Formal sensitivity, specificity, and positive/negative predictive value calculations with exact binomial confidence intervals could also be reported for each drug — Quantitative accuracy metrics would allow direct comparison of pipeline performance across software versions and against other tools, which is standard practice in diagnostic accuracy reporting (STARD framework)
-
Two samples were excluded based on low percentage of reads mapping to M. tuberculosis (9.63% and 1.47%), with Kraken2 used to characterize the dominant organisms↳ Could also: A pre-specified quantitative QC threshold (e.g., minimum breadth of coverage, minimum mean depth, or minimum fraction of reads mapped) stated a priori could also be applied uniformly to all samples — Explicit, pre-defined QC criteria improve reproducibility and allow other users of the workbench to apply consistent filtering; the current description leaves the exclusion threshold implicit
-
Pipeline run times are reported as single observed values (Table 2) on one virtual machine configuration↳ Could also: Repeated timing measurements with a measure of variability (e.g., mean ± SD across multiple runs) could also be reported — Run times on shared virtual machines can vary due to system load; replicated measurements would give readers a more reliable estimate of expected performance and its variability
-
The phylogenetic visualization was used to assess the relationship between phylogenetic clustering and hospital collection site by visual inspection↳ Could also: A formal phylogenetic-association test such as the association index (AI), parsimony score (PS), or BaTS (Bayesian tip-association significance testing) could also be applied — Quantitative tests for phylogenetic clustering by discrete trait (e.g., hospital site) provide a significance estimate and would complement the visual observation that no clear clustering by site was apparent
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-35138128 (COMBAT-TB Workbench)
Paper: van Heusden P, Mashologu Z, Lose T, Warren R, Christoffels A. "The COMBAT-TB Workbench: Making Powerful Mycobacterium tuberculosis Bioinformatics Accessible." mSphere 2022. DOI 10.1128/msphere.00991-21.
This is a software/workbench paper (a Galaxy + IRIDA deployment for TB WGS analysis). Most of the paper describes the platform, not numeric results. The demonstration analyses re-run two previously published TB WGS datasets through the workbench's two Galaxy workflows. The reproducible, pipeline-derived numbers are few but specific.
In scope (pipeline-derived, deterministic → attempted)
The "M. tuberculosis Sample Report" workflow = Trimmomatic v0.38.1 →
snippy v4.4.5 (bwa-mem + freebayes) → SnpEff → tb_variant_filter →
TB-Profiler v2.8.4 → tb_vcf_report. Mapping is against the inferred ancestral
M. tuberculosis genome (Comas et al. 2010), Zenodo 10.5281/zenodo.3497110,
file MTB_ancestor_reference.fasta.
| id | reported result | paper location | how to reproduce |
|---|---|---|---|
| C1 | SRR12416824: 9.63% of reads mapped to the M. tuberculosis genome | Results, "Examining the read mapping outputs…" | Trimmomatic → bwa-mem to ancestral ref → samtools flagstat % mapped |
| C2 | SRR12416842: 1.47% of reads mapped | same sentence | same |
| C3 | 30 samples uploaded → 2 excluded (low mapping) → 28 to phylogeny | Results | map all 30 runs; show exactly these 2 fall far below the rest (QC outcome) |
Dataset: BioProject PRJNA633244 (Tania et al., Indonesia TB WGS), 30 runs, all WGS M. tuberculosis, paired-end, public on SRA/ENA.
Out of scope (not pipeline-derived / not deterministic → not attempted)
- Runtime / performance numbers (upload 9 min, QC 3 min, sample processing 3 h 23 min, phylogeny 3 h 43 min; Xu/Spain: 117 samples 42 GB, 6 h, 8 h 4 min). These are Galaxy-server wall-clock metrics tied to the authors' specific hardware/queue — not reproducible 1:1 and not a property of the data.
- Workbench software/deployment (irida-galaxy-deploy): infrastructure, not a numeric result.
- Xu et al. (Spain) 117-sample dataset: a performance/scale demo on a different BioProject, not the assigned accession; no specific per-sample numeric claims to compare. Not attempted.
- Per-sample lineage / drug-resistance calls: the paper demonstrates the workflow produces these but reports no specific per-sample values to grade against (no table of lineages/DR calls). Nothing pinnable → not graded. (TB-Profiler lineage for the 28 passing samples could be run as a bonus if compute budget allows, but there is no reported value to compare to.)
Notes on faithfulness
- snippy v4.4.5 reports mapped-read stats from samtools after bwa-mem. The reported % almost certainly reflects post-Trimmomatic reads (Trimmomatic runs before snippy in the workflow). We replicate that order.
- The total mapping rate is dominated by contamination (these are poor-quality samples where ~90–98% of reads are non-TB), so the choice of TB reference (ancestral vs H37Rv NC_000962.3, which differ by <0.1% of sites) changes the rate negligibly. We use the ancestral genome as the paper did, and can run H37Rv as a sensitivity check.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.