Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

The COMBAT-TB Workbench: Making Powerful Mycobacterium tuberculosis Bioinformatics Accessible.

mSphere · 2022
L1 90/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
  • No authors-side cause for any deviation
What did not (or only partly)
  • 🟡A deviation arose in the data or preprocessing
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
90/100
Reproducibility score
0.9 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 79% of all assessed papers rank 211 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

REPRODUCED. The COMBAT-TB Workbench is a software/Galaxy-deployment paper; its only deterministic pipeline-derived numbers are the two read-mapping percentages and the 30->28 QC count from running PRJNA633244 (30 Indonesian M.tb WGS runs) through the 'Sample Report' mapping step (Trimmomatic -> snippy 4.4.5 [bwa mem -Y -M | samclip --max 10] -> samtools flagstat) against the Comas inferred-ancestral reference (Zenodo 3497110). All three reproduced on «our HPC» (SLURM «job» mapped all 30 runs; 2248035 ran real snippy 4.4.5 on the 2 flagged runs): C3 EXACT (30 uploaded, exactly SRR12416842+SRR12416824 are the low-mapping outliers -> 2 excluded -> 28 to phylogeny, identities match the paper); C1 within-tol (9.37% vs 9.63%, abs 0.26pp); C2 within-tol (1.35% vs 1.47%, abs 0.12pp). KEY METHOD POINT: plain bwa-mem over-counts (29.88%/10.17%) on these contamination-heavy samples; snippy's samclip filter (drops heavily soft-clipped spurious alignments) is what produces the paper's low values - reproducing it brought us to within 0.1-0.3pp. No fabrication indicators: every reported value is derivable from the shipped public data via the described pipeline. NOT attempted (out of scope): server wall-clock/performance numbers (hardware-dependent), the workbench deployment (irida-galaxy-deploy, infrastructure not a numeric result), the Xu/Spain 117-sample performance demo (different BioProject, no pinnable per-sample claim), and per-sample lineage/drug-resistance calls (the paper reports no specific values to grade against). Both datasets profiled: complete, deliver exactly what the paper describes, grade A.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 50
    assessed: 2026-06-19 ⛓ 20152de6a1a4
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-25
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

Can combining the IRIDA data-management platform with the Galaxy workflow platform into a single Docker-deployable application make Mycobacterium tuberculosis whole-genome sequencing bioinformatics (variant calling, drug resistance prediction, lineage typing, phylogenetics) accessible to public health laboratories in low- and middle-income countries that lack dedicated bioinformatics expertise?

Core claims
  • The COMBAT-TB Workbench combines the IRIDA web platform and the Galaxy workflow platform into a single easy-to-install, Docker-based application resource
  • Two Galaxy workflows were implemented: M. tuberculosis sample analysis and M. tuberculosis phylogeny method
  • Existing Galaxy tools (Trimmomatic, snippy, snp-sites) were updated and new Galaxy tools (snp-dists, TB-Profiler, tb_variant_filter, TB Variant Report) were written to build the workflows method
  • irida-wf-ga2xml was updated for recent Galaxy versions and used to build IRIDA plugins for both workflows, including metadata update integration for the sample analysis plugin method
  • The whole Workbench (IRIDA, IRIDA plugins, MariaDB, Galaxy) can be deployed with a single docker-compose command, unlike comparator platforms Innuendo and IRIDA alone resource
  • Reanalysis of a 30-sample Indonesian MDR-TB dataset with the Workbench was broadly concordant with the original published results finding
  • Advanced phylogeny visualization showed no clear relationship between phylogenetic clustering (lineage) and hospital collection site among the Indonesian isolates finding
  • Updating TB-Profiler from version 2.8.4 to 3.0.6 improved concordance of streptomycin resistance prediction with phenotypic MGIT results finding
Experimental setups
Assay System Perturbation Readout Platform
Whole-genome sequencing variant calling and drug resistance/lineage prediction (Galaxy workflow: snippy mapping/variant calling + TB-Profiler) M. tuberculosis clinical isolates, 30 samples, Java, Indonesia (Tania et al. dataset) none variants, drug resistance phenotype prediction, lineage assignment, % reads mapped snippy; TB-Profiler; COMBAT-TB Workbench/Galaxy
Maximum-likelihood phylogenetic analysis from SNVs M. tuberculosis isolates, 28 samples (Indonesia) and 117 samples (Xu et al., Valencia, Spain) none Newick phylogenetic tree, lineage clustering, association with metadata (hospital site) snippy, snp-sites, snp-dists via Galaxy/IRIDA phylogeny pipeline
Read quality control M. tuberculosis WGS reads, uploaded samples none sequence quality metrics FastQC
Taxonomic classification of sequencing reads Two outlier samples (SRR12416824, SRR12416842) with poor M. tuberculosis mapping none species/taxon assignment of reads kraken2 (standard database, 14 April 2020)
Phenotypic drug susceptibility testing (from source studies, reanalyzed for comparison) M. tuberculosis cultured clinical isolates none MDR/sensitive/mono-resistant phenotype MGIT
Key results
  • Two Indonesian samples had very low proportions of reads mapping to M. tuberculosis 9.63% (SRR12416824) and 1.47% (SRR12416842)
  • Majority of reads from SRR12416824 classified as Mycobacterium avium complex by kraken2 63.55%
  • Majority of reads from SRR12416842 classified as Mycolicibacterium fortuitum by kraken2 75.39%
  • Upgrading TB-Profiler version reduced discordance between predicted and phenotypic streptomycin resistance reduced discordance by 2 samples
  • No clear relationship found between phylogenetic clustering and hospital collection site despite site distances up to 107 km 10 km to 107 km between sites
  • Phylogeny of Spanish (Xu et al.) samples showed clustering by M. tuberculosis lineage as expected
  • Sample processing runtimes scaled with dataset size 30 samples/25GB: upload 9 min, processing 3h23min, phylogeny 3h43min; 117 samples/42GB: upload 21 min, processing 6h, phylogeny 8h4min
Key statistics
  • other 9.63% (reads mapped to M. tuberculosis genome for sample SRR12416824)
  • other 1.47% (reads mapped to M. tuberculosis genome for sample SRR12416842)
  • other 63.55% (SRR12416824 reads classified as Mycobacterium avium complex by kraken2)
  • other 75.39% (SRR12416842 reads classified as Mycolicibacterium fortuitum by kraken2)
  • count 30 samples, 25 GB (Indonesian MDR-TB dataset (Tania et al.) analyzed)
  • count 117 samples, 42 GB (Spanish transmission dataset (Xu et al.) analyzed)
  • count reduced discordance by two samples (effect of updating TB-Profiler from v2.8.4 to v3.0.6 on streptomycin resistance concordance with MGIT)
  • other 10 km to 107 km (range of distances between hospital collection sites of Indonesian isolates)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This is a software resource paper describing the COMBAT-TB Workbench, a bioinformatics platform for M. tuberculosis WGS analysis. The paper contains no inferential statistics; analytical outputs consist of bioinformatics pipeline results (drug resistance prediction via TB-Profiler, maximum-likelihood phylogeny, and per-sample QC metrics). Results from two published datasets are reported descriptively, primarily as pipeline run times, percentage of reads mapped, and qualitative concordance with previously published findings.

Replicationunclear Sample sizeSample sizes are inherited from two previously published datasets (Tania et al., n=30; Xu et al., n=117); no power calculation or prospective sample size justification is provided, as the paper is a software demonstration rather than a hypothesis-driven study GroupsNo formal group comparisons; pipeline outputs were compared qualitatively to published results from original study authors Pairingna Randomization/blindingna Dispersionnone
Statistical tests used
Test Applied to n Assumptions
Maximum-likelihood phylogeny (via SNV-based pipeline, tool not named in extracted text) Phylogeny of 28 M. tuberculosis samples (Tania et al. dataset) and 117 samples (Xu et al. dataset) 28 samples (after QC exclusion) and 117 samples respectively not stated
Taxonomic read classification (Kraken2) Quality-control investigation of two samples with low M. tuberculosis read mapping rates (SRR12416824 and SRR12416842) 2 samples na
Drug resistance phenotype prediction (TB-Profiler computational pipeline) All 30 samples from Tania et al. dataset; results compared qualitatively to published MGIT phenotypic DST 30 samples (28 after exclusions) not stated
Approaches that could also have been used
  • Phylogenetic trees were inferred using a maximum-likelihood approach from SNV data
    Could also: Bayesian phylogenetic inference (e.g., BEAST2, MrBayes) could also be applied to the same SNV data — Bayesian methods additionally produce posterior probability support values and can incorporate molecular clock models for dated phylogenies, which may be informative for transmission dynamics analyses; the trade-off is substantially higher computational cost
  • Concordance between pipeline drug resistance predictions and published phenotypic DST results was described qualitatively (e.g., 'broadly concordant', difference of two samples after version update)
    Could also: Formal sensitivity, specificity, and positive/negative predictive value calculations with exact binomial confidence intervals could also be reported for each drug — Quantitative accuracy metrics would allow direct comparison of pipeline performance across software versions and against other tools, which is standard practice in diagnostic accuracy reporting (STARD framework)
  • Two samples were excluded based on low percentage of reads mapping to M. tuberculosis (9.63% and 1.47%), with Kraken2 used to characterize the dominant organisms
    Could also: A pre-specified quantitative QC threshold (e.g., minimum breadth of coverage, minimum mean depth, or minimum fraction of reads mapped) stated a priori could also be applied uniformly to all samples — Explicit, pre-defined QC criteria improve reproducibility and allow other users of the workbench to apply consistent filtering; the current description leaves the exclusion threshold implicit
  • Pipeline run times are reported as single observed values (Table 2) on one virtual machine configuration
    Could also: Repeated timing measurements with a measure of variability (e.g., mean ± SD across multiple runs) could also be reported — Run times on shared virtual machines can vary due to system load; replicated measurements would give readers a more reliable estimate of expected performance and its variability
  • The phylogenetic visualization was used to assess the relationship between phylogenetic clustering and hospital collection site by visual inspection
    Could also: A formal phylogenetic-association test such as the association index (AI), parsimony score (PS), or BaTS (Bayesian tip-association significance testing) could also be applied — Quantitative tests for phylogenetic clustering by discrete trait (e.g., hospital site) provide a significance estimate and would complement the visual observation that no clear clustering by site was apparent
Software: TB-Profiler 2.8.4 and 3.0.6 (both mentioned) · snippy · snp-sites · snp-dists · Trimmomatic · FastQC · Kraken2 · Galaxy 21.05 · IRIDA 21.05 · Docker 19.03.14 · docker-compose 1.27.4

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-35138128 (COMBAT-TB Workbench)

Paper: van Heusden P, Mashologu Z, Lose T, Warren R, Christoffels A. "The COMBAT-TB Workbench: Making Powerful Mycobacterium tuberculosis Bioinformatics Accessible." mSphere 2022. DOI 10.1128/msphere.00991-21.

This is a software/workbench paper (a Galaxy + IRIDA deployment for TB WGS analysis). Most of the paper describes the platform, not numeric results. The demonstration analyses re-run two previously published TB WGS datasets through the workbench's two Galaxy workflows. The reproducible, pipeline-derived numbers are few but specific.

In scope (pipeline-derived, deterministic → attempted)

The "M. tuberculosis Sample Report" workflow = Trimmomatic v0.38.1 → snippy v4.4.5 (bwa-mem + freebayes) → SnpEff → tb_variant_filter → TB-Profiler v2.8.4 → tb_vcf_report. Mapping is against the inferred ancestral M. tuberculosis genome (Comas et al. 2010), Zenodo 10.5281/zenodo.3497110, file MTB_ancestor_reference.fasta.

id reported result paper location how to reproduce
C1 SRR12416824: 9.63% of reads mapped to the M. tuberculosis genome Results, "Examining the read mapping outputs…" Trimmomatic → bwa-mem to ancestral ref → samtools flagstat % mapped
C2 SRR12416842: 1.47% of reads mapped same sentence same
C3 30 samples uploaded → 2 excluded (low mapping) → 28 to phylogeny Results map all 30 runs; show exactly these 2 fall far below the rest (QC outcome)

Dataset: BioProject PRJNA633244 (Tania et al., Indonesia TB WGS), 30 runs, all WGS M. tuberculosis, paired-end, public on SRA/ENA.

Out of scope (not pipeline-derived / not deterministic → not attempted)

  • Runtime / performance numbers (upload 9 min, QC 3 min, sample processing 3 h 23 min, phylogeny 3 h 43 min; Xu/Spain: 117 samples 42 GB, 6 h, 8 h 4 min). These are Galaxy-server wall-clock metrics tied to the authors' specific hardware/queue — not reproducible 1:1 and not a property of the data.
  • Workbench software/deployment (irida-galaxy-deploy): infrastructure, not a numeric result.
  • Xu et al. (Spain) 117-sample dataset: a performance/scale demo on a different BioProject, not the assigned accession; no specific per-sample numeric claims to compare. Not attempted.
  • Per-sample lineage / drug-resistance calls: the paper demonstrates the workflow produces these but reports no specific per-sample values to grade against (no table of lineages/DR calls). Nothing pinnable → not graded. (TB-Profiler lineage for the 28 passing samples could be run as a bonus if compute budget allows, but there is no reported value to compare to.)

Notes on faithfulness

  • snippy v4.4.5 reports mapped-read stats from samtools after bwa-mem. The reported % almost certainly reflects post-Trimmomatic reads (Trimmomatic runs before snippy in the workflow). We replicate that order.
  • The total mapping rate is dominated by contamination (these are poor-quality samples where ~90–98% of reads are non-TB), so the choice of TB reference (ancestral vs H37Rv NC_000962.3, which differ by <0.1% of sites) changes the rate negligibly. We use the ancestral genome as the paper did, and can run H37Rv as a sensitivity check.
C1
Reported
SRR12416824: 9.63% of reads mapped to M. tuberculosis genome
Reproduced
9.37% (real snippy 4.4.5 snps.bam, samtools flagstat % primary mapped)
within tolerance
C2
Reported
SRR12416842: 1.47% of reads mapped to M. tuberculosis genome
Reproduced
1.35% (real snippy 4.4.5 snps.bam, samtools flagstat)
within tolerance
C3
Reported
30 samples uploaded, 2 excluded (low mapping), 28 to phylogeny
Reproduced
30 mapped; exactly SRR12416842 (10.2%) + SRR12416824 (29.9%) are the low-mapping outliers, both far below the other 28 (all >=67%) -> 30 / 2 / 28; excluded samples match the paper exactly
exact

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 90/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟡3. Location of the main deviation
🟢4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

76.2 k
tokens (I/O) · 2.9 M incl. cache
12 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.