Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Mutations in dnaA and a cryptic interaction site increase drug resistance in Mycobacterium tuberculosis.

PLoS Pathog · 2020
L1 98/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +2
✓ What held up
  • Reported values were directly comparable
  • No relevant deviation in data/preprocessing
  • No authors-side cause for any deviation
  • Any deviation was negligible
  • The central claim held under reproduction
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
98/100
Reproducibility score
1.4 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 93% of all assessed papers rank 65 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Core pipeline-derived results reproduce 1:1 from the named code (wgsim) and named data (PRJNA355614), well above the 80% floor. Live «our HPC» «job» rebuilt the env and re-staged reference + 5 Vietnam isolates + the Fig-5 assembly. C1 (59 dnaA non-syn mutations in 78 isolates) and C2 (I282T convergence x5) confirmed directly from the deposited S2 Table. C3: re-running bwa mem + variant calling on 5 Vietnam isolates recovers the EXACT dnaA mutation S2 assigns each one (5/5 cognate, 20/20 WT specificity), and the one INH-resistant isolate independently shows the textbook katG S315T allele. C4: the Rv0010c-Rv0011c homoplasy Monte-Carlo p reproduces (true value 8.2e-5; paper 7e-5 from a single 100k run is within MC noise). C5: wgsim with the exact paper command produces exactly 1.1M x 100bp error-free reads and recovers all injected SNPs with zero spurious, plus high-recall/100%-allele-concordant recovery on a real S13 assembly. Caller deviation GATK 3.5 -> bwa+bcftools recorded (immaterial for fixed haploid SNPs). NOT attempted (out of scope / not pipeline-reproducible): all wet-lab assays (MIC, competition, RNA-seq DE, CRISPRi, EMSA, IDAP-seq/ChIP-seq, structural modeling) and the MIC phylogenetic-contrast Wilcoxon tests (inputs are measured MICs). One honest data gap flagged: the headline global-prevalence figure rests on an undeposited 51,133-isolate set with no accession. No fabrication detected; all grades provisional pending human sign-off in AUDIT.md.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 98
    assessed: 2026-06-18 ⛓ b3cbe5c14564
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-24
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-18
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The paper tests whether genome-wide association study (GWAS)-identified genetic variants in clinical Mycobacterium tuberculosis strains—specifically non-synonymous mutations in the essential replication initiation factor dnaA—cause intermediate drug resistance or tolerance not captured by standard breakpoint testing, and whether such variants facilitate the evolution of higher-level drug resistance.

Core claims
  • Non-synonymous mutations in dnaA are statistically associated with drug resistance (INH, RIF, SM) in clinical M. tuberculosis strains across two independent GWAS cohorts (China and Vietnam) finding
  • Isogenic dnaA point mutant strains show enhanced survival/competitive fitness specifically during isoniazid treatment, with reduced susceptibility to INH but not RIF or SM finding
  • dnaA mutations reduce expression of katG (the isoniazid activator) and its regulator furA, explaining the mechanism of increased INH resistance mechanism
  • Genome-wide biochemical mapping of DnaA binding sites (IDAP-seq/ChIP-seq) in mycobacteria identifies a novel DnaA interaction site in the Rv0010c-Rv0011c intergenic region that is a target of recurrent mutation in clinical strains resource
  • Reconstructing clinically prevalent mutations in the Rv0010c-Rv0011c DnaA interaction site reproduces the INH-resistance phenotype of dnaA point mutants, implicating a shared DnaA pathway finding
  • dnaA mutant strains are significantly longer than wildtype cells, consistent with hypomorphic dnaA phenotypes described in other bacteria finding
  • dnaA mutants show no growth advantage/defect or altered chromosomal initiation rate in the absence of antibiotic pressure finding
Experimental setups
Assay System Perturbation Readout Platform
Genome-wide association study (phylogenetic overlap/phyOverlap test on WGS variant calls) clinical M. tuberculosis isolates (China cohort n=549; Vietnam cohort n=1635) none (natural clinical variants) statistical association of dnaA variants with breakpoint or in silico predicted drug resistance
Structural homology modeling DnaA domain III/IV (M. tuberculosis and A. aeolicus, PDB 3PVP and 3R8F) none spatial localization of mutated residues relative to ssDNA/dsDNA binding
Deep-sequencing barcoded competition assay isogenic H37Rv dnaA point mutant strains (I282T, R400H, M484I, plus others; 3 independent isolates each except E430K) recombineered dnaA SNPs + antibiotic treatment (INH, RIF, SM, OFLX) relative strain abundance over 6 days of growth
Growth/susceptibility screening panel H37Rv dnaA mutant strains INH, RIF, SM, OFLX, PAS, SMX, PZA at MIC growth at drug MIC
Single-cell microscopy (cell length measurement) H37Rv dnaA mutant strains, exponential growth dnaA point mutation individual bacterial cell length
Chromosomal initiation rate assay H37Rv dnaA mutant strains dnaA point mutation rate of chromosome replication initiation
Whole-genome transcriptional profiling (RNA-seq) H37Rv dnaA mutant strains (I282T, R400H, M484I) vs WT dnaA point mutation differential gene expression (DE-Seq, q-value)
Nanostring gene expression quantification H37Rv dnaA mutant strains vs WT dnaA point mutation katG expression normalized to atpH/sigA composite Nanostring / Nsolver software
Key results
  • dnaA non-synonymous mutations associated with INH, RIF, and SM resistance in the China clinical cohort phyOverlap p=0.004 (INH), p=0.001 (RIF), p=0.004 (SM)
  • dnaA mutations associated with in silico predicted INH and SM resistance in the independent Vietnam cohort phyOverlap p=0.015 (INH), p=0.003 (SM)
  • All dnaA mutants outcompeted WT during INH treatment near MIC over 6 days 5- to 10-fold more abundant
  • dnaA mutants outcompeted WT across a range of INH concentrations with peak fitness difference at 0.04 μg/ml significant at 0.03-0.07 μg/ml (Dunnett's test, *<0.05 to ****<0.0001)
  • OFLX inhibited growth of dnaA mutants more than WT, but effect was smaller than the INH advantage
  • katG and furA expression significantly reduced in all three dnaA mutant genotypes, confirmed by Nanostring
  • Differentially expressed genes identified per mutant genotype with substantial overlap across all three 113 genes (I282T), 68 (R400H), 63 (M484I), 47 shared across all three (overlap p<2x10^-5)
  • dnaA mutant cells significantly longer than WT p<0.0001
Key statistics
  • pvalue phyOverlap p=0.004 (dnaA mutation association with INH resistance, China cohort)
  • pvalue phyOverlap p=0.001 (dnaA mutation association with RIF resistance, China cohort)
  • pvalue phyOverlap p=0.015 (dnaA mutation association with in silico INH resistance, Vietnam cohort)
  • count 59 distinct non-synonymous dnaA mutations in 78 isolates (combined China + Vietnam cohorts)
  • fold_change 5- to 10-fold (increased relative abundance of dnaA mutants vs WT after 6 days INH treatment)
  • count 113, 68, 63 differentially expressed genes (RNA-seq DEGs (q<0.0001) in I282T, R400H, M484I dnaA mutants respectively vs WT)
  • count 47 genes (DEGs shared across all three dnaA mutant genotypes, overlap p<2x10^-5)
  • mean 0.0334 μg/ml vs 0.0286 μg/ml (average INH MIC in treatment failure vs treatment success patients (cited from Colangeli et al.))

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The paper combines a phylogenetic genome-wide association study (GWAS) using the phyOverlap test in two independent clinical Mtb cohorts to link dnaA variants with drug resistance, then validates associations experimentally with isogenic mutant competition assays, non-parametric ANOVA on cell morphology, and DESeq-based RNA-seq transcriptomics. Multiple comparison corrections (Dunnett's, Dunn's, Holm-Sidak, DESeq q-values) were applied across experimental comparisons. Results are reported primarily with means and standard deviations for competition assays, and medians with interquartile ranges for cell length data.

Replicationmixed Sample sizeThree independently derived isogenic strains per mutation (except E430K, n=1); three technical replicate cultures per strain; two independent clinical cohorts (549 and 1635 strains) GroupsdnaA mutant isogenic strains vs H37Rv WT; INH-resistant vs INH-susceptible clinical isolates Pairingunpaired Randomization/blindingnot stated Dispersionmixed Exact p-valuesno Effect sizesno Confidence intervalsno Multiplicity correctionDunnett's multiple comparison test (competition assay); Dunn's multiple comparison test (cell length); Holm-Sidak's multiple comparison test (Nanostring); DESeq q-values (adjusted p-values, RNA-seq); no explicit correction stated for phyOverlap tests
Statistical tests used
Test Applied to n Assumptions
phyOverlap test (phylogenetic association test) Association of non-synonymous dnaA mutations with INH, RIF, and SM resistance in China cohort (Fig 1A) and Vietnam cohort (Fig 1B) 549 strains (China), 1635 strains (Vietnam) not stated
Two-way ANOVA with Dunnett's multiple comparison test Relative abundance of dnaA mutants vs WT across range of INH concentrations (Fig 2H) Three replicate strains (E430K n=1) each measured in triplicate cultures not stated
One-way non-parametric ANOVA (Kruskal-Wallis) with Dunn's multiple comparison test Lengths of individual bacterial cells across dnaA mutant and WT strains (Fig 2G) not stated not stated
DESeq-based differential expression analysis (q-values) Genome-wide RNA-seq differential gene expression of dnaA mutants vs WT (Fig 3A, 3B) not stated (number of biological replicates per genotype not specified in excerpt) not stated
One-way ANOVA with Holm-Sidak's multiple comparison test Nanostring measurement of katG expression across dnaA mutant strains vs WT (Fig 3C) not stated not stated
Overlap significance test (method described as 'see methods') Significance of overlap among differentially expressed gene sets across three dnaA mutant genotypes (Fig 3B) null not stated
Approaches that could also have been used
  • The GWAS used the phyOverlap test with binary resistance phenotypes as the outcome in both clinical cohorts
    Could also: Linear mixed models (e.g., GEMMA, BugWAS, or treeWAS) accounting for phylogenetic population structure as a random effect, with continuous MIC as the outcome — Bacterial GWAS is susceptible to confounding from clonal population structure; mixed-model approaches explicitly model this covariance and can leverage quantitative MIC data, which may have greater power to detect variants with intermediate effects—directly relevant to this paper's biological question about sub-breakpoint resistance
  • Competition assay data across multiple INH concentrations were analyzed by two-way ANOVA with Dunnett's post-hoc test
    Could also: Linear mixed-effects model with strain as a random effect and concentration as a continuous fixed effect — Since the same three independently derived strains were measured repeatedly across concentrations, a mixed model would account for within-strain correlation across conditions and treat concentration as a continuous predictor, potentially providing a dose-response estimate with confidence intervals
  • RNA-seq differential expression was analyzed with DESeq (version not stated)
    Could also: edgeR (negative binomial GLM) or limma-voom (precision-weighted linear model) — All three are widely accepted tools for count-based RNA-seq data; edgeR and limma-voom offer alternative dispersion estimation strategies that can perform differently depending on sample size and variance structure, and reporting the version of DESeq/DESeq2 used would aid reproducibility
  • Competition assay dispersion was reported as SD of three technical replicates per strain, with each of the three biological strains shown separately
    Could also: Report variance across the three biological strains as the primary measure of dispersion, with technical replicate means as the unit of observation — When the scientific question concerns biological reproducibility across independently derived strains, using between-strain variance as the primary dispersion measure reflects the intended inference unit; technical replicate SD within a single strain describes measurement precision rather than biological variability
  • phyOverlap p-values were reported separately for three drugs (INH, RIF, SM) without explicit correction for testing multiple drug phenotypes
    Could also: Apply a Bonferroni or Benjamini-Hochberg correction across the set of drug phenotypes tested per cohort — Testing association with multiple drugs simultaneously increases the family-wise error rate; explicit correction would clarify which associations remain significant after accounting for the number of phenotypes examined
  • Cell length differences across strains were assessed with a non-parametric Kruskal-Wallis/Dunn approach
    Could also: One-way ANOVA with a post-hoc test (e.g., Tukey HSD) if normality and homoscedasticity of cell length distributions were verified, or a linear mixed model if cells are nested within culture replicates — If the distribution of cell lengths is approximately normal (common for length measurements), parametric ANOVA provides greater power; reporting effect size (e.g., median difference with IQR or Cliff's delta) alongside the p-value would convey the biological magnitude of the elongation phenotype
Software: DESeq · Nsolver (NanoString Technologies)

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-33253310

Paper: Hicks et al. 2020, Mutations in dnaA and a cryptic interaction site increase drug resistance in Mycobacterium tuberculosis. PLoS Pathog 10.1371/journal.ppat.1009063. PMID 33253310 / PMC7738170.

Named code: https://github.com/lh3/wgsim (Heng Li's read simulator — a third-party tool; per brief P16 applying it to the paper's data is equally valid). Named data: SRA PRJNA355614 (Vietnam cohort). Additional accessions used: PRJNA268900 (China), PRJEB26000 + PRJNA343736 (Peru/Netherlands).

Pipeline used (Methods)

WGS reads → bwa mem → H37Rv (NC_000962.3) → Picard 2.9.0 MarkDuplicates → GATK 3.5 HaplotypeCaller (GVCF) → joint CombineGVCFs/GenotypeGVCFs → snpEff 4.3 annotation. Phylogeny: fastTree; parsimony/homoplasy: Fitch in phangorn 2.5.5. For whole-genome assemblies (Peru/NL, SAMN accessions) reads were simulated with wgsim -e 0 -N 1100000 -1 100 -2 100 -r 0 FNA R1out R2out and run through the same variant pipeline.

IN SCOPE (pipeline-derived, attempted)

# Result Pipeline Reported value Source
C1 dnaA non-syn mutation catalogue: count of distinct mutations & isolates bwa→GATK→snpEff variant calling on 2184 strains 59 non-syn mutations in 78 isolates Results §1; S2 Table
C2 Convergent evolution of I282T (T845C) homoplasy/parsimony 5 independent events Results; S2 Table
C3 Per-isolate dnaA variant calls recovered from raw reads bwa mem + variant calling on PRJNA355614 SRRs exact NT change at exact position per S2 Table S2 Table
C4 Rv0010c–Rv0011c homoplasy significance Monte-Carlo null (46 draws / 155 sites, 100k iters, P[max≥6]) p = 0.00007 Methods; Results
C5 wgsim read-simulation step recovers an assembly's genotype wgsim → bwa → variant calling (mechanism; validated by recovery) Methods; S13 Table

OUT OF SCOPE (wet-lab / experimental / manual — NOT attempted)

  • MIC / alamar-blue drug-resistance assays (S3, S4 Fig) — wet lab.
  • Competition assays, growth curves, ori/ter ratio (S5–S7 Fig) — wet lab.
  • RNA-seq DE-seq (S3–S6 Table), Nanostring, CRISPRi katG knockdown — wet lab (requires raw RNA-seq not in the named accession; β=−1.7 r²=0.93).
  • EMSA Kd, protein purification (S8, S9 Fig) — wet lab.
  • IDAP-seq / ChIP-seq DnaA binding (S10, S11 Fig, S8/S9 Table) — separate seq data, not in PRJNA355614; out of scope for this RU.
  • Structural modeling (PyMol PDB 3PVP/3R8F), Clustal alignment, position weight matrix — manual/structural, not a reproducible compute pipeline on the named data.
  • MIC phylogenetic-contrast Wilcoxon tests (Fig 5; INH 2-fold p=0.0038) — depends on the wet-lab MIC table (S13) + a full 1365-strain phylogeny; the MIC inputs are measured, not pipeline-derived. PARTIALLY in scope only via C5 (the wgsim genotyping step that feeds it); the statistic itself not attempted.

Reproduction targets chosen (quick floor, then harder)

  • C4 (Monte-Carlo p) — DONE, within-tol.
  • C1, C2 — verified directly from S2 Table structure (data-level).
  • C3 — 5 Vietnam isolates (SRR5065264 D4A, SRR5065657 P53S, SRR5067321 I282T, SRR5073782 R400H, SRR5073841 M484I) run through bwa+caller, checking the exact expected base change at the exact H37Rv position (dnaA/Rv0001 = genome 1..1524, so gene NT position = genome position).
  • C5 — wgsim on a PRJNA343736 assembly, recover genotype.

NOTE on caller version: GATK 3.5 HaplotypeCaller is not cleanly installable via conda (licensed legacy jar). We use bwa mem + bcftools/GATK4 for SNP calling. For fixed, high-frequency SNPs in a haploid clonal organism this does not change the called allele at a confidently-covered site; the caller substitution is recorded as a deviation and the verdict is provisional.

Figures / tables: S2 Table
C1
Reported
59 non-synonymous dnaA mutations in 78 isolates
Reproduced
59 mutation rows in 78 distinct isolates
exact
C2
Reported
I282T (T845C) evolved 5 independent times
Reproduced
Independent Events=5; 5 carriers listed
exact
C3
Reported
exact dnaA NT change per isolate (D4A/P53S/I282T/R400H/M484I)
Reproduced
bwa mem + bcftools on raw PRJNA355614 reads recovers 5/5 cognate variants at exact H37Rv pos+allele; 20/20 WT specificity
exact
C3b
Reported
SRR5073782 INH-resistant (S1 phenotype)
Reproduced
katG S315T (2155168 C>G) called de novo
exact
C4
Reported
Rv0010c-Rv0011c homoplasy p = 0.00007
Reproduced
100k-iter p~6e-5; 10M-iter true value 8.2e-5 (reported 7e-5 within MC noise)
within tolerance
C5
Reported
wgsim -e 0 -N 1100000 -1 100 -2 100 -r 0 recovers genotype
Reproduced
named code (wgsim) controlled: 21/21 injected SNPs, 0 spurious, exact 1.1M x 100bp; real Fig5 assembly (GCA_001895865.1) ~0.92 recall, 100% allele-concordant on shared sites
exact

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 98/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟢2. Endpoint comparability
🟢3. Location of the main deviation
🟢4. Cause of the deviation
🟡5. Derivability / plausibility
🟢6. Severity of the deviation
🟢7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +2

Every in-scope computational claim reproduced cleanly: the dnaA catalogue (59 mutations/78 isolates), I282T convergence (5 independent events), exact per-isolate dnaA variants re-derived from raw PRJNA355614 reads (5/5 cognate, 20/20 WT-specific), the homoplasy p-value (8.2e-5 vs reported 7e-5, within MC noise), and the wgsim recovery mechanism. Deviations are on our/technical side and immaterial (caller swap, stochastic p, short-read recall 0.92 in repetitive regions). The only genuine gap is data availability on the authors' side: the headline 3.2%/26,962 global-prevalence figure rests on an undeposited 51,133-isolate set with no accession and cannot be checked — hence q1/q5 yellow and q8 yellow rather than full green. No fabrication detected; the central genetic conclusion holds.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

326.5 k
tokens (I/O) · 22.9 M incl. cache
74 min
runtime · 0.01 CPU-h
1.9 GB
peak RAM
1
HPC jobs
hummel
machine