High-resolution transcriptome and genome-wide dynamics of RNA polymerase and NusA in Mycobacterium tuberculosis.
The main results reproduced, with only marginal, non-material deviations.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Same input data as the authors
- ✓The central claim held under reproduction
- 🟡Reported values were only indirectly comparable
- 🟡A deviation arose in the data or preprocessing
- 🟡A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The deviation was non-trivial in magnitude
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
This is a genuine partial reproduction, not a 1:1 replication: using a third-party Bowtie1+bedtools+DESeq pipeline (bbcfutils-style, since paper-specific EPFL scripts are not in the linked repo) on all 13 GSE40862 runs, we exactly reproduced the stated alignment methodology (Bowtie1 -v1/-m5 for RNA-seq, -v3/-m10 for ChIP-seq, both with 38bp trimming) and the ChIP/Input duplicate-replicate design. Several downstream numeric claims land within-tolerance (Stat rRNA%, Stat total-enriched count, Exp NusA-only count, both prophage-locus NusA ERs, Rv1584c RNAP ER) while others show real, unresolved gaps (Exp rRNA% ~12pp higher than reported, Exp total-enriched count ~12% low, Rv2657c Stat ER far higher than the paper's single reported value, and RNA-seq replicate r2=0.88 vs the paper's stated >0.95). Differential expression (Exp vs Stat) could only be computed with classic DESeq's most conservative blind/pooled-dispersion mode because the design is asymmetric (2 Exp vs 1 Stat replicate) -- this yielded 0 significant genes at padj<0.05, which is a genuine methodological-power finding, not comparable to any explicit paper-reported DE count (none was found in the accessible main text). Not attempted at all: TU/antisense/TR/replication-bias/codon-usage analyses, which require paper-specific custom scripts absent from the linked repo. All compute ran successfully end-to-end on «our HPC» across all 13 samples (after diagnosing and fixing two real infrastructure bugs ourselves: a bedtools coverage OOM on large BAMs, fixed via the -sorted streaming algorithm plus higher per-job memory; and a container R environment bug where the $ data.frame accessor silently returns the whole object instead of a column, fixed by switching to [[ ]] indexing) plus one genuine external system finding (severe multi-tenant SLURM queue congestion, resolved by switching from 12h to 2h walltime requests per the account owner's updated policy).
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-07-29
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-07-31no human curator yet
- Last updated
- 2026-07-31
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusThe study asks how the transcriptional complex is assembled, distributed and active across the Mycobacterium tuberculosis genome, testing whether the anti-terminator NusA associates with RNA polymerase genome-wide and whether their occupancy reflects transcriptional activity in exponential versus stationary growth phases.
- ★ NusA interacts with RNAP ubiquitously throughout the M. tuberculosis chromosome and its ChIP-seq profile mirrors RNAP distribution in both exponential and stationary phase, despite NusA not binding DNA directly. finding
- ★ Promoter-proximal peaks for RNAP and NusA are generally observed, followed by a decrease in signal strength reflecting transcriptional polarity. finding
- ★ Differential binding of RNAP and NusA between the two growth conditions correlates with transcriptional activity as reflected by RNA abundance. finding
- ★ A significant association exists between expression levels and the presence of NusA throughout the gene body, confirming a transcription-promoting role for NusA. mechanism
- ★ Integration of ChIP-seq and RNA-seq data pinpointed transcriptional units, mapped promoters and uncovered new anti-sense and non-coding transcripts. resource
- ★ Highly expressed transcriptional units are situated mainly on the leading strand despite a relatively unbiased distribution of genes across the genome, helping replicative and transcriptional complexes to align. finding
- ★ Combining strand-specific RNA-seq with ChIP-seq of RNAP (RpoB) and NusA at single-nucleotide resolution provides a genome-wide regulatory map of M. tuberculosis. method
- Most ChIP-seq enrichment signals lie in intergenic regions where promoter sequences are most likely present; regions enriched for NusA alone are a small minority. finding
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| ChIP-seq (anti-RpoB, beta subunit of RNAP) | Mycobacterium tuberculosis H37Rv, exponential (OD600 0.4-0.6) and stationary (4 weeks) phase cultures | none (growth phase comparison) | genome-wide RNAP occupancy as normalized read counts and enrichment ratio versus Input | Illumina Genome Analyzer IIx; ChIP-Seq Sample Preparation Kit (Illumina); SR Cluster Generation Kit v4 and SBS 36 Cycle Kit; Neoclone monoclonal anti-RpoB clone 8RB13; Illumina Pipeline Software v1.60 |
| ChIP-seq (anti-NusA) | Mycobacterium tuberculosis H37Rv, exponential and stationary phase cultures | none (growth phase comparison) | genome-wide NusA occupancy as normalized read counts and enrichment ratio versus Input | Illumina Genome Analyzer IIx; TruSeq SR Cluster Generation Kit v3 and TruSeq SBS Kit v3; rat polyclonal anti-NusA (Statens Serum Institut); Illumina Pipeline Software v1.80 |
| Input DNA sequencing (ChIP-seq control) | Mycobacterium tuberculosis H37Rv, exponential and stationary phase cultures | none | background read coverage used as denominator for enrichment ratios | Illumina Genome Analyzer IIx; Illumina Pipeline Software v1.70 |
| Strand-specific RNA-seq | Mycobacterium tuberculosis H37Rv, exponential and stationary phase cultures | none (growth phase comparison) | transcript abundance as reads per million (RPM) per feature on forward and reverse strands; differential expression | Illumina Genome Analyzer IIx; TruSeq SR Cluster Generation Kit v3 and TruSeq SBS Kit v3; v1.5 sRNA adapters; SuperScript III RT; Illumina Pipeline Software v1.82; DESeq |
| Quantitative PCR (qPCR) validation of ChIP-seq | Mycobacterium tuberculosis H37Rv IP DNA from exponential and stationary phase, subset of CDSs and intergenic regions | none; mock-IP (no antibody) background control | log2 enrichment ratio of IP DNA versus Input, correlated with ChIP-seq ER | Applied Biosystems 7900HT Sequence Detection System; Sybr Green PCR Master Mix |
| Reverse transcription quantitative PCR (RT-qPCR) validation of RNA-seq | Mycobacterium tuberculosis H37Rv total RNA, exponential and stationary phase | none (growth phase comparison); no-RT control | log2 ratio of target molecule number in exponential versus stationary phase, correlated with log2 RPM ratio | Applied Biosystems 7900HT Sequence Detection System; SuperScript III reverse transcriptase; Sybr Green PCR Master Mix |
| Sequence read alignment and genome-wide occupancy/expression quantification (bioinformatics) | M. tuberculosis H37Rv genome (NCBI NC_000962.2), TubercuList annotation: 4019 CDS, 73 RNA genes, 3080 intergenic regions (7172 features) | none | per-base normalized coverage, per-feature read counts, enrichment ratios, RPM, MA plot, transcriptional unit promoter versus body quantification | Bowtie, samtools, custom Perl/Python scripts, DESeq, R/Bioconductor, UCSC Genome Browser, GraphPad Prism 5 |
- ▲ RNAP ChIP-seq signal was present along the entire genome with prominent accumulation at putative promoter regions (promoter-proximal peaks) in both exponential and stationary phase.
- – NusA-binding profile qualitatively matched and extensively overlapped RNAP distribution in both growth phases, demonstrating NusA association with the transcriptional complex.
- – Approximately equal numbers of features were enriched (ER >= 2) in exponential and stationary phase. 1117 (Exp) versus 1062 (Stat) features
- ▲ In both phases most enrichment signals were located in intergenic regions rather than within CDSs and RNA-encoding genes. >60% in IGs versus <40% inside CDSs/RNA genes
- ▼ Regions enriched for NusA alone (without RNAP) represented only a small fraction of signals. 193 (Exp) and 107 (Stat) features
- ▲ Genomic regions encoding the putative excisionases of prophages phiRv1 and phiRv2 showed exceptionally high enrichment for both RNAP and NusA, mainly in exponential phase. rv1584c: ER 21 (RNAP) and 7 (NusA); rv2657c: ER 17 (RNAP) and 6.7 (NusA)
- ▲ Strongest NusA peaks in exponential phase were at the ribosomal RNA operon locus, the rpsM-rpmJ ribosomal protein operon, the rv1534-rv1535 intergenic region and upstream of sRNA/tRNA genes; RNAP was mainly enriched at the putative rrn promoter (mcr3 locus), the B11 locus and between rv1733c and rv1734c.
- ▲ Unexpectedly high RNAP and NusA densities at ends of convergent genes or inside CDSs (e.g. between rv3661 and rv3662c, inside ino1) corresponded to recently identified new sRNAs; RNAP but not NusA was detected at some DosR regulon genes.
- count 1117 and 1062 features enriched with ER >= 2 (ChIP-seq enrichment in exponential and stationary phase, respectively, out of 7172 annotated features)
- count 193 and 107 features (features enriched for NusA only (no RNAP) in exponential and stationary phase)
- fold_change ER 21 (RNAP) and 7 (NusA) for rv1584c; ER 17 (RNAP) and 6.7 (NusA) for rv2657c (ChIP-seq enrichment ratios at prophage phiRv1/phiRv2 putative excisionase genes, mainly exponential phase)
- other >60% of signals in intergenic regions versus <40% inside CDSs and RNA-encoding genes (genomic distribution of enriched ChIP-seq features in both growth phases)
- other >95% of reads mapping to the ribosomal RNA operon (preliminary RNA-seq analysis; data sets were subsequently normalized to non-rRNA reads)
- other >98% of sequence tags mapped to annotated CDSs in the sense orientation (RNA-seq mapping statistics)
- count 4019 protein CDS, 73 genes encoding stable RNAs/sRNAs/tRNAs, 3080 intergenic regions, 7172 total features (TubercuList H37Rv annotation used for all quantification)
- count 2283 intergenic regions >= 30 bases in length (subset of annotated intergenic regions)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
The paper compares RNA polymerase and NusA ChIP-seq enrichment and RNA-seq transcript abundance between exponential and stationary growth phases of M. tuberculosis, using duplicate biological replicates per condition. Reproducibility between replicates was assessed with Pearson correlation, differences between Exp and Stat phase values for matched genomic features were assessed with the Wilcoxon signed-rank (Mann-Whitney U) test, and RNA-seq differential expression was analyzed with the DESeq package. ChIP-seq enrichment was quantified as a ratio to input (ER), with an ER ≥ 2 threshold used to call features enriched, and both ChIP-seq and RNA-seq findings were cross-checked by qPCR, correlating qPCR-derived log2 ratios with the sequencing-derived values. Analyses were run in R/Bioconductor, with GraphPad Prism 5 used for plotting.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Pearson's product moment correlation | reproducibility between biological replicates for RNAP, NusA and Input ChIP-seq data sets (Supplementary Table S2, Supplementary Figure S1); qPCR versus sequencing-derived log2 ratios | duplicate biological replicates per condition (as stated) | not stated |
| Wilcoxon signed-rank test (referred to in text as Mann-Whitney U test) | differences between Exp and Stat phase values for the same set of genomic features | — | not stated |
| DESeq (negative-binomial-based differential expression) | RNA-seq differential expression between Exp and Stat phase (RPM values), visualized via MA plot | mean RPM from experimental replicates (as stated) | not stated |
-
Differences between Exp and Stat phase values for matched features were assessed with the Wilcoxon signed-rank test (described in text alongside the Mann-Whitney U test).↳ Could also: A paired Student's t-test — If the paired differences are approximately normally distributed, a paired t-test would also be applicable and additionally yields a mean difference with a confidence interval, complementing the rank-based test.
-
RNA-seq differential expression between two conditions with a small number of replicates was analyzed using the DESeq package.↳ Could also: edgeR, or the later DESeq2 package with its shrinkage-based dispersion and fold-change estimation — These related count-based frameworks were developed to improve variance estimation when replicate numbers are low, and could also be applied to this type of duplicate-replicate design.
-
ChIP-seq features were classified as enriched using a fixed enrichment ratio cutoff (ER ≥ 2) relative to input.↳ Could also: A dedicated peak-calling algorithm (e.g. MACS) that assigns a statistical significance value or FDR to each peak — Such tools incorporate a background model and provide a probabilistic threshold, which could also be used alongside or instead of a fixed fold-enrichment cutoff.
-
Reproducibility between biological replicates was quantified using Pearson's correlation coefficient.↳ Could also: Spearman's rank correlation coefficient — Spearman's correlation is based on ranks and is often used as a complementary or alternative measure for sequencing count data, which can be skewed or contain outliers.
-
Comparisons were made across thousands of genomic features (CDSs and intergenic regions) without an explicitly stated multiple-testing correction in the text.↳ Could also: An explicitly reported false-discovery-rate procedure such as Benjamini-Hochberg — Reporting an FDR-controlling method explicitly for genome-wide, many-feature comparisons is a standard way to communicate how the family-wise error rate was managed across the large number of simultaneous tests.
-
ChIP-seq and RNA-seq experiments were each based on duplicate (n=2) biological replicates per condition.↳ Could also: A larger number of biological replicates (e.g. n=3 or more) — Additional replicates would also increase statistical power to detect differential binding or expression and allow more precise estimation of biological variability between conditions.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
Raw data identity is excellent — all 13 GSE40862 runs are open, downloadable and were processed end-to-end — and the two purely methodological claims (Bowtie1 -v1/-m5 for RNA-seq, -v3/-m10 for ChIP, 38bp trim; duplicate ChIP/Input design) reproduced exactly. The numeric endpoints drift moderately: Exp rRNA fraction 82% reported vs 94.22%/93.08% observed, enriched features 1117->981 (Exp) and 1062->998 (Stat), NusA-only 193->176 and 107->128, prophage ERs ~20-25% lower in Exp (Rv1584c RNAP 21->16.46, Rv2657c RNAP 17->13.72). These deviations sit on our side: the paper's feature set, rRNA operon boundaries and ER pseudocount are undisclosed and the linked bbcfutils repo lacks the paper-specific EPFL scripts, so we had to self-define them — a classic underspecified-method/annotation-boundary case rather than an authors' defect. Two claims were structurally untestable (the r2>0.95 replicate correlation, since stationary phase has only one RNA-seq run; the DE gene count, which lives only in the inaccessible Table S5), but every qualitative core conclusion — RNAP/NusA co-occupancy at ~1000 features per phase, a distinct NusA-only subset, strong prophage enrichment with RNAP ER >> NusA ER — held up.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.