Dynamics of the compartmentalized Streptomyces chromosome during metabolic differentiation.
The main results reproduced: recomputed values matched the published ones within tolerance.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Same input data as the authors
- ✓Reported values were directly comparable
- ✓No relevant deviation in data/preprocessing
- ✓No authors-side cause for any deviation
- ✓Reported values are derivable from the shared data
- ✓Any deviation was negligible
- ✓The central claim held under reproduction
- ✓Overall, the reproduction was clean
- Every checked point held up.
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Described well enough to reproduce 1:1. The paper's named tool (SARTools, github.com/PF2-pasteur-fr/SARTools) is its RNA-seq differential-expression pipeline; I checked out the paper's exact tag v1.6.3 and ran it (run.DESeq2 + exportResults.DESeq2, all defaults, reference C1) on the authors' own deposited featureCounts sense count tables (GSE162865, 21 libs C1-C7). Results match the authors' own Supplementary Data 4 SARTools report to <=0.3% per comparison (per-comparison DEG totals: 1839/3276/4299/5002/996/4967 vs reported 1844/3273/4298/5002/996/4964; two exact). DESeq2-normalized reads-per-kb (normRPK, Supp Data 3) reproduce to full precision (Pearson r=1.00000, median ratio 1.0). The headline '>93% of genes DE in at least one condition' is confirmed: the authors' own DIFF005 flag = 93.78%, bracketed by my vs-C1 union (92.14%) and all-pairwise union (96.46%). The '~90% expressed' (CAT) claim checks at 88.4%. Small DEG deltas are explained by DESeq2 version (1.50.2 vs paper-era ~1.30); normalization is byte-stable. NOT attempted (out of scope): the 3C-seq/Hi-C contact-map analysis (different code artifact = koszullab tools; raw reads not in the processed deposit, only normalized *_corr.txt matrices) and all wet-lab/genome-mining results (antiSMASH SMBGCs, genomic islands, persistence indices, antibacterial assays).
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 97assessed: 2026-06-18 ⛓ 5a816d0edd40
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-18
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusDoes the linear chromosome of Streptomyces ambofaciens fold into structurally distinct compartments that correlate with its genetic compartmentalization and gene expression, and is chromosome architecture rearranged during metabolic differentiation?
- ★ Chromosome 3D structure in S. ambofaciens correlates with genetic compartmentalization during exponential phase, with the central region segmented into domains and quiescent terminal compartments finding
- ★ Onset of metabolic differentiation is accompanied by rearrangement of chromosome architecture from an 'open' to a 'closed' conformation in which highly expressed SMBGCs form new boundaries finding
- ★ Conserved, large, highly transcribed genes (especially rDNA operons) form boundaries that segment the central chromosome into domains (CIDs) mechanism
- ★ Genetic compartmentalization (core/persistent genes central; SMBGCs/GIs/unique genes terminal) correlates with compartmentalized sense and antisense transcription finding
- Gene persistence index based on 125 complete Streptomyces genomes used to map core-genome and chromosome organization method
- 3C-seq coupled with RNA-seq across growth conditions maps chromosome folding dynamics during differentiation method
- Antisense index is high for GIs, mobile elements, and SMBGCs but decreases during metabolic differentiation, suggesting antisense-transcription is linked to regulation of differentiation genes mechanism
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| 3C-seq (chromosome conformation capture coupled to deep sequencing) | Streptomyces ambofaciens ATCC 23877 | none (growth phase / media comparison: exponential vs differentiation, YEME10 vs MP5) | frequency of contacts between genome loci; boundaries via frontier index | — |
| RNA-seq (transcriptome) | Streptomyces ambofaciens ATCC 23877, liquid culture (MP5 and YEME10 media) | none (growth time course C1-C7 conditions) | DESeq2 normalized reads per kb (normRPK) per gene; sense and antisense transcription | — |
| Comparative genomics / gene persistence analysis | 125 complete Streptomyces genomes vs S. ambofaciens ATCC 23877 reference | none | gene persistence index, gene order conservation (GOC) synteny scores, core-genome definition | — |
| Genome mining for biosynthetic gene clusters | S. ambofaciens ATCC 23877 genome | none | identification of putative SMBGCs and genomic islands | antiSMASH v5.1.0 |
| Antibacterial activity assay | S. ambofaciens cultures (MP5 and YEME10 media) | none (growth conditions) | size of inhibition halo (antibiotic production) | — |
- ▲ Genomic islands are enriched in SMBGC genes relative to the rest of the genome 3.5-fold
- – ~90% of S. ambofaciens genes significantly expressed in at least one condition ~90%
- – >93% of genes differentially expressed in at least one condition vs C1 reference >93%
- – In exponential phase (YEME10, C6) the central region contains 10 boundaries defining 9 domains ranging 240-700 kb 9 domains, 240-700 kb
- ▲ Transcription of genes within boundaries tends to be oriented in the direction of continuous replication odds ratio 1.8
- ▲ Terminal regions poorly expressed at early time points, with transcription gradually increasing toward terminal ends over growth
- ▼ Up to 19% of SMBGC genes and 23% of GI genes are silent/poorly expressed in all tested conditions 19% SMBGC; 23% GI
- ▼ Antisense indices of GIs, pSAM1, prophage and especially SMBGCs decrease during metabolic differentiation
- fold_change 3.5-fold (GI enrichment in SMBGC genes)
- pvalue < 2.2 × 10^-6 (two-sided Fisher's exact test for GI enrichment in SMBGC genes)
- other odds ratio 1.8 (boundary gene transcription oriented with continuous replication)
- pvalue 1.4 × 10^-8 (two-sided test for boundary transcription orientation toward continuous replication)
- count 125 complete Streptomyces genomes (panel used to calculate gene persistence index)
- count 332 chromosomal genes expressed at very high level in all conditions (highly expressed genes enriched in central region)
- count 10 boundaries / 9 domains (240-700 kb) (central region structure in exponential phase YEME10 (C6))
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This multi-omics study of Streptomyces ambofaciens combined RNA-seq (7 growth conditions across two liquid media) and 3C chromosome conformation capture sequencing to characterize transcriptome and chromosome architecture dynamics during metabolic differentiation. Differential gene expression was assessed with DESeq2, genome-feature enrichments with Fisher's exact tests, and expression distributions across genomic compartments with Wilcoxon rank-sum tests. Results were visualized primarily as IQR-based boxplots and contact maps, with exact p-values reported for key comparisons and one odds ratio noted.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| DESeq2 (Wald test, negative binomial GLM) with adjusted p-values | Differential gene expression across all growth conditions, using C1 (24 h MP5) as reference; Fig. 2a, b and Supplementary Data 3 | 7 growth conditions (C1–C7); number of biological replicates per condition not stated in excerpt | not stated |
| Two-sided Wilcoxon rank-sum test with continuity correction | Comparison of normalized read counts (normRPK) per genomic feature category versus whole chromosome; Fig. 2c | Gene counts per feature indicated in figure (e.g., 332 highly expressed genes); exact per-group n not given in excerpt | not stated |
| Two-sided Wilcoxon rank-sum test with continuity correction | Antisense index comparisons: (i) 24 h vs 48 h within each genome feature; (ii) each genome feature vs whole chromosome at 24 h and 48 h; Fig. 2d | Gene counts per feature indicated in figure; exact per-group n not given in excerpt | not stated |
| Two-sided Fisher's exact test for count data | Enrichment of SMBGC genes within genomic islands (GIs); reported 3.5-fold enrichment, p < 2.2 × 10^-6 | Not stated in excerpt | not stated |
| Unspecified test yielding an odds ratio | Association between transcription orientation of boundary genes and direction of continuous replication; odds ratio 1.8, p = 1.4 × 10^-8 | Not stated | not stated |
| Hierarchical classification (clustering) | Grouping of the 7 growth conditions by transcriptome similarity; Fig. 2a, b and Supplementary Fig. 3a | 7 conditions | na |
-
Multiple pairwise Wilcoxon rank-sum tests were used to compare expression distributions for each of several genomic feature categories against the whole chromosome↳ Could also: A Kruskal-Wallis omnibus test followed by Dunn's post-hoc test (or Steel-Dwass) with FDR or Bonferroni correction across all pairwise comparisons — An omnibus test followed by a corrected post-hoc step makes the experiment-wide false-positive rate explicit and quantified when many simultaneous non-parametric comparisons are performed
-
DESeq2 was used for RNA-seq differential expression analysis↳ Could also: edgeR (negative binomial GLM with quasi-likelihood F-test or exact test) or limma-voom (linear model on log-CPM with empirical Bayes shrinkage) — All three are widely accepted for count-based RNA-seq; edgeR and limma-voom offer complementary dispersion-estimation strategies and are commonly applied in parallel as a cross-validation step or when per-condition sample sizes are very small
-
Expression distributions were summarized using IQR-based boxplots without reporting SD, SEM, or individual data points↳ Could also: Overlaying individual data points (strip or jitter plot) on the boxplot, or additionally reporting mean ± SD — Individual-point overlays reveal sample size, clustering, and distributional shape directly; SD conveys absolute spread on the original scale—both are frequently preferred when the number of observations per group is small or varies across groups
-
Hierarchical classification was used to group the seven transcriptomic conditions into two clusters↳ Could also: Principal component analysis (PCA) for dimensionality reduction, optionally complemented by t-SNE or UMAP — PCA provides a continuous low-dimensional projection of total variance, making the magnitude of separation between conditions interpretable and helping identify outlier samples; it is a standard complement to hierarchical clustering in transcriptomics workflows
-
A Fisher's exact test was used to assess enrichment of SMBGC genes within genomic islands↳ Could also: A permutation-based enrichment test or a logistic regression model that can incorporate covariates such as gene length, local GC content, or chromosomal position — When genomic features like gene length or chromosomal arm position covary with island annotation, covariate-aware approaches can disentangle confounded enrichment signals that a simple 2×2 contingency test cannot
-
3C-seq domain boundaries were called with the 'frontier index' multiscale algorithm↳ Could also: Alternative boundary/TAD callers such as the insulation score method, HiCExplorer's hicFindTADs, or TADtool — Different callers use distinct mathematical formulations and sensitivity thresholds; comparing boundary positions across two or more methods (consensus calling) is a common practice to assess robustness and reduce algorithm-specific artifacts
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-34471117
Paper: Lioy et al. 2021, Dynamics of the compartmentalized Streptomyces chromosome during metabolic differentiation, Nat Commun 12:5221. DOI 10.1038/s41467-021-25462-1 · PMCID PMC8410849 · GEO GSE162865.
The paper combines two computational pipelines on Streptomyces ambofaciens ATCC 23877: (A) RNA-seq differential expression (featureCounts → SARTools/ DESeq2) and (B) 3C-seq / Hi-C contact-map analysis (koszullab tools). The named code artifact in the brief is SARTools (https://github.com/PF2-pasteur-fr/SARTools), the RNA-seq DE pipeline.
IN SCOPE (pipeline-derived, attempted)
| # | Result | Pipeline | Reproducible? |
|---|---|---|---|
| R1 | ">93% of genes differentially expressed in ≥1 condition (DESeq2 adj-p<0.05, C1=MP5_24h as reference)" — Results §"Transcriptome dynamics", refers to Supplementary Data 3/4 | featureCounts (deposited counts) → SARTools v1.6.3 DESeq2 | YES — deposited per-sample sense count tables + SARTools defaults fully specify it |
| R2 | Per-comparison DEG counts (Cx vs C1) — the SARTools statistical report = Supplementary Data 4 | SARTools v1.6.3 DESeq2 | YES (re-run; compare to Supp Data 4 if obtainable) |
| R3 | DESeq2-normalized counts per kb (normRPK) per gene/condition — Supplementary Data 3, Fig 2a heatmaps | SARTools DESeq2 size-factor normalization + gene-length scaling | PARTIAL (normalized counts reproducible; normRPK = norm.counts/genelen×1000) |
Methods (verbatim essentials): featureCounts v2.0.1, -t gene -g ID, sense &
antisense, annotation GCF_001267885.1_ASM126788v1 (06/15/2020). SARTools v1.6.3
DESeq2-based; reference condition C1 (MP5_24h); BH adjustment, significance
threshold 0.05; VST for PCA/clustering; all SARTools default parameters
unchanged; genes with null counts in all samples excluded; second-TIR genes
(not used for mapping) and rRNAs (RiboZero) excluded.
OUT OF SCOPE (not the named artifact / not attempted)
- 3C-seq / Hi-C contact maps, boundaries, domains, compartments (Figs 1,3–6):
analyzed with koszullab tools (E_coli_analysis, hicstuff-type), NOT SARTools.
GEO ships only the normalized contact matrices (
*_corr.txt.gz), not raw reads → not re-derivable from the deposit; different code artifact. Not attempted. - Wet-lab / manual results: antibacterial halo assays (Table 1), antiSMASH SMBGC identification, genomic-island definition, 125-genome persistence/conservation indices — manual/external, not the SARTools pipeline. Not attempted.
Reproduction strategy
Apply SARTools v1.6.3 (the paper's exact tool + version, third-party-OK per P16) to the paper's own deposited featureCounts sense count tables (21 libraries, conditions C1–C7), reference = C1, defaults; count genes with padj≤0.05 & betaConv in any Cx-vs-C1 comparison ÷ analyzed genes → compare to the ">93%" claim. Heavy compute on «our HPC» SLURM; data on «infra».
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
The paper's RNA-seq differential-expression pipeline reproduces essentially 1:1 from the complete GSE162865 deposit using the authors' own named tool (SARTools v1.6.3/DESeq2) at the matching tag and parameters. DESeq2-normalized counts regenerate to full precision (r=1.00000) and per-comparison DEG totals agree within <=5 genes (<0.3%, two exact), with the '>93% DE' headline bracketed by the reproduced union definitions. The only deviations are technical version-drift numerics on our side; data and method are fully available and the central conclusion holds. The Hi-C/3C-seq and wet-lab results are out of scope and not adjudicated here.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.