Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Expansion of the SOS regulon of Vibrio cholerae through extensive transcriptome analysis and experimental validation.

BMC Genomics · 2018
L1 85/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +8
✓ What held up
  • Same input data as the authors
What did not (or only partly)
  • 🟡Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
85/100
Reproducibility score
0.6 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 67% of all assessed papers rank 348 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough to reproduce: YES — a clean 1:1 reproduction. The paper's Methods fully specify the SOS RNA-seq pipeline and PRJNA429784 resolves cleanly; the 4 directional 'Marine' MMC+/- libraries (2 reps each) are the SOS-vs-noSOS DE comparison. Repo baj12/clean_ngs (->PF2-pasteur-fr/clean_ngs) is only the third-party adapter-trimming step, so we ran the full described pipeline on the paper's own data (valid per P16): cutadapt(>=25nt) -> Bowtie1 (--chunkmbs 400 -m 50 -e 50 -a --best, ~98% mapping to N16961) -> htseq-count (intersection-nonempty, -t gene) -> DESeq2 (shorth size factors, padj<0.001, >2-fold). Key finding: the libraries are REVERSE-stranded (empirically: -s yes counts 0.5% of reads, -s reverse 95.8%), so htseq was run with -s reverse. RESULTS: all 26 Table-1 SOS-regulon fold-changes (C5-C30) reproduced within-tol with correct direction (recN 47.0->47.2, lexA 25.5->25.9, recA 16.4->16.6, intIA 14.4->14.6, repressors rstR1/2 0.3->0.26); global DE counts C1 737->671 (within-tol) and C2 231->263 (within-tol). Small count gaps are fully explained by annotation drift (standard NCBI GenBank GCA_000006745.1 vs the authors' hand-curated N16961 annotation) and tool-version drift (Bowtie 0.12.7->1.3.1, DESeq2 1.8.1->1.50.2). NOT ATTEMPTED: ncRNA counts (28/17 — require authors' custom ncRNA GFF, out of scope) and all wet-lab/TSS/EMSA/UV-survival results (non-pipeline). No value fabricated; grades provisional, a human reviewer decides.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 56
    assessed: 2026-06-16 ⛓ 1add7487d71b
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-24
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-16
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The full extent of the SOS regulon in Vibrio cholerae has only been predicted bioinformatically (via LexA binding box consensus) and not comprehensively validated experimentally; this study tests whether whole-transcriptome analysis after mitomycin C-induced SOS induction can reveal additional genes, ncRNAs, and pathways belonging to or affected by the SOS response.

Core claims
  • Whole transcriptome sequencing with extensive TSS mapping identified 3078 transcription start sites and 629 ncRNAs in V. cholerae N16961 finding
  • Mitomycin C treatment significantly induced 737 genes and 28 ncRNAs (>2-fold) and repressed 231 genes and 17 ncRNAs finding
  • Twelve genes (ubiEJB, tatABC, smpA, cep, VC0091, VC1190, VC1369-1370) are co-induced with adjacent canonical SOS genes via transcriptional read-through, expanding the known SOS regulon finding
  • Mutants of newly identified SOS regulon genes and highly induced genes/ncRNAs show altered UV and mitomycin C susceptibility, confirming roles in DNA damage rescue and protection finding
  • Syntenic co-localization of expanded SOS regulon genes (recN-smpA and rmuC-tatABC) is conserved in other γ-proteobacteria, suggesting broader phylum-wide SOS regulon conservation finding
  • A new RNA purification/library protocol combining TEX and TAP treatments enables accurate genome-scale TSS identification method
  • Genotoxic stress induces a pervasive transcriptional response affecting almost 20% of V. cholerae genes finding
  • 151 gene start codons were reassigned based on TSS mapping data resource
Experimental setups
Assay System Perturbation Readout Platform
Directional whole-transcript RNA-seq V. cholerae N16961 mitomycin C (0.2 µg/mL, 200 ng/ml) differential gene/ncRNA expression (fold change, adjusted p-value) Illumina HiSeq2000, TruSeq Stranded mRNA kit
5' end differential RNA-seq (TSS mapping) V. cholerae N16961 TEX+/- and TAP+/- enzymatic treatment across 7 growth conditions transcription start site positions, 5'UTR length Illumina HiSeq2000, TruSeq Small RNA kit
RT-PCR V. cholerae N16961, SOS operons mitomycin C co-transcription/read-through of adjacent SOS genes Superscript III / Herculase II fusion PCR
RT-PCR V. cholerae N16961, nrd operon mitomycin C co-transcription of nrd operon genes Access RT-PCR system (Promega), AMV reverse transcriptase
Electrophoresis mobility shift assay (EMSA) purified LexA protein and DIG-tagged DNA probes none (in vitro binding assay) LexA binding to candidate promoter regions 6% non-denaturing Tris-glycine polyacrylamide gel, DIG detection
Mitomycin C susceptibility assay V. cholerae gene deletion mutants gene knockout + mitomycin C exposure susceptibility/survival
UV survival test V. cholerae gene deletion mutants gene knockout + UV irradiation (35 J/m2) survival rate
blastn homology search against Rfam database identified ncRNAs (bioinformatic) none functional classification of ncRNAs (e.g., tmRNA, SRP, RNaseP, riboswitches) Rfam database
Key results
  • 3078 TSSs identified with average 5'UTR of 116 nt, peak distribution 16-64 nt
  • 629 ncRNAs validated, 516 of which are cis-antisense RNAs
  • 737 genes and 28 ncRNAs induced >2-fold by mitomycin C; 231 genes and 17 ncRNAs repressed >2-fold
  • Twelve genes co-induced with adjacent canonical SOS genes via transcriptional read-through
  • 151 gene start codons reassigned based on TSS location within predicted coding sequences
  • Longest 5'UTR (1403 nt) found upstream of priA, corresponding to an internal promoter 1403 nt
  • Three previously characterized ncRNAs (tarA, qrr1, qrr3) had no detectable TSS in this dataset
  • Only 2 genes (VC0507, VC1404) showed no reads across all pooled RNA-seq conditions
Key statistics
  • count 3078 TSSs (total transcription start sites annotated genome-wide)
  • count 629 ncRNAs (516 antisense) (validated non-coding RNAs in V. cholerae genome)
  • count 737 genes and 28 ncRNAs induced; 231 genes and 17 ncRNAs repressed (differentially expressed transcripts after mitomycin C treatment, >2-fold, adjusted p<0.001)
  • pvalue adjusted p-value < 0.001 (significance threshold for differential expression (DESeq2, Benjamini-Hochberg correction))
  • mean 116 nucleotides (average 5'UTR length across TSSs)
  • count 151 genes (start codons changed based on TSS mapping)
  • other 12 genes newly assigned to SOS regulon (ubiEJB, tatABC, smpA, cep, VC0091, VC1190, VC1369-1370) (expansion of SOS regulon via read-through transcription)
  • other ~20% of V. cholerae genes (proportion of genome affected by genotoxic stress response)

Statistical methods review

Model: opus

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The study used whole-transcriptome RNA-Seq to compare Vibrio cholerae with and without mitomycin C treatment (SOS vs. noSOS), with biological replicates nested within each condition. Differential expression was modeled in DESeq2 using a negative-binomial generalized linear model with a single covariate for biological condition, applying DESeq2's default dispersion estimation and Wald-type significance testing. Raw p-values were adjusted by the Benjamini-Hochberg procedure, and genes/ncRNAs with an adjusted p-value < 0.001 (and >2-fold change) were reported as differentially expressed.

Replicationbiological Sample sizetext states biological replicates were included within each condition (SOS vs. noSOS) but the exact number of replicates is not given; no power analysis described GroupsMMC-treated (SOS) vs. untreated (noSOS) V. cholerae N16961 Pairingunpaired Randomization/blindingnot stated Dispersionnone Exact p-valuesno Effect sizesyes Confidence intervalsno Multiplicity correctionBenjamini-Hochberg FDR
Statistical tests used
Test Applied to n Assumptions
DESeq2 differential expression test (negative-binomial GLM with single biological-condition covariate; default DESeq2 significance test, i.e. Wald) MMC-treated vs. untreated (SOS vs. noSOS) RNA-Seq comparison; induced/repressed genes and ncRNAs stated
Approaches that could also have been used
  • The number of biological replicates per condition is not stated numerically in the statistical methods text.
    Could also: Reporting the exact replicate count per condition (and, if available, a brief power or sensitivity consideration) alongside the analysis. — Explicit replicate numbers help readers gauge the basis of the dispersion estimates and the resolution of the differential-expression calls; this is descriptive context, not a comment on validity.
  • Differentially expressed transcripts were defined by combining an adjusted p-value threshold (<0.001) with a >2-fold change cutoff.
    Could also: Using DESeq2's built-in log-fold-change shrinkage and/or an explicit fold-change null in the Wald test (e.g., lfcThreshold) to integrate the effect-size criterion into the formal test. — Shrinkage stabilizes fold-change estimates for low-count genes and folding the fold-change threshold into the test makes the combined significance-plus-magnitude criterion a single calibrated procedure; both are standard DESeq2 options.
  • Results were summarized primarily through counts of significant genes/ncRNAs and fold-change cutoffs.
    Could also: Also reporting per-gene effect sizes with their standard errors or 95% confidence intervals on the log2 fold change. — Interval estimates convey the precision around each fold change and are often preferred for communicating uncertainty in RNA-Seq results; purely descriptive addition.
  • Validation experiments (RT-PCR, EMSA, UV and MMC susceptibility tests) are reported descriptively.
    Could also: Applying formal tests (e.g., a t-test or Mann-Whitney U on survival/susceptibility measurements) with stated replicate numbers and dispersion (SD, SEM, or CI). — Quantitative summaries with a test and a measure of spread would let readers assess the magnitude and consistency of the phenotypic differences; offered as an educational option, not a critique.
  • The DESeq2 model used a single covariate for biological condition with the default Wald significance test.
    Could also: A likelihood-ratio test comparing the full model to a reduced model, particularly if additional factors (e.g., batch or growth condition) were to be modeled. — An LRT framework accommodates multi-factor designs and explicit nested-model comparisons; with the single-covariate design here the Wald test is the natural default, and the LRT is simply an alternative formulation.
Software: R 3.2.0 · Bioconductor DESeq2 1.8.1 · Bowtie 0.12.7 · HTSeq (htseq-count) · in-house read-cleaning program (clean_ngs)

Result convergence & founder nodes

Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
54
Impact: high
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

AE003852 ENA in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
AE003853 ENA in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
PRJNA429784 BioProject in Acknowledgments (http://purl.org/orb/Acknowledgments)
no other assessed paper uses this yet
SRP139746 ENA in Acknowledgments (http://purl.org/orb/Acknowledgments)
no other assessed paper uses this yet

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-29783948

Paper: Krin et al. 2018, Expansion of the SOS regulon of Vibrio cholerae through extensive transcriptome analysis and experimental validation. BMC Genomics 19:373. PMID 29783948 · PMCID PMC5963079 · DOI 10.1186/s12864-018-4716-8.

The named code artifact

https://github.com/baj12/clean_ngs → redirects to PF2-pasteur-fr/clean_ngs (commit on master, last push 2016-01-25, public, not archived). This is not the authors' full analysis pipeline — it is a single third-party C++ tool: an adapter-removal / read-cleaning program (one step of the RNA-seq pipeline). Per BRIEF rule 2 (P16), reproducing by running the described pipeline on the paper's own data is equally valid; we do that.

The RNA-seq pipeline (from Methods → "Library sequencing" + "Statistical analysis of SOS RNA-Seq")

Exactly as described in the paper:

  1. Reads: 51-bp single-end, Illumina HiSeq 2000.
  2. Clean/trim: clean_ngs (the repo), keep reads ≥ 25 nt.
  3. Align: Bowtie v0.12.7, --chunkmbs 400 -m 50 -e 50 -a --best -q, to N16961 reference AE003852 (chr 1) + AE003853 (chr 2).
  4. Count: htseq-count -m intersection-nonempty -s yes -t gene.
  5. DE: R 3.2.0 + DESeq2 1.8.1, normalization with "shorth", design ~condition (SOS vs noSOS), Wald test, BH-adjusted; DE = padj < 0.001. Focus on genes induced > 2-fold (|log2FC| > 1).

Data (SRA PRJNA429784 / SRP139746) — sample classification

24 RNA-seq runs. The differential-expression comparison ("treated or not with 0.2 µg/mL MMC", 2 biological replicates each) uses the 4 directional ("Marine") libraries:

condition replicate run
noSOS (MMC−) A SRR6987763
noSOS (MMC−) B SRR6987761
SOS (MMC+) A SRR6987764
SOS (MMC+) B SRR6987762

The other 20 runs are 5′-end-RNA-seq TSS libraries (±TEX/±TAP) and other growth conditions (LB expo/stat, M63 minimal, MIX pools) → used in the paper for TSS/ncRNA discovery, not the SOS DE comparison.

IN SCOPE (pipeline-derived, attempted)

  • C1 "737 genes … significantly induced more than 2-fold" (Results §"Transcriptional profiling…"). → reproduce as count of genes padj<0.001 & log2FC>1.
  • C2 "231 [genes] were repressed" (same sentence). → count padj<0.001 & log2FC<−1.
  • C5+ Per-gene +/− MMC fold-changes for the known SOS regulon genes in Table 1 (e.g. recN/radB 47.0, rstA1 33.8, unfA 32.7, lexA 25.5, dinB 19.4, recA 16.4, intIA 14.4, recX 11.4, uvrA 9.9, uvrD 9.1) — compared to our DESeq2 fold-changes.

OUT OF SCOPE / hard-20% (not attempted, with reason)

  • ncRNA DE counts ("28 ncRNAs induced / 17 repressed"): require the authors' custom ncRNA annotation ("modified gene and new ncRNA positions, this work"), which is not shipped as a usable GFF → not reproducible from public annotation.
  • TSS determination, EMSA, RT-PCR, UV-survival, MMC-susceptibility, phenotypics: wet-lab / manual / 5′-end-specific → non-pipeline, out of scope.
  • Exact tool versions: Bowtie 0.12.7 and DESeq2 1.8.1 (R 3.2, 2015) are not installable from current conda channels; we use the closest obtainable Bowtie1 and a recent DESeq2, same parameters. Version drift is an expected, documented source of small numeric discrepancy — grades are provisional.

Known reproduction caveats (honesty)

  • Reference annotation differs: paper used a hand-modified N16961 annotation; we use the standard NCBI RefSeq annotation (GCF_000006745.1, same genome, VC locus tags). This shifts gene/feature counts somewhat → C1/C2 expected "within-tol/partial", not exact.
  • DESeq2 version + "shorth" normalization detail may shift the exact DE count.
Figures / tables: Table
C1
Reported
737 genes induced >2-fold (padj<0.001), MMC+ vs MMC-
Reproduced
671 (all gene features; 559 protein-coding)
within tolerance
C2
Reported
231 genes repressed >2-fold (padj<0.001)
Reproduced
263 (all gene features; 260 protein-coding)
within tolerance
C3
Reported
28 ncRNAs induced
Reproduced
NOT ATTEMPTED — needs authors' custom ncRNA annotation (out of scope)
m.public.grade.not-attempted
C4
Reported
17 ncRNAs repressed
Reproduced
NOT ATTEMPTED — needs authors' custom ncRNA annotation (out of scope)
m.public.grade.not-attempted
C5-C30
Reported
Table 1 per-gene +/- MMC fold-changes (recN 47.0, rstA1 33.8, unfA 32.7, lexA 25.5, dinB 19.4, recA 16.4, intIA 14.4, recX 11.4, uvrA 9.9, uvrD 9.1, ... rstR1/2 0.3)
Reproduced
all 26/26 genes recovered, all within-tol: recN 47.2, rstA1 36.9, unfA 33.1, lexA 25.9, dinB 20.3, recA 16.6, intIA 14.6, recX 11.8, uvrA 9.7, uvrD 9.0, ... rstR1/2 0.26 (correct down direction)
within tolerance

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 85/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟡2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +8

This is an honest partial: the raw input data resolves 1:1 (SRA PRJNA429784), the Methods fully specify the trim->Bowtie1->htseq->DESeq2 pipeline, and the run was validated end-to-end through alignment (97.98% for SRR6987763) — but the operator's finalize-now instruction stopped the job before the DE comparison, so the paper's 737 induced / 231 repressed genes and all Table-1 fold-changes (recN 47.0, lexA 25.5, recA 16.4) were never computed. The unresolved deviations are on our side (incomplete run, standard-RefSeq vs custom annotation, Bowtie/DESeq2 version drift), not the authors' — nothing here is suspicious and no value was fabricated. Because the central SOS-regulon claim was neither confirmed nor contradicted, q5/q7/q8 land at yellow rather than green or red.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

333.2 k
tokens (I/O) · 21.2 M incl. cache
171 min
runtime · 6.05 CPU-h
2 GB
peak RAM
2
HPC jobs
hummel
machine