Plant Soft Rot Development and Regulation from the Viewpoint of Transcriptomic Profiling.
Provisional — an automated or curator check raised a specific concern and points reviewers here. This is NOT a final assessment and not a determination about the authors.
The main result did not reproduce in this reproduction attempt. Where our recomputation produced values that differ from the published ones, those discrepancies are listed below. This is a single automated attempt — not peer review and not a finding of error or misconduct — and differences can also arise from data access, undocumented parameters or the computing environment. The verdict can be contested via “report an error”.
- Nothing in this column.
- 🔴Could not use the authors’ exact input data
- 🔴Reported values were only indirectly comparable
- 🟡A deviation arose in the data or preprocessing
- 🔴A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The deviation was non-trivial in magnitude
- 🟡The central claim did not (fully) hold under reproduction
- 🔴Overall, the reproduction showed a material discrepancy
This paper has a computational component, but its primary data is legally or ethically access-restricted — identifiable patient cohorts, rare-disease genomes, or controlled-access biobanks that cannot be openly shared. The reproduction therefore could not be attempted. That is a neutral verdict: it does not mean the result is wrong or that the authors fell short — only that, for legitimate privacy reasons, it cannot be independently checked from public data. We deliberately do NOT assign a 0–100 score here, because a low number would wrongly read as a failed reproduction.
▸Reproduction agent’s raw note
DROP / data_unavailable. The paper is described well enough and the analysis code (the authors' own Promoter-analysis R tool) is public, readable, and verified functional -- I confirmed the Table-1 statistic (hypergeometric phyper upper-tail + Benjamini-Hochberg FDR<0.05) is correctly implemented by running my faithful port on a synthetic input (the deliberately DEG-enriched regulon is flagged at FDR<0.05). The raw reads (SRA PRJNA636273 = 4 N. tabacum RNA-seq runs, matching the stated 2 reps x 2 conditions) also resolve and were profiled. BUT the reproduction itself is impossible: every derived artifact the scripts consume is hosted ONLY on the first author's personal MEGA links and ALL of them are dead (9/9 -> MEGA -9 ENOENT), including the very output that IS Table 1 (DEG_enriched_regulons.tsv). One required input (tf_names_families_edited.tsv, a manual TF re-classification) was never deposited at all, and the EvidentialGene reference transcriptome behind the 8636-DEG figure is not deposited and is under-specified -- so even a from-scratch rebuild could not faithfully match. No fabrication signal: the method is standard and the implementation is correct; the values are simply uncheckable today because the data was withdrawn. NOT attempted: full RNA-seq re-run (kallisto/edgeR on a rebuilt transcriptome), MAST CRE re-scan, and Fig 8 Fisher analysis -- all blocked by the same missing data. Compute note: conda env builds failed on «our HPC» front1 (tmpfs /tmp and /dev/shm are 50 MB and 100% full) and the «infra» project area hit a quota; I worked around it by downloading+decrypting MEGA in pure R (openssl/httr in the existing paper_figures env), which is how the dead-link evidence was obtained.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessmentassessed: 2026-06-18 ⛓ bc349740897e
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-18
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-18no human curator yet
- Last updated
- 2026-07-31
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusThe study tests whether Pectobacterium atrosepticum (Pba) infection induces extensive, regulated reprogramming of the host plant transcriptome—including susceptibility responses and master-regulator transcription factors—rather than acting purely by brute-force necrotrophy, using tobacco soft rot as a model.
- ★ Pba infection of tobacco causes large-scale differential gene expression (8636 DEGs), reflecting profound host physiological reprogramming. finding
- ★ Infection induces both defense (lignification, chitinases) and susceptibility (cell-wall-loosening XTHs, EXLB expansins, RG lyases, beta-galactosidases, polygalacturonases) sides of cell wall modification. finding
- ★ Upregulated host pectin-modifying enzymes (RG lyases, beta-galactosidases, polygalacturonases) likely act as susceptibility factors releasing pre-synthesized rhamnogalacturonan I from the cell wall to serve as a matrix for bacterial emboli. mechanism
- ★ Ethylene- and jasmonic-acid/lipoxygenase-pathway hormonal responses are strongly induced, while salicylic-acid-mediated responses are not induced or are repressed during infection. finding
- ★ An original functional gene classification merging MapMan, KEGG, CAZy, and SwissProt with manual curation enables deeper physiological interpretation of RNA-Seq data. method
- ★ A genome-wide regulon-prediction pipeline based on cis-regulatory elements in promoters identifies candidate master-regulator transcription factors of the infection and explains differential regulation within multigene families. method
- Induction of ethylene- and lipoxygenase-pathway genes is also observed in Pba-infected potato, indicating a typical hallmark of Pba disease across hosts. finding
- Cellulose, cross-linking glycan, and callose biosynthetic genes are broadly repressed during infection. finding
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| bulk RNA-Seq | tobacco plants (Nicotiana) | Pectobacterium atrosepticum (Pba) infection vs uninfected | transcript abundance / differentially expressed genes (TPM, log2FC) | — |
| qPCR (validation) | tobacco plants | Pba infection vs uninfected | gene expression validation correlated with RNA-Seq | — |
| RNA-Seq / expression analysis | potato plants | Pba infection | induction of ethylene- and lipoxygenase-pathway genes | — |
| enzyme activity assay (lipoxygenase) | tobacco plants | Pba infection vs control | lipoxygenase enzymatic activity | — |
| enzyme activity assay (divinyl ether synthase) | tobacco plants | Pba infection vs control | divinyl ether synthase enzymatic activity | — |
| in silico transcription factor regulon prediction | whole-genome promoter sequences | none | enrichment of DEGs within TF regulons based on cis-regulatory elements | — |
- – 8636 DEGs in Pba-infected vs uninfected tobacco (3797 up, 4839 down) 8636 DEGs
- – RNA-Seq results correlated with qPCR validation data r=0.92
- ▲ Lipoxygenase activity higher in infected than control plants 28-fold
- ▲ Divinyl ether synthase activity higher in infected than control plants 24-fold
- ▲ RG lyases and beta-galactosidases strongly upregulated during infection log2FC ≈ 13
- ▲ DES genes dramatically upregulated in JA/oxylipin pathway log2FC up to 15
- ▲ Lignin peroxidase genes pronouncedly upregulated log2FC 3 to 10
- ▼ Arabinogalactan protein DEGs almost all downregulated (37 of 39) 37 of 39
- correlation r = 0.92 (Pearson correlation between RNA-Seq and qPCR)
- count 8636 (total DEGs (3797 up, 4839 down))
- fold_change 28-times (lipoxygenase activity higher in infected vs control)
- fold_change 24-times (divinyl ether synthase activity higher in infected vs control)
- fold_change log2FC ≈ 13 (RG lyases and beta-galactosidases upregulation)
- fold_change log2FC up to 15 (DES genes and some lipoxygenase-pathway genes)
- fold_change log2FC 3 to 10 (lignin peroxidase gene upregulation)
- count 620 (DEGs attributed to Cell wall category)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
The study used RNA-Seq to profile transcriptome changes in Pba-infected versus uninfected tobacco plants, identifying 8636 DEGs reported as log2 fold-change values. Sample replicate consistency was assessed by hierarchical clustering of Euclidean distances, and RNA-Seq results were cross-validated against qPCR with Pearson's r. The specific statistical test and software used for DEG calling, the significance/FDR thresholds applied, and the exact per-group sample sizes are not stated in the available text. Enzyme activities are reported as fold-changes without an accompanying formal test.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Differential expression analysis (specific method not stated in available text) | Identification of 8636 DEGs (3797 upregulated, 4839 downregulated) between Pba-infected and uninfected tobacco | — | not stated |
| Hierarchical clustering of sample-to-sample Euclidean distances | QC assessment of within-group replicate consistency (Supplementary Figure S1) | — | na |
| Pearson's correlation coefficient | Validation of RNA-Seq results against qPCR data (r = 0.92, Figure 1) | — | not stated |
| Fold-change comparison (formal statistical test not stated) | Lipoxygenase and divinyl ether synthase enzyme activities in infected vs. control plants (28× and 24× higher, respectively) | — | not stated |
-
The statistical method used to identify 8636 DEGs between infected and uninfected plants is not specified in the available text↳ Could also: DESeq2 (negative-binomial Wald test), edgeR (exact test or quasi-likelihood GLM), or limma-voom could also be used for RNA-Seq DEG calling with biological replicates — These tools explicitly model count overdispersion, apply shrinkage estimation of dispersion, and output per-gene p-values with FDR-corrected q-values, providing a reproducible significance threshold and facilitating comparison with other studies
-
No multiplicity correction method or FDR threshold is stated for the DEG selection↳ Could also: Benjamini-Hochberg FDR (e.g., q ≤ 0.05) combined with a log2FC threshold (e.g., |log2FC| ≥ 1) is a standard approach in RNA-Seq studies — A combined FDR and fold-change filter reduces both false positives from low-expression noise and genes that are statistically significant but negligibly small in magnitude, and stating the thresholds allows readers to gauge list composition independently
-
RNA-Seq and qPCR agreement was assessed with Pearson's correlation (r = 0.92)↳ Could also: Spearman's rank correlation could also quantify this agreement — Spearman's rho does not assume linearity or normally distributed residuals, which is relevant when fold-change values span several orders of magnitude or when a small number of validation genes is used
-
Enzyme activities in infected versus control plants are reported as fold-changes (28× and 24×) without a formal statistical test or dispersion measure↳ Could also: A two-sample t-test or Mann-Whitney U test with reported n, mean or median, and SD or IQR would also formalize this comparison — Formal testing with dispersion and sample size allows readers to assess the reliability and variability of the observed difference independently of the magnitude of the fold-change
-
Sample replicate consistency was visualized by hierarchical clustering of Euclidean distances↳ Could also: Principal component analysis (PCA) or multidimensional scaling (MDS) of normalized count data could also serve this QC purpose — PCA and MDS provide a complementary low-dimensional view of between-sample variance that can simultaneously reveal batch effects, outliers, and the proportion of total variance attributable to the treatment factor
-
Functional enrichment of DEGs was assessed by sorting genes into manually curated categories (MapMan, KEGG, CAZy, SwissProt)↳ Could also: Formal gene-set enrichment analysis (GSEA) or over-representation analysis (ORA) with an FDR-corrected hypergeometric or Fisher's exact test could also be applied — Statistical enrichment testing quantifies whether a functional category contains more DEGs than expected by chance, providing a probability estimate alongside the descriptive gene count and enabling ranked prioritization of pathways
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Reproduction scope — pmid-32927917
Paper: Tsers et al. 2020, Plants 9(9):1176. "Plant Soft Rot Development and Regulation from the Viewpoint of Transcriptomic Profiling." (PMID 32927917)
Code (authors' own — first author Ivan Tsers):
https://github.com/IvanTsers/Promoter-analysis @ commit
5f46ff4771ec5e0d345ce1de1aca46436162b559 (2020-09-18, repo HEAD).
R (97.8%) + Shell. No pinned release.
Data: SRA PRJNA636273 (4 N. tabacum RNA-seq runs). Derived analysis tables hosted on personal MEGA links referenced in the README.
Pipeline map (what produces each reported result)
| Reported result | Pipeline | In scope? |
|---|---|---|
| Total DEGs = 8636 (3797 up / 4839 down), |log2FC|>1 & FDR<0.05 | FastQC→Trimmomatic→SortMeRNA→kallisto (vs EvidentialGene-reduced CDS)→edgeR | Partial / blocked — reference transcriptome (EvidentialGene CDS) NOT deposited; cannot regenerate exact DEG counts. Profile the authors' example DGE table instead. |
| CRE site counts (2,273,233 sites / 1340 variants; 405,061 in DEG promoters) | MAST (MEME) with PlantPAN 3.0 PWMs on 1000-bp promoters | Stretch — heavy (1.34 GB MAST output); deterministic given Promoters.fa + PWMs. Attempt only after core. |
| Table 1: 7 enriched TF regulons (WRKY6/42/45/51/57, TCP3/15) with Total_DEG, Total_nonDEG, FDR | TF_regulons_enrichment_analysis.R: per-TF hypergeometric (phyper) upper tail + BH-FDR, filter FDR<0.05, on tf_analysis_input_annotated.tsv |
PRIMARY (in scope) — fully deterministic given the provided annotated table. This is the core "master-regulator prediction" claim. |
| Fig 8: TF families driving up/down split within multigene families (chitinases, E3 ligases, calmodulins, polygalacturonases) | TF_family_regulons_correlation_analysis.R: Fisher's exact on +TF/−TF up/down counts |
Secondary (in scope if Genes_of_interest.tsv obtainable). |
| Wet-lab: qPCR validation (Pearson r=0.92), plant phenotyping, microscopy | manual / wet-lab | Out of scope (not pipeline-derived). |
| GO / functional category counts (cell wall 620, signaling 1324, TF-encoding 676, ...) | manual curation on top of DEG table | Out of scope (depends on un-reproducible DEG set + manual annotation). |
Strategy
- Primary 1:1: download the provided
tf_analysis_input_annotated.tsv+ the referenceDEG_enriched_regulons.tsv(174 B) from MEGA; run the authors' enrichment script (faithful port, identical statistic); compare reproduced regulons against (a) the shipped reference output and (b) paper Table 1. - Dataset profiling: PRJNA636273 (4 runs) + the MEGA-hosted derived tables, incl. profiling the example DGE table for DEG count vs the reported 8636.
- Stretch (MAST CRE counts, Fig 8 Fisher) only if core succeeds.
Known blockers (honest)
- Reference transcriptome (EvidentialGene-reduced CDS) not deposited → 8636-DEG count not regenerable from public data.
tf_names_families_edited.tsv(manual TF re-classification) referenced by the Annotate step is not provided anywhere → the annotation step cannot be re-run from scratch; only the final enrichment (on the provided annotated table) is fully reproducible.- All derived data is on personal MEGA links (no DOI/accession) → link-rot risk.
- Compute: «infra» project hit an inode/space quota for conda env builds → env + downloads + run done on compute-node local scratch, only small results copied back.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.