Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Integrative Transcriptomic and Evolutionary Analysis of Drought and Heat Stress Responses in Solanum tuberosum and Solanum lycopersi

Plants (Basel) · 2025
84/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
How its reproducibility compares
84/100
Reproducibility score
0.6 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 63% of all assessed papers rank 392 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

1:1-style reproduction of the paper's described upstream pipeline on its own data. Paper is an integrative RNA-seq meta-analysis (450 samples / 21 experiments); it ships NO authors' own code (registry code = NCBI sra-tools), so per BRIEF P16 we reproduced the fully-specified third-party pipeline (SRA-tools 3.1.1 -> HISAT2 2.2.1 --no-unal --no-mixed on Ensembl Plants rel-57 SolTub_3.0/SL3.0 -> featureCounts -O -Q 10 -t exon -g Parent on GFF3) on a representative multi-dataset/multi-stress subset (n=28: 14 potato + 14 tomato, 4 of the 21 series). RESULT (n=14 each, from raw logs): C2 tomato align 91.93% vs 92% EXACT; C3 potato assigned 69.11% vs 70% EXACT; C1 potato align 73.40% vs 77% within-tol (-3.6pp; our subset is heat-PE-weighted: GSE158644 aligns ~56%); C4 tomato assigned 90.70% vs 85% partial (+5.7pp; our 2 tomato series are both high-quality HEAT data, whereas the paper averages 12 tomato experiments incl. drought that assign lower). All gaps are subset-composition effects, not pipeline mismatches; no fabrication concern (every value is derivable from the deposited data via the described pipeline). Audit note: Methods name a GTF for potato but set -g Parent (a GFF3 attribute) — only GFF3 works and reproduces 70% assigned, confirming GFF3 is the intended annotation. Dataset N checks: GSE77826=48, GSE158644=12, GSE174607=12 match Table 1; GSE151277 deposit holds 39 runs vs the 33 the paper attributes (heat-15 split unlabelled/uncheckable). NOT attempted (out of 80/20 scope): combined DEG counts, WGCNA, OrthoFinder/Ka-Ks/motifs, phylogenetics — each has unspecified parameters that would confound a model-choice mismatch with a pipeline mismatch.

💻 Code ↗ 🗄 Data: GSE77826

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 84
    assessed: 2026-06-20 ⛓ b81be6c0003e
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-20
Rubric version
not recorded
Assessed by
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

Whether an integrative meta-analysis of publicly available transcriptomic datasets, combined with comparative/evolutionary analysis across the Solanum genus, can reveal conserved regulatory mechanisms (gene co-expression networks, transcription factors) underlying drought and heat stress adaptation shared between potato (Solanum tuberosum) and tomato (Solanum lycopersicum).

Core claims
  • Drought and heat stress induce coordinated transcriptional reprogramming in potato and tomato: induction of molecular chaperone activity, oxidative stress responses, and immune signaling, with repression of photosynthetic and primary metabolic pathways reflecting energy reallocation. finding
  • bZIP, bHLH, DOF, and BBR/BPC transcription factor families are implicated as central regulators of drought- and heat-induced transcriptional programs based on promoter motif and TF enrichment analysis. finding
  • Orthogroup inference and Ka/Ks analysis across representative Solanum species show a predominance of purifying selection, indicating evolutionary conservation of regulatory network architecture. finding
  • Potato and tomato share distinct but conserved transcriptional signatures for heat and drought stress, indicative of conserved abiotic stress strategies across Solanaceae. finding
  • Integration of motif occurrence, co-expression profiles, and protein-protein interaction data enabled reconstruction of regulatory networks and identification of conserved hub transcription factors coordinating stress responses. method
  • A curated meta-analysis dataset of 450 RNA-seq samples across drought and heat stress experiments in potato and tomato was compiled from public SRA/BioProject data. resource
  • Combining combined-analysis DESeq2 results with per-experiment reproducibility filtering yields a validated high-confidence set of stress-responsive DEGs. method
Experimental setups
Assay System Perturbation Readout Platform
bulk RNA-seq meta-analysis (DESeq2, Hisat2, featureCounts) Solanum tuberosum (potato) heat stress differentially expressed genes Hisat2/featureCounts/DESeq2
bulk RNA-seq meta-analysis (DESeq2, Hisat2, featureCounts) Solanum tuberosum (potato) drought/water deficit stress differentially expressed genes Hisat2/featureCounts/DESeq2
bulk RNA-seq meta-analysis (DESeq2, Hisat2, featureCounts) Solanum lycopersicum (tomato) heat stress differentially expressed genes Hisat2/featureCounts/DESeq2
bulk RNA-seq meta-analysis (DESeq2, Hisat2, featureCounts) Solanum lycopersicum (tomato) drought/water deficit stress differentially expressed genes Hisat2/featureCounts/DESeq2
GO term functional enrichment analysis Solanum tuberosum and Solanum lycopersicum (top DEG sets) drought and heat stress overrepresented biological process GO terms among up/downregulated DEGs clusterProfiler (R v4.4.1)
de novo promoter motif discovery and TF family annotation Solanum tuberosum and Solanum lycopersicum (promoters of top 1000 DEGs per species-stress combination) drought and heat stress enriched sequence motifs mapped to TF families/binding sites bedtools, XSTREME and TOMTOM (MEME Suite)
comparative genomics / orthogroup inference and Ka/Ks selection analysis multiple Solanum species genomes (including wild potato genomes) none (evolutionary/comparative analysis) orthologous gene groups, selection pressure (Ka/Ks) on stress-responsive genes
principal component analysis of normalized expression profiles Solanum tuberosum and Solanum lycopersicum RNA-seq samples drought and heat stress (batch-corrected) clustering of biological replicates limma (removeBatchEffect)
Key results
  • 4536 significant DEGs identified in S. tuberosum under heat stress (combined DESeq2 analysis) 4536 genes
  • 5484 significant DEGs identified in S. tuberosum under drought stress 5484 genes
  • 5527 significant DEGs identified in S. lycopersicum under heat stress 5527 genes
  • 440 DEGs identified in S. lycopersicum under drought stress 440 genes
  • 305 significant motif-set associations retained after q<0.05 filtering, comprising 119 unique motifs classified into 14 TF families, with bHLH and bZIP most represented 305 associations; 119 motifs; 14 TF families
  • Upregulated DEGs enriched for protein folding/chaperone activity, oxidative stress response, and immune/defense signaling; downregulated DEGs enriched for photosynthesis and primary metabolism
  • Ka/Ks analysis across Solanum species shows predominance of purifying selection in orthogroups
  • Average RNA-seq alignment rate (Hisat2) was 92% for tomato and 77% for potato; mean assigned-read proportion (featureCounts) was 85% for tomato and 70% for potato 92%/77% alignment; 85%/70% assignment
Key statistics
  • count 4536 (significant DEGs in S. tuberosum under heat stress (combined analysis))
  • count 5484 (significant DEGs in S. tuberosum under drought stress (combined analysis))
  • count 5527 (significant DEGs in S. lycopersicum under heat stress (combined analysis))
  • count 440 (significant DEGs in S. lycopersicum under drought stress (combined analysis))
  • count 450 (total RNA-seq samples curated across all species-stress datasets)
  • mean 92% (tomato), 77% (potato) (average Hisat2 alignment rate by species)
  • mean 85% (tomato), 70% (potato) (mean proportion of reads assigned by featureCounts by species)
  • pvalue q-value < 0.05 (significance threshold for motif-set association filtering, yielding 305 associations and 119 unique motifs)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The study performed an integrative meta-analysis of 450 publicly available RNA-seq samples from potato and tomato under drought and heat stress using a dual-track differential expression strategy: a combined-dataset DESeq2 analysis and parallel per-experiment analyses, with the final validated DEG set defined by their intersection. Functional enrichment (GO over-representation), promoter motif discovery (XSTREME/TOMTOM), co-expression profiling, orthogroup inference, and Ka/Ks evolutionary analysis were then applied to prioritized DEG sets, with motif significance controlled by q-value < 0.05.

Replicationbiological Sample size450 RNA-seq samples from 21 independent public experiments (3 potato heat, 6 potato drought, 6 tomato heat, 6 tomato drought); no formal prospective power analysis described GroupsStressed vs. control within each species-stress combination (potato/tomato × drought/heat); cross-species and cross-stress comparisons of DEG overlap Pairingunpaired Randomization/blindingnot stated Dispersionnone Exact p-valuesno Effect sizesyes Confidence intervalsno Multiplicity correctionBenjamini-Hochberg FDR (DESeq2 default) for DEGs; q-value < 0.05 (FDR-based) for motif enrichment via XSTREME
Statistical tests used
Test Applied to n Assumptions
DESeq2 (negative binomial Wald test with Benjamini-Hochberg FDR correction, default) Combined differential expression analysis across all samples for each species-stress combination (potato heat, potato drought, tomato heat, tomato drought) 450 total RNA-seq samples (48 potato heat, 184 potato drought, 143 tomato heat, 75 tomato drought) not stated
Per-experiment DESeq2 pairwise comparisons with frequency-based occurrence ranking Individual-experiment DEG analysis used to identify reproducible stress-responsive genes; thresholds: ≥10/17 comparisons (tomato heat), ≥6/8 (potato heat), ≥13/18 (potato drought) Number of pairwise comparisons per condition stated (8 potato heat, 18 potato drought, 17 tomato heat); per-experiment n not stated in extracted text not stated
Gene Ontology over-representation analysis (clusterProfiler) Functional enrichment of top DEG sets (up- and downregulated) for each species-stress combination Top 1000 DEGs per species-stress combination (500 up, 500 down); all 440 DEGs for tomato drought not stated
De novo motif enrichment (XSTREME, MEME Suite) with q-value < 0.05 threshold, followed by TOMTOM motif-to-TF annotation Promoter regions of top 1000 DEGs per species-stress group (500 up, 500 down) 1000 DEG promoters per group (440 for tomato drought) not stated
Ka/Ks (dN/dS) ratio analysis Evolutionary selection analysis across orthogroups of representative Solanum species null not stated
Principal Component Analysis (PCA) Quality control visualization of normalized expression profiles after batch-effect correction, per species and stress condition 450 total samples na
Approaches that could also have been used
  • Batch-effect correction was applied to normalized expression values using removeBatchEffect from the limma package (designed for continuous microarray-like data) prior to PCA visualization, while downstream DEG testing was performed on raw counts with DESeq2
    Could also: ComBat-seq, which operates directly on raw count matrices and accounts for the discrete negative-binomial distribution of RNA-seq data, could also be used for count-level batch adjustment before combined DESeq2 analysis — removeBatchEffect was designed for continuous data; ComBat-seq preserves integer count properties and the variance structure relevant to negative-binomial modeling, which may improve concordance between the corrected visualization and the count-based statistical test
  • Cross-experiment integration was implemented as a frequency-based occurrence filter — genes were retained if detected as DEGs in a defined threshold fraction of pairwise comparisons
    Could also: Formal meta-analysis frameworks such as Fisher's combined p-value method, the inverse-variance weighted approach, or random-effects meta-regression (e.g., metafor in R) could also aggregate evidence across experiments — Frequency-based voting is transparent and robust to outlier experiments, but formal meta-analytic methods provide a single pooled effect estimate with confidence interval and a statistically controlled combined false discovery rate, enabling direct comparison of effect magnitudes across conditions
  • Functional enrichment was performed as GO over-representation analysis (ORA) on a fixed list of top DEGs (up to 1000 per condition) using clusterProfiler
    Could also: Gene Set Enrichment Analysis (GSEA) applied to the full ranked list of expressed genes ordered by DESeq2 LFC or Wald statistic could also be used — GSEA uses the complete ranked gene list without an arbitrary count cutoff and can detect coordinated but moderate shifts in pathway activity that may fall below an ORA membership threshold; it also avoids the sensitivity of ORA results to the chosen top-N cutoff
  • The top 1000 DEGs per species-stress group were selected by prioritizing occurrence frequency across experiments and then ranking within each occurrence bin by absolute log2 fold change
    Could also: A continuous composite scoring approach (e.g., the pi-value combining adjusted p-value and LFC, or using all statistically significant DEGs without a fixed count cap) could also define the input set for downstream analyses — Fixed-count selection (500 per direction) introduces an arbitrary boundary; a data-driven criterion makes the cutoff scale with the number of significant findings per dataset and avoids differential exclusion of genes across conditions of different statistical power
  • Ka/Ks ratios were computed across orthogroups to assess the mode of evolutionary selection acting on stress-responsive genes across Solanum species
    Could also: Branch-site models implemented in PAML or episodic diversifying selection tests in HyPhy (BUSTED, aBSREL) could also be applied to test for selection on specific lineages within the orthogroup phylogeny — A single pairwise or orthogroup-averaged Ka/Ks value can obscure lineage-specific accelerations; branch-site models provide formal statistical tests for positive selection on particular branches and can identify species-specific evolutionary events relevant to stress adaptation
  • edgeR or limma-voom are established alternatives for negative-binomial RNA-seq differential expression, while DESeq2 was the sole method used here
    Could also: Requiring DEGs to be significant by two independent methods (e.g., DESeq2 and edgeR) — a consensus DEG approach — could also be applied to the combined dataset — Each method makes different dispersion estimation assumptions; requiring concordance between two methods is a commonly used strategy in large meta-analyses to further reduce false positives beyond what a single method's FDR threshold provides
Software: DESeq2 · R/limma (removeBatchEffect) · R/clusterProfiler 4.4.1 · MEME Suite (XSTREME, TOMTOM) · HISAT2 · featureCounts (Subread) · bedtools

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-41470732

Paper: Bondar EI, Zubairova US, Bobrovskikh AV, Doroshkov AV. Integrative Transcriptomic and Evolutionary Analysis of Drought and Heat Stress Responses in Solanum tuberosum and Solanum lycopersicum. Plants (Basel) 2025. DOI 10.3390/plants14243851 · PMC12736803.

Nature of the study

An integrative meta-analysis that re-processes publicly archived RNA-seq datasets (GEO/SRA/ENA, 450 samples / 21 experiments, Table 1) through one uniform pipeline. There is no authors' own code repository; the registry "code" link is the generic third-party tool NCBI sra-tools (used for data retrieval). Per BRIEF rule 2 (P16), applying the described third-party pipeline to the paper's own data is a fully valid reproduction.

Pipeline (Methods, fully specified)

  1. Retrieval: SRA Toolkit 3.1.1 (prefetch, fasterq-dump)
  2. QC: FastQC v0.11.9, MultiQC v1.27 (no trimming — raw quality adequate)
  3. Alignment: HISAT2 v2.2.1, params --no-unal --no-mixed
    • Potato ref: Ensembl Plants rel-57 Solanum_tuberosum.SolTub_3.0 (.dna_sm.toplevel)
    • Tomato ref: Ensembl Plants rel-57 Solanum_lycopersicum.SL3.0 (.dna.toplevel)
  4. Quantification: featureCounts (subread), -O -Q 10 -g Parent (GFF3 rel-57)
  5. DE: DESeq2, ashr LFC shrinkage, s-value < 0.005, |log2FC| ≥ 0.32

Downstream: WGCNA networks, OrthoFinder, Ka/Ks_Calculator, MEME/XSTREME motifs.

IN SCOPE (reproduced) — the cleanly-specified pipeline numbers

The paper states verbatim: "the average alignment rate obtained with Hisat2 was 92% for tomato and 77% for potato. The mean proportion of successfully assigned reads, as calculated by featureCounts, reached 85% for tomato and 70% for potato." These are the four cleanest 1:1 numbers the paper reports for the upstream pipeline and they are exactly what an independent re-run regenerates.

id claim reported how we test
C1 HISAT2 align rate, potato 77% mean over 14 potato samples (GSE77826 drought SE + GSE158644 heat PE)
C2 HISAT2 align rate, tomato 92% mean over 14 tomato samples (GSE151277 + GSE174607, heat PE)
C3 featureCounts assigned %, potato 70% -O -Q 10 -g Parent, per-sample assigned% averaged
C4 featureCounts assigned %, tomato 85% -O -Q 10 -g Parent, per-sample assigned% averaged

Design note (vs. the earlier n=3 attempt): the reported rates are averages across many datasets/stresses. To make the reproduced mean a like-for-like cross-dataset average — not a single cherry-picked series — we process a representative multi-dataset, multi-stress subset (n=28: 14 potato + 14 tomato) spanning both a drought and a heat study per species, taking the first-N runs in accession order (unbiased). A subset of the full 450 cannot reproduce the global mean to the decimal, but it tests whether the described pipeline lands in the reported regime.

OUT OF SCOPE — the hard ~20% (NOT attempted; BRIEF rule 3)

  • Combined per-species/stress DEG counts (potato heat 4536, potato drought 5484, tomato heat 5527, tomato drought 440). Each combines 3–9 independent experiments with an unspecified batch/design model in DESeq2; the exact combine/covariate structure is not given → a model-choice mismatch would be confounded with a pipeline mismatch. Skipped honestly.
  • WGCNA gene-regulatory networks (gene/edge counts) — many free thresholds (0.4–0.9, soft power) not pinned by the text.
  • OrthoFinder orthogroups, Ka/Ks, XSTREME/TOMTOM motifs, 11-genome phylogenetics — downstream of the above, out of the 80/20 budget.
  • All wet-lab / manual / interpretive statements — out of scope by definition.

Datasets we process (representative subset, by design — NOT all 450)

  • Potato drought: GSE77826 / SRP069961 (single-end) — 8 runs (first-8 in order).
  • Potato heat: GSE158644 / SRP285602 (paired-end) — 6 runs.
Figures / tables: Fig 1ATable
C1
Reported
HISAT2 average alignment rate, potato = 77% (Results 2.1, Fig 1A / Suppl Table S2)
Reproduced
73.40% (mean n=14: 8 drought-SE GSE77826 mean 86.06% + 6 heat-PE GSE158644 mean 56.53%)
within tolerance
C2
Reported
HISAT2 average alignment rate, tomato = 92% (Results 2.1, Fig 1A / Suppl Table S2)
Reproduced
91.93% (mean n=14: 8 GSE151277 mean ~88.3% + 6 GSE174607 mean 96.73%)
exact
C3
Reported
featureCounts mean assigned reads, potato = 70% (Results 2.1, Fig 1A / Suppl Table S2)
Reproduced
69.11% (mean n=14)
exact
C4
Reported
featureCounts mean assigned reads, tomato = 85% (Results 2.1, Fig 1A / Suppl Table S2)
Reproduced
90.70% (mean n=14; both tomato HEAT series assign 89.8-91.8%)
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

No assessment has been recorded yet.
🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

981.1 k
tokens (I/O) · 69.2 M incl. cache
508 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.