Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Comprehensive Transcriptome Analysis Reveals Genome-Wide Changes Associated with Endoplasmic Reticulum (ER) Stress in Potato (Solanum tuberosum L.).

Int J Mol Sci · 2022
L1 62/100 3/4
Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
What did not (or only partly)
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
62/100
Reproducibility score
0.7 SD below mean
vs. all fields · 1173 studies
🎯 Scores higher than 23% of all assessed papers rank 891 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

PARTIAL reproduction (honest 1:1 attempt, well-documented forced deviation). Potato ER-stress RNA-seq (tunicamycin TM vs Mock, 2h/5h, 3 reps = 12 libs, PRJNA865435), authors' own repo (HISAT2 2.2.1 -> htseq-count 2.0.1 intersection-strict -> edgeR exactTest). KEY DATA FINDING: the SRA deposit is DE-PAIRED single-end - paper states PE150 but each pair is stored as 2 single-end spots with an empty mate (sra-stat: 75.6M spots, 11.3Gbp = 37.8M pairs x2; fastq-dump --split-files yields only read-2). All bases present, pairing irrecoverable -> FORCED single-end alignment (htseq --stranded no; strand probe confirms de-paired pool is 50/50 sense/antisense). Reference Castle Russet v2.0 (SpudDB), authors' _clean GTF not shipped so GFF3->GTF via gffread. RESULTS: alignment reproduces cleanly (overall 94.06-94.73% vs >94%; 37-38M pairs-equiv reads). DEG counts in same magnitude AND same up/down asymmetry direction but ~20-40% higher (2h 248/370 vs 204/278; 5h 197/180 vs 157/129; total 966 vs 806) - attributable to single-end+stranded-no+converted-GTF. Cross-timepoint overlap reproduced closely (8/8 vs paper 9/10). NOT attempted: isoform-switching (optional). Also flagged: shipped edgeR script uses decideTestsDGE lfc=0 (FDR-only) not the |log2FC|>=1 the paper text states; we report both. Honest verdict: pipeline + reported magnitudes are reproducible; exact DEG counts are not, primarily due to a real SRA deposit defect (de-pairing) outside our control.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 50
    assessed: 2026-06-19 ⛓ 865ef64a05c6
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-22
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The study tests whether treating potato (Solanum tuberosum) leaves with tunicamycin (TM) to induce ER stress reveals genome-wide transcriptional and post-transcriptional changes underlying the unfolded protein response (UPR), including responses not previously reported in Arabidopsis.

Core claims
  • TM treatment of potato leaves produces widespread, time-dependent differential gene expression associated with ER stress and the UPR finding
  • Chromatin remodeling and transcriptional reprogramming (histone methyltransferases, SWI/SNF components, NAP proteins, histone acetyltransferase complex, DNA methylation/demethylation factors) occur as an early ER stress response finding
  • Limited genome-wide changes in alternative RNA splicing/isoform usage of protein-coding transcripts occur following TM treatment finding
  • RNA metabolism, translation machinery components, and protein folding/maturation factors are extensively affected by TM treatment, involving a broader set of genes than reported in Arabidopsis finding
  • Antioxidant defense and oxygen metabolic enzymes are differentially regulated, consistent with oxidative stress during ER stress finding
  • Surges in protein kinase gene expression indicate early signal transduction events during ER stress finding
  • Potato has an expanded set of ER stress-responsive StbZIP genes (StbZIP17/28/33/60/67/70/71) relative to Arabidopsis orthologs, suggesting unique stress-adaptation factors resource
  • At least 39 differentially expressed transcripts contribute to structure/function of the ER, Golgi, endocytic, and vacuolar networks finding
Experimental setups
Assay System Perturbation Readout Platform
RNA-seq (BGI platform) Potato (Solanum tuberosum cv. Russet Norkotah) leaves Tunicamycin (TM) treatment vs DMSO (solvent) control Differentially expressed genes (DEGs) at 2 h and 5 h post-treatment, log2 fold change, adjusted p-value, FDR BGI RNA-seq; reads aligned to Castle Russet potato genome
Gene ontology (GO) annotation and functional enrichment analysis Potato leaf transcriptome DEG dataset (TM vs mock, 2 h and 5 h) TM vs DMSO GO term distribution and enrichment (Biological Process, Molecular Function, Cellular Component) via Fisher's exact test BLAST2GO tool within OmicsBox
Alternative splicing / transcript isoform usage analysis Potato leaf RNA-seq transcriptome TM vs DMSO Loci showing shifts in RNA isoform usage (alternative TSS/termination) at 2 h and 5 h
Key results
  • TM treatment produced 806 unique DEGs total across 2 h and 5 h (fold change threshold 1.0, p<0.05) 806 genes
  • At 2 h: 204 uniquely upregulated and 278 uniquely downregulated genes 204 up / 278 down
  • At 5 h: 157 uniquely upregulated and 129 uniquely downregulated genes 157 up / 129 down
  • Overlap classes: 9 genes up at both 2&5h, 9 down-to-up (2h to 5h), 9 up-to-down (2h to 5h), 10 down at both timepoints 9/9/9/10 genes
  • RNA-seq reads aligned to Castle Russet potato genome with high overall alignment rate; majority uniquely mapped >94% overall alignment; 56-62% unique
  • Enrichment for endomembrane system (ER, Golgi) increased from 2 h to 5 h post-TM treatment ER-associated sequences ~2-7%; Golgi elevated to 4% at 5h
  • 9 loci at 2 h and 15 loci at 5 h showed changes in RNA isoform usage, largely without overall change in gene expression 9 loci (2h), 15 loci (5h)
  • At least 28 factors in mRNA synthetic processes, 21 in rRNA processing/ribosome biogenesis, and 32 in tRNA modification/translation were differentially expressed 28/21/32 factors
Key statistics
  • count 806 unique DEGs (Total DEGs across 2h and 5h TM treatment vs DMSO)
  • pvalue p < 0.05 (Significance threshold for DEG calling, with fold change threshold of 1.0)
  • count 37-38 million quality read pairs (RNA-seq sequencing depth per sample)
  • other >94% overall alignment rate; 21.3-22.8 million (56-62%) uniquely aligned read pairs (RNA-seq alignment statistics to Castle Russet genome)
  • other 14.1 million (39%) to 15.8 million (43%) read pairs returned multiple hits (RNA-seq multi-mapping reads)
  • other 703,240 (1.9%) to 886,837 (2.4%) read pairs unaligned (RNA-seq unmapped reads)
  • count 204 up / 278 down (2h); 157 up / 129 down (5h) (Uniquely regulated DEGs per timepoint)
  • count 39 transcripts (Genes contributing to ER, Golgi, endocytic, and vacuolar network structure/function)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The study used RNA-seq (BGI platform) to compare gene expression in tunicamycin (TM)-treated versus DMSO (mock)-treated potato leaves at 2 and 5 hours, identifying differentially expressed genes (DEGs) using a fold-change threshold of 1.0 and p < 0.05, with adjusted p-values and false discovery rate (FDR) reported in supplementary tables. Gene ontology enrichment of the DEG sets was performed using Blast2GO (within OmicsBox) with Fisher's exact test to assess enrichment of cellular component/biological process/molecular function categories. Results were visualized with volcano plots, Venn-style overlap counts, GO distribution charts, and word clouds; the provided text does not include a separate Materials and Methods statistics subsection detailing replicate numbers or the specific DE-calling algorithm.

Replicationunclear Sample sizenot stated in the provided text (read pair counts and alignment rates are given, but number of biological or technical replicates per condition is not stated) GroupsTM-treated vs DMSO (solvent-only) mock-treated potato leaves, each at 2 h and 5 h post-treatment Pairingunclear Randomization/blindingnot stated Dispersionunclear Effect sizesyes Multiplicity correctionFalse discovery rate (FDR) correction referenced for the DEG tables; the specific procedure (e.g., Benjamini-Hochberg) is not named in the provided text
Statistical tests used
Test Applied to n Assumptions
Differential gene expression analysis (specific statistical model/tool not named in provided text, e.g., count-based RNA-seq DE test) TM- vs DMSO-treated leaf transcriptomes at 2 h and 5 h (Figure 1A,B; Tables S2-S4) not stated not stated
Fisher's exact test (via Blast2GO enrichment analysis) GO term/cellular component enrichment comparing TM-treated vs mock datasets (Figure 3A,C,D) not stated not stated
Approaches that could also have been used
  • Differentially expressed genes were called using a fold-change threshold (1.0) combined with p < 0.05, without naming the underlying statistical model for the RNA-seq counts.
    Could also: A count-based differential expression model such as DESeq2 (Wald or likelihood-ratio test) or edgeR (negative binomial exact test/GLM), explicitly reporting biological replicate number and dispersion estimation — Naming the statistical model and its assumptions (e.g., negative binomial mean-variance modeling) makes explicit how significance was derived from raw counts and allows others to reproduce or reanalyze the DEG calls.
  • FDR correction is referenced for the DEG p-values, but the specific multiple-testing procedure is not stated in the provided text.
    Could also: Explicitly specifying the correction method, such as the Benjamini-Hochberg procedure — Naming the exact procedure clarifies the family-wise scope of the correction (e.g., per-comparison vs. genome-wide) and lets readers directly compare thresholds across studies.
  • GO enrichment of DEGs was assessed with Fisher's exact test via Blast2GO.
    Could also: Gene-length-aware enrichment tools such as GOseq, or hierarchy-aware tools such as topGO (elim/weight algorithms) — These approaches can additionally account for transcript-length sequencing bias or the nested structure of GO terms, which can refine which categories appear most enriched.
  • DEGs were defined using a fixed fold-change cutoff (1.0) together with a nominal p-value threshold (p < 0.05).
    Could also: Ranking/selecting genes primarily by the adjusted p-value (FDR) with fold-change reported alongside, or using shrinkage-based effect-size estimates (e.g., DESeq2's apeglm/ashr shrinkage) — This can reduce reliance on an unadjusted p-value threshold for calling significance and can stabilize fold-change estimates for lower-count genes.
  • Comparisons were made independently at 2 h and 5 h, with genes then categorized post hoc by their pattern across the two time points (Figure 1B).
    Could also: A joint time-course differential expression framework (e.g., maSigPro, ImpulseDE2, or a time-by-treatment interaction term in a single model) — Modeling both time points jointly can directly test for time-by-treatment interaction and may increase power to detect temporally dynamic expression patterns.
  • The text does not state the number of biological or technical replicates used for the RNA-seq comparisons.
    Could also: Reporting the replicate number and, where feasible, a power/sample-size justification — Stating replicate numbers helps readers assess the precision of the fold-change and p-value estimates for the reported DEGs.
Software: BGI RNA-seq platform/pipeline · Blast2GO (within OmicsBox)

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-36430273

Paper: Herath V, Verchot J. "Comprehensive Transcriptome Analysis Reveals Genome-Wide Changes Associated with ER Stress in Potato (Solanum tuberosum L.)." Int J Mol Sci 2022. DOI 10.3390/ijms232213795. PMCID PMC9696714.

Design: Potato (cv. not relevant; ref = Castle Russet) treated with tunicamycin (TM, 5 µg/mL in DMSO) vs Mock (DMSO) to induce ER stress / UPR. Two timepoints (2 h, 5 h), 3 biological replicates each → 12 bulk RNA-seq libraries (BGISEQ-500, PE150, strand-specific RF). SRA: PRJNA865435 (SRR20760923–SRR20760934).

Repo: https://github.com/venuraherath/TM-Transcriptome-Potato — authors' own SLURM bash + R scripts (Grace_HISAT2.sh, Grace_HISAT2_SAMTools.sh, Grace_HTSeq_Count_Clean.sh, Grace_DE_edgeR_2hr.r / _5hr.R, Grace_IsoformSwitch_Analysis.R).

IN SCOPE (pipeline-derived, attempted)

Result Pipeline Reproducible?
Alignment / mapping rates (>94% overall, 56–62% unique, 37–38M pairs) HISAT2 2.2.1 --dta --rna-strandness RF → samtools sort YES
Per-gene counts htseq-count 2.0.1 --mode intersection-strict --stranded reverse --minaqual 1 --type exon --idattr gene_id YES
DEG counts per timepoint (2 h 204↑/278↓, 5 h 157↑/129↓) edgeR exactTest (classic, TMM + tagwise disp), BH FDR≤0.05, paper applies |log2FC|≥1 YES (see threshold note)
Total unique DEGs (806) + overlap (9↑/10↓/18 opp) set operations on DEG lists YES

THRESHOLD NOTE (important for honest comparison)

The shipped Grace_DE_edgeR_*.R calls decideTestsDGE(et, adjust.method="BH", p=.05) with the default lfc=0 — i.e. the script's printed up/down counts are FDR-only, not the |log2FC|≥1 the paper text states. To match the paper's 204/278/157/129 the logFC cutoff must be applied downstream. Our edgeR.R reports BOTH (fdronly and fdr_lfc1) so the discrepancy is visible and auditable.

REFERENCE NOTE

Paper/repo annotation = cr.working_models.pm.locus_assign_clean.gtf (a "clean" locus-assigned working-models GTF). The cleaned GTF is not shipped; we download the public cr.working_models.pm.locus_assign.gff3 from SpudDB and convert with gffread -T. Minor count differences from the unshipped cleaning step are possible.

OUT OF SCOPE (not attempted / harder)

  • Isoform switching (9 loci 2 h / 15 loci 5 h): IsoformSwitchAnalyzeR + Kallisto + external CPC2/PFAM/SignalP5/IUPred2A — many external deps; OPTIONAL, attempt only after core DEG/alignment results land.
  • GO/KEGG functional enrichment, heatmaps, wet-lab qRT-PCR validation — manual / external, not pipeline-deterministic.
Figures / tables: Fig 1BTable
align_overall
Reported
>94% overall
Reproduced
94.06-94.73% (all 12)
within tolerance
qc_readpairs
Reported
37-38M read pairs/sample
Reproduced
74.3-76.1M SE reads = 37.2-38.0M pairs-equiv
within tolerance
align_unique
Reported
56-62% unique (21.3-22.8M pairs)
Reproduced
48.2-48.9% of SE reads unique
partial
align_multi
Reported
39-43% multi-mapped
Reproduced
45.2-46.4% of SE reads multi
partial
deg_2hr
Reported
204 up / 278 down (FDR & |log2FC|>=1)
Reproduced
248 up / 370 down (FDR-only 562/464)
partial
deg_5hr
Reported
157 up / 129 down (FDR & |log2FC|>=1)
Reproduced
197 up / 180 down (FDR-only 210/217)
partial
deg_total
Reported
806 unique DEGs
Reproduced
966 union unique
partial
deg_shared
Reported
9 up both / 10 down both
Reproduced
8 up both / 8 down both
within tolerance
isoform_switch
Reported
9 loci (2h) / 15 loci (5h)
Reproduced
not attempted (out of scope this pass)
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 62/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

587.5 k
tokens (I/O) · 51.3 M incl. cache
216 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.