Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Identification of Key Differentially Expressed Genes in Arabidopsis thaliana Under Short- and Long-Term High Light Stress.

Int J Mol Sci · 2025
L1 75/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score 0
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
  • No relevant deviation in data/preprocessing
  • No authors-side cause for any deviation
  • Reported values are derivable from the shared data
  • Any deviation was negligible
  • Overall, the reproduction was clean
What did not (or only partly)
  • 🟡The central claim did not (fully) hold under reproduction
How its reproducibility compares
75/100
Reproducibility score
at the mean
vs. all fields · 1173 studies
🎯 Scores higher than 45% of all assessed papers rank 612 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough -> EXACT 1:1 for the headline Step-1 result. AraLightMeta (commit edbbb59) is a self-contained meta-analysis that ships its own pre-computed inputs (count matrix + per-gene DEG table + GO annotation + GENIE3 edges); it does NOT re-align FASTQ. Ran the authors' own frequent-DEG classification code (AraLightMeta.R L602-632, copied verbatim into repro_step1.R) on the authors' own shipped data on «our HPC» (env paper_figures, R 4.5.2). Reproduced EXACTLY: 978 frequent DEGs (498 up / 474 down / 6 mixed) and 99 unique conditions. One documented, value-preserving deviation: the edgeR dendrogram block was omitted because it only orders conditions and does not enter these counts (verified order-invariant). NOT attempted: (1) Step-2 short/long-specific DEG counts C3a-d (the harder ~20%, L785-1730); (2) upstream HISAT2/featureCounts/edgeR DEG calling on 280 raw libraries across 21 GEO series (out of scope - repo neither ships nor runs it; the shipped tables ARE its output); (3) GRN/WGCNA cluster counts (heavy deps; GENIE3 edges are shipped pre-computed). No fabrication signal: the reported headline figures are exactly regenerable from shipped data by shipped code.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 75
    assessed: 2026-06-14 ⛓ 7f2e531654dc
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-14
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-07-31

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

Can a uniform transcriptomic meta-analysis of accumulated high light (HL) stress datasets in Arabidopsis thaliana distinguish short- versus long-term HL transcriptional programs and identify the key differentially expressed genes (DEGs) and transcription factors governing the long-term adaptive response?

Core claims
  • Short- and long-term HL responses in Arabidopsis leaves are driven by distinct transcriptional programs, with duration of HL treatment as the primary factor separating transcriptomic clusters. finding
  • Long-term HL adaptation involves key TFs CRF3 and PTF1 (antioxidant/jasmonate signaling), ATWHY2, WHY3, and emb2746 (chloroplast/mitochondrial gene coordination), AT2G28450 (ribosome biogenesis), and AT4G12750 (methyltransferase activity). mechanism
  • The AraLightDEGs publicly accessible knowledge base and the AraLightMeta R pipeline were developed for searching, meta-analysis, clustering, and gene regulatory network reconstruction of HL-responsive DEGs. resource
  • A meta-analysis aggregating DEG frequency across many independent experiments (rather than fold changes from individual experiments) yields a robust core set of consistently regulated HL-response DEGs. method
  • The core set of frequently upregulated HL DEGs is enriched for ROS/oxidative stress, general stress responses, and jasmonate/hormonal pathways, while downregulated genes are linked to auxin signaling and growth/photosynthesis suppression. finding
  • Leaf and seedling tissues show tissue-specific transcriptomic effects comparable to or exceeding stress-induced changes, requiring separate analysis. finding
  • Two experiments (GSE251796 and PRJNA699408) exhibited pronounced batch effects and were excluded as outliers. method
Experimental setups
Assay System Perturbation Readout Platform
Bulk RNA-seq transcriptomic meta-analysis (re-processed from raw sequencing data) Arabidopsis thaliana leaves and seedlings (photosynthetic tissues) High light stress (500–2000 µmol·m−2·s−1; durations from minutes to 3 days) Differentially expressed genes (DEGs), frequency of DEG occurrence, log2 fold change
Hierarchical clustering / dendrogram of transcriptomic libraries Arabidopsis thaliana, 99 unique transcriptomic conditions (21 experiments, 280 libraries) High light across varying intensity and duration vs control Clustering of conditions by similarity, tissue and duration grouping AraLightMeta R pipeline
Correlation analysis of aggregated transcriptomic libraries Arabidopsis thaliana control, short-HL, and long-HL leaf and seedling samples Short HL (≤90 min) vs long HL (2 h–3 days) vs control Pairwise correlation coefficients between tissue/treatment groups
GO functional enrichment analysis Arabidopsis thaliana, most frequent DEGs (≥50% of experiments) High light Enriched GO biological process terms for up- and down-regulated gene sets
Gene regulatory network (GRN) reconstruction Arabidopsis thaliana leaf short- and long-term HL DEGs High light Key transcription factors and their target genes AraLightMeta R pipeline
Key results
  • Meta-analysis of 21 experiments / 58 HL conditions yielded ~218,000 DEG instances corresponding to ~19,000 unique genes. 218,000 DEG instances; 19,000 unique genes
  • Relatively low correlation between short- and long-term HL responses in leaves revealed a dynamic temporal shift in gene expression. r=0.88
  • Control leaves vs control seedlings correlation was lower than control leaves vs short-HL leaves, indicating strong tissue-specific effects. 0.94 (control leaf vs seedling) vs 0.97 (control vs short-HL leaf)
  • 978 most frequent DEGs identified (in ≥50% of experiments): 498 upregulated (median log2FC ≥0.5) and 474 downregulated (median log2FC ≤−0.5). 978 DEGs; 498 up / 474 down
  • Filtering for DEGs found in ≥5 leaf-specific experiments produced 8510 DEGs for downstream short-/long-term analysis. 8510 DEGs
  • Upregulated core DEGs strongly enriched for ROS/oxidative stress, wounding/heat/cold stress, and jasmonic acid response pathways.
  • Downregulated core DEGs enriched for auxin pathway, cell communication, and signaling, linked to growth and photosynthesis suppression.
  • Two experiments (GSE251796, PRJNA699408; 8 conditions) clustered as outliers with pronounced batch effects and were excluded. 2 experiments / 8 conditions
Key statistics
  • count 21 experiments covering 58 HL conditions (39 time/intensity combinations) (Curated dataset scope)
  • count ~218,000 instances of DEGs, ~19,000 unique genes (Total DEGs identified in meta-analysis)
  • count 280 transcriptomic libraries; 99 unique conditions (Libraries after QC used for clustering)
  • correlation 0.88 (Short- vs long-term HL response in leaves)
  • correlation 0.94 (Control leaves vs control seedlings)
  • correlation 0.97 (Control leaves vs short-HL treated leaves)
  • count 978 frequent DEGs (498 up, 474 down) (DEGs detected in ≥50% (27) of conditions)
  • fold_change median log2FC ≥0.5 (up) / ≤−0.5 (down) (Thresholds classifying up/down-regulated DEGs)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The paper presents a transcriptomic meta-analysis of Arabidopsis thaliana high-light stress responses by uniformly reprocessing raw RNA-seq data from 21 independent experiments (280 libraries, 58 conditions) using a custom R pipeline (AraLightMeta). Experimental conditions were classified into short-term and long-term high-light groups via hierarchical clustering on aggregated transcriptomic libraries. DEGs were identified through frequency-based filtering (detected in at least 50% of experiments) combined with median log2 fold-change thresholding, and GO enrichment analysis was used to characterize functional categories; gene regulatory networks were then reconstructed to identify key transcription factors.

Replicationbiological Sample size21 independent experiments; 280 transcriptomic libraries; 58 total HL conditions (54 after outlier removal); 25 short-term and 11 long-term leaf conditions retained for main analysis GroupsShort-term HL (2–60 min) vs. long-term HL (2–30 h) in leaf tissue; leaf vs. seedling as a secondary cross-tissue comparison Pairingunpaired Randomization/blindingnot stated Dispersionnone Exact p-valuesno Effect sizesyes Confidence intervalsno Multiplicity correctionnot stated in available text
Statistical tests used
Test Applied to n Assumptions
Hierarchical clustering (linkage method not stated in available text) Classification of 99 unique transcriptomic conditions into short-term, medium-term, and long-term high-light groups (Figure 1 dendrogram) 99 transcriptomic conditions from 21 experiments not stated
Correlation analysis (Pearson or Spearman — type not stated) Assessment of similarity between leaf and seedling transcriptomes under control, short-HL, and long-HL conditions (Figure 2C) Aggregated transcriptomic libraries; exact n not stated not stated
GO enrichment analysis (specific statistical test and background set not stated in available text) Functional characterization of upregulated and downregulated genes among the 978 most-frequent DEGs (Figure 3) 498 upregulated and 474 downregulated genes from 978 most-frequent DEGs not stated
Frequency-based filtering with median log2FC thresholding (non-inferential criterion) Identification of the 978 most-frequent DEGs (detected in at least 50% of 54 non-outlier conditions; median log2FC threshold ±0.5) and the 8510 DEGs retained for downstream analysis (detected in at least 5 leaf-specific experiments) 27 experimental conditions for the 50% threshold; 5 leaf-specific experiments for the secondary filter na
Approaches that could also have been used
  • DEGs were identified using frequency-based filtering (detected in ≥50% of experiments) and median log2FC thresholding as the meta-analytic summary across 21 studies
    Could also: Formal statistical meta-analysis methods such as RankProd, metaRNAseq, or Fisher's combined probability test could also integrate effect sizes and p-values from individual studies into a unified test statistic — Formal meta-analysis methods yield a single p-value with explicit heterogeneity quantification (e.g., Cochran's Q or I²), which allows researchers to distinguish genes that are consistently regulated from those showing high cross-study variability
  • Correlation coefficients between tissue or condition groups (0.88, 0.94, 0.97) were reported as point estimates without measures of uncertainty
    Could also: Bootstrap confidence intervals or Fisher z-transformation intervals could also be reported alongside the correlation point estimates — Confidence intervals around correlations convey estimation precision and allow formal comparison of whether, for example, the short-vs-long-HL correlation (0.88) is meaningfully lower than the within-tissue correlation (0.97)
  • Outlier experiments were identified and excluded based on visual inspection of the dendrogram (described as 'pronounced batch effects')
    Could also: Quantitative outlier-detection approaches such as principal variance component analysis (PVCA), median absolute deviation on inter-sample distances, or RLE plot inspection could also be used to guide and document exclusion decisions — A quantitative threshold for exclusion makes the decision reproducible and provides a numerical criterion that other researchers can apply consistently when extending the dataset
  • Short-term and long-term HL groups were defined by a binary time threshold (≤90 min vs. ≥2 h), with the boundary informed by the clustering results
    Could also: A continuous time-course model such as spline regression implemented in limma/voom or the maSigPro package could also characterize the temporal trajectory of expression without requiring a single cutoff — Continuous time-course models capture the dynamic shape of expression changes across all treatment durations, which may reveal gradual transitions that are obscured by a binary grouping
  • GO enrichment results were described qualitatively as 'statistically significant' without reporting exact or adjusted p-values, enrichment statistics, or the background gene set
    Could also: Reporting Benjamini–Hochberg FDR-adjusted p-values, fold-enrichment or odds-ratio statistics, and the explicit background gene universe (e.g., all expressed genes) would also be standard practice for enrichment analyses — Explicit FDR thresholds and effect-size measures allow readers to calibrate the strength of enrichment signals and to reproduce the analysis with alternative thresholds or background sets
  • Hierarchical clustering was the sole method used to assess global structure among the 99 transcriptomic conditions, with group boundaries interpreted visually from the dendrogram
    Could also: Complementary dimensionality-reduction methods such as PCA or UMAP, optionally combined with a quantitative cluster-validity index (e.g., silhouette width), could also be used alongside the dendrogram — PCA or UMAP provides a continuous low-dimensional layout of inter-sample distances that makes the degree of separation between temporal groups more directly quantifiable and easier to inspect for gradients or intermediate-state samples
Software: R (custom AraLightMeta pipeline) · AraLightDEGs (custom web knowledge base)

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
2
Impact: low
Foundation confidence
Built on 1 assessed reference(s) · mean reproducibility 63/100
partly built on non-reproducible work
Topics

Assessed papers, coloured by verdict. Click a node to open it.

Built on (assessed references) (1)
Cited by (assessed papers) (0)
  • No assessed neighbours yet — the network grows as more papers are assessed.

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-40869111

Paper: Bobrovskikh, Zubairova, Doroshkov (2025). Identification of Key Differentially Expressed Genes in Arabidopsis thaliana Under Short- and Long-Term High Light Stress. Int J Mol Sci 26(16):7790. DOI 10.3390/ijms26167790. Code: https://github.com/av-bobrovskikh/AraLightMeta @ edbbb59 (AraLightMeta.R, 2867 lines, R). Data: repo ships ALL inputs as zips (self-contained). GEO GSE251796 is only one of 21 series the upstream AraLightDEGs resource aggregated.

Pipeline structure (what the repo actually does)

The repo is a meta-analysis that starts from PRE-COMPUTED inputs shipped in the repo, NOT from raw FASTQ:

  • aralightdegs_counts_metadata.csv — aggregated count matrix + 6 metadata rows (280 libraries / 99 conditions across 21 GEO series).
  • aralightdegs_all_degs.csv — precomputed per-gene DEG table from the AraLightDEGs database (14,869 genes; cols incl. total_numbers = frequency, upregulated_fc/downregulated_fc = ';'-joined log2FC values).
  • ATH_GO_GOSLIM.txt — TAIR GO annotation (external, dated 2023-12-31).
  • edges_test_{short,long}_{leaves,seedlings}.rds — precomputed GENIE3/DIANE edges.

AraLightMeta.R does: (1) condition clustering/dendrogram, (2) frequent-DEG classification + GO enrichment, (3) short/long-specific gene classification, (4) WGCNA + GENIE3 gene-regulatory-network reconstruction.

IN SCOPE (pipeline-derived, reproducible from shipped data)

id reported result paper loc pipeline difficulty
C1 978 frequent DEGs (≥50% of HL conditions): 498 up, 474 down (+mixed) Results / Fig (GO input set) AraLightMeta.R L602–632, from all_degs.csv LOW — verbatim author code, light env
C2 total HL conditions / threshold used for "frequent" filter Methods L297–306, L624 LOW
C3 short/long-specific DEG counts: ST up 2450 / down 2954, LT up 2408 / down 2333; 4234 unique (60%), 2398 shared (34%) Results Step 2 (L785–1730), leaves+seedlings MED — more complex classification

OUT OF SCOPE (not attempted — and why)

  • Upstream RNA-seq (HISAT2 2.2.1 → featureCounts 2.0.3 → edgeR 4.4.1 exactTest, FDR≤0.05 |log2FC|≥0.5) on the 280 raw libraries from 21 GEO series. The repo does NOT ship or run this; it ships the already-computed count/DEG tables. Re-deriving it would require downloading/aligning 280 libraries — far beyond the 80/20 line and not the published code's job. The shipped tables ARE the authors' output of this step.
  • GRN reconstruction (WGCNA + GENIE3) detailed cluster numbers (162/64 edges, GO clusters). Heavy deps (WGCNA, GENIE3, biomaRt online); GENIE3 edges are shipped pre-computed (.rds) — so the network is partly an input, not freshly derived. Optional last-20%; attempted only if C1–C3 land with budget to spare.
  • GO enrichment terms (clusterProfiler/org.At.tair.db): derivable from C1 gene set but term-level output is large/qualitative; we reproduce the C1 INPUT counts, which are the auditable numbers.

Primary target: C1 (+ C2, C3 if cheap). Honest 1:1 from the authors' own code on their own data.

C1a
Reported
978
Reproduced
978
exact
C1b
Reported
498
Reproduced
498
exact
C1c
Reported
474
Reproduced
474
exact
C2
Reported
99
Reproduced
99
exact
C3a
Reported
2450
Reproduced
partial
C3b
Reported
2954
Reproduced
partial
C3c
Reported
2408
Reproduced
partial
C3d
Reported
2333
Reproduced
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 75/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟢3. Location of the main deviation
🟢4. Cause of the deviation
🟢5. Derivability / plausibility
🟢6. Severity of the deviation
🟡7. Core claim
🟢8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score 0

The reproduced headline values — 978 frequent DEGs (498 up / 474 down / 6 mixed) and 99 unique conditions — match the paper exactly, obtained by running the authors' own code (AraLightMeta.R L602-632) verbatim on the authors' own shipped input tables; this is a clean 1:1 reproduction with strong anti-fabrication evidence and one documented, value-preserving deviation (omitted order-invariant dendrogram block). The deviation here is coverage, not correctness: the short/long-term-specific DEG counts (C3a-d) and GRN/WGCNA numbers — central to the title's short-vs-long framing — were not attempted. No issue lies on the authors' or data side for what was tested; the limitation is our scope choice (80/20). Overall a solid, exact reproduction whose central conclusion is fully confirmed for Step-1 but only partially verified for the time-specific arm.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

122.9 k
tokens (I/O) · 8.8 M incl. cache
13 min
runtime · 0 CPU-h
0.3 GB
peak RAM
3 (2 failed)
HPC jobs
hummel
machine