A Meta-Analysis of the Effects of Chronic Stress on the Prefrontal Transcriptome in Animal Models and Convergence With Existing Human Data.
The main results reproduced, with only marginal, non-material deviations.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Reported values were directly comparable
- ✓No authors-side cause for any deviation
- ✓Reported values are derivable from the shared data
- ✓The central claim held under reproduction
- 🟡Could not use the authors’ exact input data
- 🟡A deviation arose in the data or preprocessing
- 🟡The deviation was non-trivial in magnitude
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Described well enough to reproduce: YES. The paper's primary pipeline is a metafor REML random-effects meta-analysis of 8 chronic-stress-vs-control Gemma DE contrasts from 6 mouse-PFC GEO datasets, followed by BH-FDR. I re-implemented the two repo R scripts faithfully and fetched the same Gemma result sets live from the Gemma REST API (5/6 datasets retained the exact result-set IDs + probe counts the paper used). Result is a faithful PARTIAL reproduction: method, pipeline, and gene-level findings reproduce well (21321 vs 21379 genes analyzed = within-tol; 91/104 DEGs in the paper's reported top-100, all top-25 hits match, Xdh upregulated as reported, downregulation dominant as reported), but the exact DEG count is lower (104 vs 133; 76 vs 97 down; 28 vs 36 up). The most likely causes are documented input drift (Gemma re-analysed GSE84572, 1 of 6 datasets: rsId 480898->519601) and metafor version advance (3.4.0->5.0.1), both of which move borderline genes across the FDR<0.05 cutoff. No fabrication concern: every reported number is regenerable from the shipped code on public data, and the repo's own hard-coded intermediate counts matched the live Gemma data. NOT attempted (the hard 20%): fGSEA/Brain.GMT enrichment (53 gene sets; curated .gmt not shipped), human-psychiatric convergence correlations (external DE tables), and forest/heatmap figures (visual).
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 64assessed: 2026-06-14 ⛓ d1a260b6348b
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-14
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15no human curator yet
- Last updated
- 2026-09-19
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetThe study tests whether chronic stress (via chronic social defeat stress or chronic unpredictable/variable mild stress paradigms) produces a consistent, convergent transcriptional signature in the rodent prefrontal cortex across different studies and stress paradigms, and whether this signature overlaps with transcriptional changes seen in human psychiatric disorders.
- ★ A meta-analysis of 6 public PFC transcriptional profiling datasets (n=117 mice, 8 stress vs control contrasts) identified 133 genes consistently differentially expressed across chronic stress studies and paradigms (FDR<0.05) finding
- ★ fGSEA identified 53 gene sets enriched with differential expression (FDR<0.05), dominated by glial and neurovascular markers (oligodendrocyte, astrocyte, endothelial/vascular) and stress-related signatures (MDD, hormonal responses) finding
- ★ Immediate-early gene markers of neuronal activity (Fos, Junb, Arc, Dusp1) were consistently suppressed by chronic stress finding
- ★ Chronic stress induces a robust, cross-paradigm PFC signature characterized by downregulation of glia/myelin and vascular pathways and suppression of immediate-early gene activity mechanism
- ★ Meta-analysis results resembled findings from prior independent chronic stress meta-analyses (CSDS, early life stress) despite minimal sample overlap finding
- ★ Meta-analysis results resembled transcriptional signatures from human psychiatric disorders including alcohol abuse disorder, MDD, bipolar disorder, and schizophrenia finding
- The Brain Data Alchemy Project standardized pipeline (leveraging the Gemma database) can be used to systematically identify, extract, and meta-analyze public transcriptional profiling datasets method
- Analysis code and curated Brain.GMT gene set database are released as public resources for the research community resource
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Microarray (Affymetrix Mouse GeneChip 1.0 ST) | Adult male C57/BL6 mice, medial PFC | CSDS (14 days, 5 min/day) | Differential gene expression (log2FC), susceptible vs control | Affymetrix Mouse GeneChip 1.0 ST |
| RNA-seq (96 bp single end) | Adult male C57/BL6NCrl and DBA/2NCrl mice, medial PFC | CSDS (10 days, 10 min/day) | Differential gene expression (log2FC), susceptible/resilient vs control | — |
| RNA-seq (50 bp paired-end) | Adult male C57BL/6J mice, PFC | CSDS (10 days, 5 min/day) | Differential gene expression (log2FC), susceptible/resilient vs control | — |
| RNA-seq (100 bp paired-end) | Post-weaning male C57BL/6J mice, medial PFC | CUMS (3 weeks, P21-P51) | Differential gene expression (log2FC), CUMS susceptible vs control | — |
| RNA-seq (50 bp paired-end) | Adult male and female C57BL/6J mice, ventromedial PFC | CUMS (3 weeks) | Differential gene expression (log2FC), CUMS vs control | — |
| Microarray (Affymetrix GeneChip Mouse Genome 430 2.0) | Adult male C57BL/6 mice, PFC | CUMS (4 weeks) | Differential gene expression (log2FC), CUMS vs control | Affymetrix GeneChip Mouse Genome 430 2.0 |
| Random-effects meta-analysis (metafor, REML) | Pooled across 6 mouse PFC datasets | Chronic stress (CSDS/CUMS) vs control (pooled) | Meta-analytic log2FC and FDR-corrected p-value per gene (n=21,379 genes) | metafor v.3.4.0 |
| Fast gene set enrichment analysis (fGSEA) | Meta-analysis results ranked by estimated log2FC, mouse PFC | Chronic stress (CSDS/CUMS) | Enrichment of gene sets (Brain.GMT, MSigDB C5) among differentially expressed genes | fGSEA (Sergushichev 2016) |
- – 133 genes consistently differentially expressed across chronic stress studies/paradigms FDR<0.05
- – 53 gene sets enriched with differential expression, dominated by glial/neurovascular and stress-related signatures FDR<0.05
- ▼ Immediate-early genes Fos, Junb, Arc, Dusp1 suppressed
- – Overlap with 10-day CSDS meta-analysis (Reshetnikov et al. 2022) 15% overlap (n=18 of 117)
- – Overlap with stress susceptibility vs resilience meta-analysis (Reshetnikov et al. 2022) 15% overlap (n=18 of 117)
- – Overlap with broader chronic stress meta-analysis (Gururajan 2022) 32% overlap (37 of 117)
- – No overlap with 30-day CSDS meta-analysis (Reshetnikov et al. 2022) n=20, 0% overlap
- count n=117 subjects total across 6 datasets/8 contrasts (Final meta-analysis sample size)
- count n=21,379 genes analyzed (Genes represented in at least 5 stress vs control contrasts, included in random effects model)
- count 133 differentially expressed genes at FDR<0.05 (Meta-analysis result of consistently DE genes)
- count 53 enriched gene sets at FDR<0.05 (fGSEA using Brain.GMT and MSigDB C5)
- count 174 datasets initially identified via Gemma search terms (Initial dataset search before inclusion/exclusion filtering)
- count 6 datasets / 8 stress vs control contrasts included (Final included studies after PRISMA-style screening)
- other 80% power at alpha=0.05 to detect medium effect sizes (Post hoc power assessment of final sample)
- other log2 expression range RNA-seq -5 to 12, microarray 4 to 15 (GSE114224 excluded for range 2.5-4) (Dataset quality control criterion for inclusion)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This paper performed a random effects meta-analysis (metafor, REML, inverse-variance weighting) of log(2) fold change (Log2FC) effect sizes across six publicly available transcriptional profiling datasets (n=117 mice, 8 stress-vs-control contrasts, 21,379 genes) to characterize chronic stress effects on the prefrontal cortex transcriptome. Individual-study differential expression was computed via limma (microarray) or limma-voom (RNA-Seq) within the Gemma database, and meta-analytic p-values were corrected for false discovery rate using the Benjamini-Hochberg method. Functional patterns were identified via fast gene set enrichment analysis (fGSEA), and results were benchmarked against published meta-analyses of chronic stress and psychiatric disorder cortical transcriptomics.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| limma (microarray) / limma-voom (RNA-Seq) differential expression (moderated t-statistic) | Individual study stress-vs-control contrasts, preprocessed and run within the Gemma database | Per-study sample sizes n=4 to n=38; total n=117 across all studies | not stated |
| Random effects meta-analysis rma() with REML estimation and inverse-variance weighting | Pooling Log2FC effect sizes across 8 stress-vs-control contrasts for each of 21,379 genes | 21,379 genes each represented in at least 5 of 8 contrasts; n=117 animals total | not stated |
| Fast gene set enrichment analysis (fGSEA) | Gene set enrichment across Brain.GMT and MSigDB C5 gene sets, gene list ranked by meta-analytic Log2FC | 21,379 ranked genes | na |
| Benjamini-Hochberg FDR correction | Meta-analytic p-values for all 21,379 genes and separately for fGSEA gene set p-values | 21,379 genes (differential expression); gene set count not stated (fGSEA) | na |
-
An intercept-only random effects model was used, pooling across stress paradigms (CSDS and CUMS) without moderator variables↳ Could also: A mixed-effects meta-regression model including moderators such as stress paradigm, stress duration, or sex could also have been fit — Moderator analysis partitions the estimated between-study heterogeneity (tau²) into explained and residual components, allowing assessment of whether paradigm type or duration systematically shifts gene-level effect sizes — though the small number of datasets (k=6) limits statistical power for such analyses
-
Log2 fold changes from limma/limma-voom were used as the effect size metric pooled across microarray and RNA-Seq platforms↳ Could also: Standardized mean differences (Hedges' g or Cohen's d) could also have been computed and pooled — Standardized effect sizes explicitly account for within-study variance and are scale-free across measurement platforms, making them more directly comparable when pooling microarray and RNA-Seq data that differ in dynamic range and noise structure
-
A minimum of 5 out of 8 contrasts per gene was required for inclusion, yielding 21,379 genes↳ Could also: A lower inclusion threshold (e.g., ≥3 contrasts, ≥2 studies) could also have been applied — A lower threshold would increase coverage of less-frequently measured transcripts and improve power for those genes, at the cost of greater uncertainty in the pooled estimate; the chosen threshold ensures at least three independent studies per gene
-
fGSEA was applied using all genes ranked continuously by meta-analytic Log2FC↳ Could also: CAMERA (which models intra-gene-set correlation) or overrepresentation analysis (ORA) on the discrete set of FDR-significant genes could also have been used — CAMERA explicitly accounts for inter-gene correlations within gene sets that can inflate enrichment statistics; ORA provides a complementary threshold-based view; using multiple methods in parallel can improve confidence in enriched pathways
-
Benjamini-Hochberg FDR was applied to control the false discovery rate across all tested genes↳ Could also: Storey's q-value method or a permutation-based FDR could also have been used — Storey's q-value estimates the proportion of true nulls empirically from the observed p-value distribution and can be more powerful when many genes are truly differentially expressed; permutation-based FDR is distribution-free and may better reflect the null in correlated transcriptomic data
-
Comparisons between the meta-analytic results and published psychiatric disorder datasets were performed (method details partially outside the excerpted text)↳ Could also: Rank-rank hypergeometric overlap (RRHO) analysis or permutation-based Spearman correlation could also have been used for cross-dataset comparisons — RRHO characterizes threshold-independent, directional concordance between two ranked gene lists and visualizes the overlap across the full effect-size spectrum, revealing partial overlaps that fixed-threshold comparisons may miss
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
Assessed papers, coloured by verdict. Click a node to open it.
- No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-41566898
Paper: Xiong et al. 2026, Brain & Behavior 10.1002/brb3.71197 — "A Meta-Analysis of the Effects of Chronic Stress on the Prefrontal Transcriptome in Animal Models…" Repo: https://github.com/Jinglin0320/Metaanalysis_GemmaOutput_CSDS_PFC_Jinglin (R, 100%)
Pipeline (what produces the numbers)
The analysis is a random-effects meta-analysis of Gemma-precomputed differential expression (DE) results across 6 mouse PFC datasets (8 stress-vs-control contrasts):
DEextraction-CSDS+CUMS-PFC.R— reads, per dataset, the Gemma DE result-set tables (Log2FC, Tstat per gene), drops rows with missing/ambiguous gene symbol, collapses to one value per gene symbol (mean Log2FC, mean Tstat; SE = mean(Log2FC/Tstat), sampling variance SV = SE²).Meta-Analysis of CSDS and CUMS.R— aligns all 8 contrast columns by gene symbol (full outer join), keeps genes with <3 NAs (present in ≥6 of 8 contrasts) → 21,379 genes, runsmetafor::rma(yi=Log2FC, vi=SV)(REML random-effects) per gene, BH-FDR corrects the p-values, counts FDR<0.05.
In scope (pipeline-derived, attempted)
| id | reported result | paper location |
|---|---|---|
| C1 | 21,379 genes analyzed (≥5/≥6 contrasts) | Results; script line 396 |
| C2 | 133 DEGs at FDR<0.05 | Results; script line 599 |
| C3 | 97 downregulated | Results; script line 615 |
| C4 | 36 upregulated | Results; script line 608 |
| C5 | top DEGs (Hepacam, Elovl5, Dusp1, Fa2h, S100b, Aqp4, P2rx7, Xdh↑, Fos, Arc…) | Results; script lines 525-549 |
The 8 contrasts (Gemma result-set + contrast value IDs, current as of 2026-06-14)
| dataset | rsId | contrast valueId | label |
|---|---|---|---|
| GSE84572 | 519601 | 142768 | CSDS (social defeat) — re-analysed by Gemma; paper used now-deleted rsId 480898 |
| GSE109315 | 498170 | 173376 | resilient to CSDS |
| GSE109315 | 498170 | 173377 | susceptible to CSDS |
| GSE81672 | 492712 | 154643 | resistant (resilient) |
| GSE81672 | 492712 | 154646 | susceptible |
| GSE81587 | 538108 | 176084 | CUMS |
| GSE102556.1 | 506248 | 181946 | chronic variable stress 21d |
| GSE151807 | 520045 | 170598 | chronic mild stress |
5 of 6 datasets retain the exact result-set IDs + probe counts the paper used; only GSE84572 was re-analysed (probe count 35531→35496), a small input drift we flag.
Out of scope (not attempted, with reason)
- fGSEA / Brain.GMT enrichment (53 gene sets) — depends on a curated Brain.GMT
.gmtnot in the repo; deferred as the hard 20%. - Convergence with human psychiatric datasets (correlations R=0.21..0.58) — uses external published DE tables (Duan 2025, Reshetnikov 2022, etc.) not shipped; out of scope.
- Forest plots / heatmaps — visual, derived from the same metaOutput; not graded numerically.
Hard rules compliance
All compute on «our HPC» SLURM; Gemma TSVs downloaded inside the compute job («infra» cwd); «host» holds only small results + pointers.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
This is a faithful partial reproduction of a metafor REML random-effects meta-analysis: 21321 vs 21379 genes analyzed (within-tol), all reproduced top-25 DEGs match with Xdh up and downregulation dominant. The only gap is the exact DEG count (104 vs 133; 76 vs 97 down; 28 vs 28... 36 up), ~22% lower, sitting at the FDR<0.05 boundary. The deviation is on the technical/data-availability side, not the authors': Gemma re-analysed 1 of 6 datasets and metafor advanced 3.4.0->5.0.1 — every reported value remains regenerable from the shipped code on public data, with no fabrication concern. The central biological conclusion holds; overall a solid yellow.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at [email protected].
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.