A Meta-Analysis of the Effects of Acute Sleep Deprivation on the Cortical Transcriptome in Rodent Models.
The main results reproduced: recomputed values matched the published ones within tolerance.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Reported values were directly comparable
- ✓No relevant deviation in data/preprocessing
- ✓No authors-side cause for any deviation
- ✓Reported values are derivable from the shared data
- ✓Any deviation was negligible
- ✓The central claim held under reproduction
- ✓Overall, the reproduction was clean
- 🟡Could not use the authors’ exact input data
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
DESCRIBED WELL ENOUGH + 1:1 on the deposited-output stage. The paper is a Gemma-database random-effects meta-analysis (metafor) of acute sleep deprivation across 8 mouse-cortex GEO datasets (18 SD-vs-control contrasts); the repo is a single self-contained R script and ships no data, but the paper deposits the full pipeline OUTPUT as supplementary Table S1 (per-gene Log2FC/p/FDR, 16,290 genes) and the GSE114845 validation DE as Table S2. I re-ran the authors' DE-calling stage on their own deposited estimates on «our HPC» («job»): all 12 in-scope anchors reproduce EXACTLY -- 16,290/16,255 genes, 182 DEGs (104 up / 78 down), the named Table 2/3 effect sizes (Nr3c1, Nr2e1, Cdc42ep3, B3gnt3, Hspa12b, Gmeb1 all match to 3 dp), and the validation 115/182 (criterion = same direction AND nominal p<0.05 in GSE114845). An INDEPENDENT Benjamini-Hochberg recomputation from the deposited p-values reproduces the deposited FDR column (max abs diff 0.0021) and the same 182 cutoff -> NO fabrication: the reported numbers are genuine BH outputs of the deposited data. WHAT I DID NOT ATTEMPT (the ~20%): re-running metafor from the RAW 2022 Gemma DE files. Those inputs are unretrievable -- the analysis/resultSet IDs in the script's folder names return HTTP 404, and Gemma has re-curated and restructured the experimental designs (SD timepoints now a separate factor), so the 18 contrasts cannot be rebuilt 1:1. A Gemma-derived meta-analysis is not byte-reproducible once the upstream DB moves on; the authors' deposit of Table S1 is exactly the mitigation that keeps the result fully auditable. Data lived on «infra»; only small results pulled to «host».
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 96assessed: 2026-06-15 ⛓ d53aa07302ae
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-15
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-16no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusThe study tests whether a meta-analysis of public transcriptional profiling datasets can identify consistent, reproducible effects of acute sleep deprivation (SD) on gene expression in the rodent cerebral cortex across diverse SD paradigms and studies.
- ★ Meta-analysis of 18 SD-vs-control contrasts identified 182 genes differentially expressed in the murine cortex in response to sleep deprivation (FDR < 0.05). finding
- ★ Most meta-analysis DEGs (115/182) replicated with concordant effects (FDR < 0.05) in an independent large RNA-Seq validation dataset (GSE114845). finding
- ★ SD down-regulates pathways related to stress response (e.g., glucocorticoid receptor Nr3c1), vasculature, growth and development, and up-regulates pathways related to stress, inflammation, and neuropeptide signalling. mechanism
- ★ Exploratory analyses suggest recovery sleep (1–18 h) could reverse the impact of SD on gene expression. finding
- ★ A random-effects meta-analysis model fit to gene-level Log2 fold changes (and sampling variances) using the metafor rma() function with REML is an effective approach to detect consistent SD effects across heterogeneous datasets. method
- The meta-analysis provides a reference database (full results) illustrating the diverse molecular impact of SD on the rodent cerebral cortex. resource
- The majority of the cortical transcriptome was differentially expressed in response to SD in GSE114845 (63% FDR < 0.05 in re-analysis). finding
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Affymetrix microarray (GeneChip Mouse Genome 430 2.0) | Mouse (C57BL/6J) cerebral cortex | Gentle handling sleep deprivation (6/9/12 h SD vs Ctrl) | Gene expression (Log2 fold change) | Affymetrix GeneChip Mouse Genome 430 2.0 Array (GSE6514) |
| Affymetrix microarray (Mouse Exon/Gene ST arrays) | Mouse (C57BL/6J) cerebral cortex / anterior cingulate cortex | SD (gentle handling, constant movement, unknown) with/without recovery sleep | Gene expression (Log2 fold change) | Affymetrix Mouse Exon 1.0 ST / Gene 2.1 ST / Gene 2.0 ST Arrays (GSE33491, GSE78215, GSE93041) |
| RNA-Seq (paired-end and single-end) | Mouse (C57BL/6J) prefrontal/medial prefrontal/frontal/cerebral cortex | Gentle handling SD (3–12 h) with/without recovery sleep | Gene expression (aligned read counts, Log2 fold change) | Illumina HiSeq 2500/4000, NextSeq 500 (GSE113754, GSE128770, GSE132076, GSE144957) |
| RNA-Seq validation re-analysis | Mouse (full genetic panel) cerebral cortex (86 samples pooled from 222 mice) | 6 h gentle handling SD vs Ctrl | Differential expression (log2 cpm, Log2FC) via limma-voom | Illumina HiSeq 2500 (HiSeq SBS Kit v3), GSE114845; re-aligned with ARCHS4/Kallisto |
| Random-effects meta-analysis (in silico) | 16,290 genes across 18 SD contrasts (n = 293 mice) | none | Pooled SD effect size (Log2FC), FDR | metafor rma() REML (R) |
| Gene-set enrichment analysis (fGSEA) | Meta-analysis gene ranking | none | Pathway enrichment (directional and non-directional) | fGSEA with Brain.GMT gene set database |
- – 182 differentially expressed genes detected in cortex after SD (104 upregulated, 78 downregulated)
- – 115 of 182 meta-analysis DEGs showed similar effects (FDR < 0.05) in independent validation dataset GSE114845 115/182
- ▼ Glucocorticoid receptor Nr3c1 consistently down-regulated across SD paradigms and also down-regulated after 6 h SD in GSE114845
- ▲ CD7 (Cd7) immunoglobulin superfamily gene consistently upregulated across SD paradigms and experiments
- – Down-regulation in stress response, vasculature, growth/development pathways; up-regulation in stress, inflammation, neuropeptide signalling pathways
- – 63% of cortical transcriptome differentially expressed in response to SD in GSE114845 re-analysis (78% in original publication) 63%
- – 16,255 of 16,290 genes produced stable meta-analysis estimates 16,255/16,290
- count 182 DEGs (FDR < 0.05) (differentially expressed genes from meta-analysis)
- count 115/182 validated (FDR < 0.05) (DEGs replicated in GSE114845 validation)
- count 16,290 genes (genes included in meta-analysis (present in ≥13 of 18 contrasts))
- count 8 datasets, 18 SD contrasts (datasets meeting inclusion criteria; collective sample size)
- count n = 222 mice (86 pooled RNA-Seq samples; 43 SD, 43 CTRL) (GSE114845 validation dataset)
- other 63% FDR < 0.05 (78% in original publication) (proportion of cortical transcriptome differentially expressed in GSE114845)
- count 104 upregulated, 78 downregulated (direction of the 182 meta-analysis DEGs)
- other 80% power to detect medium effect sizes at alpha 0.05 (statistical power of collective meta-analysis sample)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This meta-analysis pooled per-gene differential-expression effect sizes (log2 fold changes) from 18 sleep-deprivation-vs-control contrasts drawn from 8 publicly available rodent cortical transcriptomics datasets (microarray and RNA-Seq; collective n = 293 mice) identified via systematic search in the Gemma database. A random effects model was fit per gene (n = 16,290) using inverse-variance weighting with REML estimation (metafor package), and gene-level results were corrected for false discovery rate using the Benjamini–Hochberg method. Functional gene-set enrichment was characterised with fGSEA, and meta-analysis DEGs were validated by comparison with an independent RNA-Seq dataset (GSE114845; n = 86 RNA pools from 222 mice) re-analysed with the limma-voom pipeline, with cross-dataset consistency assessed via Spearman's rank correlation and simple linear regression.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Random effects meta-analysis, inverse-variance weighting, REML heterogeneity estimation (metafor::rma(), intercept-only model) | Per-gene pooling of log2 fold changes across 18 SD-vs-control contrasts; main planned outcome | 16,290 genes; 18 contrasts; collective n = 293 mice | not stated |
| Random effects meta-analysis with SD duration (numeric, centered) and recovery sleep (factor) as predictors (metafor::rma()) | Exploratory secondary analysis examining moderating effects of SD duration and presence of recovery sleep on gene expression | same 18 contrasts and 16,290 genes as main analysis | not stated |
| limma-voom pipeline with empirical Bayes correction (eBayes()) | Differential expression analysis within each individual dataset via Gemma's standardised pipeline (effect sizes extracted for meta-analysis input), and re-analysis of validation dataset GSE114845 | varies by dataset (n = 4 to 108 per individual study); GSE114845 re-analysis: n = 86 RNA-Seq samples | not stated |
| Simple linear regression (stats::lm()) | Comparison of per-gene Log2FC estimates between meta-analysis output and GSE114845 re-analysis, for both top DEGs (FDR < 0.05) and all shared genes | number of genes present in both outputs; not explicitly stated | not stated |
| Spearman's rank correlation (stats::cor.test()) | Rank-based comparison of per-gene Log2FC between meta-analysis and GSE114845, for top DEGs and all shared genes | number of genes present in both outputs; not explicitly stated | not stated |
| fast Gene Set Enrichment Analysis (fGSEA) | Functional pathway/ontology characterisation of meta-analysis results ranked by Log2FC (directional) and |Log2FC| (non-directional) using Brain.GMT v1 | 16,255 genes with stable meta-analysis estimates | na |
-
A random effects model with REML estimation was used to pool per-gene effect sizes, with heterogeneity treated as a single random component across all 18 contrasts↳ Could also: A three-level (multilevel) random effects model explicitly nesting contrasts within studies could also be used, given that several datasets contributed multiple contrasts (e.g., GSE78215 contributed 4 contrasts, GSE128770 contributed 4) — Contrasts from the same study share biological material, lab environment, and platform, introducing within-study correlation; a three-level model partitions variance into within-study and between-study components, which may produce better-calibrated standard errors for the pooled estimate
-
Between-study heterogeneity (τ²) was estimated via REML with the intercept-only model↳ Could also: The DerSimonian–Laird method-of-moments estimator or a fully Bayesian approach (e.g., bayesmeta or brms) could also estimate τ² — DL is the most widely used alternative and eases comparison with prior meta-analyses; Bayesian approaches propagate uncertainty in τ² into gene-level summaries rather than treating it as a fixed point estimate, which can matter when the number of studies is small (here k = 8–18 per gene)
-
Validation agreement between meta-analysis Log2FC and GSE114845 Log2FC was assessed with Spearman's rank correlation and simple linear regression↳ Could also: A concordance correlation coefficient (CCC) or Bland–Altman analysis could also quantify the agreement between the two sets of effect-size estimates — Spearman correlation captures rank agreement regardless of scale or systematic offset; CCC jointly rewards both precision and accuracy, and Bland–Altman plots reveal whether any bias or proportional differences in effect magnitude exist between the meta-analysis and validation estimates
-
Genotype was not included as a covariate in the limma-voom re-analysis of GSE114845 (a genetic reference population) to avoid model overfitting↳ Could also: A linear mixed model treating genotype as a random blocking factor (e.g., using limma's duplicateCorrelation or lme4) could also be used — Modelling genotype as a random effect may reduce residual variance and improve power to detect SD effects without the degree-of-freedom cost of a fixed-effects parameterisation for each genotype; it also more accurately reflects the design of a genetic reference population study
-
fGSEA was applied to the full ranked gene list from the meta-analysis for pathway enrichment↳ Could also: Over-representation analysis (ORA) applied to the 182 FDR-significant DEGs, or variance-weighted approaches such as camera (limma) or GSVA, could also characterise pathway-level signals — ORA offers a more direct interpretation when a discrete DEG list is the primary deliverable; camera explicitly accounts for inter-gene correlation within gene sets, which is common in transcriptomics data and can affect type I error for enrichment tests
-
Dispersion in validation box plots was displayed as IQR (box) with range or 1.5×IQR whiskers, accompanied by jittered individual data points↳ Could also: Notched box plots or overlaid 95% CIs on group means could also communicate the precision of group-level expression estimates — With n = 86 samples per condition, 95% CIs on means directly convey inferential uncertainty about the group mean, complementing the distributional information already provided by the IQR and individual data points
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
Assessed papers, coloured by verdict. Click a node to open it.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-41031900
Paper: Rhoads et al. 2025, A Meta-Analysis of the Effects of Acute Sleep Deprivation on the Cortical Transcriptome in Rodent Models, J Sleep Res, DOI 10.1111/jsr.70205. PMCID PMC13131251.
Code: https://github.com/rhoadsco/Sleep-Deprivation-MetaAnalysis — a single
self-contained R script (Meta Analysis of Sleep Deprivation Studies.R, 715 lines).
Functions originate from M. Hagenauer's BrainDataAlchemy. (P16: even though this
re-uses a lab template, it is the authors' own analysis script for this paper.)
Pipeline (as described in Methods + repo)
- Input: per-study differential-expression (DE) results pulled from Gemma (Pavlidis lab DB) for 8 mouse cortex datasets → 18 SD-vs-control contrasts (GSE6514, GSE33491, GSE78215, GSE93041, GSE113754, GSE128770, GSE132076, GSE144957). Each contrast gives a Log2FC + t-stat per probe.
- Per-study collapse: filter to good gene annotation, average probes →
one Log2FC + SE per gene symbol (
CollapsingDEResults_OneResultPerGene). - Meta-analysis: per gene, random-effects inverse-variance model via
metafor::rma(effect, var)across the 18 contrasts (genes with ≥5 NA dropped) → estimate Log2FC, SE, p-value (RunBasicMetaAnalysis). - DE calling: Benjamini-Hochberg FDR on the meta p-values
(
multtest::mt.rawp2adjp), threshold FDR<0.05 (FalseDiscoveryCorrection). - Validation: regress meta Log2FC against an independent GSE114845 re-analysis.
- Enrichment: fgsea over Brain.GMT gene sets.
Reported results (anchors to reproduce)
- R1 16,290 genes analysed; 16,255 produced stable meta estimates (Results 3.1).
- R2 182 DEGs at FDR<0.05; 104 up, 78 down (Results 3.1).
- R3 Top DEGs (Tables 2 & 3): down Nr3c1 (−0.122, FDR 7.73e-5), Nr2e1 (−0.163, 8.46e-5), Cdc42ep3 (−0.093, 1.11e-4); up B3gnt3 (0.155, 1.08e-7), Hspa12b (0.153, 7.73e-5), Gmeb1 (0.135, 8.35e-5).
- R4 115/182 DEGs (64%) validated in GSE114845 (Fig 3C).
- R5 Exploratory covariate meta-analysis: 16,248 stable estimates (Table S6).
Deposited artifacts that enable reproduction
The repo deposits no data, but the paper deposits the pipeline OUTPUT as supplementary tables (PMC OA package, CC BY):
- Table S1 (
*-s001.xlsx): full meta-analysis output, 16,290 genes / 16,255 stable estimates, per-gene Log2FC + p-value + FDR. → enables R1, R2, R3. - Table S2 (
*-s005.xlsxper caption): GSE114845 re-analysis DE, 19,798 genes → enables R4 validation. - Table S6 (
*-s002.xlsxper caption): exploratory covariate meta (16,248) → R5. - Table S5 (
*-s004.xlsx): fGSEA over 10,436 gene sets.
IN SCOPE (attempted)
Reproduce the DE-calling + counting stage of the authors' own pipeline by
re-running their BH-FDR/threshold/sign-split/ranking on their deposited per-gene
meta estimates (Table S1), plus the validation overlap (S1×S2). This is a
faithful, deterministic re-execution of FalseDiscoveryCorrection() and the
downstream counting on the authors' own data, and is a direct internal-consistency
/ possible-fabrication check of R1–R4.
OUT OF SCOPE (not attempted — documented blocker, the hard ~20%)
Re-running the upstream metafor meta-analysis from raw Gemma DE files is not reproducible:
- The exact 2022 Gemma inputs are gone: the analysis/resultSet IDs embedded in the script's folder names (e.g. analysis 92880, resultSet 477377 for GSE6514) now return HTTP 404 on the Gemma REST API.
- Gemma continuously re-curates: the current GSE6514 analysis is 268577 (build 1.32.7, mm10 annot. updated 2022-06-30) and has restructured the experimental design — SD timepoints are now a separate factor from treatment — so the 18 per-timepoint SD-vs-control contrasts the paper meta-analysed cannot be reconstructed 1:1 from today's Gemma. Current REST also exposes per-contrast fold-changes only for 2-level factors.
- Net: a Gemma-DB-derived meta-analysis is **not byt
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
All 12 in-scope anchors reproduce exactly from the authors' deposited supplementary tables — 16,290/16,255 genes, 182 DEGs (104 up/78 down), the named Table 2/3 effect sizes (to 3 dp), and the 115/182 GSE114845 validation — and an independent Benjamini-Hochberg recompute (max abs diff 0.0021) confirms the reported FDRs are genuine, clearing fabrication. The only deviations are 4th-decimal rounding. The single gap is external: the raw 2022 Gemma input files are unretrievable (404; DB re-curated), so the upstream metafor step is not byte-reproducible — but the authors mitigated this by depositing the full pipeline output, keeping the result auditable. This is a strong reproduction; the caveat sits on data availability / a moving-target database, not on the authors' integrity or our method.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.