Exploring Gene Expression Patterns in Alzheimer's Disease Using a Human Microarray Data Meta-Analysis.
The main results reproduced: recomputed values matched the published ones within tolerance.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Same input data as the authors
- ✓Reported values were directly comparable
- ✓No relevant deviation in data/preprocessing
- ✓No authors-side cause for any deviation
- ✓Reported values are derivable from the shared data
- ✓Any deviation was negligible
- ✓The central claim held under reproduction
- ✓Overall, the reproduction was clean
- Every checked point held up.
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Described well enough; EXACT 1:1 on the paper's headline result. Ran the authors' own shipped code (Parsers/Meta-Analysis.php + Mosteller_Rosenthal.php + commands.R, pinned commit 8a418c3) on the authors' own shipped per-study limma DEG lists ('Metanalysis Demo/'), exactly as the README directs to replicate the meta-analysis. The Mosteller-Bush weighted-Stouffer meta-analysis + BH FDR at adjP<0.001 over 10 sub-studies reproduced 4218 total DEGs, 1944 up-regulated (O), 2274 down-regulated (U) -- all three matching the paper to the digit (output SHA256 0eff0303...d89f8). Compute on «our HPC» (SLURM 2176311, php 8.5.7 + r-base 4.5.2 via conda on «infra»). NOTE: the room brief's data accession GSE28146 was actually EXCLUDED by the paper (paraffin samples); the real inputs are 8 studies/10 sub-studies shipped in the repo, which is what was used. NOT attempted (80/20): upstream RMA+BrainArray+SVA+limma from raw CEL (the per-study DEG lists are shipped, so unnecessary for the headline number) and downstream enrichment/PPI via WebGestalt/STRING/Cytoscape (external non-deterministic web tools). No fabrication indicators -- the headline counts are fully regenerable from shipped data+code.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 100assessed: 2026-06-14 ⛓ d38b8d1deab2
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-14
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusThis meta-analysis investigates the differentially expressed genes between Alzheimer's disease and healthy brain to identify genes that can serve as risk factors or biomarkers of diagnostic, prognostic, or pharmacological value, hypothesizing that AD brains display a distinct transcriptomic profile.
- ★ AD brains show a distinct transcriptomic profile with up-regulation of immune/inflammation genes and down-regulation of synapse/neuronal-signaling genes finding
- ★ A meta-analysis of eight microarray studies yields a combined list of 4218 differentially expressed genes (1944 up-regulated, 2274 down-regulated) finding
- ★ Up-regulated DEGs are enriched for immune response processes while down-regulated DEGs are enriched for synapse-related pathways finding
- ★ A Mosteller–Bush weighted meta-analysis approach combines per-study DEG p-values weighted by sample size to identify consistent DEGs across studies method
- The resulting DEG list provides candidate genes that may serve as diagnostic/prognostic biomarkers for early AD detection resource
- PRISMA 2020 guidelines were followed to systematically collect Affymetrix microarray datasets from GEO and ArrayExpress method
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Affymetrix microarray gene expression profiling (re-analysis of public datasets) | human AD and healthy brain tissue (multiple regions: hippocampus, temporal/frontal lobes, posterior cingulate cortex, etc.) | none (AD vs. healthy disease state comparison) | per-probe-set/gene expression intensities | Affymetrix platform chips (CEL files); BrainArray custom CDF v25 |
| Quality control of microarray samples (NUSE and RLE plots) | human brain microarray samples per study | none | normalized unscaled standard error and relative log expression after MAS5 normalization | R v4.30, oligo/Bioconductor |
| Normalization and batch correction | human brain microarray gene expression matrices | none | quantile-normalized gene expression matrix; batch-corrected matrix | RMA algorithm; SVA algorithm; BrainArray CDF |
| Differential gene expression analysis | human brain microarray sub-studies | AD vs. healthy | log2 fold change, p-value, BH-adjusted p-value per gene | limma; org.Hs.eg.db; HGNC annotation |
| Statistical meta-analysis (combination of per-study DEG lists) | 10 sub-studies from 8 human brain microarray studies | AD vs. healthy | meta-analysis z-score and FDR-adjusted p-value per gene | Mosteller–Bush weighted Stouffer method (custom) |
| Over-representation enrichment analysis (ORA) | up- and down-regulated human DEG sets | none | enriched GO/KEGG/Reactome/network/disease/cytogenetic terms | WebGestalt 2024 |
| Protein–protein interaction network construction and hub gene identification | up- and down-regulated human DEG sets | none | PPI networks and hub genes (most interactions) | STRING v12; Cytoscape stringApp v2.2.0 / Cytoscape 3.10.4 |
- – Combined meta-analysis produced 4218 statistically significant DEGs at adjP < 0.001 4218 genes
- ▲ 1944 DEGs were up-regulated and enriched for immune response processes 1944 genes
- ▼ 2274 DEGs were down-regulated and enriched for synapse-related pathways 2274 genes
- ▲ Up-regulated DEG GO:BP enrichment terms (immune response and regulation, cytokine production, cell population proliferation) all had adjP < 10^-10 adjP < 10^-10
- – Eight microarray studies (10 sub-studies) remained for quantitative meta-analysis after quality control and filtering 8 studies / 10 sub-studies
- count 4218 DEGs (total significant DEGs at adjP cut-off of 0.001)
- count 1944 up-regulated (immune response enriched DEGs)
- count 2274 down-regulated (synapse-related enriched DEGs)
- pvalue adjP < 0.001 (significance cut-off for final DEG list)
- pvalue adjP < 10^-10 (GO:BP immune-response terms for up-regulated DEGs)
- count 104 studies initially (35 ArrayExpress + 69 GEO) (PRISMA database search results)
- count 8 studies (final studies included in meta-analysis)
- other NUSE = 1.05 ± 0.10; RLE = 0.0 ± 0.2 (quality control thresholds for sample removal)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This study performed a systematic meta-analysis of eight publicly available human brain microarray datasets comparing Alzheimer's disease (AD) to healthy controls, following PRISMA 2020 guidelines. Each dataset was individually pre-processed using RMA normalization and SVA batch correction, then subjected to per-study differential expression analysis with limma. Study-level p-values were combined across 10 sub-studies using the Mosteller–Bush weighted Stouffer z-score method (weighting by sample size), and meta-analysis p-values were FDR-adjusted at adjP < 0.001, yielding 4218 DEGs whose functional context was explored via over-representation analysis in WebGestalt 2024 and protein–protein interaction network construction in STRING v12/Cytoscape.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| limma moderated t-test (empirical Bayes linear model for microarray data) | Per-study/sub-study differential expression: AD vs. healthy brain samples for each of 10 sub-studies | Varies per sub-study; 10 sub-studies from 8 datasets; per-sub-study n referenced in Table 1 (not reproduced in full text) | not stated |
| Mosteller–Bush weighted Stouffer z-score combination | Meta-analysis combining per-study one-tailed p-values (signed by log2FC direction) across all 10 sub-studies for each gene | 10 sub-studies; weights proportional to sqrt(n_i − 2) per sub-study | not stated |
| Benjamini–Hochberg FDR adjustment | Applied at two stages: (1) within each sub-study on limma output; (2) at meta-analysis level on combined p-values; adjP < 0.001 cut-off applied for final DEG list | All genes detected across respective platform CDFs; 4218 genes passed the meta-analysis filter | na |
| Over-representation analysis (ORA; hypergeometric test) | Enrichment of up-regulated (n = 1944) and down-regulated (n = 2274) DEG sets across GO (BP/CC/MF), KEGG, Reactome, DisGeNET, transcription factor targets, microRNA targets, and cytogenetic band databases via WebGestalt 2024 | 1944 up-regulated and 2274 down-regulated DEGs; background = union of genes represented across all included Affymetrix platform CDFs | not stated |
-
Per-study p-values were combined using the Mosteller–Bush weighted Stouffer method, a fixed-effects p-value combination approach that treats between-study variation as sampling noise↳ Could also: A random-effects meta-analysis applied to per-study log2FC estimates and their standard errors (e.g., DerSimonian–Laird model via the metafor R package) could also have been used — A random-effects model explicitly estimates between-study heterogeneity (τ²) and reports it as I², producing pooled effect estimates with confidence intervals; this is informative when studies differ in brain region, patient demographics, or platform, as it quantifies how consistent the effect is across studies rather than only whether it reaches significance
-
Enrichment was performed using ORA, which dichotomizes genes into a selected DEG list versus a background based on the adjP < 0.001 threshold↳ Could also: Gene Set Enrichment Analysis (GSEA) using the full ranked gene list — for example ranked by meta-analysis z-score — could also have been applied — GSEA uses the continuous ranking of all genes rather than a binary cut-off, making results less sensitive to the choice of significance threshold and capable of detecting coordinated but moderate expression shifts across a pathway; it is particularly useful when many genes show small but consistent directional changes
-
DEGs were selected using only an adjusted p-value cut-off (adjP < 0.001), without an additional minimum fold-change filter↳ Could also: A dual threshold combining adjP with a minimum absolute log2 fold change (e.g., |log2FC| ≥ 1) could also have been applied — With aggregated sample sizes across many studies, very small expression differences can achieve high statistical significance; adding a fold-change filter retains genes whose effect magnitude is more likely to be biologically relevant, and is a common practice in transcriptomic meta-analyses to improve downstream interpretability
-
Studies containing samples from multiple brain tissues were split into tissue-specific sub-studies, each analyzed independently before meta-analysis combination↳ Could also: Tissue type could also have been included as a covariate or random effect within a single linear mixed-effects model per study (e.g., using lme4 or limma's duplicate-correlation approach) — Modeling tissue as a covariate within a unified model preserves statistical power by using all samples jointly while still adjusting for tissue-driven variance, rather than reducing effective sample sizes through splitting; it also allows a formal test of tissue-by-diagnosis interaction
-
Batch effects were addressed by applying SVA separately within each study or sub-study to identify and regress out latent technical factors↳ Could also: ComBat (empirical Bayes batch correction) treating study-of-origin as the known batch variable, applied after combining studies into a single expression matrix, could also have been used — Cross-study ComBat correction explicitly harmonizes systematic inter-study differences (e.g., scanner, protocol, lab) by treating each study as a known batch, which can reduce cross-study variance in a more transparent and interpretable way than latent-factor removal; the corrected matrix can then be used for downstream joint analysis
-
Hub genes in protein–protein interaction networks were identified based on the highest number of interactions (node degree centrality)↳ Could also: Other centrality metrics such as betweenness centrality, closeness centrality, or eigenvector centrality could also have been used to identify hub genes — Degree centrality captures the most highly connected nodes but can favor promiscuously interacting proteins; betweenness centrality identifies nodes that serve as bridges between modules and may pinpoint functionally critical regulatory genes not captured by degree alone, which can be particularly informative in disease network analyses
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-41744654
Paper: Dermitzaki et al. (2026), Exploring Gene Expression Patterns in
Alzheimer's Disease Using a Human Microarray Data Meta-Analysis. Biology (Basel),
DOI 10.3390/biology15040345. PMID 41744654 / PMC12938635.
Code: https://github.com/imichalop/Meta-Analysis (own code, P16 own-repo)
pinned commit 8a418c313171667538479db2c452996de37afa2c (pushed 2026-01-30).
NOTE ON THE BRIEF'S DATA ACCESSION
The room brief lists geo:GSE28146 as the dataset. The paper explicitly EXCLUDED
GSE28146 ("excluded after screening the meta-data, as the samples were stored in
paraffin blocks", Methods). GSE28146 was a harvest artifact, not an input. The
paper's actual inputs are 8 studies / 10 sub-studies:
GSE48350 (Hip), GSE39420, GSE36980 (FrontCort/Hip/TempCort), GSE16759, GSE1297,
GSE5281 (PCC), GSE12685, E-MEXP-2280.
Pipeline stages (Methods)
- Preprocessing — Affymetrix Power Tools + RMA with BrainArray custom CDFs; GEOquery/oligo (R). [raw CEL stage]
- Batch correction — SVA. QC by NUSE (1.05±0.10) / RLE (0.0±0.2). [manual QC]
- Per-study DEG — limma topTables; org.Hs.eg.db annotation; BH FDR. [pipeline]
- Meta-analysis — Mosteller–Bush (weighted Stouffer variant) combining the
per-study limma p-values across the 10 sub-studies, then BH FDR; significance
cutoff adjP < 0.001. → headline DEG list. [pipeline —
Parsers/Meta-Analysis.php] - Enrichment / networks — WebGestalt 2024 (ORA), STRING v12, Cytoscape stringApp. [external web tools — OUT OF SCOPE, not scriptable/deterministic here]
IN SCOPE (attempted) — the meta-analysis stage (stage 4)
The repo ships Metanalysis Demo/ containing the exact per-study limma DEG
lists used in the article plus studies.txt. The README states: "the meta-analysis
script can be run there, to replicate the meta-analysis results." This makes stage 4
a fully self-contained, deterministic 1:1 reproduction with no raw-data download,
no APT, no SVA — the cleanest low-hanging pipeline output and the paper's headline
number.
Target claims (paper Section 3 / abstract):
- Total DEGs at adjP<0.001: 4218
- Up-regulated (over, "O"): 1944
- Down-regulated (under, "U"): 2274
OUT OF SCOPE (not attempted) — and why
- Stages 1–3 (RMA/BrainArray/SVA/limma from raw CEL): the per-study DEG lists are shipped as the demo inputs, so regenerating them from CEL is the hard, optional last ~20% and is unnecessary to reproduce the headline meta-analysis number. Not attempted (80/20 rule). The shipped lists are taken as the pipeline's stage-3 output.
- Stage 5 enrichment/PPI (WebGestalt/STRING/Cytoscape hub genes MYC, GAPDH; node/edge counts): external interactive web services, non-deterministic versions, not a scriptable reproduction. Out of scope.
- Comparison-with-prior-meta-analyses Venn counts (380 common, etc.): derived from stage-4 output + external gene lists; out of scope for the 1:1 stage-4 check.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
The paper's headline meta-analysis result — 4218 significant DEGs at adjP<0.001 (1944 up, 2274 down) — reproduced exactly to the digit by running the authors' own shipped code on their own shipped per-study limma DEG lists, with a verified output hash. There is no deviation on any side: data, method, and endpoint are all 1:1. Only the upstream raw-CEL→limma stage and downstream WebGestalt/STRING enrichment were not attempted, but those are not needed for the headline counts since the per-study lists are shipped. A clean, fabrication-free exact reproduction.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.