Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Exploring Gene Expression Patterns in Alzheimer's Disease Using a Human Microarray Data Meta-Analysis.

Biology (Basel) · 2026
L1 100/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Every question reproduced
-1 pts
From: “every question reproduced”
Total score -7
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
  • No relevant deviation in data/preprocessing
  • No authors-side cause for any deviation
  • Reported values are derivable from the shared data
  • Any deviation was negligible
  • The central claim held under reproduction
  • Overall, the reproduction was clean
What did not (or only partly)
  • Every checked point held up.
How its reproducibility compares
100/100
Reproducibility score
1.5 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 95% of all assessed papers rank 1 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough; EXACT 1:1 on the paper's headline result. Ran the authors' own shipped code (Parsers/Meta-Analysis.php + Mosteller_Rosenthal.php + commands.R, pinned commit 8a418c3) on the authors' own shipped per-study limma DEG lists ('Metanalysis Demo/'), exactly as the README directs to replicate the meta-analysis. The Mosteller-Bush weighted-Stouffer meta-analysis + BH FDR at adjP<0.001 over 10 sub-studies reproduced 4218 total DEGs, 1944 up-regulated (O), 2274 down-regulated (U) -- all three matching the paper to the digit (output SHA256 0eff0303...d89f8). Compute on «our HPC» (SLURM 2176311, php 8.5.7 + r-base 4.5.2 via conda on «infra»). NOTE: the room brief's data accession GSE28146 was actually EXCLUDED by the paper (paraffin samples); the real inputs are 8 studies/10 sub-studies shipped in the repo, which is what was used. NOT attempted (80/20): upstream RMA+BrainArray+SVA+limma from raw CEL (the per-study DEG lists are shipped, so unnecessary for the headline number) and downstream enrichment/PPI via WebGestalt/STRING/Cytoscape (external non-deterministic web tools). No fabrication indicators -- the headline counts are fully regenerable from shipped data+code.

💻 Code ↗ 🗄 Data: GSE28146

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 100
    assessed: 2026-06-14 ⛓ d38b8d1deab2
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-14
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

This meta-analysis investigates the differentially expressed genes between Alzheimer's disease and healthy brain to identify genes that can serve as risk factors or biomarkers of diagnostic, prognostic, or pharmacological value, hypothesizing that AD brains display a distinct transcriptomic profile.

Core claims
  • AD brains show a distinct transcriptomic profile with up-regulation of immune/inflammation genes and down-regulation of synapse/neuronal-signaling genes finding
  • A meta-analysis of eight microarray studies yields a combined list of 4218 differentially expressed genes (1944 up-regulated, 2274 down-regulated) finding
  • Up-regulated DEGs are enriched for immune response processes while down-regulated DEGs are enriched for synapse-related pathways finding
  • A Mosteller–Bush weighted meta-analysis approach combines per-study DEG p-values weighted by sample size to identify consistent DEGs across studies method
  • The resulting DEG list provides candidate genes that may serve as diagnostic/prognostic biomarkers for early AD detection resource
  • PRISMA 2020 guidelines were followed to systematically collect Affymetrix microarray datasets from GEO and ArrayExpress method
Experimental setups
Assay System Perturbation Readout Platform
Affymetrix microarray gene expression profiling (re-analysis of public datasets) human AD and healthy brain tissue (multiple regions: hippocampus, temporal/frontal lobes, posterior cingulate cortex, etc.) none (AD vs. healthy disease state comparison) per-probe-set/gene expression intensities Affymetrix platform chips (CEL files); BrainArray custom CDF v25
Quality control of microarray samples (NUSE and RLE plots) human brain microarray samples per study none normalized unscaled standard error and relative log expression after MAS5 normalization R v4.30, oligo/Bioconductor
Normalization and batch correction human brain microarray gene expression matrices none quantile-normalized gene expression matrix; batch-corrected matrix RMA algorithm; SVA algorithm; BrainArray CDF
Differential gene expression analysis human brain microarray sub-studies AD vs. healthy log2 fold change, p-value, BH-adjusted p-value per gene limma; org.Hs.eg.db; HGNC annotation
Statistical meta-analysis (combination of per-study DEG lists) 10 sub-studies from 8 human brain microarray studies AD vs. healthy meta-analysis z-score and FDR-adjusted p-value per gene Mosteller–Bush weighted Stouffer method (custom)
Over-representation enrichment analysis (ORA) up- and down-regulated human DEG sets none enriched GO/KEGG/Reactome/network/disease/cytogenetic terms WebGestalt 2024
Protein–protein interaction network construction and hub gene identification up- and down-regulated human DEG sets none PPI networks and hub genes (most interactions) STRING v12; Cytoscape stringApp v2.2.0 / Cytoscape 3.10.4
Key results
  • Combined meta-analysis produced 4218 statistically significant DEGs at adjP < 0.001 4218 genes
  • 1944 DEGs were up-regulated and enriched for immune response processes 1944 genes
  • 2274 DEGs were down-regulated and enriched for synapse-related pathways 2274 genes
  • Up-regulated DEG GO:BP enrichment terms (immune response and regulation, cytokine production, cell population proliferation) all had adjP < 10^-10 adjP < 10^-10
  • Eight microarray studies (10 sub-studies) remained for quantitative meta-analysis after quality control and filtering 8 studies / 10 sub-studies
Key statistics
  • count 4218 DEGs (total significant DEGs at adjP cut-off of 0.001)
  • count 1944 up-regulated (immune response enriched DEGs)
  • count 2274 down-regulated (synapse-related enriched DEGs)
  • pvalue adjP < 0.001 (significance cut-off for final DEG list)
  • pvalue adjP < 10^-10 (GO:BP immune-response terms for up-regulated DEGs)
  • count 104 studies initially (35 ArrayExpress + 69 GEO) (PRISMA database search results)
  • count 8 studies (final studies included in meta-analysis)
  • other NUSE = 1.05 ± 0.10; RLE = 0.0 ± 0.2 (quality control thresholds for sample removal)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This study performed a systematic meta-analysis of eight publicly available human brain microarray datasets comparing Alzheimer's disease (AD) to healthy controls, following PRISMA 2020 guidelines. Each dataset was individually pre-processed using RMA normalization and SVA batch correction, then subjected to per-study differential expression analysis with limma. Study-level p-values were combined across 10 sub-studies using the Mosteller–Bush weighted Stouffer z-score method (weighting by sample size), and meta-analysis p-values were FDR-adjusted at adjP < 0.001, yielding 4218 DEGs whose functional context was explored via over-representation analysis in WebGestalt 2024 and protein–protein interaction network construction in STRING v12/Cytoscape.

Replicationbiological Sample sizeEight published microarray studies (split into 10 tissue-specific sub-studies) selected via PRISMA 2020 systematic search; per-sub-study sample sizes in Table 1; no a priori power calculation described GroupsAlzheimer's disease brain tissue vs. healthy/control brain tissue across multiple brain regions Pairingunpaired Randomization/blindingnot stated Dispersionnone Exact p-valuesno Effect sizesyes Confidence intervalsno Multiplicity correctionBenjamini–Hochberg FDR
Statistical tests used
Test Applied to n Assumptions
limma moderated t-test (empirical Bayes linear model for microarray data) Per-study/sub-study differential expression: AD vs. healthy brain samples for each of 10 sub-studies Varies per sub-study; 10 sub-studies from 8 datasets; per-sub-study n referenced in Table 1 (not reproduced in full text) not stated
Mosteller–Bush weighted Stouffer z-score combination Meta-analysis combining per-study one-tailed p-values (signed by log2FC direction) across all 10 sub-studies for each gene 10 sub-studies; weights proportional to sqrt(n_i − 2) per sub-study not stated
Benjamini–Hochberg FDR adjustment Applied at two stages: (1) within each sub-study on limma output; (2) at meta-analysis level on combined p-values; adjP < 0.001 cut-off applied for final DEG list All genes detected across respective platform CDFs; 4218 genes passed the meta-analysis filter na
Over-representation analysis (ORA; hypergeometric test) Enrichment of up-regulated (n = 1944) and down-regulated (n = 2274) DEG sets across GO (BP/CC/MF), KEGG, Reactome, DisGeNET, transcription factor targets, microRNA targets, and cytogenetic band databases via WebGestalt 2024 1944 up-regulated and 2274 down-regulated DEGs; background = union of genes represented across all included Affymetrix platform CDFs not stated
Approaches that could also have been used
  • Per-study p-values were combined using the Mosteller–Bush weighted Stouffer method, a fixed-effects p-value combination approach that treats between-study variation as sampling noise
    Could also: A random-effects meta-analysis applied to per-study log2FC estimates and their standard errors (e.g., DerSimonian–Laird model via the metafor R package) could also have been used — A random-effects model explicitly estimates between-study heterogeneity (τ²) and reports it as I², producing pooled effect estimates with confidence intervals; this is informative when studies differ in brain region, patient demographics, or platform, as it quantifies how consistent the effect is across studies rather than only whether it reaches significance
  • Enrichment was performed using ORA, which dichotomizes genes into a selected DEG list versus a background based on the adjP < 0.001 threshold
    Could also: Gene Set Enrichment Analysis (GSEA) using the full ranked gene list — for example ranked by meta-analysis z-score — could also have been applied — GSEA uses the continuous ranking of all genes rather than a binary cut-off, making results less sensitive to the choice of significance threshold and capable of detecting coordinated but moderate expression shifts across a pathway; it is particularly useful when many genes show small but consistent directional changes
  • DEGs were selected using only an adjusted p-value cut-off (adjP < 0.001), without an additional minimum fold-change filter
    Could also: A dual threshold combining adjP with a minimum absolute log2 fold change (e.g., |log2FC| ≥ 1) could also have been applied — With aggregated sample sizes across many studies, very small expression differences can achieve high statistical significance; adding a fold-change filter retains genes whose effect magnitude is more likely to be biologically relevant, and is a common practice in transcriptomic meta-analyses to improve downstream interpretability
  • Studies containing samples from multiple brain tissues were split into tissue-specific sub-studies, each analyzed independently before meta-analysis combination
    Could also: Tissue type could also have been included as a covariate or random effect within a single linear mixed-effects model per study (e.g., using lme4 or limma's duplicate-correlation approach) — Modeling tissue as a covariate within a unified model preserves statistical power by using all samples jointly while still adjusting for tissue-driven variance, rather than reducing effective sample sizes through splitting; it also allows a formal test of tissue-by-diagnosis interaction
  • Batch effects were addressed by applying SVA separately within each study or sub-study to identify and regress out latent technical factors
    Could also: ComBat (empirical Bayes batch correction) treating study-of-origin as the known batch variable, applied after combining studies into a single expression matrix, could also have been used — Cross-study ComBat correction explicitly harmonizes systematic inter-study differences (e.g., scanner, protocol, lab) by treating each study as a known batch, which can reduce cross-study variance in a more transparent and interpretable way than latent-factor removal; the corrected matrix can then be used for downstream joint analysis
  • Hub genes in protein–protein interaction networks were identified based on the highest number of interactions (node degree centrality)
    Could also: Other centrality metrics such as betweenness centrality, closeness centrality, or eigenvector centrality could also have been used to identify hub genes — Degree centrality captures the most highly connected nodes but can favor promiscuously interacting proteins; betweenness centrality identifies nodes that serve as bridges between modules and may pinpoint functionally critical regulatory genes not captured by degree alone, which can be particularly informative in disease network analyses
Software: R 4.30 (as stated; likely 4.3.0) · oligo (Bioconductor) · limma (Bioconductor) · SVA (Bioconductor) · org.Hs.eg.db (Bioconductor) · PHP CLI 7.4 · Array Power Tools (APT) · WebGestalt 2024 · STRING v12 · Cytoscape 3.10.4 · Cytoscape stringApp v2.2.0

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
0
Impact: low
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-41744654

Paper: Dermitzaki et al. (2026), Exploring Gene Expression Patterns in Alzheimer's Disease Using a Human Microarray Data Meta-Analysis. Biology (Basel), DOI 10.3390/biology15040345. PMID 41744654 / PMC12938635. Code: https://github.com/imichalop/Meta-Analysis (own code, P16 own-repo) pinned commit 8a418c313171667538479db2c452996de37afa2c (pushed 2026-01-30).

NOTE ON THE BRIEF'S DATA ACCESSION

The room brief lists geo:GSE28146 as the dataset. The paper explicitly EXCLUDED GSE28146 ("excluded after screening the meta-data, as the samples were stored in paraffin blocks", Methods). GSE28146 was a harvest artifact, not an input. The paper's actual inputs are 8 studies / 10 sub-studies: GSE48350 (Hip), GSE39420, GSE36980 (FrontCort/Hip/TempCort), GSE16759, GSE1297, GSE5281 (PCC), GSE12685, E-MEXP-2280.

Pipeline stages (Methods)

  1. Preprocessing — Affymetrix Power Tools + RMA with BrainArray custom CDFs; GEOquery/oligo (R). [raw CEL stage]
  2. Batch correction — SVA. QC by NUSE (1.05±0.10) / RLE (0.0±0.2). [manual QC]
  3. Per-study DEG — limma topTables; org.Hs.eg.db annotation; BH FDR. [pipeline]
  4. Meta-analysis — Mosteller–Bush (weighted Stouffer variant) combining the per-study limma p-values across the 10 sub-studies, then BH FDR; significance cutoff adjP < 0.001. → headline DEG list. [pipeline — Parsers/Meta-Analysis.php]
  5. Enrichment / networks — WebGestalt 2024 (ORA), STRING v12, Cytoscape stringApp. [external web tools — OUT OF SCOPE, not scriptable/deterministic here]

IN SCOPE (attempted) — the meta-analysis stage (stage 4)

The repo ships Metanalysis Demo/ containing the exact per-study limma DEG lists used in the article plus studies.txt. The README states: "the meta-analysis script can be run there, to replicate the meta-analysis results." This makes stage 4 a fully self-contained, deterministic 1:1 reproduction with no raw-data download, no APT, no SVA — the cleanest low-hanging pipeline output and the paper's headline number.

Target claims (paper Section 3 / abstract):

  • Total DEGs at adjP<0.001: 4218
  • Up-regulated (over, "O"): 1944
  • Down-regulated (under, "U"): 2274

OUT OF SCOPE (not attempted) — and why

  • Stages 1–3 (RMA/BrainArray/SVA/limma from raw CEL): the per-study DEG lists are shipped as the demo inputs, so regenerating them from CEL is the hard, optional last ~20% and is unnecessary to reproduce the headline meta-analysis number. Not attempted (80/20 rule). The shipped lists are taken as the pipeline's stage-3 output.
  • Stage 5 enrichment/PPI (WebGestalt/STRING/Cytoscape hub genes MYC, GAPDH; node/edge counts): external interactive web services, non-deterministic versions, not a scriptable reproduction. Out of scope.
  • Comparison-with-prior-meta-analyses Venn counts (380 common, etc.): derived from stage-4 output + external gene lists; out of scope for the 1:1 stage-4 check.
C1
Reported
4218
Reproduced
4218
exact
C2
Reported
1944
Reproduced
1944
exact
C3
Reported
2274
Reproduced
2274
exact

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 100/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟢3. Location of the main deviation
🟢4. Cause of the deviation
🟢5. Derivability / plausibility
🟢6. Severity of the deviation
🟢7. Core claim
🟢8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Concordant (toward reproduced)
All content-critical questions reproduced
-4 pts
From: Q7 · Core claim 🟢
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Every question reproduced
-1 pts
From: “every question reproduced”
Total score -7

The paper's headline meta-analysis result — 4218 significant DEGs at adjP<0.001 (1944 up, 2274 down) — reproduced exactly to the digit by running the authors' own shipped code on their own shipped per-study limma DEG lists, with a verified output hash. There is no deviation on any side: data, method, and endpoint are all 1:1. Only the upstream raw-CEL→limma stage and downstream WebGestalt/STRING enrichment were not attempted, but those are not needed for the headline counts since the per-study lists are shipped. A clean, fabrication-free exact reproduction.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

91.2 k
tokens (I/O) · 6.8 M incl. cache
11 min
runtime · 0.01 CPU-h
1.9 GB
peak RAM
1
HPC jobs
hummel
machine