Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Systematic review of human post-mortem immunohistochemical studies and bioinformatics analyses unveil the complexity of astrocyte reaction in Alzheimer's diseas

Neuropathol Appl Neurobiol · 2021
L1 92/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
✓ What held up
  • Nothing in this column.
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🟡Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
92/100
Reproducibility score
1.0 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 83% of all assessed papers rank 179 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Repo (github serrano-pozo-lab/astrocyte-review, commit a6cf893, GPL-3.0) ships analysis code only, NO Data/ folder. The key input - the ADRA protein set - was recovered from paper Table S1 sheet 'ADRA Protein Set' (196 proteins, 18 functional categories: Inflammation=26, Cytoskeleton=5, etc. - all matching the paper). EXACT 1:1 reproduction of the STRING v11.0b PPI network: 196 input -> 193 mapped (3 unmapped immunoglobulins IGHA1/IGHG1/IGHM, exactly as stated), 193 nodes, 2331 edges, avg degree 24.2, clustering 0.563, 836 expected edges, PPI-enrichment p<1e-16. The 15 eigen-centrality hub genes reproduced 15/15 identical including order. Enrichr ENCODE/ChEA TF enrichment puts ESR1 (#1) and CTCF (#4) in the top-10, matching the highlighted TFs. The Simpson GSE29652 microarray cross-validation hypergeometric reproduced same-significance (p=5.2e-3 vs reported 1.55e-2; ~3x off due to hgu133plus2.db annotation + RMA version drift since 2021, which the repo did not pin). NOT reproducible 1:1: Johnson (2.25e-13) & Grubman (3.45e-12) hypergeometric (authors' DEG CSVs unshipped; Johnson data Synapse-restricted), pathway enrichment Table S2 (manual GSEA web step), TFEA.ChIP Table S3 (unshipped Rdata). The systematic review itself is manual IHC curation - out of scope. Verdict: a DIFFERENT-but-faithful reproduction - the central bioinformatic pipeline (PPI network + hubs + TF enrichment) is described well enough to reproduce exactly; the omics cross-validation is only partially reproducible because intermediate artifacts were not deposited.

💻 Code ↗ 🗄 Data: 10.5281/zenodo.5140749

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 50
    assessed: 2026-06-18 ⛓ 745e6993408b
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-24
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The authors hypothesised that systematically compiling the post-mortem human neuropathological immunohistochemistry literature could produce a catalogue of dysregulated proteins in Alzheimer's disease reactive astrocytes (ADRA) around plaques and tangles, reveal the complexity of their functional changes, and inform development of fluid and PET biomarkers of astrocyte reaction.

Core claims
  • Systematic review of 306 eligible articles identified 196 distinct proteins constituting the ADRA (AD reactive astrocyte) protein set finding
  • Astrocyte reaction in AD is complex and heterogeneous, spanning 18 functional categories beyond cytoskeletal remodelling (e.g., inflammation, oxidative stress, lipid metabolism, proteostasis, ECM, neurotransmission, BBB integrity) finding
  • Increased GFAP immunoreactivity is the most frequently reported hallmark of astrocyte reaction in AD finding
  • CTCF and ESR1 emerged as potential transcription factors driving expression changes of the ADRA protein set finding
  • The ADRA protein set significantly overlaps with published transcriptomic and proteomic changes reported in AD brain and/or CSF finding
  • Findings were catalogued into a new public online resource, www.astrocyteatlas.org resource
  • Immunohistochemistry remains the gold-standard technique for capturing spatial expression patterns of astrocytes in post-mortem tissue, complementing transcriptomic/proteomic approaches that lack spatial information mechanism
  • A bioinformatics pipeline combining functional categorisation, PPI network analysis, pathway enrichment (GO/Reactome via MSigDB), and transcription factor enrichment (TFEA.ChIP, Enrichr) was applied to the ADRA protein set method
Experimental setups
Assay System Perturbation Readout Platform
Immunohistochemistry (literature-derived, systematic review) Post-mortem human Alzheimer's disease brain tissue AD vs control (disease state) Immunoreactivity/expression of candidate astrocyte marker proteins
Protein-protein interaction network analysis ADRA protein set (in silico, Homo sapiens) none Direct and indirect protein-protein interaction network structure STRING database v11.0
Pathway enrichment analysis (PEA) ADRA protein set (in silico) none Enriched Gene Ontology and Reactome pathways Molecular Signatures Database (MSigDB)
Transcription factor enrichment analysis ADRA protein set (in silico, ChIP-seq database-derived) none Enriched transcription factors potentially regulating ADRA markers TFEA.ChIP and Enrichr
Comparative transcriptomic/proteomic enrichment analysis Human control and AD brain and/or CSF ('omics datasets) AD vs control Overlap of ADRA set with differentially expressed genes/proteins (Fisher's exact test) and expression-level heatmaps
Key results
  • 1237 records identified via PubMed, APA PsycInfo and WoS-SCIE searches, narrowed through PRISMA screening to 306 eligible original articles
  • 306 articles rendered 196 proteins (ADRA protein set), most reported as upregulated in AD vs control brains
  • Increased GFAP immunoreactivity was the most frequently described hallmark of astrocyte reaction
  • Inflammation was the largest functional category with 26 markers (e.g., IL6, MAPK1/3/8, TNF), predominantly increased
  • Oxidative stress/antioxidant defence category (19 markers) showed mixed direction: most pro-oxidant/antioxidant enzymes increased (e.g., MT1A/2A, SOD1/2, PRDX6), but NFE2L2, SLC40A1 and HAMP decreased
  • CTCF and ESR1 identified as potential transcription factors regulating ADRA marker expression via TF enrichment analysis
  • ADRA protein set showed significant overlap with published transcriptomic and proteomic AD brain/CSF datasets
  • Lipid metabolism markers, especially APOE, CLU and LRP1, were reported as increased in ADRA by the majority of studies
Key statistics
  • count 306 (Original articles meeting eligibility criteria included in the systematic review)
  • count 196 (Distinct proteins identified as ADRA markers across included studies)
  • count 1237 (Records initially identified from PubMed, APA PsycInfo and WoS-SCIE database searches)
  • count 1067 (Unique records after deduplication screened by title/abstract)
  • count 391 (Records assessed for full eligibility after title/abstract screening)
  • count 26 (Number of protein markers classified under the Inflammation functional category)
  • count 22 (Number of protein markers classified under the Proliferation/apoptosis functional category)
  • count 19 (Number of protein markers classified under the Oxidative stress functional category)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This paper combines a PRISMA-guided systematic review of human post-mortem immunohistochemical studies with downstream bioinformatics analyses. The 196 AD reactive astrocyte (ADRA) proteins extracted from 306 eligible articles were subjected to pathway enrichment analysis against GO and Reactome databases via MSigDB, protein–protein interaction network construction via STRING v11.0, and transcription factor enrichment analysis via TFEA.ChIP and Enrichr. Overlap between the ADRA protein set and published transcriptomic and proteomic AD datasets was assessed with Fisher's exact test, and results were visualised as heatmaps.

Replicationunclear Sample size306 eligible articles identified via PRISMA screening; 196 proteins extracted; no formal a priori power calculation reported GroupsAD vs. control human post-mortem brains, as reported within each included primary study Pairingna Randomization/blindingnot stated Dispersionnone Exact p-valuesno Effect sizesno Confidence intervalsno Multiplicity correctionnot stated
Statistical tests used
Test Applied to n Assumptions
Fisher's exact test (over-representation/overlap test) Comparison of ADRA protein set against differentially expressed genes or proteins from published AD transcriptomic and proteomic datasets 196 ADRA proteins vs. external omics gene/protein sets; background universe size not stated in extracted text not stated
Pathway enrichment analysis (MSigDB; GO and Reactome gene sets) Functional annotation and validation of the 196 ADRA proteins against Gene Ontology and Reactome databases 196 ADRA proteins not stated
Transcription factor enrichment analysis (TFEA.ChIP and Enrichr; ChIP-seq-based) Identification of transcription factors potentially regulating expression of ADRA markers 196 ADRA proteins not stated
Protein–protein interaction network analysis (STRING v11.0) Mapping direct and indirect interactions among the 196 ADRA proteins 196 ADRA proteins na
Approaches that could also have been used
  • Overlap between the ADRA protein set and external omics datasets was assessed with Fisher's exact test, treating ADRA membership as a binary yes/no
    Could also: Gene set enrichment analysis (GSEA) or a permutation-based over-representation test ranked on differential expression statistics could also be applied to the same overlap question — Ranking-based methods such as GSEA use the full continuum of differential expression effect sizes rather than a binary membership cut-off, which can increase sensitivity to distributed, moderate-magnitude signals and avoids dependence on an arbitrary significance threshold for defining the reference gene set
  • The systematic review extracted directional information (increased/decreased/unchanged) for each marker but did not pool quantitative effect sizes across studies
    Could also: A formal meta-analysis pooling standardised mean differences in immunoreactive area or optical density across studies that reported quantitative data could also be performed where sufficient primary data existed — Quantitative synthesis would provide a pooled magnitude estimate and confidence interval for each marker, allow formal assessment of between-study heterogeneity, and weight studies by their precision—none of which is recoverable from direction-of-change summaries alone
  • Transcription factor enrichment was conducted with two ChIP-seq-based tools (TFEA.ChIP and Enrichr)
    Could also: Sequence-based motif enrichment tools such as HOMER, AME (MEME-Suite), or JASPAR-based promoter scanning could also be applied to the same gene list — Motif-based approaches assess whether known transcription factor binding motifs are statistically over-represented in promoter regions of the query gene set, providing complementary evidence—particularly for transcription factors that lack publicly available ChIP-seq data in the databases used
  • The 196 ADRA proteins were manually assigned to 18 functional categories based on published literature by the authors
    Could also: Data-driven clustering approaches such as community detection on the STRING network, hierarchical clustering on GO semantic similarity matrices, or topic modelling on gene–pathway membership could also generate functional groupings — Data-driven groupings reduce the influence of prior categorisation assumptions and may surface functional modules not immediately apparent from established literature, providing a complementary view alongside the expert-curated classification
  • The PPI network was constructed using STRING, which integrates both experimentally validated and computationally predicted interaction evidence
    Could also: Restricting the network to experimentally validated interactions only (e.g., BioGRID or IntAct filtered to co-immunoprecipitation or yeast two-hybrid evidence) could also be applied — Limiting to experimentally validated edges reduces the inclusion of predicted interactions that may not reflect physical binding, potentially improving specificity in hub and module identification at the cost of lower network coverage
  • Pathway enrichment was conducted against GO and Reactome gene sets via MSigDB
    Could also: Additional databases such as KEGG, WikiPathways, or disease-specific resources (e.g., DisGeNET) could also be interrogated alongside GO and Reactome — Different databases vary in coverage, curation philosophy, and pathway granularity; querying complementary resources can surface biologically relevant pathways that are absent or poorly represented in GO and Reactome and provides a cross-database consistency check on the enrichment results
Software: STRING (protein–protein interaction database) 11.0 · TFEA.ChIP · Enrichr · MSigDB (Molecular Signatures Database)

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-34297416

Paper: Viejo, Noori, Merrill, Das, Hyman, Serrano-Pozo (2021/2022). Systematic review of human post-mortem immunohistochemical studies and bioinformatics analyses unveil the complexity of astrocyte reaction in Alzheimer's disease. Neuropathol Appl Neurobiol 48(1):e12753. PMID 34297416 · PMC8766893 · DOI 10.1111/nan.12753

Code: https://github.com/serrano-pozo-lab/astrocyte-review (commit pinned at run time; default branch main, last push 2021-08-20, GPL-3.0). Zenodo 10.5281/zenodo.5140749 is a snapshot of the same repo (no separate data).

Repo content: 5 R-Markdown analysis scripts (R 4.1.0), rendered with eval=FALSE (so the shipped docs/*.html contain code only, no computed output values):

  • network-analysis.Rmd — STRING v11.0b PPI network of the ADRA protein set
  • omics-comparison.Rmd — Simpson microarray (GSE29652) DE + Johnson/Grubman + hypergeometric tests
  • pathway-enrichment.Rmd — MSigDB (GO/Reactome) overlap done on the GSEA web tool, then Jaccard clustering
  • tf-enrichment.Rmd — TFEA.ChIP + Enrichr ENCODE/ChEA TF enrichment

Critical data gap & resolution

The repo ships NO Data/ folder. Every script reads Data/ADRA Protein Set.csv (+ per-analysis input CSVs / CEL files / Rdata) that were never committed.

  • ADRA Protein Set RECOVERED from paper Table S1 (sheet "ADRA Protein Set"): 196 proteins × {Category, Symbol, Name, UniProtKB}. Verified: 196 proteins, 18 categories (cytoskeleton=5 … inflammation=26 — matches paper text), immunoglobulins IGHA1/IGHG1/IGHM and KIF21B present (the markers the code says STRING excludes). Rebuilt as data/ADRA_Protein_Set.csv (Symbol,Group,UniProtKB).
  • Supplementary tables downloaded (PMC PoW-gated): Table S1/S2/S3 → «infra».

IN SCOPE (pipeline-derived, attempted)

id result pipeline input availability
R1 STRING network: input 196, mapped 193/196, excluded IGHA1/IGHG1/IGHM; 193 nodes / 2331 edges; avg degree 24.2; clustering 0.563; PPI-enrichment p<1e-16 STRING v11.0b API + igraph ADRA set (recovered) + frozen STRING version → reproducible
R2 Network hub genes (eigen-centrality): IL6, TP53, CASP3, TNF, MAPK3, MAPK8, MAPK1, MYC, PTGS2, IGF1, APP, IL1B, CCL2, FGF2, ESR1 igraph eigen_centrality as R1
R3 Cross-validation hypergeometric p, Simpson et al. = 1.55e-2 GSE29652 RMA+limma DE → ∩ ADRA → phyper GSE29652 CEL (public, downloaded) + ADRA set
R4 TF enrichment (Enrichr ENCODE/ChEA): top TFs incl. CTCF, ESR1 Enrichr API on ADRA set ADRA set + Enrichr API

OUT OF SCOPE / BLOCKED (not attempted or partial — honest)

  • Systematic review itself (196 markers from 306 articles, IHC): manual literature curation, wet-lab/manual → out of scope.
  • R3 Johnson (p=2.25e-13) & Grubman (p=3.45e-12) hypergeometric: the precomputed DEG CSVs (Johnson DEGs.csv, Grubman DEGs ...csv) were NOT shipped. Johnson bulk/CSF proteomics = AMP-AD/Synapse (registered/restricted access). Grubman snRNA = GSE138852 (raw public) but the authors' DEG-calling pipeline is not in the repo. → not reproducible 1:1.
  • Pathway enrichment (Table S2): the enrichment step was run manually on the GSEA/MSigDB web tool (annotate.jsp); only the downstream Jaccard clustering is scripted and it needs the unshipped intermediate CSVs. → not reproducible 1:1.
  • TFEA.ChIP (Table S3): needs unshipped ReMap+GH_doubleElite.Rdata. Enrichr half (R4) is reproducible.

Datasets to profile

  • GSE29652 (Simpson microarray, 18 CEL) — public, downloaded.
  • AMP-AD/Synapse Johnson proteomics — restricted.
  • GSE138852 (Grubman snRNA) — public but DEG pipeline absent.
  • Zenodo 5140749 — code snapshot only.
  • Paper Table S1 (ADRA set + review) — the recovered key input.
Figures / tables: Table
R1-network-nodes-edges
Reported
193 nodes / 2331 edges (196 input, 193 mapped, excluded IGHA1/IGHG1/IGHM)
Reproduced
193 nodes / 2331 edges; 196 input, 193 mapped, unmapped = IGHA1,IGHG1,IGHM
exact
R1-avg-degree
Reported
24.2
Reproduced
24.16 (rounds to 24.2)
exact
R1-expected-edges
Reported
836 expected
Reproduced
836
exact
R1-clustering
Reported
0.563 (local clustering coefficient)
Reproduced
0.563
exact
R1-ppi-enrichment-p
Reported
< 1.0e-16
Reproduced
0 (below STRING reporting floor, i.e. < 1e-16)
exact
R2-hub-genes
Reported
IL6,TP53,CASP3,TNF,MAPK3,MAPK8,MAPK1,MYC,PTGS2,IGF1,APP,IL1B,CCL2,FGF2,ESR1
Reproduced
15/15 identical, same order
exact
R4-enrichr-tf
Reported
CTCF, ESR1 highlighted as top ENCODE/ChEA TFs
Reproduced
ESR1 (#1), CTCF (#4) in ENCODE_and_ChEA_Consensus_TFs top-10
within tolerance
R3-simpson-hyper
Reported
1.55e-2
Reproduced
5.21e-3 (phyper(q=34,m=2760,21306,196)); same significance call, ~3x lower; annotation/RMA version drift
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 92/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟡2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

183.2 k
tokens (I/O) · 15.9 M incl. cache
19 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.