Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Integrating transcriptomic datasets across neurological disease identifies unique myeloid subpopulations driving disease-specific signatures.

Glia · 2022
L1 26/100 PQI 80
⚑ Flagged for review — a reproduced result did not match the reported value

Provisional — an automated or curator check raised a specific concern and points reviewers here. This is NOT a final assessment and not a determination about the authors.

Why this verdict

The main result did not reproduce in this reproduction attempt. Where our recomputation produced values that differ from the published ones, those discrepancies are listed below. This is a single automated attempt — not peer review and not a finding of error or misconduct — and differences can also arise from data access, undocumented parameters or the computing environment. The verdict can be contested via “report an error”.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Main result did not reproduce
Decisive
From: Q5 · Derivability / plausibility 🔴
✓ What held up
  • Reported values were directly comparable
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🟡A deviation arose in the data or preprocessing
  • 🔴A deviation was attributed to the published material
  • 🔴Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
26/100
Reproducibility score
2.7 SD below mean
vs. all fields · 1173 studies
🎯 Scores higher than 1% of all assessed papers rank 1158 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Paper is a re-analysis of 10 public GEO datasets with NO authors' code repo (Data Availability: no new data generated; the auto-mined ncbi/TPMCalculator is only a tool the bulk arm mentions). Per P16 I reproduced the described Seurat v4 scRNA QC (250-25000 genes/cell, <5% mito) exactly on the paper's own deposited GSE98969 (Keren-Shaul MARS-seq) tables on «our HPC» («job»), targeting Table-2 QC cell counts -- the clearest deterministic pipeline number. Result: PARTIAL/different. (1) The stated <5% mito criterion is NON-BINDING -- the deposited MARS-seq tables have ~0% mito reads (median 0.0%), so it removes nothing and cannot explain the reported ~5% QC loss [FLAG]. (2) The gene-count filter on deposited data gives ~67% survival, not the reported ~95%, and the reproduced after-QC counts (8930; 2008) are BELOW the reported after-QC (9636; 2693) -- the paper's pipeline detects more genes/cell than the deposited UMI tables yield, so the reported counts are not reconstructable from the deposited data under the Methods as written. (3) The 'before QC' cell selection is under-specified: no genotype/timepoint subset reproduces 10146/2820 (closest: 5xFAD >=200-gene cells = 10034, ~1%). NOT a fabrication but a documented reproducibility gap. NOT ATTEMPTED (out of scope, the hard >20%): the 6-dataset LIGER integration -> 8 microglial/4 monocyte clusters (Fig 4), downsampled 2814+1407 cells, Cd81 25x fold-change (Fig 7c), gene sets A=319/B=7804, 192/119 cross-disease DE genes, 12/22 membrane markers, and the bulk RNA-seq DESeq2/TPMCalculator arm -- all non-deterministic or under-specified with no authors' code.

💻 Code ↗ 🗄 Data: GSE98969

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 26
    assessed: 2026-06-15 ⛓ 72a2b978426e
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-15
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

Is there a universal microglial/myeloid disease signature conserved across CNS pathologies, or are myeloid responses disease-specific? The paper tests this by meta-integrating transcriptomic datasets of microglia and monocyte-derived cells across diverse neurological disease models.

Core claims
  • The bulk microglial and monocyte transcriptomic program is highly contingent on the disease environment, challenging the notion of a universal microglial disease signature finding
  • Disease-specific bulk signatures are driven by differing proportions of unique myeloid subpopulations that are individually expanded in different disease settings finding
  • scRNA-seq integration defines a conserved core set of functional myeloid states: neurodegeneration-associated, inflammatory, interferon-responsive, phagocytic, antigen-presenting, and LPS-responsive finding
  • Cd81 is a neuroinflammatory-stable, microglia-enriched gene that distinguishes microglia from monocyte-derived cells across all examined CNS disease models at bulk and single-cell levels finding
  • A meta-analysis resource integrating bulk and single-cell transcriptomic datasets of microglia and monocytes across CNS disease (autoimmunity, neurodegeneration, sterile injury, infection) resource
  • High-parameter data integration of transcriptomes using clustering and dimensionality reduction to compare myeloid profiles across disease models and cell types method
Experimental setups
Assay System Perturbation Readout Platform
Meta-analysis of bulk RNA-seq / microarray transcriptomic datasets Acutely isolated microglia and monocyte-derived cells from adult mouse brain/spinal cord across disease models (demyelinating, ischemic, neurodegenerative, traumatic injury, infectious) various disease models (none of own; reanalysis) Gene expression / differential expression between cell types and diseases Galaxy: HISAT2, featureCounts, DESeq2, CutAdapt, FastQC, TPMCalculator
Integration of single-cell RNA-sequencing datasets Microglia and monocytes from six/multiple scRNA-seq datasets across CNS disease (mouse) various disease models (reanalysis) Single-cell transcriptomes; identification of myeloid subpopulations/states Seurat v4, LIGER v0.5.0, Spectre
Spectral flow cytometry C57BL/6J mouse brain, spleen, bone marrow West Nile virus (Sarafend) intranasal infection vs non-infected Surface/intracellular marker expression on myeloid cells (e.g., CD81, P2RY12, CD45, CD11b) 5-laser Aurora Spectral cytometer (Cytek Biosciences); FlowJo v10.8
Key results
  • Microglia and monocyte transcriptomes are highly divergent across pathologies, with no single universal disease signature
  • Unique myeloid subpopulations are differentially/proportionally expanded in different disease contexts, generating disease-specific bulk signatures
  • Six functionally-defined myeloid cellular states conserved across CNS pathology identified from scRNA-seq integration 6 states
  • Cd81 accurately identifies microglia and distinguishes them from monocyte-derived cells across all experimental models at bulk and single-cell level
  • DAM/MGnD signatures characteristically downregulate homeostatic genes (P2ry12, Tmem119, Cx3cr1) and upregulate inflammatory genes (Trem2, Apoe, Axl, Lpl, Itgax, Clec7a)
Key statistics
  • count Ten transcriptomic datasets from eight separate studies integrated (Datasets integrated for signature analysis and differential expression)
  • count Six single-cell RNA-sequencing datasets integrated (scRNA-seq integration revealing disease-specific subpopulations)
  • count minimum of 600 single cells passing quality-control (scRNA-seq dataset inclusion threshold)
  • count 1.2 × 10^5 plaque-forming units (WNV intranasal infection dose in 10 μl PBS)
  • count 19 samples excluded from Keren-Shaul et al. 2017 (Dataset sample exclusions for meta-analysis)
  • count at least three mice per group (Flow cytometry experimental groups (infected/non-infected))

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This meta-analysis integrates 10 bulk RNA-seq and 6 scRNA-seq transcriptomic datasets of microglia and monocyte-derived cells from multiple CNS disease models; differential expression was assessed with DESeq2 (Wald test) and cross-study gene ranking with RankProd, while six scRNA-seq datasets were jointly integrated using LIGER and analyzed with Seurat to define myeloid subpopulations. GO enrichment was evaluated via topGO and ViSEAGO, and gene-set overlap significance via the SuperExact test. A small experimental flow cytometry validation was performed in a West Nile virus mouse model (n ≥ 3 per group, randomly assigned) to corroborate transcriptomic findings at the protein level.

Replicationbiological Sample sizeAt least 3 mice per group for flow cytometry validation; scRNA-seq datasets required minimum 600 cells post-QC; 10 bulk and 6 scRNA-seq datasets from 8 published studies integrated GroupsMicroglia vs. monocyte-derived cells; disease conditions vs. homeostatic controls across autoimmunity, neurodegeneration, sterile injury, and infection models Pairingunpaired Randomization/blindingstated Dispersionunclear
Statistical tests used
Test Applied to n Assumptions
DESeq2 Wald test (negative binomial GLM) Differential expression analysis between microglia and monocyte-derived cells across bulk RNA-seq disease datasets 10 bulk RNA-seq datasets from 8 studies; per-comparison sample n not stated in provided text not stated
RankProd (rank product meta-analysis) Cross-study aggregation of transcriptional signatures to identify conserved and disease-specific genes not stated
SuperExact test (multi-set intersection significance) Testing significance of overlap between gene sets across datasets or disease conditions not stated
topGO / ViSEAGO (GO term enrichment; Fisher's exact or Kolmogorov-Smirnov, method not stated in provided text) Gene ontology enrichment of differentially expressed or cluster-marker gene lists not stated
LIGER iNMF integration + Seurat Louvain/Leiden clustering with UMAP/tSNE dimensionality reduction Integration and unsupervised clustering of 6 scRNA-seq datasets to define myeloid subpopulations Minimum 600 cells per dataset (post-QC inclusion threshold) not stated
Approaches that could also have been used
  • Cross-study aggregation of differential expression results was performed using RankProd applied to multiple separate DESeq2 outputs
    Could also: A mixed-effects meta-regression model (e.g., via the metafor R package) applied to log2 fold-change estimates and their standard errors from each study could also serve as a cross-study aggregation framework — Mixed-effects meta-regression explicitly models and reports between-study heterogeneity (I², τ²), making it straightforward to quantify how much variability in a gene's effect is attributable to between-study differences vs. sampling error, and to formally test whether disease type or cell-sorting method explains heterogeneity
  • Six scRNA-seq datasets were integrated using LIGER (integrative NMF)
    Could also: Harmony, Scanorama, or Seurat's RPCA/WNN integration could also be applied for multi-dataset scRNA-seq batch correction and integration — Different integration methods rest on different assumptions (linear vs. nonlinear correction, shared vs. dataset-specific factors); cross-method comparison is a common robustness check to confirm that identified subpopulations are not specific to one integration algorithm
  • GO term enrichment was assessed via topGO and ViSEAGO using a defined gene list (threshold-selected)
    Could also: Gene Set Enrichment Analysis (GSEA) or fgsea using the full pre-ranked gene list could also be applied for pathway enrichment — GSEA operates on the complete ranked list rather than a binary significant/non-significant threshold, reducing sensitivity to the chosen cutoff and potentially detecting moderate but concordant pathway signals that threshold-based methods may miss
  • Multi-set gene-list intersections were tested for significance using the SuperExact test
    Could also: Pairwise hypergeometric or Fisher's exact tests on each pair of gene lists, with FDR correction across comparisons, could also quantify overlap significance — The SuperExact test handles simultaneous multi-set intersections; reporting pairwise hypergeometric P-values alongside would make results interpretable to readers more familiar with that simpler and widely used formulation
  • The experimental flow cytometry validation used n ≥ 3 mice per group with no stated a priori power calculation
    Could also: An a priori power analysis based on an expected effect size and variance estimate (from pilot data or literature) could also be reported to justify the chosen group size — Reporting a power calculation or the assumed effect size clarifies the study's designed sensitivity, which aids interpretation of non-significant findings and helps readers judge whether observed differences of a given magnitude were likely to be detected
  • Bulk datasets from multiple studies were normalized to TPM using TPMCalculator and then integrated for cross-study comparison
    Could also: Batch-aware normalization at the count level (e.g., ComBat-seq, or joint TMM normalization) could also be applied before cross-study analysis to explicitly model inter-study technical variation — TPM normalization corrects for library size and gene length within a sample but does not remove between-batch technical effects arising from different sequencing platforms, library preparation protocols, or cell-sorting strategies across studies; explicit batch correction would make residual technical variance more transparent
Software: DESeq2 Galaxy Version 2.11.40.6+galaxy1 · Seurat 4 · LIGER 0.5.0 · RankProd 2.0 · topGO 2.44.0 · ViSEAGO 1.6.0 · SuperExactTest (R) · Spectre (R) · FlowJo 10.8 · pheatmap (R) 1.0.12 · featureCounts Galaxy Version 2.0.1 · HISAT2 Galaxy Version 2.1.0+galaxy7 · CutAdapt Galaxy Version 3.4+galaxy0 · FastQC Galaxy Version 0.72+galaxy1 · TPMCalculator · ggplot2 (R) 3.3.5 · UpSetR (R) 1.4.0 · Python 3.8.5 · Galaxy platform

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
14
Impact: medium
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

GSE59725 GEO in Methods (http://purl.org/orb/Methods)
also used by 1 paper:
GSE101686 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE101688 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE107792 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE120701 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE121654 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE127233 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE146113 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE162610 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE175430 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE98969 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet

Downstream reach in the literature

66 downstream papers · 11 datasets

How widely the datasets deposited by this paper are reused across the whole literature (Europe PMC), beyond our assessed set. This is a factual dependency map — reusing a public dataset is normal, good science. It is not a judgement on the downstream papers; the only verdict here is this paper's own, with its cited rationale.

This paper is currently under reproducibility review (see the verdict above). The map below shows where the data in question has propagated — so reuse can be traced, not so the downstream work is presumed affected.
GSE175430 GEO reused by 4 papers in the literature
Most-cited downstream papers:
GSE120701 GEO reused by 2 papers in the literature
Most-cited downstream papers:
GSE59725 GEO reused by 2 papers in the literature 1 assessed here
Most-cited downstream papers:
GSE101686 GEO reused by 1 papers in the literature
GSE101688 GEO reused by 1 papers in the literature
GSE107792 GEO reused by 1 papers in the literature
GSE127233 GEO reused by 1 papers in the literature
GSE146113 GEO reused by 1 papers in the literature

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-36527260

Paper: Wishart CL, Spiteri AG, Locatelli G, King NJC. Integrating transcriptomic datasets across neurological disease identifies unique myeloid subpopulations driving disease-specific signatures. Glia 2022. PMID 36527260 · PMC10952672 · DOI 10.1002/glia.24314.

Nature of the study

A meta-analysis / re-analysis integrating 10 public GEO datasets of CNS myeloid cells across neurological diseases. The authors state (Data Availability): "This study did not generate new reagents or datasets." There is no authors' code repository. The brief's auto-mined github.com/ncbi/TPMCalculator is a tool the paper mentions (used to compute TPM in the bulk-RNA-seq arm via Galaxy), NOT the authors' analysis code. Per brief rule P16, reproducing with the same third-party tools on the paper's own data is equally valid.

Pipelines described (Methods)

  • scRNA-seq: Seurat v4.0 — QC filter (250–25,000 unique genes/cell, <5% mito), sctransform normalization, PCA on top 2,000 HVGs, UMAP (29 PCs), graph-based FindClusters (default). Cross-dataset integration with LIGER v0.5.0 (k=12, λ=5). Markers via FindAllMarkers (Wilcoxon, padj<0.05, log2FC>0.25).
  • Bulk RNA-seq (Galaxy, usegalaxy.org.au): CutAdapt → FastQC → HISAT2 (mm10) → featureCounts → TPMCalculator (TPM) → DESeq2.
  • Cross-study: RankProd 2.0, topGO/ViSEAGO GO enrichment, UpSetR, pheatmap.

In scope (attempted)

S1 — Seurat scRNA-seq QC cell counts (Table 2) for GSE98969 (Keren-Shaul MARS-seq). This is the clearest, lowest-ambiguity pipeline-derived numeric result: applying the exact stated QC thresholds (250–25,000 genes/cell, <5% mito) to the paper's own deposited UMI tables and counting cells before/after. GSE98969 is the brief's auto-mined accession and contributes the 5xFAD and SOD1G93A arms.

  • Reported (Table 2): 5xFAD 10,146 → 9,636 (3,958 myeloid); SOD1G93A 2,820 → 2,693 (1,091 myeloid).
  • Method: stream the 97 MARS-seq amplification-batch UMI tables, map wells via the deposited experimental_design file, compute per-cell genes-detected and mito fraction, apply the QC thresholds, count. (The "after QC given before QC" survival is a deterministic function of the thresholds + cell subset, so it is exactly reproducible if the cell subset matches.)

Out of scope (NOT attempted — documented why)

  • Full 6-dataset LIGER integration → 8 microglial / 4 monocyte clusters (Fig 4), downsampled to 2,814 microglia + 1,407 MCs. Multi-dataset LIGER (k=12,λ=5) + graph clustering is highly seed-/preprocessing-sensitive and non-deterministic; with no authors' code the exact cluster count is the hard ">20%". Not attempted.
  • Cd81 "25-fold higher median expression in microglia vs MCs" (Fig 7c) and the 319/7,804 gene sets, 192/119 cross-disease DE genes, 12/22 membrane markers — all downstream of the integrated/clustered object above. Not attempted.
  • Bulk RNA-seq DESeq2 arm — needs raw FASTQ for an under-specified sample selection across multiple studies; Galaxy-version-pinned. Not attempted.
  • GO enrichment (topGO/ViSEAGO), RankProd — downstream of the above. Not attempted.

Known reproducibility risk (flagged up front)

The paper's per-dataset "before QC" counts do not map cleanly onto any single genotype/timepoint grouping in GSE98969 (5XFAD has 13,772 sorted single cells vs the paper's 10,146; SOD1 has 5,700 vs the paper's 2,820). The sample/plate selection that yields the paper's "before QC" number is under-specified and cannot be pinned without the authors' code (which does not exist). We therefore scan candidate subsets and report which, if any, reproduces the reported counts.

Figures / tables: TableFig 4Fig 4aFig 7cFig 7a
S1a_5xFAD_beforeQC
Reported
10146
Reproduced
13241 raw / 10034 at >=200 genes
partial
S1a_5xFAD_afterQC
Reported
9636
Reproduced
8930 (67.4% survival)
did not match
S1b_SOD1_beforeQC
Reported
2820
Reproduced
3040
partial
S1b_SOD1_afterQC
Reported
2693
Reproduced
2008 (66.1% survival)
did not match
S1c_mito_filter
Reported
<5% mito removes ~5% of cells (Table 2 survival)
Reproduced
0% removed; median mito_frac 0.0%; non-binding on deposited data
did not match

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 26/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟢2. Endpoint comparability
🟡3. Location of the main deviation
🔴4. Cause of the deviation
🔴5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Main result did not reproduce
Decisive
From: Q5 · Derivability / plausibility 🔴

This is a re-analysis paper with no authors' code and a Data Availability stating no new data; we applied the described Seurat QC to the authors' own deposited GSE98969 tables targeting Table-2 counts. The reported QC numbers are not derivable from the deposited data as written: the <5% mito criterion is provably inert (median 0.0% mito), gene-count survival is ~67% not ~95%, and reported after-QC (9636) even exceeds the deposited cells passing the stated threshold — a genuine reproducibility gap on the authors'/Methods side, not fabrication. The deviation sits in input/preprocessing bookkeeping and is moderate in absolute counts (within ~7–25%); the central myeloid-subpopulation conclusion (LIGER clusters, Fig 4/7) was out-of-scope and untested, so it is neither confirmed nor refuted.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

162 k
tokens (I/O) · 8.5 M incl. cache
15 min
runtime · 0.01 CPU-h
0.5 GB
peak RAM
1
HPC jobs
hummel
machine