A single-cell compendium of human cerebrospinal fluid identifies disease-associated immune cell populations.
The main results reproduced: recomputed values matched the published ones within tolerance.
- Nothing in this column.
- 🔴Could not use the authors’ exact input data
- 🔴Reported values were only indirectly comparable
- 🟡A deviation arose in the data or preprocessing
- 🟡A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The deviation was non-trivial in magnitude
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
PARTIAL reproduction (described well enough; partly 1:1, partly gated). Paper = a Scanpy/Seurat scRNA-seq INTEGRATION compendium of human CSF (403,973-cell atlas, 51 immune subpopulations). BRIEF pointers were placeholders and are corrected here: real analysis code = github.com/rasmirnov/MS_CSF_project (NOT mojaveazure/seurat-disk, a generic I/O helper); primary newly-generated data = Synapse syn51730532 (login-gated, no creds), while GSE133028 is only 1 of 7 REUSED public datasets. WHAT REPRODUCED (open data, on «our HPC»): independently re-assembled the open GSE133028 GEO GEX matrices (38 scRNA-seq samples, 175,895 raw cells, 18 CSF + 20 PB) and ran the paper's own Scanpy 1.9.1 pipeline end-to-end -- QC -> 175,439 cells, 932 HVGs, Harmony integration on 'sample', Leiden res=1.0 -> 26 clusters. The published pipeline runs reproducibly and yields a sensible integrated decomposition on the open input (process reproduction; graded within-tol). Dataset profiling: GSE133028 open + complete + delivers-promised (grade B); Synapse deposit access-gated (uncheckable). WHAT WAS NOT ATTEMPTED: the integrated-atlas headline numbers (total cells, 51 subpops, per-lineage counts) -- these depend on the gated Synapse Counts/*.h5 (incl. dbGaP + new raw) and on stochastic upstream QC, so they are not reproducible from open data and were recorded as gated, NOT fabricated or graded as mismatch. Six SLURM jobs (env/dep iteration: scanpy 1.9.1 needs matplotlib<3.7, pandas<2, numpy<1.25; final clean run = «job»).
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 50assessed: 2026-06-20 ⛓ 7889a629b649
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-23
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-20no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetThe study tests the hypothesis that the number and transcriptional features of myeloid and lymphoid cell populations within blood and CSF reflect disease states across a range of neurologic diseases.
- ★ Integration of public and newly generated scRNA-seq datasets yields a compendium of 139 subjects (193 samples, 403,973 immune cells) spanning CSF and blood across healthy controls and multiple neurologic diseases. resource
- ★ A previously undescribed AREG+ cDC2 dendritic cell subset is present exclusively in CSF (not blood) and is increased in frequency in multiple sclerosis. finding
- ★ CSF microglia-like cells arise from peripheral CD14+ monocytes through a border-associated macrophage (BAM) intermediate, based on pseudotime trajectory inference. mechanism
- ★ The FN1+ microglia-like cell subcluster in CSF is uniquely increased in neurodegenerative diseases compared with healthy controls. finding
- ★ CD4+ T cell proportion is statistically significantly higher in CSF than PBMC overall, and is dramatically elevated in the CSF of MS subjects. finding
- B cell proportion is markedly higher and myeloid cell proportion lower in CSF of MS subjects compared with healthy controls. finding
- NK cell proportion in CSF is significantly elevated in neurodegenerative disease subjects compared with healthy controls. finding
- ★ Flow cytometry confirms a distinct HLA-DR+BDCA-2–XCR1–CLEC9A–CD1c+FCER1A+CD32B+AREG+ dendritic cell population enriched in CSF versus blood in MS subjects. method
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| scRNA-seq (integrated, mostly 10x Genomics 5' or 3') | human CSF and PBMC samples (healthy controls and multiple neurologic diseases: MS, Alzheimer's, Parkinson's, COVID-19, autoimmune encephalitis, etc.) | disease state (none/various) | immune cell subset frequency and transcriptional profile | 10x Genomics |
| trajectory/pseudotime inference (computational, on scRNA-seq data) | human CSF myeloid cells (CD14+ monocytes, BAMs, microglia-like subclusters) | none | transcriptional state as a function of pseudotime, lineage relationships | — |
| flow cytometry | human CSF and blood from 4 MS subjects | disease (MS) | surface expression of DC markers (HLA-DR, BDCA-2, XCR1, CLEC9A, CD1c, FCER1A, CD32B, AREG) and cell frequency | — |
| single-cell subclustering (computational, on scRNA-seq data) | human B cells and plasmablasts from CSF and PBMC | disease group (HC, MS, ND, INF, OID) | subset composition and frequency differences between tissue compartments and disease groups | — |
- – 403,973 total immune cells profiled (195,431 PBMC, 208,542 CSF) from 193 samples across 139 subjects.
- ▲ AREG+ cDC2 population identified transcriptomically and by flow cytometry, exclusive to CSF and increased in MS vs HC.
- – Continuous transcriptional trajectory observed from CD14+ CSF monocytes to BAMs to CCL2+/FN1+/SPP1+ microglia-like subclusters.
- ▲ FN1+ microglia-like cell frequency higher in neurodegenerative disease CSF compared with HC.
- ▲ CD4+ T cell proportion significantly higher in CSF vs PBMC overall; dramatically elevated in MS CSF.
- – B cell proportion higher and myeloid cell proportion lower in MS CSF vs HC.
- ▲ NK cell proportion significantly elevated in CSF of neurodegenerative disease subjects vs HC.
- ▼ CD16+ monocytes were virtually absent in MS CSF despite being present in blood; other myeloid subsets (CD14+ Mono, microglia-like) also underrepresented in MS CSF vs HC.
- count 403,973 cells (195,431 PBMC; 208,542 CSF) (total integrated dataset size)
- count 139 subjects; 135 CSF and 58 blood samples (compendium cohort composition)
- pvalue adjusted P = 0.056 (trend for higher CD4+ T cell frequency in MS vs HC CSF, not statistically significant)
- count 59,770 myeloid cells (36,450 PBMC; 23,320 CSF) (myeloid cell object size)
- count 204,738 CD4+ T cells (75,776 PBMC; 128,962 CSF) (CD4+ T cell cluster, largest object)
- count 25,129 B cells/plasmablasts (21,387 PBMC; 3,742 CSF) (B cell and plasmablast subclustering)
- count 1,343 plasmablasts (579 blood; 764 CSF) (plasmablast subclusters: pre-, IgA+, IgG+)
- count 6,330 CSF microglia-like cells (CCL2+, FN1+, SPP1+ microglia-like subclusters)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This is a single-cell RNA-seq compendium study integrating scRNA-Seq datasets from multiple published sources plus newly generated samples (139 subjects, 403,973 cells) comparing immune cell subset frequencies between PBMC and CSF compartments and across disease groups (HC, MS, neurodegenerative disease, infectious CNS disease, other inflammatory disease). Differences in cell-type proportions between tissues and disease groups are repeatedly described as 'statistically significant,' including at least one instance of an adjusted P value (0.056) for a CD4+ T cell comparison, indicating a multiple-comparison adjustment was applied to these frequency comparisons; the specific statistical test(s) and correction method are not stated in the portion of the text provided. A novel AREG+ cDC2 subpopulation identified transcriptomically was also validated by flow cytometry in CSF versus blood from 4 MS subjects.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| not specified in available text (described only as differences being 'statistically significant') | cell-type proportion comparisons between PBMC vs CSF, and CSF vs HC across disease groups (e.g., Figure 2D, 2E, 3F, 4E, 5E, 5F) | — | not stated |
-
Differences in immune cell-type proportions across many pairwise group comparisons (PBMC vs CSF, and each disease group vs HC) are each reported as statistically significant or not, with one adjusted p-value shown.↳ Could also: A single omnibus test (e.g., ANOVA or Kruskal-Wallis across all disease groups) followed by a post-hoc test with a family-wise or FDR correction (e.g., Tukey HSD, Dunn's test with Benjamini-Hochberg) could also be used — This approach explicitly controls the overall false-positive rate across all group comparisons within a cell type, which can be a useful complement when many disease groups and many cell populations are each compared against a common reference.
-
Cell-type proportions (compositional data that sum to 100% within a sample) are compared using what appears to be standard significance testing on proportions.↳ Could also: Methods designed specifically for single-cell compositional data, such as scCODA, propeller, or a beta-binomial/Dirichlet-multinomial regression framework, could also be applied — These approaches account for the compositional (sum-constrained) nature of cell-type frequency data and for variable cell numbers per sample, which can provide additional robustness for frequency comparisons derived from scRNA-Seq data.
-
The study integrates datasets from multiple independent published sources with newly generated samples into a shared analysis space.↳ Could also: A mixed-effects or hierarchical model with dataset/batch as a random effect could also be used for the downstream frequency comparisons — This would allow explicit modeling of between-dataset/study variability (e.g., differing sequencing chemistry noted as 5' vs 3', or cohort-specific technical factors) alongside the biological group effect of interest.
-
Some comparisons (e.g., PBMC vs CSF) involve samples that are 'almost exclusively paired,' while others (e.g., across disease groups) are between independent subjects.↳ Could also: A paired analysis approach (e.g., paired t-test/Wilcoxon signed-rank test or a mixed model with subject as a random effect) could also be used specifically for the paired PBMC-CSF comparisons — Explicitly modeling the paired structure when the same subject contributes both a PBMC and CSF sample can increase statistical power by accounting for between-subject variability.
-
One comparison (CD4+ T cell frequency, MS vs HC) is reported with an adjusted p-value of 0.056, described as a strong trend that did not reach statistical significance.↳ Could also: Reporting an effect size with a confidence interval (e.g., difference in proportions with a 95% CI) alongside the p-value could also be presented — This would let readers gauge the magnitude and precision of the observed difference independent of the specific significance threshold used.
-
Trajectory inference (pseudotime) is used to support a hypothesis about microglia-like cell ontogeny from monocytes via a BAM intermediate.↳ Could also: Complementary lineage-tracing or orthogonal validation approaches (e.g., RNA velocity, or independent fate-mapping/experimental validation) could also be used alongside pseudotime inference — Pseudotime methods infer likely trajectories from a snapshot of transcriptional states; combining them with an orthogonal method can provide converging evidence for a proposed differentiation path.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — PMID 39744938
Paper: A single-cell compendium of human cerebrospinal fluid identifies disease-associated immune cell populations. JCI 2025 (PMCID PMC11684814, DOI 10.1172/jci177793). Smirnov/Martyomov/Wu lab (WashU).
Type: scRNA-seq meta-analysis / integration compendium — combines 9 source datasets (7 public GEO + 1 "all_ris"/Tabula reference + 1 dbGaP controlled) plus newly generated raw scRNA-seq, integrated into a 403,973-cell atlas.
Corrected pointers (the BRIEF's pointers were wrong)
- Code: BRIEF said
github.com/mojaveazure/seurat-disk(a generic Seurat I/O helper — NOT this paper's code). Real analysis code:github.com/rasmirnov/MS_CSF_project(R + Python, Snakemake workflows; cloned to «infra», commit onmain25dc83b). - Data: BRIEF said
GSE133028— that is only ONE of seven reused public datasets. Primary newly-generated data: Synapse syn51730532 (login-gated; no DUA/access-requirement, but anonymous users get metadata-only → file download needs a Synapse account we do not have).
Pipeline (from repo + Methods)
Per-study processing (study.R, reanalysis.R): Seurat v4.2.0, SCTransform, miQC (adaptive
mito filter) + DoubletFinder (doublet removal), CellCycleScoring, anchor integration.
Whole-atlas integration (integration.py): Scanpy v1.9.1 — normalize_total(1e4) → log1p →
HVG(min_mean .0125, max_mean 3, min_disp .5, batch_key=sample, nbatches>5) → regress_out
(nCount_RNA, percent.mito) → scale(10) → PCA(arpack) → Harmony integrate on sample →
neighbors(n=10, pcs=30) → UMAP → Leiden (res 0.4–1.0). Trajectory: Slingshot/Monocle3.
DEA: pseudobulk.
In scope (pipeline-derived, attemptable)
| Result | Pipeline | Reproducible from open data? |
|---|---|---|
| Source datasets exist & deliver promised scRNA-seq | GEO retrieval | YES (open) — profiled |
| GSE133028 single-cell cell count / clustering (assigned dataset) | Scanpy (paper params) | YES (open GEO GEX matrices) — computed on «our HPC» |
| Per-source sample/subject accounting | metadata | PARTIAL (GEO sample lists open) |
Out of scope / blocked
- Headline atlas numbers (403,973 cells; 51 subpopulations; per-lineage counts myeloid
59,770 / CD4 204,738 / CD8 75,649 / B 25,129 / NK 29,894 / γδ 6,271 / microglia 6,330):
require the Synapse-deposited
Counts/*.h5(the authors' post-QC per-study matrices, including the dbGaP counts and the new raw data) → all login-gated on Synapse with no account available → NOT reproducible from open data. The repo's Snakefiles hardcode internal WashU paths (/storage1/fs1/martyomov/...) for these inputs; they are not shipped. - Bit-exact per-study post-QC counts: upstream QC uses miQC + DoubletFinder (data-adaptive, stochastic) → not bit-reproducible without the deposited objects.
- Wet-lab (10x library prep, sequencing), the online navigator app — out of scope.
Verdict shape
PARTIAL: open source data profiled + assigned dataset (GSE133028) independently reprocessed with the paper's own scanpy pipeline on «our HPC»; the integrated-atlas headline numbers are blocked by login-gated Synapse inputs (data_restricted-adjacent: open-with-account, no creds).
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.