Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

A single-cell compendium of human cerebrospinal fluid identifies disease-associated immune cell populations.

J Clin Invest · 2025
L1 78/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
✓ What held up
  • Nothing in this column.
What did not (or only partly)
  • 🔴Could not use the authors’ exact input data
  • 🔴Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
78/100
Reproducibility score
at the mean
vs. all fields · 1173 studies
🎯 Scores higher than 51% of all assessed papers rank 533 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

PARTIAL reproduction (described well enough; partly 1:1, partly gated). Paper = a Scanpy/Seurat scRNA-seq INTEGRATION compendium of human CSF (403,973-cell atlas, 51 immune subpopulations). BRIEF pointers were placeholders and are corrected here: real analysis code = github.com/rasmirnov/MS_CSF_project (NOT mojaveazure/seurat-disk, a generic I/O helper); primary newly-generated data = Synapse syn51730532 (login-gated, no creds), while GSE133028 is only 1 of 7 REUSED public datasets. WHAT REPRODUCED (open data, on «our HPC»): independently re-assembled the open GSE133028 GEO GEX matrices (38 scRNA-seq samples, 175,895 raw cells, 18 CSF + 20 PB) and ran the paper's own Scanpy 1.9.1 pipeline end-to-end -- QC -> 175,439 cells, 932 HVGs, Harmony integration on 'sample', Leiden res=1.0 -> 26 clusters. The published pipeline runs reproducibly and yields a sensible integrated decomposition on the open input (process reproduction; graded within-tol). Dataset profiling: GSE133028 open + complete + delivers-promised (grade B); Synapse deposit access-gated (uncheckable). WHAT WAS NOT ATTEMPTED: the integrated-atlas headline numbers (total cells, 51 subpops, per-lineage counts) -- these depend on the gated Synapse Counts/*.h5 (incl. dbGaP + new raw) and on stochastic upstream QC, so they are not reproducible from open data and were recorded as gated, NOT fabricated or graded as mismatch. Six SLURM jobs (env/dep iteration: scanpy 1.9.1 needs matplotlib<3.7, pandas<2, numpy<1.25; final clean run = «job»).

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 50
    assessed: 2026-06-20 ⛓ 7889a629b649
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-23
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-20
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The study tests the hypothesis that the number and transcriptional features of myeloid and lymphoid cell populations within blood and CSF reflect disease states across a range of neurologic diseases.

Core claims
  • Integration of public and newly generated scRNA-seq datasets yields a compendium of 139 subjects (193 samples, 403,973 immune cells) spanning CSF and blood across healthy controls and multiple neurologic diseases. resource
  • A previously undescribed AREG+ cDC2 dendritic cell subset is present exclusively in CSF (not blood) and is increased in frequency in multiple sclerosis. finding
  • CSF microglia-like cells arise from peripheral CD14+ monocytes through a border-associated macrophage (BAM) intermediate, based on pseudotime trajectory inference. mechanism
  • The FN1+ microglia-like cell subcluster in CSF is uniquely increased in neurodegenerative diseases compared with healthy controls. finding
  • CD4+ T cell proportion is statistically significantly higher in CSF than PBMC overall, and is dramatically elevated in the CSF of MS subjects. finding
  • B cell proportion is markedly higher and myeloid cell proportion lower in CSF of MS subjects compared with healthy controls. finding
  • NK cell proportion in CSF is significantly elevated in neurodegenerative disease subjects compared with healthy controls. finding
  • Flow cytometry confirms a distinct HLA-DR+BDCA-2–XCR1–CLEC9A–CD1c+FCER1A+CD32B+AREG+ dendritic cell population enriched in CSF versus blood in MS subjects. method
Experimental setups
Assay System Perturbation Readout Platform
scRNA-seq (integrated, mostly 10x Genomics 5' or 3') human CSF and PBMC samples (healthy controls and multiple neurologic diseases: MS, Alzheimer's, Parkinson's, COVID-19, autoimmune encephalitis, etc.) disease state (none/various) immune cell subset frequency and transcriptional profile 10x Genomics
trajectory/pseudotime inference (computational, on scRNA-seq data) human CSF myeloid cells (CD14+ monocytes, BAMs, microglia-like subclusters) none transcriptional state as a function of pseudotime, lineage relationships
flow cytometry human CSF and blood from 4 MS subjects disease (MS) surface expression of DC markers (HLA-DR, BDCA-2, XCR1, CLEC9A, CD1c, FCER1A, CD32B, AREG) and cell frequency
single-cell subclustering (computational, on scRNA-seq data) human B cells and plasmablasts from CSF and PBMC disease group (HC, MS, ND, INF, OID) subset composition and frequency differences between tissue compartments and disease groups
Key results
  • 403,973 total immune cells profiled (195,431 PBMC, 208,542 CSF) from 193 samples across 139 subjects.
  • AREG+ cDC2 population identified transcriptomically and by flow cytometry, exclusive to CSF and increased in MS vs HC.
  • Continuous transcriptional trajectory observed from CD14+ CSF monocytes to BAMs to CCL2+/FN1+/SPP1+ microglia-like subclusters.
  • FN1+ microglia-like cell frequency higher in neurodegenerative disease CSF compared with HC.
  • CD4+ T cell proportion significantly higher in CSF vs PBMC overall; dramatically elevated in MS CSF.
  • B cell proportion higher and myeloid cell proportion lower in MS CSF vs HC.
  • NK cell proportion significantly elevated in CSF of neurodegenerative disease subjects vs HC.
  • CD16+ monocytes were virtually absent in MS CSF despite being present in blood; other myeloid subsets (CD14+ Mono, microglia-like) also underrepresented in MS CSF vs HC.
Key statistics
  • count 403,973 cells (195,431 PBMC; 208,542 CSF) (total integrated dataset size)
  • count 139 subjects; 135 CSF and 58 blood samples (compendium cohort composition)
  • pvalue adjusted P = 0.056 (trend for higher CD4+ T cell frequency in MS vs HC CSF, not statistically significant)
  • count 59,770 myeloid cells (36,450 PBMC; 23,320 CSF) (myeloid cell object size)
  • count 204,738 CD4+ T cells (75,776 PBMC; 128,962 CSF) (CD4+ T cell cluster, largest object)
  • count 25,129 B cells/plasmablasts (21,387 PBMC; 3,742 CSF) (B cell and plasmablast subclustering)
  • count 1,343 plasmablasts (579 blood; 764 CSF) (plasmablast subclusters: pre-, IgA+, IgG+)
  • count 6,330 CSF microglia-like cells (CCL2+, FN1+, SPP1+ microglia-like subclusters)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This is a single-cell RNA-seq compendium study integrating scRNA-Seq datasets from multiple published sources plus newly generated samples (139 subjects, 403,973 cells) comparing immune cell subset frequencies between PBMC and CSF compartments and across disease groups (HC, MS, neurodegenerative disease, infectious CNS disease, other inflammatory disease). Differences in cell-type proportions between tissues and disease groups are repeatedly described as 'statistically significant,' including at least one instance of an adjusted P value (0.056) for a CD4+ T cell comparison, indicating a multiple-comparison adjustment was applied to these frequency comparisons; the specific statistical test(s) and correction method are not stated in the portion of the text provided. A novel AREG+ cDC2 subpopulation identified transcriptomically was also validated by flow cytometry in CSF versus blood from 4 MS subjects.

Replicationbiological Sample sizeCohort described as 139 subjects (135 CSF and 58 blood samples in one integration; elsewhere 193 samples/403,973 cells with 195,431 PBMCs and 208,542 CSF cells); per-disease-group sample counts referenced via Table 1 but not given as numbers in the provided text; flow cytometry validation used 4 MS subjects GroupsPBMC vs CSF; and HC vs MS vs neurodegenerative disease (ND) vs infectious CNS disease (INF) vs other inflammatory disease (OID) Pairingmixed Randomization/blindingnot stated Dispersionunclear Multiplicity correctionnot stated (method name not given in available text)
Statistical tests used
Test Applied to n Assumptions
not specified in available text (described only as differences being 'statistically significant') cell-type proportion comparisons between PBMC vs CSF, and CSF vs HC across disease groups (e.g., Figure 2D, 2E, 3F, 4E, 5E, 5F) not stated
Approaches that could also have been used
  • Differences in immune cell-type proportions across many pairwise group comparisons (PBMC vs CSF, and each disease group vs HC) are each reported as statistically significant or not, with one adjusted p-value shown.
    Could also: A single omnibus test (e.g., ANOVA or Kruskal-Wallis across all disease groups) followed by a post-hoc test with a family-wise or FDR correction (e.g., Tukey HSD, Dunn's test with Benjamini-Hochberg) could also be used — This approach explicitly controls the overall false-positive rate across all group comparisons within a cell type, which can be a useful complement when many disease groups and many cell populations are each compared against a common reference.
  • Cell-type proportions (compositional data that sum to 100% within a sample) are compared using what appears to be standard significance testing on proportions.
    Could also: Methods designed specifically for single-cell compositional data, such as scCODA, propeller, or a beta-binomial/Dirichlet-multinomial regression framework, could also be applied — These approaches account for the compositional (sum-constrained) nature of cell-type frequency data and for variable cell numbers per sample, which can provide additional robustness for frequency comparisons derived from scRNA-Seq data.
  • The study integrates datasets from multiple independent published sources with newly generated samples into a shared analysis space.
    Could also: A mixed-effects or hierarchical model with dataset/batch as a random effect could also be used for the downstream frequency comparisons — This would allow explicit modeling of between-dataset/study variability (e.g., differing sequencing chemistry noted as 5' vs 3', or cohort-specific technical factors) alongside the biological group effect of interest.
  • Some comparisons (e.g., PBMC vs CSF) involve samples that are 'almost exclusively paired,' while others (e.g., across disease groups) are between independent subjects.
    Could also: A paired analysis approach (e.g., paired t-test/Wilcoxon signed-rank test or a mixed model with subject as a random effect) could also be used specifically for the paired PBMC-CSF comparisons — Explicitly modeling the paired structure when the same subject contributes both a PBMC and CSF sample can increase statistical power by accounting for between-subject variability.
  • One comparison (CD4+ T cell frequency, MS vs HC) is reported with an adjusted p-value of 0.056, described as a strong trend that did not reach statistical significance.
    Could also: Reporting an effect size with a confidence interval (e.g., difference in proportions with a 95% CI) alongside the p-value could also be presented — This would let readers gauge the magnitude and precision of the observed difference independent of the specific significance threshold used.
  • Trajectory inference (pseudotime) is used to support a hypothesis about microglia-like cell ontogeny from monocytes via a BAM intermediate.
    Could also: Complementary lineage-tracing or orthogonal validation approaches (e.g., RNA velocity, or independent fate-mapping/experimental validation) could also be used alongside pseudotime inference — Pseudotime methods infer likely trajectories from a snapshot of transcriptional states; combining them with an orthogonal method can provide converging evidence for a proposed differentiation path.
Software: not specified in available text (single-cell integration/analysis pipeline implied, e.g. for UMAP, clustering, trajectory inference, but no package/tool named)

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — PMID 39744938

Paper: A single-cell compendium of human cerebrospinal fluid identifies disease-associated immune cell populations. JCI 2025 (PMCID PMC11684814, DOI 10.1172/jci177793). Smirnov/Martyomov/Wu lab (WashU).

Type: scRNA-seq meta-analysis / integration compendium — combines 9 source datasets (7 public GEO + 1 "all_ris"/Tabula reference + 1 dbGaP controlled) plus newly generated raw scRNA-seq, integrated into a 403,973-cell atlas.

Corrected pointers (the BRIEF's pointers were wrong)

  • Code: BRIEF said github.com/mojaveazure/seurat-disk (a generic Seurat I/O helper — NOT this paper's code). Real analysis code: github.com/rasmirnov/MS_CSF_project (R + Python, Snakemake workflows; cloned to «infra», commit on main 25dc83b).
  • Data: BRIEF said GSE133028 — that is only ONE of seven reused public datasets. Primary newly-generated data: Synapse syn51730532 (login-gated; no DUA/access-requirement, but anonymous users get metadata-only → file download needs a Synapse account we do not have).

Pipeline (from repo + Methods)

Per-study processing (study.R, reanalysis.R): Seurat v4.2.0, SCTransform, miQC (adaptive mito filter) + DoubletFinder (doublet removal), CellCycleScoring, anchor integration. Whole-atlas integration (integration.py): Scanpy v1.9.1 — normalize_total(1e4) → log1p → HVG(min_mean .0125, max_mean 3, min_disp .5, batch_key=sample, nbatches>5) → regress_out (nCount_RNA, percent.mito) → scale(10) → PCA(arpack) → Harmony integrate on sample → neighbors(n=10, pcs=30) → UMAP → Leiden (res 0.4–1.0). Trajectory: Slingshot/Monocle3. DEA: pseudobulk.

In scope (pipeline-derived, attemptable)

Result Pipeline Reproducible from open data?
Source datasets exist & deliver promised scRNA-seq GEO retrieval YES (open) — profiled
GSE133028 single-cell cell count / clustering (assigned dataset) Scanpy (paper params) YES (open GEO GEX matrices) — computed on «our HPC»
Per-source sample/subject accounting metadata PARTIAL (GEO sample lists open)

Out of scope / blocked

  • Headline atlas numbers (403,973 cells; 51 subpopulations; per-lineage counts myeloid 59,770 / CD4 204,738 / CD8 75,649 / B 25,129 / NK 29,894 / γδ 6,271 / microglia 6,330): require the Synapse-deposited Counts/*.h5 (the authors' post-QC per-study matrices, including the dbGaP counts and the new raw data) → all login-gated on Synapse with no account available → NOT reproducible from open data. The repo's Snakefiles hardcode internal WashU paths (/storage1/fs1/martyomov/...) for these inputs; they are not shipped.
  • Bit-exact per-study post-QC counts: upstream QC uses miQC + DoubletFinder (data-adaptive, stochastic) → not bit-reproducible without the deposited objects.
  • Wet-lab (10x library prep, sequencing), the online navigator app — out of scope.

Verdict shape

PARTIAL: open source data profiled + assigned dataset (GSE133028) independently reprocessed with the paper's own scanpy pipeline on «our HPC»; the integrated-atlas headline numbers are blocked by login-gated Synapse inputs (data_restricted-adjacent: open-with-account, no creds).

gse133028_reused_open
Reported
GSE133028 = 1 of 7 reused public GEO datasets (scRNA-seq human CSF/blood); 98 GEO samples
Reproduced
Confirmed: GSE133028 retrieved + parsed on «our HPC»; 38 scRNA-seq GEX matrices (18 CSF + 20 PB), 175,895 raw cells x 33,538 genes
exact
gse133028_sc_pipeline
Reported
Paper's Scanpy 1.9.1 integration pipeline (normalize1e4->log1p->HVG->regress->scale->PCA->Harmony(sample)->Leiden)
Reproduced
Ran end-to-end on «our HPC»: 175,895 raw -> 175,439 QC-pass cells -> 932 HVGs -> Harmony integration on 'sample' (succeeded) -> 26 Leiden clusters (res 1.0)
within tolerance
total_cells_403973
Reported
403,973 cells in integrated atlas (51 subpopulations; per-lineage counts)
Reproduced
NOT reproducible from open data: inputs login-gated on Synapse syn51730532 (incl. dbGaP + new raw data); repo Snakefiles hardcode internal WashU paths
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 78/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🔴1. Data identity
🔴2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

341.1 k
tokens (I/O) · 28 M incl. cache
142 min
runtime · 1.53 CPU-h
35.3 GB
peak RAM
6 (3 failed)
HPC jobs
hummel
machine