Spatial transcriptomics demonstrates the role of CD4 T cells in effector CD8 T cell differentiation during chronic viral infection.
The main results reproduced, with only marginal, non-material deviations.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Reported values were directly comparable
- 🟡Could not use the authors’ exact input data
- 🟡A deviation arose in the data or preprocessing
- 🔴A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The deviation was non-trivial in magnitude
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
P16 third-party-tool / own-data reproduction of Topchyan 2022 Cell Rep (CD4 T cells in CD8 differentiation, chronic LCMV). BRIEF metadata was wrong: the paper's own deposit is GSE200721 (scRNA CD4) + GSE200720 (Visium); GSE129139 is only one of five SPOTlight reference sets, and the named code (MarcElosua/SPOTlight 0.1.7) is a third-party tool ('this paper does not report original code'). Reproduced the scRNA-seq QC pipeline on GSE200721 with Seurat 4.3.0.1 on «our HPC» (SLURM 2177774): per-condition total cell counts reproduce in the RIGHT DIRECTION and 'comparable' relationship (control 16,014 vs depleted 17,715 after QC), but run ~12% above the reported 14,286/15,629. Root cause established and is itself the key finding: GSE200721 deposits ONLY Gene-Expression matrices (32,285 features, no Antibody-Capture/HTO matrix), so the authors' HTO singlet demultiplexing cannot be reproduced from the deposit and its thresholds are unstated -> the exact headline counts are NOT regenerable from the shipped data + stated methods. Described well enough to reproduce the PIPELINE but NOT the exact numbers => PARTIAL, honest 1:1 attempt. NOT attempted (documented): cluster counts depend on an unstated resolution and hit an r-matrix/SeuratObject env incompatibility (C2/C3, soft targets); the Visium 7-cluster result (C3, GSE200720); and the SPOTlight colocalization analysis (C4) which needs an assembled 5-dataset integrated reference (the hard 20%). Also flagged: the Data/Code-availability statement contains an unfilled 'GSE' placeholder.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 50assessed: 2026-06-15 ⛓ f215d73b0803
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-15
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusWhy do CD4 T cells that repopulate after transient CD4 depletion fail to rescue functional effector CD8 T cell responses during chronic LCMV Cl13 infection — are the repopulating CD4 subsets phenotypically/functionally inferior, numerically insufficient, or spatially mislocalized to provide help?
- ★ Following transient CD4 depletion, IL-21-producing Tfh cells do repopulate but are outnumbered by immunomodulatory CD4 T cells (Tregs, Th17), and IL-21+ frequency remains low within activated CD4 T cells. finding
- ★ Loss of CD4 T cell help disrupts splenic architecture, decreasing white pulp regions and causing germinal center losses. finding
- ★ Disrupted splenic architecture diminishes colocalization of Tfh and progenitor CD8 T cells, providing a potential mechanism for impaired progenitor-to-effector CD8 T cell differentiation under un-helped conditions. mechanism
- ★ Adoptive transfer of in-vitro-activated LCMV-specific IL-21-producing CD4 (SMARTA) cells fails to rescue effector CX3CR1+ CD8 T cell formation in CD4-depleted chronically infected mice. finding
- Combining spatial transcriptomics (10x Visium) with scRNA-seq characterizes CD4 heterogeneity and cellular colocalization during chronic infection. method
- ★ CD4 depletion reduces Tfh cells and CD95+GL7+ germinal center B cells in both frequency and absolute number. finding
- scRNA-seq of CD44+ CD4 T cells at 21 dpi identifies 13 clusters; Il21-expressing clusters (GC Tfh, pre-Tfh, activated memory) are reduced in the depleted group. finding
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Flow cytometry (IL-21-tRFP reporter, blood and splenocytes) | IL-21-tRFP reporter mice, LCMV Cl13 chronic infection | anti-CD4 depletion vs control | frequency of IL-21-tRFP+ cells within CD44+ CD4 T cells | — |
| scRNA-seq | CD44+ CD4 T cells from mouse spleen, LCMV Cl13, 21 dpi | CD4 depletion vs control | transcriptional clusters / cell-type frequencies (UMAP, DEGs) | — |
| Flow cytometry (Tfh and GC B cells) | mouse splenocytes, LCMV Cl13 | CD4 depletion vs control | CXCR5+BCL6+ Tfh and CD95+GL7+ GC B cell frequency and number | — |
| Flow cytometry (Treg and Th17) | mouse splenocytes, LCMV Cl13 | CD4 depletion vs control | Foxp3+CD44+ and RORγt+CD44+ CD4 T cell frequency and number | — |
| Adoptive transfer + flow cytometry | chronically infected CD4-depleted mice receiving in-vitro-activated SMARTA TCR-transgenic LCMV-specific CD4 T cells (IL-21-tRFP) | adoptive CD4 transfer into CD4-depleted hosts | IL-21-tRFP+ CD4 detection and effector CX3CR1+ CD8 T cell formation | — |
| Spatial transcriptomics (ST) | mouse spleens, LCMV Cl13, 7 and 21 dpi | CD4 depletion vs control | 55 μm spatial spot clusters, dominant cell-type marker gene expression, GC/Treg/Tfh gene localization | 10x Genomics Visium |
| Spatial deconvolution (SPOTlight) | mouse spleen ST spots, LCMV Cl13 | CD4 depletion vs control | colocalization of Tfh, B, and progenitor CD8 T cells | SPOTlight |
- ▼ At 8 dpi, repopulated CD4 T cells in depleted mice did not express IL-21-tRFP vs ~5% of CD44+ CD4 T cells in control ~5% in control vs none in depleted
- ▼ At 35 dpi, IL-21-tRFP+ frequency within splenic CD44+ CD4 T cells significantly lower in depleted group
- ▲ Th17 cluster nearly unique to depleted condition, ~7% of CD44+ CD4 T cells vs ~1% in control ~7% vs ~1%
- ▲ Majority of antigen-experienced CD4 T cells found in control mice 68% in control
- ▼ CD4-depleted mice had significantly reduced frequency and number of CXCR5+BCL6+ Tfh cells
- ▼ Significant reduction in frequency and absolute number of CD95+GL7+ GC B cells in depleted mice
- – Higher frequency of Foxp3+CD44+ Tregs repopulate after CD4 depletion (though lower total number per spleen)
- ▼ GC B cell genes Fas and Bcl6 significantly reduced in CD4-depleted spleens by ST; Il21 increased over time in control but unchanged in depleted
- count 13 clusters (CD4 T cell clusters identified by scRNA-seq UMAP)
- count 14,286 cells (control) and 15,629 cells (CD4 depleted) (total cells per condition in scRNA-seq)
- percent 68% (antigen-experienced CD4 T cells present in control mice)
- percent ~7% vs ~1% (Th17 cluster in depleted vs control CD44+ CD4 T cells)
- percent ~5% (CD44+ CD4 T cells expressing IL-21-tRFP in control at 8 dpi)
- count 7 distinct clusters of 55 μm spatial spots (Visium ST UMAP analysis)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
The paper combines scRNA-seq (UMAP-based clustering with differential gene expression to characterize CD4 T cell subsets), flow cytometry validation of key populations, and 10x Genomics Visium spatial transcriptomics with SPOTlight deconvolution to study CD4 T cell heterogeneity and splenic architecture in control versus CD4-depleted LCMV Cl13-infected mice. Two-group comparisons (control vs. CD4-depleted) are made across multiple assays and time points (7 and 21 dpi). Results are described as 'significantly' different throughout, but the statistical tests, exact p-values, and dispersion measures are not named in the provided text excerpt, which does not include a dedicated statistical-analysis section.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Not stated — differential gene expression for scRNA-seq cluster marker identification | Identification of top DEGs defining the 13 CD4 T cell clusters and 7 spatial clusters (Figures 1C–1D, 4C–4E) | 14,286 cells (control) and 15,629 cells (CD4-depleted) for scRNA-seq; not stated for spatial spots | not stated |
| Not stated — two-group comparison of cluster frequencies | Comparison of Il21-expressing cluster frequencies between control and CD4-depleted conditions (Figure 1E, 2C, 3C) | null | not stated |
| Not stated — two-group comparison of flow cytometry frequencies and absolute numbers | Tfh cell frequency and number (Figures 2D–2F), GC B cell frequency and number (Figures 2G–2I), Treg frequency and number (Figures 3D–3F), Th17 frequency and number (Figures 3G–3I), IL-21-tRFP+ cell frequency (Figures S1B–S1F) | null | not stated |
| Not stated — spatial gene expression comparison between conditions and time points | Comparison of Bcl6, Fas, Cd19, Ighd, Foxp3, Il21 expression across control vs. CD4-depleted spleens at days 7 and 21 (Figures 5A–5C, S3A–S3D) | null | not stated |
-
Differential gene expression for scRNA-seq cluster markers was performed, but the specific test is not stated in the available text↳ Could also: Commonly used alternatives include the Wilcoxon rank-sum test (Seurat default), MAST (mixed-effects model accounting for dropout), or DESeq2/edgeR on pseudobulk aggregates per biological replicate — Pseudobulk approaches (DESeq2/edgeR applied after aggregating cells per sample) treat the biological replicate as the statistical unit, reducing false-positive inflation that can arise when cells from the same animal are treated as independent observations; MAST explicitly models the bimodal distribution of single-cell expression data
-
Two-group flow cytometry comparisons (frequency and absolute number) between control and CD4-depleted mice are reported as significant, but no test is named↳ Could also: A two-tailed Student's t-test (if normality holds) or a Mann-Whitney U test (nonparametric) are standard choices for two-group unpaired comparisons in small-n mouse experiments — Naming the test allows readers to judge whether the assumptions (normality, equal variance) are met for the observed sample size and to reproduce the analysis; with small n (typical in mouse work), nonparametric tests are often preferred
-
Multiple flow cytometry populations are compared between conditions without a stated multiple-comparisons correction↳ Could also: A Benjamini-Hochberg FDR correction applied across all pairwise comparisons, or a Bonferroni correction for a smaller family, would also be applicable — When many cell populations are tested across the same experiment, controlling the false discovery rate reduces the probability that some reported differences are chance findings; reporting the correction scope makes the inference transparent
-
Dispersion and spread of flow cytometry measurements are not reported in the available text↳ Could also: Reporting SD (for normally distributed data) or IQR/individual data points (for small n) alongside group means or medians is standard practice — Dispersion measures allow readers to assess variability within groups and to judge effect magnitude relative to within-group spread, which is especially informative when biological n is small
-
Spatial transcriptomics spots (55 µm) were clustered using UMAP and compared between conditions by gene expression level↳ Could also: Spatial deconvolution methods such as RCTD (Robust Cell-Type Decomposition) or CARD (Conditional Autoregressive Deconvolution), in addition to or instead of SPOTlight, could also be applied to estimate cell-type composition per spot — Different deconvolution algorithms make different assumptions about cell-type reference profiles and spot mixing; using two methods and comparing concordance can strengthen confidence in colocalization findings given the acknowledged resolution limitation of 55 µm spots
-
Cluster frequencies in scRNA-seq are compared between conditions by proportion, without a stated statistical test for compositional differences↳ Could also: Compositional analysis methods such as scCODA (Bayesian) or a negative-binomial generalized linear model on cell counts per sample (with mouse as the statistical unit) could also be used — Cell-type proportions are compositional data (they sum to 1), so standard tests that assume independence can be misleading; dedicated compositional or count-based models account for this constraint and propagate uncertainty from small biological n
Result convergence & founder nodes
Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.
-
CD95+GL7+ germinal center B cells are significantly reduced in frequency and absolute number in CD4-depleted mice during chronic LCMV infectionflow-cytometry mouse spleen down 2022×1papers★ This paper is the founder (earliest)
-
IL21-expressing CD4 T cells fail to re-emerge at 8 dpi in CD4-depleted mice during chronic LCMV infection, absent compared to ~5% in controlsflow-cytometry mouse spleen down 2022×1papers★ This paper is the founder (earliest)
-
CXCR5+BCL6+ follicular helper T cells are significantly reduced in frequency and absolute number in CD4-depleted mice during chronic LCMV infectionflow-cytometry mouse spleen down 2022×1papers★ This paper is the founder (earliest)
-
FOXP3+ regulatory T cells show higher relative frequency but lower absolute number in CD4-depleted mice during chronic LCMV infectionflow-cytometry mouse spleen mixed 2022×1papers★ This paper is the founder (earliest)
-
Il21 expression increases over time in control spleens but remains unchanged in CD4-depleted mice during chronic LCMV infection by spatial transcriptomicsother mouse spleen down 2022×1papers★ This paper is the founder (earliest)
-
68% of antigen-experienced CD44+ CD4 T cells in a combined scRNA-seq atlas are derived from control mice at 21 dpi of chronic LCMV infectionscRNA-seq mouse spleen up 2022×1papers★ This paper is the founder (earliest)
-
Th17 CD4 T cell cluster is expanded (~7%) in CD4-depleted mice versus controls (~1%) at 21 dpi of chronic LCMV infection by scRNA-seqscRNA-seq mouse spleen up 2022×1papers★ This paper is the founder (earliest)
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-36450262
Paper: Topchyan, Zander, Kasmani, Nguyen, Brown, Lin, Burns, Cui (2022). "Spatial transcriptomics demonstrates the role of CD4 T cells in effector CD8 T cell differentiation during chronic viral infection." Cell Reports 41:111736. DOI 10.1016/j.celrep.2022.111736 · PMID 36450262 · PMCID PMC9792173.
Metadata corrections (vs the room BRIEF)
- BRIEF lists data = GSE129139 and code = github.com/MarcElosua/SPOTlight.
Neither is the paper's own deposit:
- The paper's own data: GSE200721 (scRNA-seq CD44+ CD4 T cells, control vs CD4-depleted, 4 samples) and GSE200720 (Visium spatial, 8 samples, day7/day21 × control/depleted). GSE129139 is only one of five reference scRNA-seq datasets borrowed from prior work to build the SPOTlight reference (GP33+ CD8 T cells, day 30 LCMV Cl13).
- The Data/Code-availability statement says verbatim: "This paper does not report original code." SPOTlight (Elosua-Bayes et al., v0.1.7) is the third-party tool the authors applied. Per P16 this is a fully valid reproduction target.
In scope (pipeline-derived)
| # | Reported result | Pipeline | Feasibility |
|---|---|---|---|
| C1 | Total cells per condition: 14,286 control / 15,629 CD4-depleted (Fig 2 text) | Cell Ranger filtered matrices + Seurat QC (nFeature 200–2500, %mt<10) | HIGH — resolution-independent → primary 1:1 anchor |
| C2 | 13 clusters of CD4 T cells (Fig 1C) | Seurat SCTransform + 30 PCs + Louvain | MEDIUM — clustering resolution not stated → soft target |
| C3 | 7 distinct clusters of 55 µm spatial spots (Fig 4C–D) | Seurat SCTransform + 30 PCs on Visium | MEDIUM — resolution not stated; needs Visium parse |
| C4 | SPOTlight cell–cell colocalization Pearson correlations (Fig 4–5) | SPOTlight 0.1.7 deconvolution of Visium spots vs integrated 5-dataset reference | LOW — the hard 20%: requires assembling a 5-source reference; attempt only if budget allows |
Out of scope (not pipeline-derived)
- Flow cytometry validation (CXCR5+BCL6+ Tfh, GC B cells CD95+GL7+).
- Wet-lab: CD4 depletion, FACS sorting, library prep, sequencing.
- Marker-gene biological interpretation / cluster naming (manual annotation).
- 68% antigen-experienced split and 7% vs 1% Th17 — depend on manual cluster labelling on top of C2 → reported but not primary.
Plan (80/20)
- C1 first (smallest, cleanest): GSE200721 → QC cell counts. «our HPC» SLURM + Seurat conda.
- C2 same job: default-resolution cluster count (report as approx).
- C3 if time: GSE200720 Visium clustering.
- C4 documented as the hard 20%, attempted only if C1–C3 land with budget to spare.
All heavy compute on «our HPC» («infra» workdir). «host» holds results only.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
This is an honest partial reproduction: per-condition total cell counts reproduce in the right direction and 'comparable' relationship (16,014 control < 17,715 depleted) but run ~12% above the reported 14,286/15,629. The offset sits on the authors'/data-availability side — GSE200721 ships Gene-Expression-only matrices, so the HTO singlet demultiplexing that prunes the ~12% cannot be reproduced and its thresholds are unstated (the availability statement even has an unfilled 'GSE' placeholder). Severity is moderate (magnitude/direction preserved, no fabrication signal), but the paper's central spatial/SPOTlight claims (C2-C4) were deferred per 80/20 and remain untested, so confidence in the core conclusion is limited rather than confirmed.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.