Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Circulating mucosal-associated invariant T cells identify patients responding to anti-PD-1 therapy.

Nat Commun · 2021
L1 70/100 PQI 93
Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +6
✓ What held up
  • Any deviation was negligible
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🟡Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🔴A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
70/100
Reproducibility score
at the mean
vs. all fields · 1173 studies
🎯 Scores higher than 37% of all assessed papers rank 732 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough to do an honest PARTIAL 1:1 on the auditable deposited values, NOT a full pipeline rerun. KEY CORRECTION: the BRIEF's accession GSE148190 is wrong for this paper (it is a third-party 10x tumor-infiltrating-lymphocyte set, PMID 32539073, used only for secondary validation); the paper's own ddSEQ scRNA-seq data is GSE166181, which I used. From the GSE166181 deposit (metadata + normalized/raw UMI matrices, fetched + analysed on «our HPC» «infra») two headline numbers reproduce EXACTLY: 51,701 purified CD8+ T cells and 20 patients. The central biological claim 'MAIT cells more abundant in responders' reproduces in direction AND magnitude (2.4-2.7x higher in R) via a transparent SLC4A10/KLRB1 marker proxy, though that proxy is weaker than the paper's Seurat-cluster MAIT definition. NOT attempted (and why): pre-purification counts (56,142 QC cells; 4,210 NK + 231 monocytes removed) are not in the public deposit (raw per-sample ddSEQ inputs were never uploaded), and the 8-cluster figure cannot be faithfully regenerated because the authors' repo is exploratory, not runnable as-is (hardcoded Windows paths, several syntax errors/placeholders) and has an unseeded-RNG bug ('set.seed <- 123' never calls set.seed), making UMAP/Louvain non-deterministic. Fabrication check: the checkable counts are internally consistent (56142-4210-231 = 51701 = deposited cell count, 20 patients) => no fabrication signal; unverifiable items flagged unverifiable, not fabricated.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 70
    assessed: 2026-06-15 ⛓ f3ec46fc6ba4
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-15
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

The study tests whether circulating immune cell populations, particularly subsets of CD8+ T cells, can serve as peripheral blood biomarkers to identify patients with metastatic melanoma who will respond to anti-PD-1 immune checkpoint therapy.

Core claims
  • Mucosal-associated invariant T (MAIT) cells are more abundant in the circulation of metastatic melanoma patients who respond to anti-PD-1 therapy than in non-responders, before and during therapy. finding
  • Patients with >1.7% MAIT cells among peripheral CD8+ T cells show a better response to anti-PD-1 treatment, supporting MAIT frequency as a predictive biomarker. finding
  • Responders harbor a higher proportion of activated, proliferating effector memory CD8+ T cells (cluster C16; Ki67+CD71+GNLY+) than non-responders. finding
  • MAIT cells from responders express higher levels of CXCR4 (homing receptor) and produce more granzyme B. finding
  • In silico analysis of public single-cell datasets supports the presence of CXCR4-expressing MAIT cells in the melanoma tumor microenvironment and their increase in regressing metastatic lesions after ICI. mechanism
  • Combining single-cell RNA-seq with high-dimensional/polychromatic flow cytometry resolves circulating CD8+ T-cell states (maturation, activation, exhaustion) longitudinally during PD-1 blockade. method
  • An activated MAIT subcluster with homing properties (expressing CXCR4, CD69, TNFAIP3, FOS, JUN) is enriched in responders across time points. finding
Experimental setups
Assay System Perturbation Readout Platform
High-dimensional (polychromatic) flow cytometry Peripheral blood CD8+ T cells from metastatic melanoma patients anti-PD-1 therapy (longitudinal T0/T1/T2) Phenograph clustering of T-cell subsets; iMFI and frequency of surface/intracellular markers
Single-cell RNA sequencing (scRNA-seq) Isolated CD3+CD8+ T cells from 20 metastatic melanoma patients anti-PD-1 therapy (T0, T1, T2) Cluster proportions and differential gene expression (e.g., MAIT, activated EM); UMAP/pseudotime
Flow cytometry MAIT identification Peripheral blood CD3+CD8+ T cells from melanoma patients anti-PD-1 therapy Proportion of MAIT cells (TCRα7.2+ CD161+) and CXCR4 expression
Intracellular cytokine staining / polyfunctionality flow cytometry PBMC-derived MAIT cells from melanoma patients in vitro stimulation with IL-12, IL-18, CD3/CD28 Combinatorial production of granzyme B, IFN-γ, and TNF SPICE software analysis
In silico scRNA-seq / scTCR-seq re-analysis (public dataset) PBMC, lymph node metastasis, and tumor from melanoma patients K383/K409/K411 none (untreated) MAIT proportion and CXCR4/KLRB1/CD69 expression across tissues GEO GSE148190
In silico scRNA-seq re-analysis (public ICI dataset) CD8 T cells from melanoma patients treated with ICI immune checkpoint inhibitor therapy MAIT cell abundance in regressing vs non-regressing metastatic lesions
cTP-net surface-protein imputation scRNA-seq of CD8+ T cells from melanoma patients none (computational) Imputed surface protein abundances to confirm T-cell phenotype cTP-net deep neural network
Key results
  • Activated/proliferating effector memory CD8+ T-cell cluster C16 higher in responders before therapy and remained higher after treatment
  • MAIT cell proportion higher in responders before therapy and after first cycle (scRNA-seq)
  • Activated MAIT cell subcluster significantly higher in responders at T0, T1, and T2
  • MAIT cells expanded in circulation of responders vs non-responders before therapy by flow cytometry
  • Proportion of MAIT cells expressing CXCR4 increased after two cycles of therapy in responders but not non-responders
  • Before therapy, proportion of MAIT cells producing only granzyme B higher in responders than non-responders
  • About 3% of cells in lymph node metastasis and tumor identified as CXCR4-expressing MAIT cells ~3%
  • Patients with MAIT >1.7% of CD8+ T cells had better/increased probability of response to therapy
Key statistics
  • pvalue p = 0.0363 (Log-rank Mantel-Cox) (MAIT >1.7% vs <1.7% stratification predicts response; N<1.7%=4, N>1.7%=8)
  • count 28 patients (17 responders, 11 non-responders) (Metastatic melanoma cohort starting anti-PD-1, followed 6 months)
  • pvalue p < 0.001 (T0); p < 0.01 after treatment (C16 activated proliferating effector cells higher in responders (NR=9, R=8))
  • count 51,701 purified CD8+ T cells (scRNA-seq after QC (56,142 cells; 4210 NK and 231 monocytes removed))
  • pvalue p = 0.04 (T0), p = 0.005 (T1), p = 0.007 (T2) (Activated MAIT proportion R vs NR; NR=8, R=11)
  • pvalue p = 0.023; p = 0.0012 (MAIT proportion differences R vs NR by scRNA-seq across time points)
  • pvalue p = 0.016 (MAIT cell proportion R vs NR at T0 by flow cytometry (NR=4, R=8))
  • pvalue p = 0.028 (Mann–Whitney); p = 0.041 (Wilcoxon) (MAIT granzyme B production R vs NR (NR=4, R=6))

Statistical methods review

Model: opus

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This observational longitudinal cohort study compared circulating CD8+ T-cell subsets between anti-PD-1 responders (R) and non-responders (NR) among 28 metastatic melanoma patients sampled before therapy (T0) and after one (T1) and two (T2) cycles, using high-dimensional flow cytometry, scRNA-seq, and functional assays. Between-group comparisons of cell-cluster proportions and marker expression were made predominantly with the two-sided Mann–Whitney nonparametric test (with Bonferroni multiple-comparisons adjustment noted in figure legends), polyfunctionality combinations were compared by permutation testing (SPICE) and Wilcoxon rank test, and the prognostic MAIT-cell cutoff (1.7%) was evaluated by log-rank (Mantel-Cox) analysis. Results were generally reported as individual values with mean ± SEM and exact p-values for significant comparisons.

Replicationbiological Sample size28 patients total (17 R, 11 NR); subset n's stated per figure (e.g. scRNA-seq from 20 patients), but no formal power/sample-size calculation described Groupsanti-PD-1 responders vs non-responders, cross-sectional and longitudinal (T0/T1/T2) Pairingmixed Randomization/blindingnot stated DispersionSEM Exact p-valuesyes Effect sizesno Confidence intervalsno Multiplicity correctionBonferroni's multiple comparisons test (BH/FDR for DE not specified); permutation testing used for polyfunctionality
Statistical tests used
Test Applied to n Assumptions
two-sided Mann–Whitney U (nonparametric) R vs NR cluster frequencies and marker iMFI/expression across flow cytometry and scRNA-seq figures (Fig. 1A, 1B, 2C, 2D, 3C, 4A, 4D) e.g. NR=9, R=8 (Fig.1B); NR=8, R=11 (Fig.2D, 3C); NR=4, R=8 (Fig.4A); NR=4, R=6 (Fig.4D) not stated
Bonferroni multiple-comparisons test (applied alongside Mann–Whitney) R vs NR comparisons across time points/clusters in Figs. 1-4 not stated
permutation test (10,000 permutations, SPICE) combinatorial GRZM-B/IFN-γ/TNF polyfunctionality of MAIT cells, R vs NR at T0 (Fig. 4C) NR=4, R=6 na
Wilcoxon rank test frequency of MAIT cells producing cytokine combinations after stimulation at T0 (Fig. 4C, right) NR=4, R=6 not stated
log-rank (Mantel-Cox) test response probability comparing patients with MAIT >1.7% vs <1.7% of CD3+CD8+ cells (Fig. 5) N<1.7%=4, N>1.7%=8 na
differential gene expression analysis (scRNA-seq; method/test not specified in text) genes between R and NR within EM and MAIT clusters (Fig. 2C, 2D); p-values in source tables not stated
Approaches that could also have been used
  • Spread was summarized as mean ± SEM throughout the figures.
    Could also: Reporting standard deviation or a 95% confidence interval alongside individual data points. — SD describes the variability of the observations themselves and CIs convey precision of the estimate; for small group sizes these are often preferred because SEM can appear to understate dispersion.
  • Many R-vs-NR comparisons were made across multiple clusters and three time points using repeated two-sided Mann–Whitney tests with Bonferroni adjustment.
    Could also: A single mixed-effects or repeated-measures model (or Friedman/aligned-rank approaches) for the longitudinal structure, with a unified correction. — A single model can incorporate the within-patient time structure and the full family of comparisons simultaneously, which can improve efficiency and make the multiplicity scope explicit.
  • Multiplicity was addressed with Bonferroni's multiple comparisons test.
    Could also: A Benjamini–Hochberg false-discovery-rate procedure for the larger families (e.g., per-cluster or differential-expression tests). — FDR control is often favored in high-dimensional/omics settings because it offers more power than family-wise Bonferroni when many features are tested.
  • The prognostic cutoff used the cohort median MAIT value (1.7%) and groups were compared by log-rank test.
    Could also: A Cox proportional-hazards model treating MAIT% as a continuous variable, optionally with a data-driven optimal-cutoff or ROC analysis. — Modeling the marker continuously avoids dichotomization at the median and can provide a hazard ratio with a confidence interval as an effect-size estimate.
  • Significance was reported via p-values, with effect sizes and confidence intervals not generally provided.
    Could also: Reporting effect-size measures (e.g., rank-biserial correlation for Mann–Whitney, or median differences) with confidence intervals. — Effect sizes with intervals convey the magnitude and precision of differences, complementing significance testing especially in small cohorts.
  • Differential gene expression between R and NR within clusters was assessed (test/method not specified in main text).
    Could also: A specified scRNA-seq DE framework such as a Wilcoxon test with BH-FDR, MAST, or pseudobulk DESeq2/edgeR aggregating cells per patient. — Pseudobulk and mixed approaches account for within-patient correlation of cells and can reduce false positives that arise when individual cells are treated as independent replicates.
Software: SPICE (permutation testing for polyfunctionality) · Phenograph (high-dimensional flow cytometry clustering) · cTP-net (deep neural network for surface-protein imputation) · UMAP / scRNA-seq clustering and pseudotime pipeline (specific package not stated in provided text)

Result convergence & founder nodes

Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
86
Impact: high
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

GSE120575 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE148190 GEO in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GSM4455931 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSM4455932 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSM4455933 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSM4455935 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSM4455937 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSM4455938 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet

Downstream reach in the literature

107 downstream papers · 8 datasets

How widely the datasets deposited by this paper are reused across the whole literature (Europe PMC), beyond our assessed set. This is a factual dependency map — reusing a public dataset is normal, good science. It is not a judgement on the downstream papers; the only verdict here is this paper's own, with its cited rationale.

GSM4455931 GEO reused by 1 papers in the literature
GSM4455932 GEO reused by 1 papers in the literature
GSM4455933 GEO reused by 1 papers in the literature
GSM4455935 GEO reused by 1 papers in the literature
GSM4455937 GEO reused by 1 papers in the literature
GSM4455938 GEO reused by 1 papers in the literature

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Figures / tables: tableFig. 2A
cd8_purified
Reported
51,701 purified CD8+ T cells
Reproduced
51,701 cells (GSE166181 normalized & raw matrices = 51701 columns; metadata = 51701 rows; 51701/51701 cell-IDs matched)
exact
n_patients_scrna
Reported
20 patients (scRNA-seq cohort)
Reproduced
20 unique patients ('sample' column, Mela3..Mela31)
exact
mait_higher_R
Reported
MAIT cells more abundant in responders (Fig. 2; direction)
Reproduced
SLC4A10+ marker proxy: R 2.81% (760/27037) vs NR 1.18% (292/24664) = 2.4x; SLC4A10+&KLRB1+: R 1.81% vs NR 0.67% = 2.7x
partial
cells_qc
Reported
56,142 cells passing QC
Reproduced
not in deposit (pre-purification; raw per-sample ddSEQ files not on GEO)
partial
n_clusters
Reported
8 unsupervised CD8 clusters (Fig. 2A)
Reproduced
not attempted (no cluster labels deposited; repo reclustering unseeded/non-deterministic, needs non-deposited raw inputs)
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 70/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟡2. Endpoint comparability
🟡3. Location of the main deviation
🔴4. Cause of the deviation
🟡5. Derivability / plausibility
🟢6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +6

On the auditable deposited values this reproduces 1:1: 51,701 purified CD8+ T cells and 20 patients match exactly, and internal consistency (56142−4210−231=51701) shows no fabrication signal. The central claim 'MAIT cells more abundant in responders' holds in direction and magnitude (2.4–2.7x) but only through a self-chosen SLC4A10/KLRB1 marker proxy, since the paper's Seurat-cluster definition and labels were not deposited. The unreproducible items — pre-purification QC counts and the 8 clusters — are limited by authors'-side gaps (raw ddSEQ inputs never uploaded, a non-runnable repo with a set.seed bug), not by our method. Overall a solid partial reproduction with explainable, non-critical deviations.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

127.4 k
tokens (I/O) · 8 M incl. cache
13 min
runtime · 0.02 CPU-h
2.5 GB
peak RAM
2
HPC jobs
hummel
machine