Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

PD-1 and TIGIT coexpression identifies a circulating CD8 T cell subset predictive of response to anti-PD-1 therapy.

J Immunother Cancer · 2020
L1 75/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +5
✓ What held up
  • Reported values were directly comparable
  • The central claim held under reproduction
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
75/100
Reproducibility score
at the mean
vs. all fields · 1173 studies
🎯 Scores higher than 45% of all assessed papers rank 612 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough to reproduce the CORE pipeline 1:1 in method, with a clear, honest gap on exact sample filtering. The paper's own workflowr limma-voom DE pipeline (valentinvoillet/Simon_et_al_2020 @c176826) was re-run on the deposited GEO count matrix GSE141119 (140 samples x 975 genes). RESULTS: DE_1 temporal DEGs EXACT (1, in DNEG T0-vs-M1); DE_2 (Fig 3B-C DEG counts) at M2 essentially exact (all 6 contrasts within +/-3) and at T0 within +/-4 for 5/6; M1 shows a systematic undercount (16-23 fewer on 4/6 contrasts); DE_3 (NR vs R) 2 vs 1 (both ~null); DPOS-specific signature (Fig 4B/Table S5) reproduced as 9 of ~12 genes with biologically on-theme content (CXCL13, CXCR5, ENTPD1/CD39, MKI67). The non-exact matches all trace to ONE root cause: the per-sample QC metric '% reads dropped <55 bp' used (OR library.size<70000, plus an author-acknowledged 'subjective threshold' manual step) to filter 140->121 samples is NOT in the GEO deposit (only in un-deposited Qiagen Summary.xlsx) -> our coded rule keeps 130, the 9 extra lower-quality samples reduce DE power (mostly at M1). NO fabrication: M2 near-exact + DE_1 exact confirm reported counts are genuine pipeline outputs. NOT ATTEMPTED: TCGA-SKCM OE survival (Fig 4C, needs external TCGA data + the ImmuneResistance OE code); GSEA gene-set counts (gene-set RDS not shipped in repo); TCR-seq (GSE141120, Fig 6); all flow/ELISPOT/ROC/KM (Fig 1,2,5 = wet-lab, out of scope). Discrepancies flagged in profiling: paper says 13 patients / deposit has 12; GEO lists 96 GSMs but matrix has 140 columns; 8-sample 'TO'->'T0' header typo.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 75
    assessed: 2026-06-18 ⛓ de71ebd56083
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-18
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-18
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

Can peripheral CD8 T cell subsets defined by PD-1 and TIGIT coexpression serve as a circulating, non-invasive biomarker predictive of clinical response to anti-PD-1 therapy in melanoma and Merkel cell carcinoma patients?

Core claims
  • The frequency of circulating PD-1+TIGIT+ (DPOS) CD8+ T cells after 1 month of anti-PD-1 therapy is associated with clinical response and overall survival across three independent cohorts. finding
  • The DPOS CD8 T-cell population is enriched in highly activated, proliferating, tumor-specific and emerging T-cell clonotypes. finding
  • The DPOS population overexpresses CXCR5, a key marker of the CD8 cytotoxic follicular T cell (Tfc) population. finding
  • Transcriptomic profiling defines a specific gene signature for the DPOS population, sharing features with the tumor PD-1high CXCL13+ subset predictive of PD-1 blockade efficacy. finding
  • TCR repertoire analysis identifies a cluster of emerging T-cell clonotypes in the DPOS population as a key feature of PD-1 therapeutic efficacy. mechanism
  • Monitoring the PD-1+TIGIT+ circulating CD8 population can serve as an early cellular-based marker for clinical decision support in anti-PD-1 therapy. resource
Experimental setups
Assay System Perturbation Readout Platform
Multiparameter flow cytometry and FACS cell sorting PBMCs from melanoma (cohorts 1 & 2) and MCC patients anti-PD-1 therapy (Nivolumab/Pembrolizumab) Frequency/phenotype of four CD8 subsets (DNEG, PD-1, DPOS, TIGIT) based on PD-1/TIGIT expression BD ARIA II cell sorter, BD FACS Diva Software
Targeted transcriptomic RNA sequencing Sorted CD8 T-cell subsets from patient PBMCs none (ex vivo sorted) Gene expression / differentially expressed genes and gene signatures QIAseq Targeted RNA Custom Panel (QIAGEN), Illumina NextSeq High Output V2
TCR sequencing (repertoire analysis) Sorted CD8 T-cell subsets from patient PBMCs none Clonotype abundance, clonality (Shannon entropy), emerging/expanding/contracting clusters Human TCR Panel QIAseq Immune Repertoire RNA Library Kit (QIAGEN), Illumina NextSeq Mid Output V2
IFN-γ ELISPOT antitumor reactivity assay Sorted and in vitro amplified T cells from HLA-A*0201 patients stimulation with tumor peptide, viral peptide or anti-CD3 IFNγ spot forming units (SFU)/10^5 cells Mabtech precoated ELISPOT plates, Bio-Sys Bioreader
Luminex multiplex cytokine assay Sorted and amplified T cells from cohort 1 patients (baseline and M1, n=5) stimulation with coated anti-CD3 (OKT3) Production of 25 cytokines Luminex EPX250-12166-901, ThermoFisher Scientific
Bulk RNA-seq signature analysis (TCGA) 473 skin cutaneous melanoma tumors (TCGA) none Overall expression of DPOS-upregulated DEG signature and predicted overall survival
Key results
  • Frequency of DPOS subset significantly higher in responding patients after 1 month of therapy, while other subsets not significantly different by clinical outcome
  • DPOS subset frequency predictive of anti-PD-1 efficacy at month 1 across melanoma and MCC cohorts
  • DNEG population was the most represented subset in all cohorts 46%±14% and 61%±19% (melanoma cohorts); 35%±13% (MCC)
  • DPOS population enriched in CXCR5-overexpressing T lymphocytes (Tfc marker)
  • DPOS and TIGIT fractions equally represented (~20%) in melanoma cohorts ~20%
Key statistics
  • mean 46%±14% (DNEG fraction frequency, melanoma cohort 1)
  • mean 61%±19% (DNEG fraction frequency, melanoma cohort 2)
  • mean 35%±13% (DNEG fraction frequency, MCC cohort)
  • mean 26%±14% (PD-1 fraction frequency, MCC cohort)
  • mean 24%±13% (DPOS fraction frequency, MCC cohort)
  • mean 14%±10% (TIGIT fraction frequency, MCC cohort)
  • count 473 (TCGA skin cutaneous melanoma tumors used for signature/survival analysis)
  • pvalue p<0.05 (Threshold for statistical significance; DPOS frequency higher in responders (Holm-Sidak))

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

Three independent patient cohorts (melanoma cohort 1, n=13; melanoma cohort 2, n=13; MCC cohort, n=15) were studied longitudinally with multiparameter flow cytometry to characterize peripheral CD8+ T-cell subsets defined by PD-1/TIGIT co-expression. Subset frequencies between clinical responders and non-responders were compared using Mann-Whitney tests or one-way ANOVA with Holm-Sidak correction; differential gene expression was assessed via the limma/edgeR framework with empirical Bayes moderated t-statistics and FDR control at 5%; and TCR clonotype differential abundance was tested with Fisher exact tests (FDR 5%). Results were displayed as box-and-whisker plots (median, IQR, min/max) with only significant p values shown.

Replicationbiological Sample sizeThree cohorts defined by enrollment: melanoma cohort 1 (n=13), melanoma cohort 2 (n=13), MCC cohort (n=15); no formal power calculation reported GroupsResponders (CR+PR) vs non-responders (SD+PD); four CD8 T-cell subsets (DNEG, PD-1-single-positive, TIGIT-single-positive, DPOS); baseline vs on-treatment time points (T0, M1, M2) Pairingmixed Randomization/blindingnot stated DispersionIQR Exact p-valuesno Effect sizesno Confidence intervalsno Multiplicity correctionHolm-Sidak (flow cytometry multiple comparisons); Benjamini-Hochberg FDR at 5% (DEGs, GSEA, TCR differential abundance)
Statistical tests used
Test Applied to n Assumptions
Mann-Whitney U test Comparisons of CD8 T-cell subset frequencies between responders and non-responders across cohorts, as indicated per figure n=13 (cohort 1), n=13 (cohort 2), n=15 (MCC cohort) not stated
One-way ANOVA with Holm-Sidak multiple comparison correction Comparisons of T-cell subset frequencies across time points and clinical response groups (e.g., figure 1B) n=13 (cohort 1) not stated
Empirical Bayes moderated t-statistics (two-tailed) via limma R package Differential expression analysis of targeted RNA-seq data comparing T-cell subsets, time points, and clinical outcomes; intraclass correlations used to account for repeated-measures structure not stated
Fisher exact test Differential abundance of individual TCR alpha/beta clonotype sequences between T0 and M1 within each T-cell fraction not stated
Gene set enrichment analysis (GSEA) via Camera (limma R package) Pathway enrichment using KEGG, Hallmark, immunological signatures (c7), and curated pathway gene sets from MSigDB not stated
Kaplan-Meier survival curves Overall survival prediction in TCGA skin cutaneous melanoma patients (n=473) stratified by high (top 20%), low (bottom 20%), or intermediate DPOS gene signature expression n=473 not stated
Approaches that could also have been used
  • Dispersion in box-and-whisker plots is shown as IQR (median, IQR, min/max) for flow cytometry frequency data
    Could also: 95% confidence intervals around the median or mean could also be displayed alongside individual data points — CIs directly communicate estimation uncertainty and facilitate visual inference about group differences, which can be particularly informative in small clinical cohorts (n=13–15)
  • Only statistically significant p values are displayed for flow cytometry comparisons across groups and time points
    Could also: All p values could be reported, together with effect-size estimates (e.g., rank-biserial correlation for Mann-Whitney, eta-squared for ANOVA) — Reporting all p values and effect sizes allows readers to evaluate the magnitude and direction of non-significant results and assess the complete pattern of group differences
  • Responders and non-responders were compared at individual time points using separate Mann-Whitney or ANOVA tests across three independent cohorts
    Could also: A linear mixed-effects model (patient as random effect) or generalized estimating equations could jointly model repeated measures across time points — Such models account explicitly for within-patient correlation over longitudinal measurements, can assess response-by-time interactions, and increase statistical efficiency relative to separate per-timepoint tests
  • TCGA survival analysis stratified patients into top 20%, bottom 20%, and intermediate groups based on gene signature expression
    Could also: A continuous Cox proportional hazards regression using the gene signature score as a predictor could also be applied — Continuous Cox regression avoids the arbitrary choice of percentile cut-points, retains all available prognostic information, and provides a hazard ratio with confidence interval as a quantified effect estimate
  • TCR clonotype differential abundance between time points was tested with Fisher exact tests applied to each sequence independently
    Could also: Negative binomial or quasi-likelihood models (e.g., edgeR or DESeq2) are also used for count-based differential abundance testing in immune repertoire data — Count-based models designed for overdispersed sequencing data can better accommodate the distributional properties of clonotype counts and incorporate library-size normalization directly into the statistical model
  • Intraclass correlations were estimated within the limma framework to account for repeated measures from the same patients in RNA-seq analyses
    Could also: A fully specified linear mixed-effects model (e.g., via lme4 or nlme in R) with patient as a random effect could also be used — Mixed-effects models offer a more flexible framework for unbalanced longitudinal designs, allow patient-level random slopes, and can incorporate additional covariates such as prior treatment history
Software: GraphPad Prism 8 · R 3.5.2 · edgeR (R package) · limma (R package) · QIAGEN GeneGlobe Data Analysis Center

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-33188038

Paper: Simon S, Voillet V, et al. PD-1 and TIGIT coexpression identifies a circulating CD8 T cell subset predictive of response to anti-PD-1 therapy. J Immunother Cancer 2020. PMID 33188038 · PMCID PMC7668369 · DOI 10.1136/jitc-2020-001631.

Two code artifacts

  1. Authors' own analysis repo (PRIMARY reproduction target): https://github.com/valentinvoillet/Simon_et_al_2020 (rendered: https://valentinvoillet.github.io/Simon_et_al_2020/). A workflowr project; the RNA-seq + TCR-seq analyses are R Markdown notebooks under analysis/, rendered to docs/*.html. This is the code that produces the paper's bioinformatic figures.
  2. Third-party tool (SECONDARY, P16): https://github.com/livnatje/ImmuneResistance (Jerby-Arnon et al.). Used by the paper for ONE step only — the TCGA SKCM validation (Fig 4C): "Using the R code and data provided by Jerby-Arnon et al. we computed the overall expression (OE) of the upregulated DEGs between PD-1+TIGIT+ and the other subsets." Applying this existing tool to the paper's own derived signature is equally valid (P16).

Data

  • GSE141121 = SuperSeries. SubSeries:
    • GSE141119 = bulk targeted RNA-seq (QIAseq Targeted RNA Custom Panel, ~975 immune genes), 4 FACS-sorted circulating CD8 T-cell subsets (DNEG=PD-1−TIGIT−, PD-1=PD-1+TIGIT−, TIGIT=PD-1−TIGIT+, DPOS=PD-1+TIGIT+) × 3 timepoints (T0/M1/M2) × melanoma patients. Processed count matrix shipped: GSE141119_RNAseq_MelanomaPatient_Simon.txt.gz (140 sample columns, ~974 genes, "alignment+quantification by Qiagen").
    • GSE141120 = TCR-seq (QIAseq Immune Repertoire), 180 samples.

IN SCOPE (pipeline-derived from GSE141119 + the shipped R code)

# Result Pipeline Where
R1 QC: 140 samples → 121 after filtering; remove TCF7 & ITGAE genes custom R filter (library.size<70000 OR %reads<55bp>0.55) QC.Rmd / Fig design
R2 DPOS-vs-subset DEG counts per timepoint (T0/M1/M2; 6 contrasts each) limma-voom + duplicateCorrelation(block=patient) DE_2.Rmd → Fig 3B–C
R3 Temporal DEG counts within each fraction (T0vsM1 etc.) = 1 DEG total limma-voom DE_1.Rmd
R4 Responder-vs-Non-responder DEG counts (~1 DEG at T0) limma-voom DE_3.Rmd → Fig 4D context
R5 GSEA gene-set counts (KEGG/Hallmark/c7) per contrast limma fry/camera or fgsea DE_2/DE_3.Rmd → Fig 4A/4E
R6 12-gene DPOS-specific signature (Fig 4B / Table S5) DE intersection / OE manuscript
R7 TCGA SKCM survival of DPOS-signature OE (n=473, log-rank p=2.2e-4) ImmuneResistance OE + survival Fig 4C

Primary quick target = R2 (clear, deterministic count table). R1/R3/R4/R5 follow from the same pipeline. R6/R7 are the harder "keep going" targets.

OUT OF SCOPE (wet-lab / manual / not pipeline-derived from deposited data)

  • Flow cytometry frequencies & phenotyping (Fig 1A–E, Fig 2): DPOS ~20%, HLA-DR+CD38+, CXCR5+, MFIs — instrument data, not in GEO.
  • ROC/AUC of DPOS frequency predicting response (Fig 1F 0.96, 1G 0.78) — derived from flow frequencies, not from RNA-seq deposit.
  • Kaplan–Meier of DPOS frequency >17% (Fig 1H) — flow-derived.
  • ELISPOT tumor-specific T-cell responses (Fig 5) — wet-lab.
  • TCR repertoire clonotype-cluster outcome association (Fig 6) — TCR-seq (GSE141120) analysis; attempt only if RNA-seq targets complete (lower priority).

Notes / discrepancies to profile

  • Paper text: 13 melanoma patients (cohort 1); repo/GEO RNA-seq: 12 patients.
  • GEO SuperSeries page lists GSE141119 as 96 GSM samples, but the shipped count matrix has 140 sample columns — to reconcile in dataset_profile.json.
  • The R code's output/RNA_count.rds is NOT shipped in the repo (output/ has only a README) → must be rebuilt from the GEO count matrix before running DE.
Figures / tables: Fig 4BTable
DE1_total
Reported
1 DEG (PD-1-TIGIT-, T0 vs M1)
Reproduced
1 DEG (DNEG, T0 vs M1)
exact
QC_genes
Reported
975 genes
Reproduced
975 genes
exact
DE2_T0_DPOS_vs_TIGIT
Reported
33
Reproduced
33
exact
DE2_M2_DNEG_vs_DPOS
Reported
166
Reproduced
164
within tolerance
DE2_M2_DPOS_vs_PD1
Reported
73
Reproduced
72
within tolerance
DE2_M2_PD1_vs_TIGIT
Reported
81
Reproduced
80
within tolerance
DE2_M2_DNEG_vs_TIGIT
Reported
148
Reproduced
150
within tolerance
DE2_M2_DPOS_vs_TIGIT
Reported
42
Reproduced
41
within tolerance
DE2_M2_DNEG_vs_PD1
Reported
95
Reproduced
92
within tolerance
DE2_T0_DNEG_vs_DPOS
Reported
171
Reproduced
174
within tolerance
DE2_T0_DPOS_vs_PD1
Reported
68
Reproduced
67
within tolerance
DE2_T0_PD1_vs_TIGIT
Reported
81
Reproduced
82
within tolerance
DE2_T0_DNEG_vs_PD1
Reported
121
Reproduced
117
within tolerance
DE2_T0_DNEG_vs_TIGIT
Reported
159
Reproduced
172
partial
DE2_M1_DPOS_vs_TIGIT
Reported
40
Reproduced
41
within tolerance
DE2_M1_PD1_vs_TIGIT
Reported
58
Reproduced
52
partial
DE2_M1_DNEG_vs_DPOS
Reported
121
Reproduced
105
partial
DE2_M1_DNEG_vs_PD1
Reported
72
Reproduced
52
partial
DE2_M1_DNEG_vs_TIGIT
Reported
125
Reproduced
102
partial
DE2_M1_DPOS_vs_PD1
Reported
68
Reproduced
47
partial
DE3_total
Reported
1
Reproduced
2
within tolerance
DPOS_signature
Reported
12 genes (Fig 4B/Table S5)
Reproduced
9 genes: CD80,CXCL13,CXCR5,ENTPD1,MKI67,RGS1,ST8SIA1,VCAM1,ZBTB32
partial
QC_samples
Reported
121 of 140
Reproduced
130 of 140
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 75/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟢2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟢7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +5

A faithful method reproduction: the authors' own limma-voom workflowr pipeline rerun on the deposited GSE141119 matrix yields DE_1 exact and all six M2 contrasts within ±3, confirming the reported DEG counts are genuine pipeline outputs (no fabrication). The only material deviations — an M1 undercount of 16–23 DEGs and a 9-of-12-gene DPOS signature — trace to a single honest gap: the '% reads <55bp' QC metric (and an acknowledged subjective step) that filters 140→121 samples is not in the deposit, so our coded rule keeps 130. The discrepancy is on the data-availability/our-method side, moderate in magnitude with direction and biology preserved, so the central DPOS-subset transcriptomic claim holds within reproducible scope.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

396.5 k
tokens (I/O) · 56.1 M incl. cache
123 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.