Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

PD-1 blockade potentiates neoadjuvant chemotherapy in NSCLC via increasing CD127+ and KLRG1+ CD8 T cells.

NPJ Precis Oncol · 2023
L1 40/100 PQI 80
Why this verdict

The main result did not reproduce in this reproduction attempt. Where our recomputation produced values that differ from the published ones, those discrepancies are listed below. This is a single automated attempt — not peer review and not a finding of error or misconduct — and differences can also arise from data access, undocumented parameters or the computing environment. The verdict can be contested via “report an error”.

Reproduced on the brainbox compute brainarbeit.com
✓ What held up
  • Reported values were directly comparable
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🟡A deviation arose in the data or preprocessing
  • 🔴A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
40/100
Reproducibility score
1.9 SD below mean
vs. all fields · 1173 studies
🎯 Scores higher than 4% of all assessed papers rank 1126 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough for a PARTIAL, honest 1:1 on the primary pipeline output, not a full reproduction. Clarification: the brief's data accession GSE179994 is only a validation set (Fig.7); the paper's OWN scRNA data is GSE229353, and the named tool (CellChat) was run on it - so we reproduced against GSE229353. C1 (cell count): we recomputed per-cell QC directly from the 7 deposited 10x matrices and applied the paper's stated Seurat filters (200<=nGene<=6000, nUMI>=1000, mito<10%); 6 intact samples yield 22,509 cells and 31,203 total cellranger calls vs the reported 26,861 (extrapolated ~25k, ~7% off). An EXACT total is impossible because GSM7159185(P03)_matrix.mtx.gz is CORRUPT in the GEO deposit (gzip-invalid, ~92% truncated, byte-identical on re-download) - a real data-availability defect flagged for the reviewer - and doublet/QC-order details are under-specified. C3 (CellChat, the named tool): attempted on «our HPC»; correct 1.1.3 source identified (untagged commit af7e152) and a Seurat env built, but the install stalled on a graphics-only dependency (harfbuzz/freetype) and, more fundamentally, the exact qualitative L-R claim needs the authors' unshipped cell annotations - so we stopped per the 80/20 rule rather than fabricate. NOT attempted: mIHC/spatial K(r) imaging results (wet-lab), Cell Ranger re-alignment from FASTQ (its output - the matrices - is already deposited), and the GSE176021/GSE190265 GSEA validations (directional, multi-dataset). All heavy compute ran on «our HPC»/«infra»; «host» holds only small results + pointers.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 40
    assessed: 2026-06-15 ⛓ e15631f3c130
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-15
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

The paper investigates the mechanisms by which adding PD-1 blockade (pembrolizumab) to neoadjuvant chemotherapy augments anti-tumor immune responses in resectable NSCLC, hypothesizing that it remodels the tumor immune microenvironment by recruiting specific T and B cell subsets.

Core claims
  • Adding PD-1 blockade to neoadjuvant chemotherapy (NAPC) increases tumor infiltration of CD20+ B cells, CD4+ T cells, CD4+CD127+ T cells, CD8+ T cells, CD8+CD127+ and CD8+KLRG1+ T cells, whereas NAC alone increases only CD20+ B cells. finding
  • NAPC skews tumor-infiltrating CD8+ T cells toward CD127+ and KLRG1+ (memory/effector) phenotypes, potentially assisted by CD4+ T cells and B cells. mechanism
  • Synergistic increase in B and T cells promotes favorable pathological therapeutic response after NAPC. finding
  • NAPC improves pathological response (MPR/pCR) and survival (DFS, OS) compared to NAC in NSCLC patients. finding
  • CD8+ T cells and their CD127+/KLRG1+ subsets are in closer spatial proximity to CD4+ T/CD20+ B cells in NAPC versus NAC. finding
  • B-cell, CD4, memory, and effector CD8 signatures correlate with therapeutic responses and clinical outcomes (validated in GEO dataset). finding
  • CD8+ T cells in NAPC highly express cytotoxic genes (GZMK, IFNG) and CXCR3 and lowly express exhaustion genes (CTLA4, PDCD1, HAVCR2, TIGIT) compared with NAC. finding
  • MPR patients in NAPC have higher CD8 cytotoxic score and lower exhausted score, with pro-inflammatory cytokines enriched in MPR and anti-inflammatory cytokines in non-MPR. finding
Experimental setups
Assay System Perturbation Readout Platform
single-cell RNA sequencing (scRNA-seq) CD45+ immune cells from surgically resected fresh tumors of 7 NSCLC patients (6 NAPC, 1 NAC) neoadjuvant chemotherapy (NAC) vs neoadjuvant pembrolizumab plus chemotherapy (NAPC) single-cell transcriptomes, immune cell cluster composition, gene expression
multiplex fluorescent immunohistochemistry (mIHC) FFPE tumor tissues from 65 resectable NSCLC patients before and after treatment NAC vs NAPC, pre- vs post-treatment proportions of CD20+ B cells, CD4+/CD8+ T cell subsets (CD127+, KLRG1+, FoxP3+) infiltration; spatial proximity
GEO dataset validation (gene signature/transcriptomic analysis) public NSCLC patient cohort (GEO dataset) none correlation of B-cell, CD4, memory and effector CD8 signatures with therapeutic response and clinical outcomes
GSEA / signature scoring (AddModuleScore in Seurat) scRNA-seq CD8+ T cells from NAC and NAPC tumors NAC vs NAPC; MPR vs non-MPR lymphocyte subpopulation signature enrichment, cytotoxic and exhausted scores Seurat
Key results
  • NAPC achieved higher MPR rate than NAC (23/35 vs 6/30) 65.7% vs 20%
  • pCR achieved in NAPC group but no patients in NAC group 45.7% (16/35) vs 0%
  • NAPC group had significantly longer DFS than NAC HR=0.3216 (95% CI 0.1455–0.7105)
  • NAPC group had significantly longer OS than NAC HR=0.2368 (95% CI 0.1043–0.5376)
  • NAPC increased infiltration of CD20+ B, CD4+, CD4+CD127+, CD8+, CD8+CD127+ and CD8+KLRG1+ T cells post vs pre-treatment; NAC increased only CD20+ B cells
  • CD8+, CD8+CD127+ and CD8+KLRG1+ T cell ratios significantly higher in post-NAPC vs post-NAC
  • In NAPC, CD20+ B, CD4+, CD8+, CD8+CD127+ and CD8+KLRG1+ cells significantly elevated in MPR but not non-MPR patients after treatment
  • CD4+CD127+ T cells slightly higher in pemetrexed-treated (adenocarcinoma) vs paclitaxel-treated (squamous) patients in NAC group
Key statistics
  • pvalue P < 0.0001 (MPR rate difference between NAC and NAPC groups)
  • pvalue HR=0.3216, 95% CI 0.1455–0.7105, P = 0.0031 (DFS NAPC vs NAC)
  • pvalue HR=0.2368, 95% CI 0.1043–0.5376, P = 0.0016 (OS NAPC vs NAC)
  • count 26,861 single-cell transcriptomes (4,057 NAC, 22,804 NAPC) (scRNA-seq cells obtained after quality filtering)
  • count 29 distinct clusters (unsupervised clustering of single-cell transcriptomes)
  • count MPR 65.7% (23/35) NAPC vs 20% (6/30) NAC (major pathological response rates)
  • count pCR 45.7% (16/35) NAPC vs 0% NAC (pathological complete response rates)
  • pvalue P = 0.0116 (CD4+CD127+ T cells higher in pemetrexed vs paclitaxel patients in NAC)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This study used a multi-modal design pairing scRNA-seq (n=7 patients) with multiplex immunohistochemistry (mIHC; n=65 patients) to characterize tumor-infiltrating immune cell changes in NSCLC before and after NAC or NAPC. TIL proportion comparisons across pre/post-treatment groups used the Kruskal-Wallis test; treatment-response subgroup dynamics were evaluated with two-way ANOVA. Survival outcomes were reported with hazard ratios and 95% CIs, and scRNA-seq data were interrogated with GSEA and Seurat-based DEG analysis, with findings externally validated in a GEO dataset.

Replicationbiological Sample sizemIHC: NAC n=30 (16 paired pre/post, 14 unpaired post); NAPC n=35 (18 paired pre/post, 5 unpaired pre, 12 unpaired post). scRNA-seq: 7 patients (6 NAPC, 1 NAC). No formal sample size or power calculation reported. GroupsNAC vs NAPC; pre- vs post-treatment within each arm; MPR vs non-MPR within NAPC; CD8+KLRG1+ vs CD8+KLRG1- T cells Pairingmixed Randomization/blindingnot stated DispersionSD Exact p-valuesyes Effect sizesyes Confidence intervalsyes Multiplicity correctionMultiple comparisons adjustment applied post-Kruskal-Wallis and post-two-way ANOVA; specific method (e.g., Dunn's, Tukey HSD) not named. scRNA-seq DEG correction method not named.
Statistical tests used
Test Applied to n Assumptions
Kruskal-Wallis test with multiple comparisons Comparison of CD20+ B cells, CD4+ T cells, CD4+CD127+ T cells, CD8+ T cells, CD8+CD127+ T cells, and CD8+KLRG1+ T cells across pre/post-NAC and pre/post-NAPC groups (Fig. 3c–h) Pre-NAC n=16, Post-NAC n=30, Pre-NAPC n=23, Post-NAPC n=30 not stated
Two-way ANOVA with multiple comparisons Comparison of TIL proportions in non-MPR vs MPR patients before and after treatment in NAC and NAPC groups (Fig. 4a, b) NAC: Pre-NR n=13, Post-NR n=24, Pre-R n=3, Post-R n=6; NAPC: Pre-NR n=6, Post-NR n=10, Pre-R n=17, Post-R n=20 not stated
Cox proportional hazards regression or log-rank test (not explicitly named; implied by hazard ratio reporting) DFS and OS comparison between NAC and NAPC groups (Supplementary Fig. 1b, c) NAC n=30, NAPC n=35 not stated
Chi-square or Fisher's exact test (not explicitly named) Comparison of clinicopathological characteristics between NAC and NAPC groups (Table 1) NAC n=30, NAPC n=35 not stated
Gene Set Enrichment Analysis (GSEA) Detection of lymphocyte subpopulation signatures differentially enriched in NAC vs NAPC scRNA-seq data (Fig. 1f) 26,861 single-cell transcriptomes: 4,057 NAC, 22,804 NAPC not stated
Seurat-based DEG analysis (underlying test not named; Seurat default is Wilcoxon rank-sum); inclusion criteria: Log2FC > 0.25, p < 0.05, min.pct > 0.1 Differential expression of CD8+ T cells between NAC and NAPC, and between CD8+KLRG1+ vs CD8+KLRG1- cells (Fig. 2d; Supplementary Fig. 4) null not stated
Approaches that could also have been used
  • The Kruskal-Wallis test was applied to all four groups (pre/post × NAC/NAPC) treating every sample as independent, even though a subset of patients contributed matched pre- and post-treatment specimens.
    Could also: A linear mixed-effects model or a repeated-measures design explicitly modelling patient as a random effect would also be appropriate for the mixed paired/unpaired structure. — Accounting for within-patient pairing uses each patient as their own control, can increase power by reducing between-subject variance, and avoids treating paired observations as independent—which can affect type I error.
  • Two-way ANOVA was used to compare TIL proportions across response (MPR/non-MPR) and timepoint (pre/post) subgroups, without reporting a normality assessment; some cell counts were very small (e.g., Pre-R n=3 in NAC).
    Could also: An aligned rank transform (ART) ANOVA or a permutation-based factorial test would also be applicable to proportion data that may not meet normality assumptions. — Cell-proportion data are bounded between 0 and 1 and can be skewed, particularly in small subgroups; rank-based or permutation alternatives do not require normality and may be more appropriate when cell sizes are as small as n=3.
  • Results throughout the mIHC comparisons were summarized as mean ± SD.
    Could also: Median with interquartile range (IQR), or mean with 95% confidence intervals, would also convey central tendency and spread. — With several small subgroups (n=3 to n=6), 95% CIs communicate estimation uncertainty around the mean, while IQR is often more interpretable than SD for non-symmetric or small-sample distributions.
  • scRNA-seq DEG analysis was performed using Seurat's cell-level approach, where individual cells are treated as the statistical unit of replication.
    Could also: A pseudo-bulk approach—aggregating per-patient count profiles and then applying DESeq2 or edgeR—would also be appropriate for a two-condition (NAC vs NAPC) comparison with patient-level biological replication. — Pseudo-bulk methods use the patient (not the cell) as the unit of replication, account for overdispersion in count data, and control type I error inflation that can arise when thousands of cells from a small number of donors are treated as independent observations.
  • Survival outcomes (DFS, OS) were compared between NAC and NAPC using a model yielding unadjusted hazard ratios, without covariate adjustment.
    Could also: A multivariable Cox proportional hazards regression adjusting for key baseline covariates (e.g., T stage, N stage, histology, resection type) would also be standard practice in an observational comparison. — Because patients were not randomized, adjusted Cox regression can reduce confounding from imbalanced baseline characteristics—such as the higher pneumonectomy rate in the NAC group (23.3% vs 8.6%)—when estimating the treatment association with survival.
  • The statistical test for Table 1 clinicopathological comparisons was not named, and no formal sample size justification or power calculation was reported.
    Could also: Explicitly specifying whether chi-square or Fisher's exact test was used (depending on expected cell counts) and reporting a power calculation or effect-size-based sample size rationale would also be standard practice. — Naming the test allows readers to verify its appropriateness for sparse cells (Fisher's exact is preferred when any expected cell count is less than 5); a power calculation contextualizes whether the study was adequately sized to detect differences of a clinically meaningful magnitude.
Software: Seurat · GSEA · BioRender.com

Result convergence & founder nodes

Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
23
Impact: medium
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

Assessed papers, coloured by verdict. Click a node to open it.

Built on (assessed references) (0)
  • No assessed neighbours yet — the network grows as more papers are assessed.
Cited by (assessed papers) (1)

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

GSE176021 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE179994 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE190265 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet

Downstream reach in the literature

27 downstream papers · 3 datasets

How widely the datasets deposited by this paper are reused across the whole literature (Europe PMC), beyond our assessed set. This is a factual dependency map — reusing a public dataset is normal, good science. It is not a judgement on the downstream papers; the only verdict here is this paper's own, with its cited rationale.

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-37231145

Paper: Hui et al. 2023, NPJ Precis Oncol. "PD-1 blockade potentiates neoadjuvant chemotherapy in NSCLC via increasing CD127+ and KLRG1+ CD8 T cells." DOI 10.1038/s41698-023-00384-x · PMID 37231145.

Datasets (clarification — brief vs paper)

  • GSE229353 = the paper's OWN newly generated scRNA-seq (7 NSCLC patients, CD45+ sorted cells, 10x 5'). Ships only GSE229353_RAW.tar (per-sample MTX/TSV count matrices; no annotation / no processed object).
  • GSE179994 (named in the brief) = a validation dataset (Liu et al.; 47 GSM / 36 pts NSCLC ICB). Used only in Fig. 7 for signature-enrichment validation. Ships a processed T-cell raw-counts RDS + a T-cell metadata TSV (annotations present, but T cells only — no B/myeloid compartment).
  • Code: CellChat (github.com/sqjin/CellChat) is the named third-party tool. Authors' own analysis code is "available on reasonable request" → NOT shipped.

Pipelines named in Methods

  • Cell Ranger 3.1.0 (align/quantify, GRCh38) — needs FASTQ (SRA); out of scope (heavy, and the deposited matrices already are its output).
  • Seurat 3.2.1 — QC filtering, PCA, clustering (res 0.8), MNN integration, DEGs.
  • CellChat 1.1.3 — cell-cell communication (Supplementary Fig. 9).
  • clusterProfiler 3.14.3 (GO), GSEA, GSVA, AddModuleScore — signature scoring.

In scope (pipeline-derived, attempted)

# Reported result Pipeline Reproducibility
C1 26,861 single-cell transcriptomes total (4,057 NAC / 22,804 NAPC) after QC Cell Ranger → Seurat 3.2.1 QC (genes 200–6000, UMI≥1000, mito<10%) HIGH — deterministic filter on shipped matrices
C2 29 clusters at resolution 0.8 (Fig 1b) Seurat PCA+MNN+FindClusters MEDIUM — clustering is version/seed sensitive (within-tol at best)
C3 CellChat: CD4+ T ↔ CD8-CD127+/CD8-KLRG1+ via CXCL12-CXCR4 & IL7-IL7R; B ↔ CD8 subsets via BTLA-TNFRSF14 (Supp Fig 9) CellChat 1.1.3 LOW/PARTIAL — requires authors' unshipped cell-type annotation; best-effort qualitative check only

Out of scope (not pipeline / not attempted)

  • mIHC / multispectral imaging proportions (Fig 3), InForm software — wet-lab imaging.
  • Spatial bivariate K(r) function (Fig 5) — derived from mIHC images, not scRNA.
  • GSEA/GSVA validation in GSE176021, GSE190265 (Fig 7) — directional, multi-dataset; beyond the clean 80%. GSE179994 validation is directional (no exact number printed).
  • Cell Ranger re-alignment from FASTQ — heavy, and its output (matrices) is already deposited.

Plan

  1. C1 (primary): download GSE229353 matrices on «infra», recompute per-cell nGene/ nUMI/mito% directly from MTX, apply the stated QC, count total + per-sample. 1:1 vs 26,861. (Transparent, no Seurat-version dependence for a pure count.)
  2. C3 (secondary, named tool): best-effort CellChat 1.1.3 on the paper data after a standard cluster+annotate, checking whether the named L-R pairs are inferred. Graded partial. Skip the exact-annotation last 20% (documented).
  3. C2 reported as a clustering sanity check if cheap; not a hard 1:1 target.
Figures / tables: Fig 1bFig 9
C1a
Reported
26,861 single-cell transcriptomes after QC (GSE229353)
Reproduced
22,509 cells pass the paper's stated Seurat QC over 6 of 7 samples; 31,203 total cellranger cell-calls across all 7; extrapolating the corrupt sample at the 80.2% intact pass-rate gives ~25,000 (within ~7% of reported)
partial
C1b
Reported
4,057 NAC / 22,804 NAPC cell split
Reproduced
unresolved - all 7 GSMs titled '_C'; the single NAC patient is not mappable to a GSM from deposited metadata
partial
C2
Reported
29 clusters at resolution 0.8
Reproduced
not attempted as a 1:1 target (Seurat version/seed/MNN-integration sensitive; P03 sample missing)
partial
C3
Reported
CellChat 1.1.3 L-R signaling: CD4+T->CD8 subsets via IL7-IL7R & CXCL12-CXCR4; B->CD8 subsets via BTLA-TNFRSF14 (Supp Fig 9)
Reproduced
not completed - CellChat 1.1.3 pinned to commit af7e152 (no v1.1.3 git tag exists); env build stalled compiling a non-essential graphics dependency (harfbuzz/freetype). Exact claim also requires the authors' unshipped fine cell-type annotation. Recipe + pinned commit shipped for a reviewer to finish.
did not match

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 40/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟢2. Endpoint comparability
🟡3. Location of the main deviation
🔴4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

209.7 k
tokens (I/O) · 20.7 M incl. cache
43 min
runtime · 0.07 CPU-h
3.2 GB
peak RAM
4 (4 failed)
HPC jobs
hummel
machine