Antigen and checkpoint receptor engagement recalibrates T cell receptor signal strength.
Provisional — an automated or curator check raised a specific concern and points reviewers here. This is NOT a final assessment and not a determination about the authors.
The main result did not reproduce in this reproduction attempt. Where our recomputation produced values that differ from the published ones, those discrepancies are listed below. This is a single automated attempt — not peer review and not a finding of error or misconduct — and differences can also arise from data access, undocumented parameters or the computing environment. The verdict can be contested via “report an error”.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Reported values were directly comparable
- 🟡Could not use the authors’ exact input data
- 🟡A deviation arose in the data or preprocessing
- 🔴A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🔴The deviation was non-trivial in magnitude
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Described well enough to ATTEMPT, but reproduces only PARTIALLY (different result). Scaffold metadata was wrong: the auto-linked code (riazn/bms038_analysis = Riaz 2017) and data (GSE93018 = Efremova 2018) are text-mining false positives w.r.t. this paper; the authors' OWN analysis code is NOT public ('available from lead contact upon reasonable request'). The paper is mostly Nr4a3-Tocky flow-cytometry wet-lab (out of scope) plus QuantSeq RNA-seq (own data GSE165817/818) and re-analysis of public ICB cohorts. Per brief HARD RULE 2/P16 we reproduced the most precisely-specified PUBLIC-data pipeline result: the TCR.strong 5-gene (TNFRSF4,ICOS,IRF8,TNIP3,STAT4) geometric-mean-TPM survival stratification of the Riaz nivolumab cohort (Fig 7F/7G), using GSE91061 FPKM->TPM + the public Riaz clinical table, with R/survival on «our HPC»/«infra». RESULT: the Ipi-naive OS association DIRECTION reproduces (TCR.strong High -> better OS, median 101 vs 77 wk, matching Fig 7G) but the reported SIGNIFICANCE does not (we get p=0.37 vs <0.05); whole-cohort PFS (Fig 7F) does not reproduce at all (p=0.86, opposite direction). Our cohort is n=51 (25 naive/26 prog) vs the paper's n=50 (21/29) -- the exact author subset/QC is unrecoverable without their code. This is a REPRODUCIBILITY GAP, explicitly NOT an assertion of fabrication; flagged provisionally for human audit. NOT attempted: Gide cohort (Fig 7L, raw-fastq alignment), GSE165817/818 DESeq2 DEG (own QuantSeq, claims less pinnable), all flow-cytometry/Tocky wet-lab (non-pipeline).
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 37assessed: 2026-06-15 ⛓ ac8c934ef7b0
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-15
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator headless) · v1.0 · run #1 2026-06-15no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusThe study tests how TCR signal strength modulates CD4+ T cell function and activation thresholds in vivo, and to what extent immune checkpoint blockade (e.g., anti-PD1) alters these processes, hypothesizing that TCR signal strength can serve as a measurable metric for monitoring immunotherapy outcomes.
- ★ TCR signal strength drives dynamic, dose- and time-dependent transcriptional changes in CD4+ T cells while single-cell Nr4a3 activation dynamics remain digital/uniform. finding
- ★ Dose-dependent programming of distinct co-inhibitory receptor modules rapidly recalibrates T cell activation thresholds, reducing responsiveness to re-stimulation after strong initial TCR signaling. finding
- ★ Anti-PD1 (PD1 blockade) lowers the T cell activation threshold and leads to an increased/strong TCR signal transcriptional signature in T cells. finding
- ★ A 5-gene strong TCR signal metric (TCR.strong) upregulated by anti-PD1 stratifies melanoma patient responses to anti-PD1 therapy and is superior to a canonical T cell activation signature. resource
- ★ The Nr4a3-Tocky/Tg4 reporter system enables in vivo tracking of synchronized T cell activation and classification of TCR signals as new, persistent, or arrested. method
- Il10+ T cells reflect cells receiving the strongest TCR signaling and show higher inhibitory receptor (Lag3, Tigit) expression and greater non-responsiveness to re-stimulation. mechanism
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Flow cytometry (Nr4a3-Timer Blue/Red, Il10-GFP, surface/intracellular markers) | Tg4 Nr4a3-Tocky Il10-GFP mice, splenic CD4+ T cells | s.c. immunization with [4Y] MBP peptide (0.8, 8, 80 μg), no adjuvant | % Nr4a3-Blue+ active TCR signaling, Nr4a3-Timer Angle, Il10-GFP+ frequency | — |
| 3' mRNA RNA-seq | Sorted splenic CD4+ Tg4 T cells (synchronized by Timer maturation) | s.c. [4Y] MBP peptide 0.8 μg vs 80 μg at 4, 12, 24 h | differentially expressed genes (DESeq2), PCA clusters, KEGG pathways | 3' mRNA library prep |
| Flow cytometry of inhibitory receptors | Live CD4+ Nr4a3-Timer+ Tg4 T cells | s.c. [4Y] MBP dose range | PD1, Lag3, Tigit, CTLA-4 expression | — |
| In vivo re-stimulation / threshold recalibration assay | Tg4 Nr4a3-Tocky Il10-GFP mice, splenic CD4+ T cells | Prime 0/8/80 μg [4Y] MBP, 24 h later re-challenge 8 or 80 μg for 4 h | % arrested (Nr4a3-Blue−Red+) vs reactivated T cells | — |
| In vivo checkpoint blockade assay | Tg4 Nr4a3-Tocky T cells in vivo | Blocking antibodies (PD1, CTLA-4, Lag3) or agonistic CD28, 30 min before peptide challenge | Nr4a3-Blue re-activation / threshold modulation | — |
| In vitro T cell activation | Tg4 T cells | MBP peptide variants [4K], [4A], [4Y] | Nr4a3, CD69, CD25, CD44 expression correlation | — |
| Clinical transcriptomic stratification | Melanoma patients receiving anti-PD1 therapy | anti-PD1 (nivolumab) therapy | TCR.strong 5-gene signature vs canonical activation signature, patient outcome stratification | — |
- ▲ Proportion of responding (Nr4a3-Blue+) T cells increased with TCR signal strength, peaking at 4 h and falling to near-zero by 24 h.
- – Nr4a3-Timer Angle trajectories were highly similar regardless of immunizing dose, indicating dose-independent single-cell Nr4a3 dynamics.
- ▲ Il10-GFP expression within Nr4a3-Timer+ T cells increased directly with immunizing antigen dose.
- ▼ Most DEGs between high and low antigen occurred at 4 h and declined over time, with most genes unique to each time point.
- ▲ Lag3, Tigit, and CTLA-4 expression were strongly dose-dependent, whereas PD1 was tightly coupled to activation but only modestly dose-dependent.
- ▼ More T cells primed with 80 μg failed to re-activate Nr4a3 upon re-challenge (remained in arrested locus) than those primed with 8 μg.
- ▲ Il10-GFPhi T cells had significantly higher Lag3 and Tigit and increased non-responsiveness to re-stimulation versus Il10-lo/− cells.
- – Agonistic CD28 antibody did not alter the T cell activation threshold in the re-activation model.
- other half-life Nr4a3-Blue 4 h vs Nr4a3-Red 120 h (Timer protein maturation basis for arrested-signal detection)
- fold_change 100-fold dose range (0.8, 8, 80 μg) ([4Y] MBP peptide immunization range)
- count 5 genes (TCR.strong metric of genes upregulated by anti-PD1)
- count 7 modules (transcriptional modules undergoing time/dose-dependent regulation)
- pvalue p<0.05 threshold; ∗∗p<0.01, ∗∗∗p<0.001, ∗∗∗∗p<0.0001 (statistical thresholds across flow cytometry comparisons (n=3–6 per group))
Statistical methods review
Model: opusA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This is an experimental immunology study using Nr4a3-Tocky reporter mice, combining flow cytometry of T cell populations with 3' mRNA RNA-seq and analysis of a clinical melanoma cohort. Group comparisons of flow cytometric readouts were made with one-way or two-way ANOVA followed by Tukey's or Sidak's multiple comparisons tests, and paired/unpaired t tests for two-group comparisons; transcriptomic differences were assessed with DESeq2 with downstream PCA and KEGG pathway analysis. Results were reported as mean ± SEM with significance indicated by p-value thresholds (asterisk symbols).
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| two-way ANOVA with Tukey's multiple comparisons test | Figure 1C and 1D: percent active TCR signaling and Nr4a3-Timer Angle across 0.8/8/80 µg doses | — | not stated |
| one-way ANOVA with Tukey's multiple comparisons test | Figure 1F (Il10-GFP expressers, n=4); Figure 3B (PD1, n=3); Figure 3C Lag3 (n=6); Figure 3D Tigit (n=6) | n=3 to n=6 as stated per figure | not stated |
| two-way ANOVA with Sidak's multiple comparisons test | Figure 4C: frequency of arrested TCR signaling T cells (n=3) | n=3 | not stated |
| unpaired t test | Figure 3G: Lag3 and Tigit in Il10-hi vs Il10-lo cells (n=3) | n=3 | not stated |
| paired t test | Figure 4E: percent Nr4a3-Blue+ in Il10-GFP hi vs lo after re-challenge (n=3) | n=3 | not stated |
| DESeq2 (differential expression) | Figure 2C: DEGs between 80 µg and 0.8 µg stimulated T cells at 4, 12, 24 h | — | not stated |
-
Spread was summarized as mean ± SEM throughout the figures.↳ Could also: The same data could also be displayed with SD or a 95% confidence interval. — SD conveys the variability of the underlying observations and a CI conveys precision of the estimate; both are often favored for small-n datasets, so reporting them alongside SEM would give readers a complementary view of dispersion and effect magnitude.
-
Significance was indicated with p-value threshold symbols (e.g., *, **, ***).↳ Could also: Exact p values together with effect-size estimates (e.g., mean differences with confidence intervals) could also be reported. — Exact p values and effect sizes communicate the magnitude and precision of differences rather than only whether a threshold was crossed, which can aid interpretation and downstream meta-analysis.
-
Several comparisons across three dose groups used one-way ANOVA with Tukey's post-hoc test, while two-group comparisons used t tests.↳ Could also: For the small-n flow cytometry readouts, non-parametric tests (e.g., Kruskal-Wallis with Dunn's, or Mann-Whitney) could also be applied. — Non-parametric alternatives relax the normality assumption, which can be useful when n is small and distributional assumptions are difficult to verify; presenting them can serve as a robustness check.
-
Differential expression was assessed with DESeq2 and significance/visualization conveyed via heatmaps, PCA, and KEGG enrichment.↳ Could also: An explicit statement of the FDR method and threshold (e.g., Benjamini-Hochberg adjusted-p cutoff) could also accompany the DEG lists. — DESeq2 applies BH adjustment by default; stating the adjusted-p threshold and fold-change criteria makes the multiplicity handling fully transparent and reproducible for the genome-wide comparisons.
-
Distributional assumptions for the parametric tests were not stated.↳ Could also: Reporting assumption checks (e.g., normality/variance-homogeneity diagnostics) or pre-registering the analysis choices could also be done. — Documenting assumption checks helps readers confirm the chosen parametric models are appropriate and supports reproducibility, particularly for the small-sample comparisons.
-
Sample sizes were given per figure without a described power calculation.↳ Could also: An a priori power/sample-size justification could also be provided. — A stated power analysis clarifies the basis for the chosen n and the sensitivity of the study to detect effects of a given size, complementing the descriptive per-figure n values.
Result convergence & founder nodes
Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.
-
Agonistic CD28 antibody does not alter the TCR re-activation threshold in antigen-primed CD4+ T cells in vivo, contrasting with checkpoint receptor blockadeflow-cytometry mouse splenic-cd4-t-cell none 2021×1papers★ This paper is the founder (earliest)
-
IL10 expression within NR4A3+ antigen-responding splenic CD4+ T cells increases proportionally with immunizing antigen doseflow-cytometry mouse splenic-cd4-t-cell up 2021×1papers★ This paper is the founder (earliest)
-
LAG3, TIGIT, and CTLA4 expression on antigen-responding CD4+ T cells is strongly antigen dose-dependent; PDCD1 is activation-linked but only modestly dose-dependentflow-cytometry mouse splenic-cd4-t-cell up 2021×1papers★ This paper is the founder (earliest)
-
NR4A3 re-activation is reduced in CD4+ T cells primed with high-dose antigen (80 μg vs 8 μg), demonstrating dose-dependent upward recalibration of the TCR re-activation thresholdflow-cytometry mouse splenic-cd4-t-cell down 2021×1papers★ This paper is the founder (earliest)
-
NR4A3-Timer angle (single-cell TCR signaling dynamics) is dose-independent in antigen-responding CD4+ T cells; dose controls the responding fraction, not per-cell kineticsflow-cytometry mouse splenic-cd4-t-cell none 2021×1papers★ This paper is the founder (earliest)
-
NR4A3-Blue+ active TCR signaling fraction in splenic CD4+ T cells increases with antigen dose, peaking at 4 h post-immunizationflow-cytometry mouse splenic-cd4-t-cell up 2021×1papers★ This paper is the founder (earliest)
-
IL10-high CD4+ T cells show significantly elevated LAG3 and TIGIT and increased non-responsiveness to re-stimulation relative to IL10-low cellsflow-cytometry mouse splenic-cd4-t-cell up 2021×1papers★ This paper is the founder (earliest)
-
Antigen dose-dependent transcriptional differences in CD4+ T cells peak at 4 h post-immunization and decline over time, with most DEGs unique to individual timepointsRNA-seq mouse splenic-cd4-t-cell down 2021×1papers★ This paper is the founder (earliest)
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
Downstream reach in the literature
148 downstream papers · 5 datasetsHow widely the datasets deposited by this paper are reused across the whole literature (Europe PMC), beyond our assessed set. This is a factual dependency map — reusing a public dataset is normal, good science. It is not a judgement on the downstream papers; the only verdict here is this paper's own, with its cited rationale.
- Defining T Cell States Associated with Response to C... 2018 · 1,561 cites
- Molecular and pharmacological modulators of the tumo... 2019 · 1,199 cites
- ImmuCellAI: A Unique Method for Comprehensive T-Cell... 2020 · 728 cites
- Collagen promotes anti-PD-1/PD-L1 resistance in canc... 2020 · 335 cites
- B cells sustain inflammation and predict response to... 2019 · 319 cites
- Exaggerated false positives by popular differential... 2022 · 237 cites
- CXCR4 inhibition in human pancreatic and colorectal... 2020 · 219 cites
- Antagonistic Inflammatory Phenotypes Dictate Tumor F... 2020 · 192 cites
- 9p21 loss confers a cold tumor immune microenvironme... 2021 · 177 cites
- A gene expression signature of TREM2<sup>hi</sup> ma... 2020 · 150 cites
- Ablation of the endoplasmic reticulum stress kinase... 2022 · 121 cites
- Blockade of LAG-3 and PD-1 leads to co-expression of... 2024 · 116 cites
- Lag3 and PD-1 pathways preferentially regulate NFAT-... 2025 · 1 cites
- Interferon-γ and IL-27 positively regulate type 1 re... 2025 · 0 cites
- Targeting immune checkpoints potentiates immunoediti... 2018 · 238 cites
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-34534438
Paper: Elliot TAE et al. "Antigen and checkpoint receptor engagement recalibrates T cell receptor signal strength." Immunity 2021. PMID 34534438 · PMC8585507 · DOI 10.1016/j.immuni.2021.08.020. Lead contact: David Bending (U. Birmingham).
Metadata correction (scaffold was mis-linked)
The auto-harvested registry pointed to github.com/riazn/bms038_analysis as the
"code" and GSE93018 as the "data". Both are text-mining false positives w.r.t.
authorship:
riazn/bms038_analysisis Riaz et al. 2017 (nivolumab melanoma) — used by THIS paper only as a third-party data source (clinical outcomes + expression of the Riaz cohort), NOT as the authors' analysis code.GSE93018is Efremova et al. 2018 (MC38 immunoediting, Innsbruck, PMID 29296022) — re-analysed by this paper as one external cohort, not its own data.
What this paper actually ships (verified from STAR Methods / PMC)
- Own new data: GEO GSE165817 (Fig 2, Tocky/TCR-signal-strength QuantSeq RNA-seq) and GSE165818 (Fig 5, anti-PD1 QuantSeq RNA-seq). Public.
- Re-analysed public cohorts: GSE91061 (Riaz, FPKM), GSE93018 (Efremova MC38), ENA PRJEB23709 (Gide 2019, raw fastq).
- Authors' own analysis code: "available from the lead contact upon reasonable
request" → NOT publicly deposited. (drop-reason
no_codefor the authors' code.)
In scope (pipeline-derived, public data, standard tools) — ATTEMPTED
Per brief HARD RULE 2 (P16): applying a standard tool to the paper's own data / clearly-specified method is equally valid. We reproduce the most precisely-specified, fully-public computational result: the TCR.strong signature survival stratification.
- TCR.strong score = geometric mean of TPM of 5 genes TNFRSF4, ICOS, IRF8, TNIP3, STAT4 per patient (paper Methods, Fig 7).
- C1 (Fig 7F): Riaz/nivolumab whole cohort (n=50), split at median TCR.strong → Kaplan-Meier PFS, log-rank. Reported: High > Low, p < 0.05.
- C2 (Fig 7G): Riaz Ipi-naive subgroup (Ipi-N, n=21) → OS, log-rank. Reported: High > Low, p < 0.05.
- Data: GSE91061 FPKM (public GEO suppl) for expression; Riaz
bms038_analysis/data(public) for per-patient PFS/OS/cohort. Tools: R +survival.
Out of scope / not attempted (with reason)
- C3 (Fig 7L) Gide cohort (PFS p=0.0034, OS p=0.017): needs raw-fastq alignment of ENA PRJEB23709 (Partek/STAR) → the hard last 20%; recorded but not run.
- GSE165817/818 DESeq2 DEG (Fig 2/5): own QuantSeq counts are public and DESeq2 is reproducible, but the specific reported claims (exact DEG counts/gene lists) are less precisely pinnable than C1/C2 from text alone; deprioritised (80/20).
- All flow cytometry / Nr4a3-Tocky reporter wet-lab (Fig 1,3,4,6, most of 2/5): manual/instrument (FlowJo, GraphPad) → out of scope (not a bioinformatic pipeline).
- Authors' exact code path: not public ("on request") → cannot run their code 1:1.
Compute plan
Single SLURM job on a «our HPC» compute node (internet) with cwd on «infra»:
download GSE91061 FPKM + clone bms038 clinical data → R: FPKM→TPM, geomean(5 genes),
median split, survdiff/survfit for PFS (whole) & OS (Ipi-N). Light compute, but
run on «our HPC» + «infra» per HARD RULE 1/1b.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
This is a third-party reimplementation (authors' own analysis code is not public) of the paper's TCR.strong survival stratification on public Riaz/nivolumab data. The Fig 7G Ipi-naive OS direction reproduces (median 101.1 > 77.4 wk) but its significance does not (p=0.367 vs <0.05), and Fig 7F whole-cohort PFS fails entirely (p=0.859, direction reversed). The deviation sits mainly on the input/cohort-definition side (n=51 vs 50; 25/26 vs 21/29), with root cause on the authors' side — the reported significance is not derivable from shared public data because the exact subset/QC lives only in non-public code. Severity is high (significance + PFS-direction flip) but this is documented as a reproducibility gap, not fabrication; overall criticality lands yellow.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.