Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Antigen and checkpoint receptor engagement recalibrates T cell receptor signal strength.

Immunity · 2021
L1 37/100 3/4
⚑ Flagged for review — a reproduced result did not match the reported value

Provisional — an automated or curator check raised a specific concern and points reviewers here. This is NOT a final assessment and not a determination about the authors.

Why this verdict

The main result did not reproduce in this reproduction attempt. Where our recomputation produced values that differ from the published ones, those discrepancies are listed below. This is a single automated attempt — not peer review and not a finding of error or misconduct — and differences can also arise from data access, undocumented parameters or the computing environment. The verdict can be contested via “report an error”.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +5
✓ What held up
  • Reported values were directly comparable
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🟡A deviation arose in the data or preprocessing
  • 🔴A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🔴The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
37/100
Reproducibility score
2.1 SD below mean
vs. all fields · 1173 studies
🎯 Scores higher than 3% of all assessed papers rank 1134 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough to ATTEMPT, but reproduces only PARTIALLY (different result). Scaffold metadata was wrong: the auto-linked code (riazn/bms038_analysis = Riaz 2017) and data (GSE93018 = Efremova 2018) are text-mining false positives w.r.t. this paper; the authors' OWN analysis code is NOT public ('available from lead contact upon reasonable request'). The paper is mostly Nr4a3-Tocky flow-cytometry wet-lab (out of scope) plus QuantSeq RNA-seq (own data GSE165817/818) and re-analysis of public ICB cohorts. Per brief HARD RULE 2/P16 we reproduced the most precisely-specified PUBLIC-data pipeline result: the TCR.strong 5-gene (TNFRSF4,ICOS,IRF8,TNIP3,STAT4) geometric-mean-TPM survival stratification of the Riaz nivolumab cohort (Fig 7F/7G), using GSE91061 FPKM->TPM + the public Riaz clinical table, with R/survival on «our HPC»/«infra». RESULT: the Ipi-naive OS association DIRECTION reproduces (TCR.strong High -> better OS, median 101 vs 77 wk, matching Fig 7G) but the reported SIGNIFICANCE does not (we get p=0.37 vs <0.05); whole-cohort PFS (Fig 7F) does not reproduce at all (p=0.86, opposite direction). Our cohort is n=51 (25 naive/26 prog) vs the paper's n=50 (21/29) -- the exact author subset/QC is unrecoverable without their code. This is a REPRODUCIBILITY GAP, explicitly NOT an assertion of fabrication; flagged provisionally for human audit. NOT attempted: Gide cohort (Fig 7L, raw-fastq alignment), GSE165817/818 DESeq2 DEG (own QuantSeq, claims less pinnable), all flow-cytometry/Tocky wet-lab (non-pipeline).

💻 Code ↗ 🗄 Data: GSE93018

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 37
    assessed: 2026-06-15 ⛓ ac8c934ef7b0
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-15
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator headless) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

The study tests how TCR signal strength modulates CD4+ T cell function and activation thresholds in vivo, and to what extent immune checkpoint blockade (e.g., anti-PD1) alters these processes, hypothesizing that TCR signal strength can serve as a measurable metric for monitoring immunotherapy outcomes.

Core claims
  • TCR signal strength drives dynamic, dose- and time-dependent transcriptional changes in CD4+ T cells while single-cell Nr4a3 activation dynamics remain digital/uniform. finding
  • Dose-dependent programming of distinct co-inhibitory receptor modules rapidly recalibrates T cell activation thresholds, reducing responsiveness to re-stimulation after strong initial TCR signaling. finding
  • Anti-PD1 (PD1 blockade) lowers the T cell activation threshold and leads to an increased/strong TCR signal transcriptional signature in T cells. finding
  • A 5-gene strong TCR signal metric (TCR.strong) upregulated by anti-PD1 stratifies melanoma patient responses to anti-PD1 therapy and is superior to a canonical T cell activation signature. resource
  • The Nr4a3-Tocky/Tg4 reporter system enables in vivo tracking of synchronized T cell activation and classification of TCR signals as new, persistent, or arrested. method
  • Il10+ T cells reflect cells receiving the strongest TCR signaling and show higher inhibitory receptor (Lag3, Tigit) expression and greater non-responsiveness to re-stimulation. mechanism
Experimental setups
Assay System Perturbation Readout Platform
Flow cytometry (Nr4a3-Timer Blue/Red, Il10-GFP, surface/intracellular markers) Tg4 Nr4a3-Tocky Il10-GFP mice, splenic CD4+ T cells s.c. immunization with [4Y] MBP peptide (0.8, 8, 80 μg), no adjuvant % Nr4a3-Blue+ active TCR signaling, Nr4a3-Timer Angle, Il10-GFP+ frequency
3' mRNA RNA-seq Sorted splenic CD4+ Tg4 T cells (synchronized by Timer maturation) s.c. [4Y] MBP peptide 0.8 μg vs 80 μg at 4, 12, 24 h differentially expressed genes (DESeq2), PCA clusters, KEGG pathways 3' mRNA library prep
Flow cytometry of inhibitory receptors Live CD4+ Nr4a3-Timer+ Tg4 T cells s.c. [4Y] MBP dose range PD1, Lag3, Tigit, CTLA-4 expression
In vivo re-stimulation / threshold recalibration assay Tg4 Nr4a3-Tocky Il10-GFP mice, splenic CD4+ T cells Prime 0/8/80 μg [4Y] MBP, 24 h later re-challenge 8 or 80 μg for 4 h % arrested (Nr4a3-Blue−Red+) vs reactivated T cells
In vivo checkpoint blockade assay Tg4 Nr4a3-Tocky T cells in vivo Blocking antibodies (PD1, CTLA-4, Lag3) or agonistic CD28, 30 min before peptide challenge Nr4a3-Blue re-activation / threshold modulation
In vitro T cell activation Tg4 T cells MBP peptide variants [4K], [4A], [4Y] Nr4a3, CD69, CD25, CD44 expression correlation
Clinical transcriptomic stratification Melanoma patients receiving anti-PD1 therapy anti-PD1 (nivolumab) therapy TCR.strong 5-gene signature vs canonical activation signature, patient outcome stratification
Key results
  • Proportion of responding (Nr4a3-Blue+) T cells increased with TCR signal strength, peaking at 4 h and falling to near-zero by 24 h.
  • Nr4a3-Timer Angle trajectories were highly similar regardless of immunizing dose, indicating dose-independent single-cell Nr4a3 dynamics.
  • Il10-GFP expression within Nr4a3-Timer+ T cells increased directly with immunizing antigen dose.
  • Most DEGs between high and low antigen occurred at 4 h and declined over time, with most genes unique to each time point.
  • Lag3, Tigit, and CTLA-4 expression were strongly dose-dependent, whereas PD1 was tightly coupled to activation but only modestly dose-dependent.
  • More T cells primed with 80 μg failed to re-activate Nr4a3 upon re-challenge (remained in arrested locus) than those primed with 8 μg.
  • Il10-GFPhi T cells had significantly higher Lag3 and Tigit and increased non-responsiveness to re-stimulation versus Il10-lo/− cells.
  • Agonistic CD28 antibody did not alter the T cell activation threshold in the re-activation model.
Key statistics
  • other half-life Nr4a3-Blue 4 h vs Nr4a3-Red 120 h (Timer protein maturation basis for arrested-signal detection)
  • fold_change 100-fold dose range (0.8, 8, 80 μg) ([4Y] MBP peptide immunization range)
  • count 5 genes (TCR.strong metric of genes upregulated by anti-PD1)
  • count 7 modules (transcriptional modules undergoing time/dose-dependent regulation)
  • pvalue p<0.05 threshold; ∗∗p<0.01, ∗∗∗p<0.001, ∗∗∗∗p<0.0001 (statistical thresholds across flow cytometry comparisons (n=3–6 per group))

Statistical methods review

Model: opus

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This is an experimental immunology study using Nr4a3-Tocky reporter mice, combining flow cytometry of T cell populations with 3' mRNA RNA-seq and analysis of a clinical melanoma cohort. Group comparisons of flow cytometric readouts were made with one-way or two-way ANOVA followed by Tukey's or Sidak's multiple comparisons tests, and paired/unpaired t tests for two-group comparisons; transcriptomic differences were assessed with DESeq2 with downstream PCA and KEGG pathway analysis. Results were reported as mean ± SEM with significance indicated by p-value thresholds (asterisk symbols).

Replicationbiological Sample sizePer-figure n values stated (e.g., n=3, 4, 6); no formal power/sample-size calculation described in the available text GroupsT cells from mice immunized with graded [4Y] MBP peptide doses (0/0.8/8/80 µg) across time points; Il10-hi vs Il10-lo subsets; checkpoint-blockade conditions Pairingmixed Randomization/blindingstated DispersionSEM Exact p-valuesno Effect sizesno Confidence intervalsno Multiplicity correctionTukey's and Sidak's multiple comparisons tests (post-hoc to ANOVA); DESeq2 used for RNA-seq (BH FDR not explicitly stated in available text)
Statistical tests used
Test Applied to n Assumptions
two-way ANOVA with Tukey's multiple comparisons test Figure 1C and 1D: percent active TCR signaling and Nr4a3-Timer Angle across 0.8/8/80 µg doses not stated
one-way ANOVA with Tukey's multiple comparisons test Figure 1F (Il10-GFP expressers, n=4); Figure 3B (PD1, n=3); Figure 3C Lag3 (n=6); Figure 3D Tigit (n=6) n=3 to n=6 as stated per figure not stated
two-way ANOVA with Sidak's multiple comparisons test Figure 4C: frequency of arrested TCR signaling T cells (n=3) n=3 not stated
unpaired t test Figure 3G: Lag3 and Tigit in Il10-hi vs Il10-lo cells (n=3) n=3 not stated
paired t test Figure 4E: percent Nr4a3-Blue+ in Il10-GFP hi vs lo after re-challenge (n=3) n=3 not stated
DESeq2 (differential expression) Figure 2C: DEGs between 80 µg and 0.8 µg stimulated T cells at 4, 12, 24 h not stated
Approaches that could also have been used
  • Spread was summarized as mean ± SEM throughout the figures.
    Could also: The same data could also be displayed with SD or a 95% confidence interval. — SD conveys the variability of the underlying observations and a CI conveys precision of the estimate; both are often favored for small-n datasets, so reporting them alongside SEM would give readers a complementary view of dispersion and effect magnitude.
  • Significance was indicated with p-value threshold symbols (e.g., *, **, ***).
    Could also: Exact p values together with effect-size estimates (e.g., mean differences with confidence intervals) could also be reported. — Exact p values and effect sizes communicate the magnitude and precision of differences rather than only whether a threshold was crossed, which can aid interpretation and downstream meta-analysis.
  • Several comparisons across three dose groups used one-way ANOVA with Tukey's post-hoc test, while two-group comparisons used t tests.
    Could also: For the small-n flow cytometry readouts, non-parametric tests (e.g., Kruskal-Wallis with Dunn's, or Mann-Whitney) could also be applied. — Non-parametric alternatives relax the normality assumption, which can be useful when n is small and distributional assumptions are difficult to verify; presenting them can serve as a robustness check.
  • Differential expression was assessed with DESeq2 and significance/visualization conveyed via heatmaps, PCA, and KEGG enrichment.
    Could also: An explicit statement of the FDR method and threshold (e.g., Benjamini-Hochberg adjusted-p cutoff) could also accompany the DEG lists. — DESeq2 applies BH adjustment by default; stating the adjusted-p threshold and fold-change criteria makes the multiplicity handling fully transparent and reproducible for the genome-wide comparisons.
  • Distributional assumptions for the parametric tests were not stated.
    Could also: Reporting assumption checks (e.g., normality/variance-homogeneity diagnostics) or pre-registering the analysis choices could also be done. — Documenting assumption checks helps readers confirm the chosen parametric models are appropriate and supports reproducibility, particularly for the small-sample comparisons.
  • Sample sizes were given per figure without a described power calculation.
    Could also: An a priori power/sample-size justification could also be provided. — A stated power analysis clarifies the basis for the chosen n and the sensitivity of the study to detect effects of a given size, complementing the descriptive per-figure n values.
Software: DESeq2 (differential expression) · KEGG pathway analysis

Result convergence & founder nodes

Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
66
Impact: high
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

1x10 in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
also used by 2 papers:
GSE91061 GEO in Methods (http://purl.org/orb/Methods)
also used by 2 papers:
PRJEB23709 in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
also used by 1 paper:
4hrs in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
A16517 in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
CVCL_B288 RRID in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet
GSE165817 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE165818 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE93018 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
RRID:AB_1027648 RRID in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet
RRID:AB_10548034 RRID in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet
RRID:AB_10612935 RRID in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet
RRID:AB_10639935 RRID in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet
RRID:AB_10696422 RRID in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet
RRID:AB_2159183 RRID in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet
RRID:AB_2290801 RRID in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet
RRID:AB_2532296 RRID in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet
RRID:AB_2561723 RRID in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet
RRID:AB_2563385 RRID in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet
RRID:AB_2565648 RRID in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet
RRID:AB_2566126 RRID in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet
RRID:AB_2566284 RRID in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet
RRID:AB_2566285 RRID in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet
RRID:AB_2572741 RRID in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet
RRID:AB_2616834 RRID in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet
RRID:AB_2621911 RRID in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet
RRID:AB_2687796 RRID in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet
RRID:AB_2715763 RRID in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet
RRID:AB_2715958 RRID in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet
RRID:AB_2732919 RRID in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet
RRID:AB_2738426 RRID in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet
RRID:AB_2742163 RRID in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet
RRID:AB_2870092 RRID in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet
RRID:AB_313254 RRID in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet
RRID:AB_429727 RRID in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet
RRID:AB_492843 RRID in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet
RRID:AB_493713 RRID in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet
RRID:AB_893288 RRID in Article (http://semanticscience.org/resource/SIO_001029)
no other assessed paper uses this yet

Downstream reach in the literature

148 downstream papers · 5 datasets

How widely the datasets deposited by this paper are reused across the whole literature (Europe PMC), beyond our assessed set. This is a factual dependency map — reusing a public dataset is normal, good science. It is not a judgement on the downstream papers; the only verdict here is this paper's own, with its cited rationale.

This paper is currently under reproducibility review (see the verdict above). The map below shows where the data in question has propagated — so reuse can be traced, not so the downstream work is presumed affected.
GSE165817 GEO reused by 4 papers in the literature
Most-cited downstream papers:
GSE93018 GEO reused by 2 papers in the literature
Most-cited downstream papers:
GSE165818 GEO reused by 1 papers in the literature

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-34534438

Paper: Elliot TAE et al. "Antigen and checkpoint receptor engagement recalibrates T cell receptor signal strength." Immunity 2021. PMID 34534438 · PMC8585507 · DOI 10.1016/j.immuni.2021.08.020. Lead contact: David Bending (U. Birmingham).

Metadata correction (scaffold was mis-linked)

The auto-harvested registry pointed to github.com/riazn/bms038_analysis as the "code" and GSE93018 as the "data". Both are text-mining false positives w.r.t. authorship:

  • riazn/bms038_analysis is Riaz et al. 2017 (nivolumab melanoma) — used by THIS paper only as a third-party data source (clinical outcomes + expression of the Riaz cohort), NOT as the authors' analysis code.
  • GSE93018 is Efremova et al. 2018 (MC38 immunoediting, Innsbruck, PMID 29296022) — re-analysed by this paper as one external cohort, not its own data.

What this paper actually ships (verified from STAR Methods / PMC)

  • Own new data: GEO GSE165817 (Fig 2, Tocky/TCR-signal-strength QuantSeq RNA-seq) and GSE165818 (Fig 5, anti-PD1 QuantSeq RNA-seq). Public.
  • Re-analysed public cohorts: GSE91061 (Riaz, FPKM), GSE93018 (Efremova MC38), ENA PRJEB23709 (Gide 2019, raw fastq).
  • Authors' own analysis code: "available from the lead contact upon reasonable request" → NOT publicly deposited. (drop-reason no_code for the authors' code.)

In scope (pipeline-derived, public data, standard tools) — ATTEMPTED

Per brief HARD RULE 2 (P16): applying a standard tool to the paper's own data / clearly-specified method is equally valid. We reproduce the most precisely-specified, fully-public computational result: the TCR.strong signature survival stratification.

  • TCR.strong score = geometric mean of TPM of 5 genes TNFRSF4, ICOS, IRF8, TNIP3, STAT4 per patient (paper Methods, Fig 7).
  • C1 (Fig 7F): Riaz/nivolumab whole cohort (n=50), split at median TCR.strong → Kaplan-Meier PFS, log-rank. Reported: High > Low, p < 0.05.
  • C2 (Fig 7G): Riaz Ipi-naive subgroup (Ipi-N, n=21) → OS, log-rank. Reported: High > Low, p < 0.05.
  • Data: GSE91061 FPKM (public GEO suppl) for expression; Riaz bms038_analysis/data (public) for per-patient PFS/OS/cohort. Tools: R + survival.

Out of scope / not attempted (with reason)

  • C3 (Fig 7L) Gide cohort (PFS p=0.0034, OS p=0.017): needs raw-fastq alignment of ENA PRJEB23709 (Partek/STAR) → the hard last 20%; recorded but not run.
  • GSE165817/818 DESeq2 DEG (Fig 2/5): own QuantSeq counts are public and DESeq2 is reproducible, but the specific reported claims (exact DEG counts/gene lists) are less precisely pinnable than C1/C2 from text alone; deprioritised (80/20).
  • All flow cytometry / Nr4a3-Tocky reporter wet-lab (Fig 1,3,4,6, most of 2/5): manual/instrument (FlowJo, GraphPad) → out of scope (not a bioinformatic pipeline).
  • Authors' exact code path: not public ("on request") → cannot run their code 1:1.

Compute plan

Single SLURM job on a «our HPC» compute node (internet) with cwd on «infra»: download GSE91061 FPKM + clone bms038 clinical data → R: FPKM→TPM, geomean(5 genes), median split, survdiff/survfit for PFS (whole) & OS (Ipi-N). Light compute, but run on «our HPC» + «infra» per HARD RULE 1/1b.

Figures / tables: Figure 7FFigure 7GFigure 7L
C1
Reported
Fig 7F: Riaz/nivolumab whole cohort (n=50), TCR.strong High>Low PFS, log-rank p<0.05
Reproduced
log-rank p=0.859; median PFS High=8.1wk < Low=24.1wk (n=51, 25/26); opposite direction, non-significant
did not match
C2
Reported
Fig 7G: Riaz Ipilimumab-naive (n=21), TCR.strong High>Low OS, log-rank p<0.05
Reproduced
DIRECTION reproduced (median OS High=101.1wk > Low=77.4wk) but NOT significant: log-rank p=0.367 (n=25, 12/13)
partial
C3
Reported
Fig 7L: Gide 2019 cohort (n=18), PFS p=0.0034 / OS p=0.017
Reproduced
not attempted (requires raw-fastq alignment of ENA PRJEB23709 = hard last 20%)
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator headless) · v1.0 L1 37/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟢2. Endpoint comparability
🟡3. Location of the main deviation
🔴4. Cause of the deviation
🟡5. Derivability / plausibility
🔴6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +5

This is a third-party reimplementation (authors' own analysis code is not public) of the paper's TCR.strong survival stratification on public Riaz/nivolumab data. The Fig 7G Ipi-naive OS direction reproduces (median 101.1 > 77.4 wk) but its significance does not (p=0.367 vs <0.05), and Fig 7F whole-cohort PFS fails entirely (p=0.859, direction reversed). The deviation sits mainly on the input/cohort-definition side (n=51 vs 50; 25/26 vs 21/29), with root cause on the authors' side — the reported significance is not derivable from shared public data because the exact subset/QC lives only in non-public code. Severity is high (significance + PFS-direction flip) but this is documented as a reproducibility gap, not fabrication; overall criticality lands yellow.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

202.1 k
tokens (I/O) · 12.2 M incl. cache
20 min
runtime · 0.01 CPU-h
1.3 GB
peak RAM
5 (2 failed)
HPC jobs
hummel
machine