Differential Infiltration of Key Immune T-Cell Populations Across Malignancies Varying by Immunogenic Potential and the Likelihood of Response to Immunotherapy.
Part of the results reproduced; minor but material deviations remained.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- Nothing in this column.
- 🔴Could not use the authors’ exact input data
- 🔴Reported values were only indirectly comparable
- 🟡A deviation arose in the data or preprocessing
- 🔴A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The deviation was non-trivial in magnitude
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
This paper has a computational component, but its primary data is legally or ethically access-restricted — identifiable patient cohorts, rare-disease genomes, or controlled-access biobanks that cannot be openly shared. The reproduction therefore could not be attempted. That is a neutral verdict: it does not mean the result is wrong or that the authors fell short — only that, for legitimate privacy reasons, it cannot be independently checked from public data. We deliberately do NOT assign a 0–100 score here, because a low number would wrongly read as a failed reproduction.
▸Reproduction agent’s raw note
DROP. The paper (Cells 2024, T-cell infiltration across malignancies) is genuinely a computational-pipeline study (bbduk->STAR->RSEM TPM->ComBat->T-cell signature z-scores->TIMEx deconvolution/GSEA->Mann-Whitney/KM/AUROC), but it cannot be reproduced 1:1 from public artifacts. (1) data_restricted [primary]: every specifically-quantified result -- Fig 1 cross-malignancy infiltration p-values, Fig 3 melanoma survival, Table 3 ORIEN AUROC, and the responder/non-responder analysis -- derives from the ORIEN/Avatar cohort (1892 patients), which the Data Availability statement releases only 'upon reasonable requests to the corresponding author' (controlled/on-request, not downloadable). (2) no_code: there is no Code Availability statement; the harvested code link github.com/alxdobin/STAR is HTTP 404 (the user does not exist) -- a text-mining typo of alexdobin/STAR, the generic STAR RNA-seq aligner (confirmed via GitHub API), not the paper's analysis code. (3) docs_insufficient: the five T-cell signature gene lists are not tabulated (only loose markers in sec 2.6; z-score cited as 'per Lee et al.'), so the method cannot be reconstructed unambiguously. (4) public path not pinnable: of 12 validation cohorts (672 pts) only GSE165278 (n=22) and GSE158403 (n=81) have in-paper accessions, and the only public metric is a POOLED average AUROC (0.605-0.638) with no per-dataset reported value to compare against. P16 (third-party tool on the paper's own data) does not rescue it because the paper's own data is the restricted ORIEN cohort. NOT ATTEMPTED on «our HPC» by design -- this resolves at the control-plane screening stage, conserving compute; no result was fabricated to avoid the drop. No author fabrication detected; the bad code link is a harvesting artifact.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessmentassessed: 2026-06-15 ⛓ abfc1017a238
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-15
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15no human curator yet
- Last updated
- 2026-09-19
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetThe study investigated whether candidate T-cell populations (stem-like TILs, TRM, early/late dysfunctional T-cells, APA T-cells, and BTN3A isoforms), estimated from mRNA co-expression of their cellular markers, differ in infiltration across tumor types varying by immunogenic potential and likelihood of immunotherapy response, and whether their expression associates with survival after immune checkpoint inhibitors.
- ★ Estimated T-cell infiltration differs significantly across melanoma, bladder, ovarian, and pancreatic cancers, tracking known immunogenic potential. finding
- ★ Ovarian cancer shows the lowest median infiltration across most T-cell populations compared to the other three cancers. finding
- ★ APA T-cell infiltration is significantly higher in melanoma and bladder cancer than in ovarian and pancreatic cancer. finding
- ★ TRM T-cell infiltration is lowest in ovarian cancer and highest in bladder cancer among the four malignancies. finding
- ★ Melanoma and ovarian cancer express BTN3A isoforms more than bladder and pancreatic cancer. finding
- ★ Higher densities of stem-like TILs, TRM, early and late dysfunctional T-cells, APA T-cells, and BTN3A isoforms are associated with increased survival in melanoma patients treated with immunotherapy. finding
- ★ The TRM gene signature is a moderate predictor of immunotherapy response/survival in melanoma (AUROC 0.65), with similar predictive performance in independent public ICI-treated melanoma datasets (AUROC 0.61-0.64). finding
- ★ Among the six T-cell signatures tested, only TRM T-cells show a statistically significant difference in expression between immunotherapy responders and non-responders in melanoma. finding
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| bulk RNA sequencing (RNA-seq) | tumor samples from patients with melanoma, bladder urothelial carcinoma, ovarian cancer, and pancreatic adenocarcinoma (ORIEN/Avatar cohort, N=1892) | none (real-world observational cohort) | mRNA co-expression-derived gene signature z-scores for stem-like TILs, TRM, early/late dysfunctional T-cells, APA T-cells, and BTN3A isoforms | Bbduk, STAR aligner (hg38), RNA-SeQC, RSEM (TPM, log2 transformed), ComBat batch correction |
| Gene Set Enrichment Analysis (GSEA) | melanoma tumor transcriptomes, immunotherapy responders vs non-responders | none (comparison by clinical response status) | normalized enrichment score (NES) for MSigDB hallmark, KEGG, and REACTOME gene sets | GSEA software v20.3.4, MSigDB v7.5.1 |
| bulk transcriptome deconvolution | melanoma tumor samples, immunotherapy responders vs non-responders | none | estimated cell type composition/proportions | TIMEx web portal |
| Kaplan-Meier survival analysis / log-rank test | melanoma patients treated with immunotherapy | none (stratified by high vs low gene signature expression) | overall survival probability | SciPy 1.7.0 |
| AUROC biomarker validation | ORIEN melanoma cohort and 10 combined public ICI-treated melanoma datasets (n=672) | none | predictive power of each T-cell signature z-score for immunotherapy response | — |
- ▲ APA T-cell infiltration significantly higher in melanoma/bladder vs ovarian/pancreatic cancer p=4.67e-12 (melanoma) and p=5.80e-12 (bladder) vs ovarian/pancreatic
- ▼ TRM T-cell expression significantly lower in ovarian vs pancreatic, melanoma, and bladder cancer p=7.852e-9, p=2.232e-8, p=3.862e-28 respectively
- ▼ Stem-like TIL expression significantly lower in ovarian vs bladder and melanoma p=2e-8 and p=6.35e-8
- ▼ Late dysfunctional T-cell expression significantly lower in ovarian vs bladder, pancreatic, and melanoma; also differs between bladder and melanoma p=4.722e-7, p=5.452e-11, p=2.112e-14; bladder vs melanoma p=0.000507
- ▲ BTN3A isoform expression significantly higher in melanoma vs bladder, pancreatic, and ovarian cancer p=9.152e-8, p=4.682e-9, p=1.962e-5 respectively
- ▲ High expression of stem-like TILs, TRM, early/late dysfunctional, APA T-cells, and BTN3A isoforms associated with improved melanoma survival by log-rank test p=0.0075, 0.00059, 0.013, 0.005, 0.0016, 0.041 respectively
- – TRM gene signature predicts survival/response in melanoma with moderate accuracy, confirmed in independent public datasets AUROC=0.65 (ORIEN cohort); AUROC 0.61-0.64 (public datasets)
- – Only TRM T-cell expression significantly differs between immunotherapy responders and non-responders in melanoma adjusted p-value = 0.02
- pvalue p=4.67e-12 (APA T-cell infiltration, melanoma vs ovarian/pancreatic)
- pvalue p=5.80e-12 (APA T-cell infiltration, bladder vs ovarian/pancreatic)
- pvalue p=2.23e-8, 3.86e-28, 7.85e-9 (TRM T-cell infiltration, ovarian vs melanoma, bladder, pancreatic respectively)
- other AUROC=0.65 (discovery); AUROC 0.61-0.64 (validation) (TRM gene signature predicting survival/response in melanoma, validated across 10 public ICI-treated datasets)
- pvalue p=0.0075, 0.00059, 0.013, 0.005, 0.0016, 0.041 (log-rank test, survival by high vs low expression of six T-cell signatures in melanoma)
- count N=1892 (melanoma 232, bladder 349, ovarian 664, pancreatic 647) (total patient cohort analyzed for T-cell infiltration)
- count n=672 across 10 public datasets (external validation cohort for AUROC biomarker analysis)
- pvalue adjusted p=0.02 (TRM T-cell z-score, immunotherapy responders vs non-responders in melanoma)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This retrospective observational study analyzed mRNA co-expression-derived T-cell signature z-scores across four cancer types (melanoma, bladder, ovarian, pancreatic; N = 1892) using non-parametric pairwise comparisons, and evaluated associations between signature levels and immunotherapy outcomes in a melanoma subset (n = 123). Cross-cancer group differences were assessed with Mann-Whitney U tests, survival associations used Kaplan-Meier estimation with log-rank tests, and predictive performance was summarized as AUROC validated in both an internal cohort and 10 external public datasets (n = 672). FDR correction was applied to the responder-versus-non-responder comparisons, with two-tailed p < 0.05 used for cross-cancer comparisons.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Mann-Whitney U test (two-tailed) | Pairwise comparisons of median T-cell signature z-scores among the four cancer types (Figure 1, all six signatures) | 1892 total (melanoma n=232, bladder n=349, ovarian n=664, pancreatic n=647) | not stated |
| Mann-Whitney U test with FDR correction | Comparison of T-cell signature z-scores between ICI responders (OS ≥ 24 months) and non-responders (OS < 24 months) in melanoma (Figure 2) | n=123 melanoma patients treated with immunotherapy | not stated |
| Log-rank test | Survival differences between high vs. low T-cell signature expression groups in melanoma (Figure 3) | n=123 (implied from melanoma ICI cohort) | not stated |
| Kaplan-Meier estimation | Survival probability curves for high vs. low expression of each T-cell signature in melanoma (Figure 3) | n=123 (implied from melanoma ICI cohort) | na |
| Area under the receiver operating characteristic curve (AUROC) | Predictive value of T-cell signature z-scores for ICI response in ORIEN melanoma cohort and 10 public datasets | 672 patients from 10 public datasets; ORIEN melanoma subset n not separately stated for this analysis | not stated |
| Gene Set Enrichment Analysis (GSEA) using normalized enrichment score (NES) | Hallmark, KEGG, and REACTOME pathway enrichment comparing ICI responders vs. non-responders | null | not stated |
-
Pairwise Mann-Whitney U tests were conducted directly between cancer-type pairs without a preceding omnibus test↳ Could also: A Kruskal-Wallis omnibus test followed by Dunn's post-hoc test with multiplicity correction (e.g., Bonferroni or Benjamini-Hochberg) could also be applied when comparing k ≥ 3 independent groups on the same outcome — An omnibus-first approach provides a single decision rule for whether any group differences exist before pairwise follow-up, which is a common convention that explicitly structures the family of comparisons
-
Immunotherapy response was dichotomized as OS ≥ 24 months vs. OS < 24 months for the responder/non-responder comparisons↳ Could also: Cox proportional hazards regression using time-to-event OS as a continuous outcome with each signature score as a continuous or ranked covariate could also be used — Treating survival as a continuous time-to-event outcome retains all temporal information and avoids information loss that can arise from dichotomization at a fixed threshold
-
T-cell infiltration density was estimated using the average z-score across marker genes for each defined signature↳ Could also: Single-sample GSEA (ssGSEA) or GSVA could also generate per-sample enrichment scores for each gene signature from the same bulk RNA-seq data — ssGSEA and GSVA account for the rank ordering of all expressed genes within each sample and can reduce sensitivity to individual outlier genes relative to simple mean z-scoring
-
AUROC values were reported as point estimates without accompanying confidence intervals↳ Could also: 95% confidence intervals for each AUROC could also be reported using DeLong's method or bootstrapping — Confidence intervals for AUROC communicate estimation precision and are especially informative when cohort sizes differ across validation datasets, allowing readers to gauge uncertainty in discriminative performance
-
Survival differences between high- and low-expression groups were assessed with unadjusted log-rank tests↳ Could also: Multivariable Cox proportional hazards regression adjusting for available clinical covariates (e.g., age, disease stage, ECOG performance status) could also be used — Multivariable modeling allows the independent prognostic contribution of each signature to be estimated after accounting for established clinical prognostic factors, which vary in distribution across the cohort
-
The FDR correction method was described as controlling the false discovery rate but the specific algorithm was not named↳ Could also: Explicitly naming the FDR algorithm (e.g., Benjamini-Hochberg) and stating the full family of tests to which it was applied could also be reported — Naming the algorithm and the test family allows independent verification of adjusted p-values and aids reproducibility, since different FDR procedures (BH, BY, q-value) can yield different results under dependence structures common in transcriptomic data
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
Assessed papers, coloured by verdict. Click a node to open it.
- No assessed neighbours yet — the network grows as more papers are assessed.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-39682743
Paper: Differential Infiltration of Key Immune T-Cell Populations Across Malignancies Varying by Immunogenic Potential and the Likelihood of Response to Immunotherapy. Cells 2024;13(23):1993. PMID 39682743 · PMC11640164 · DOI 10.3390/cells13231993.
What the paper does (pipeline-derived)
RNA-seq → adapter trim (bbduk 38.96) → align (STAR 2.7.3a, hg38) → QC (RNA-SeQC 2.3.2) → quantify TPM (RSEM 1.3.1) → log2(TPM+1) → batch-correct (ComBat / sva 3.34.0) → five T-cell-population signature z-scores (stem-like TILs, TRM, APA, early-dysfunctional, late-dysfunctional; z-score per "Lee et al." method) → immune deconvolution (TIMEx portal) + GSEA (v20.3.4, MSigDB hallmark v7.5.1) → stats (Mann–Whitney U, Kaplan–Meier/log-rank, AUROC; SciPy 1.7.0; FDR<0.05/0.02).
Datasets
- Primary — ORIEN / Avatar (Total Cancer Care): 1892 patients (melanoma 232, bladder 349, ovarian 664, pancreatic 647) across 18 centers. Drives Figure 1 (cross-malignancy p-values), Figure 3 (melanoma survival), Table 3 ORIEN AUROC, responder/non-responder analysis. Availability: "Data are available upon reasonable requests to the corresponding author." → CONTROLLED-ACCESS / on-request.
- Validation — 10 public cohorts, 672 patients (Table 1): Du(50), Gide_Pre_PD-1+CTLA4(41), Gide_Pre_PD-1(50), GSE165278(22), GSE158403(81), Freeman(38), Hugo(26), Lauss(25), Lee(78), Liu(122), Riaz(98), VanAllen(41). Only 2 carry an explicit GSE accession in-paper (GSE165278, GSE158403).
In scope vs out of scope
| Result | Pipeline | In scope? | Reason |
|---|---|---|---|
| Fig 1 cross-malignancy infiltration p-values | signature z-score + Mann–Whitney | OUT | derived from restricted ORIEN data |
| Fig 3 melanoma survival (KM/log-rank) | signature z-score + KM | OUT | restricted ORIEN data |
| Table 3 ORIEN AUROC (5 signatures) | z-score + AUROC | OUT | restricted ORIEN data |
| Table 3 PUBLIC average AUROC (0.605–0.638) | z-score + AUROC on 12 cohorts | OUT | pooled-only metric; most cohorts lack in-paper accessions; signature gene lists not provided; no per-dataset expected value to compare 1:1 |
| RNA-seq alignment/quantification (STAR/RSEM) | STAR→RSEM | OUT | generic upstream step; no paper-specific expected output value; raw data is the restricted ORIEN cohort |
Blockers (why this is a drop)
- no_code — No Code Availability statement. Harvested
code_urlgithub.com/alxdobin/STARis a 404 (user/repo does not exist); it is a typo ofalexdobin/STAR, the generic STAR RNA-seq aligner — not this paper's analysis pipeline. No authors' analysis code exists to run, and no third-party tool can substitute because (see 2–4) the inputs/expected outputs are not pinnable. - data_restricted — the central quantitative results all derive from the ORIEN/Avatar cohort, which is on-request (not downloadable).
- docs_insufficient — the five T-cell signature gene lists are not tabulated (only loose marker genes in §2.6); z-scoring only cited as "per Lee et al." The signature definitions cannot be reconstructed unambiguously.
- no_expected_result (public path) — for the only fully-public, accession-backed cohort (GSE165278, n=22) the paper reports no dataset-specific value; the public metric is a pooled average AUROC across 12 cohorts.
Decision
DROP. Primary drop_reason = data_restricted (paper's own data is on-request),
compounded by no_code + docs_insufficient + no pinnable public expected
result. No «our HPC» compute spent — resolved at the control-plane screening stage.
Per the brief, drops are valid and a result must not be fabricated to avoid one.
No individual results have been recorded for this entry yet.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
This is a clean DROP (data_restricted): every reported number (Fig 1 cross-malignancy p-values, Fig 3 melanoma survival, Table 3 ORIEN AUROC, the responder analysis) derives from the controlled-access ORIEN/Avatar cohort, so no 1:1 input exists. The drop is compounded by no code (the only link is a 404 typo of the generic alexdobin/STAR aligner) and under-specified methods (signature gene lists not tabulated; pooled-only public AUROC with no per-dataset target). The limitation sits on the data-availability/authors' side, not in our methodology, and no fabrication is indicated — nothing was reproduced, so q5/q7 are scored as unestablished (yellow) rather than red per the restricted-data principle. Overall a well-justified, non-critical drop.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at [email protected].
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.