Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Differential Infiltration of Key Immune T-Cell Populations Across Malignancies Varying by Immunogenic Potential and the Likelihood of Response to Immunotherapy.

Cells · 2024
L1 No data access 2/4
Why this verdict

Part of the results reproduced; minor but material deviations remained.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Input / endpoint not comparable 1:1
+1 pts
From: Q1 · Data identity 🔴
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +7
✓ What held up
  • Nothing in this column.
What did not (or only partly)
  • 🔴Could not use the authors’ exact input data
  • 🔴Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🔴A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
No data access Data access not granted

This paper has a computational component, but its primary data is legally or ethically access-restricted — identifiable patient cohorts, rare-disease genomes, or controlled-access biobanks that cannot be openly shared. The reproduction therefore could not be attempted. That is a neutral verdict: it does not mean the result is wrong or that the authors fell short — only that, for legitimate privacy reasons, it cannot be independently checked from public data. We deliberately do NOT assign a 0–100 score here, because a low number would wrongly read as a failed reproduction.

Reproduction agent’s raw note

DROP. The paper (Cells 2024, T-cell infiltration across malignancies) is genuinely a computational-pipeline study (bbduk->STAR->RSEM TPM->ComBat->T-cell signature z-scores->TIMEx deconvolution/GSEA->Mann-Whitney/KM/AUROC), but it cannot be reproduced 1:1 from public artifacts. (1) data_restricted [primary]: every specifically-quantified result -- Fig 1 cross-malignancy infiltration p-values, Fig 3 melanoma survival, Table 3 ORIEN AUROC, and the responder/non-responder analysis -- derives from the ORIEN/Avatar cohort (1892 patients), which the Data Availability statement releases only 'upon reasonable requests to the corresponding author' (controlled/on-request, not downloadable). (2) no_code: there is no Code Availability statement; the harvested code link github.com/alxdobin/STAR is HTTP 404 (the user does not exist) -- a text-mining typo of alexdobin/STAR, the generic STAR RNA-seq aligner (confirmed via GitHub API), not the paper's analysis code. (3) docs_insufficient: the five T-cell signature gene lists are not tabulated (only loose markers in sec 2.6; z-score cited as 'per Lee et al.'), so the method cannot be reconstructed unambiguously. (4) public path not pinnable: of 12 validation cohorts (672 pts) only GSE165278 (n=22) and GSE158403 (n=81) have in-paper accessions, and the only public metric is a POOLED average AUROC (0.605-0.638) with no per-dataset reported value to compare against. P16 (third-party tool on the paper's own data) does not rescue it because the paper's own data is the restricted ORIEN cohort. NOT ATTEMPTED on «our HPC» by design -- this resolves at the control-plane screening stage, conserving compute; no result was fabricated to avoid the drop. No author fabrication detected; the bad code link is a harvesting artifact.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment
    assessed: 2026-06-15 ⛓ abfc1017a238
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-15
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

This study investigates whether key T-cell populations (stem-like TILs, tissue-resident memory T-cells, early and late dysfunctional T-cells, activated-potentially anti-tumor T-cells, and BTN3A isoforms) are differentially infiltrated across solid tumors that vary in immunogenic potential and likelihood of response to immunotherapy, and whether their estimated infiltration associates with survival following immune checkpoint inhibitors.

Core claims
  • Immune-activation-related T-cell populations (APA, TRM, stem-like, early/late dysfunctional T-cells) are more heavily infiltrated in ICI-responsive malignancies (melanoma, bladder) than in poorly responsive ones (ovarian, pancreatic). finding
  • In melanoma treated with ICIs, higher densities of stem-like TILs, TRM, early and late dysfunctional T-cells, APA T-cells, and BTN3A isoforms are associated with improved survival. finding
  • Ovarian cancer shows the lowest infiltration across nearly all T-cell populations, while melanoma shows the highest for most populations (except TRM, highest in bladder). finding
  • The TRM gene signature is a moderate predictor of immunotherapy survival/response in melanoma, validated in independent public ICI datasets. finding
  • Melanoma and ovarian cancers express BTN3A isoforms more than other malignancies. finding
  • T-cell infiltration estimated from mRNA co-expression of marker gene signatures (z-score method of Lee et al.) using bulk RNA-seq from the ORIEN/Avatar cohort. method
  • Only TRM T-cells differed significantly between melanoma immunotherapy responders and non-responders. finding
Experimental setups
Assay System Perturbation Readout Platform
bulk RNA-seq (gene expression signature scoring of T-cell marker co-expression) human melanoma, bladder urothelial, ovarian, and pancreatic carcinoma patient tumors (ORIEN/Avatar TCC cohort, N=1892) none (observational/real-world) log2(TPM+1) expression, z-score signature scores estimating T-cell infiltration STAR v2.7.3a alignment to GRCh38/hg38; RSEM v1.3.1 TPM quantification; Bbduk v38.96; RNA-SeQC v2.3.2; sva/ComBat v3.34.0
survival analysis (Kaplan–Meier / log-rank) melanoma patients treated with immune checkpoint inhibitors (n=123) ICI treatment overall survival stratified by low vs high T-cell signature expression
AUROC predictive validation melanoma patients (ORIEN cohort plus 10 public ICI-treated datasets, 672 patients) ICI treatment prediction of immunotherapy response (responder OS≥24mo vs non-responder OS<24mo) from signature z-scores SciPy 1.7.0
Gene Set Enrichment Analysis (GSEA) / bulk transcriptomic deconvolution melanoma responder vs non-responder tumor samples none normalized enrichment score (NES) of hallmark/KEGG/REACTOME gene sets; cell-type deconvolution GSEA V.20.3.4, MSigDB V.7.5.1, TIMEx web portal
Key results
  • APA T-cell infiltration in melanoma differed significantly vs ovarian/pancreatic cancers p=4.67×10^-12 (melanoma vs ovarian)
  • APA T-cell infiltration in bladder differed significantly vs ovarian/pancreatic p=5.80×10^-12
  • Ovarian cancer had lower TRM T-cell infiltration than bladder, pancreatic, and melanoma p=3.862×10^-28 (vs bladder), 7.852×10^-9 (vs pancreatic), 2.232×10^-8 (vs melanoma)
  • Ovarian cancer had lower stem-like TIL expression than bladder or melanoma p=2×10^-8 (vs bladder), 6.35×10^-8 (vs melanoma)
  • Higher T-cell signature densities associated with improved melanoma survival by log-rank (stem-like, TRM, early dys, late dys, APA, BTN3A) p=0.0075, 0.00059, 0.013, 0.005, 0.0016, 0.041 respectively
  • TRM gene signature predicted survival in melanoma cohort and independent ICI datasets AUROC=0.65 (melanoma cohort); 0.61–0.64 (public datasets)
  • TRM T-cells were the only signature significantly different between melanoma ICI responders and non-responders adjusted p=0.02
  • BTN3A isoform expression significantly higher in melanoma vs bladder, pancreatic, ovarian p=9.152×10^-8, 4.682×10^-9, 1.962×10^-5 respectively
Key statistics
  • pvalue 4.67×10^-12 (APA T-cell infiltration melanoma vs ovarian)
  • pvalue 3.862×10^-28 (TRM infiltration ovarian vs bladder)
  • pvalue 0.00059 (TRM high vs low expression survival in melanoma (log-rank))
  • other AUROC=0.65 (TRM signature predicting survival in melanoma cohort)
  • other AUROC 0.61–0.64 (TRM signature in public ICI-treated melanoma datasets)
  • count 1892 patients (melanoma 232, ovarian 664, pancreatic 647, bladder 349) (total study cohort by cancer type)
  • count n=123 (melanoma patients treated with immunotherapy analyzed for responder vs non-responder signatures)
  • mean 62 ± 13 years (mean ± SD age of cohort)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This retrospective observational study analyzed mRNA co-expression-derived T-cell signature z-scores across four cancer types (melanoma, bladder, ovarian, pancreatic; N = 1892) using non-parametric pairwise comparisons, and evaluated associations between signature levels and immunotherapy outcomes in a melanoma subset (n = 123). Cross-cancer group differences were assessed with Mann-Whitney U tests, survival associations used Kaplan-Meier estimation with log-rank tests, and predictive performance was summarized as AUROC validated in both an internal cohort and 10 external public datasets (n = 672). FDR correction was applied to the responder-versus-non-responder comparisons, with two-tailed p < 0.05 used for cross-cancer comparisons.

Replicationbiological Sample sizeTotal N per cancer type stated (melanoma 232, bladder 349, ovarian 664, pancreatic 647); no formal power calculation or sample-size justification reported GroupsFour cancer types (melanoma, bladder, ovarian, pancreatic); ICI responders vs. non-responders within melanoma Pairingunpaired Randomization/blindingnot stated Dispersionmixed Exact p-valuesyes Effect sizesno Confidence intervalsno Multiplicity correctionFDR (specific algorithm, e.g. Benjamini-Hochberg, not named)
Statistical tests used
Test Applied to n Assumptions
Mann-Whitney U test (two-tailed) Pairwise comparisons of median T-cell signature z-scores among the four cancer types (Figure 1, all six signatures) 1892 total (melanoma n=232, bladder n=349, ovarian n=664, pancreatic n=647) not stated
Mann-Whitney U test with FDR correction Comparison of T-cell signature z-scores between ICI responders (OS ≥ 24 months) and non-responders (OS < 24 months) in melanoma (Figure 2) n=123 melanoma patients treated with immunotherapy not stated
Log-rank test Survival differences between high vs. low T-cell signature expression groups in melanoma (Figure 3) n=123 (implied from melanoma ICI cohort) not stated
Kaplan-Meier estimation Survival probability curves for high vs. low expression of each T-cell signature in melanoma (Figure 3) n=123 (implied from melanoma ICI cohort) na
Area under the receiver operating characteristic curve (AUROC) Predictive value of T-cell signature z-scores for ICI response in ORIEN melanoma cohort and 10 public datasets 672 patients from 10 public datasets; ORIEN melanoma subset n not separately stated for this analysis not stated
Gene Set Enrichment Analysis (GSEA) using normalized enrichment score (NES) Hallmark, KEGG, and REACTOME pathway enrichment comparing ICI responders vs. non-responders null not stated
Approaches that could also have been used
  • Pairwise Mann-Whitney U tests were conducted directly between cancer-type pairs without a preceding omnibus test
    Could also: A Kruskal-Wallis omnibus test followed by Dunn's post-hoc test with multiplicity correction (e.g., Bonferroni or Benjamini-Hochberg) could also be applied when comparing k ≥ 3 independent groups on the same outcome — An omnibus-first approach provides a single decision rule for whether any group differences exist before pairwise follow-up, which is a common convention that explicitly structures the family of comparisons
  • Immunotherapy response was dichotomized as OS ≥ 24 months vs. OS < 24 months for the responder/non-responder comparisons
    Could also: Cox proportional hazards regression using time-to-event OS as a continuous outcome with each signature score as a continuous or ranked covariate could also be used — Treating survival as a continuous time-to-event outcome retains all temporal information and avoids information loss that can arise from dichotomization at a fixed threshold
  • T-cell infiltration density was estimated using the average z-score across marker genes for each defined signature
    Could also: Single-sample GSEA (ssGSEA) or GSVA could also generate per-sample enrichment scores for each gene signature from the same bulk RNA-seq data — ssGSEA and GSVA account for the rank ordering of all expressed genes within each sample and can reduce sensitivity to individual outlier genes relative to simple mean z-scoring
  • AUROC values were reported as point estimates without accompanying confidence intervals
    Could also: 95% confidence intervals for each AUROC could also be reported using DeLong's method or bootstrapping — Confidence intervals for AUROC communicate estimation precision and are especially informative when cohort sizes differ across validation datasets, allowing readers to gauge uncertainty in discriminative performance
  • Survival differences between high- and low-expression groups were assessed with unadjusted log-rank tests
    Could also: Multivariable Cox proportional hazards regression adjusting for available clinical covariates (e.g., age, disease stage, ECOG performance status) could also be used — Multivariable modeling allows the independent prognostic contribution of each signature to be estimated after accounting for established clinical prognostic factors, which vary in distribution across the cohort
  • The FDR correction method was described as controlling the false discovery rate but the specific algorithm was not named
    Could also: Explicitly naming the FDR algorithm (e.g., Benjamini-Hochberg) and stating the full family of tests to which it was applied could also be reported — Naming the algorithm and the test family allows independent verification of adjusted p-values and aids reproducibility, since different FDR procedures (BH, BY, q-value) can yield different results under dependence structures common in transcriptomic data
Software: SciPy 1.7.0 · GSEA 20.3.4 · RSEM 1.3.1 · STAR 2.7.3a · sva (ComBat) 3.34.0 · TIMEx (web portal) · BBduk 38.96 · RNA-SeQC 2.3.2

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
3
Impact: low
Foundation confidence
Built on 1 assessed reference(s) · mean reproducibility 100/100
stands on reproducible work
Topics

Assessed papers, coloured by verdict. Click a node to open it.

Built on (assessed references) (1)
Cited by (assessed papers) (0)
  • No assessed neighbours yet — the network grows as more papers are assessed.

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-39682743

Paper: Differential Infiltration of Key Immune T-Cell Populations Across Malignancies Varying by Immunogenic Potential and the Likelihood of Response to Immunotherapy. Cells 2024;13(23):1993. PMID 39682743 · PMC11640164 · DOI 10.3390/cells13231993.

What the paper does (pipeline-derived)

RNA-seq → adapter trim (bbduk 38.96) → align (STAR 2.7.3a, hg38) → QC (RNA-SeQC 2.3.2) → quantify TPM (RSEM 1.3.1) → log2(TPM+1) → batch-correct (ComBat / sva 3.34.0) → five T-cell-population signature z-scores (stem-like TILs, TRM, APA, early-dysfunctional, late-dysfunctional; z-score per "Lee et al." method) → immune deconvolution (TIMEx portal) + GSEA (v20.3.4, MSigDB hallmark v7.5.1) → stats (Mann–Whitney U, Kaplan–Meier/log-rank, AUROC; SciPy 1.7.0; FDR<0.05/0.02).

Datasets

  • Primary — ORIEN / Avatar (Total Cancer Care): 1892 patients (melanoma 232, bladder 349, ovarian 664, pancreatic 647) across 18 centers. Drives Figure 1 (cross-malignancy p-values), Figure 3 (melanoma survival), Table 3 ORIEN AUROC, responder/non-responder analysis. Availability: "Data are available upon reasonable requests to the corresponding author." → CONTROLLED-ACCESS / on-request.
  • Validation — 10 public cohorts, 672 patients (Table 1): Du(50), Gide_Pre_PD-1+CTLA4(41), Gide_Pre_PD-1(50), GSE165278(22), GSE158403(81), Freeman(38), Hugo(26), Lauss(25), Lee(78), Liu(122), Riaz(98), VanAllen(41). Only 2 carry an explicit GSE accession in-paper (GSE165278, GSE158403).

In scope vs out of scope

Result Pipeline In scope? Reason
Fig 1 cross-malignancy infiltration p-values signature z-score + Mann–Whitney OUT derived from restricted ORIEN data
Fig 3 melanoma survival (KM/log-rank) signature z-score + KM OUT restricted ORIEN data
Table 3 ORIEN AUROC (5 signatures) z-score + AUROC OUT restricted ORIEN data
Table 3 PUBLIC average AUROC (0.605–0.638) z-score + AUROC on 12 cohorts OUT pooled-only metric; most cohorts lack in-paper accessions; signature gene lists not provided; no per-dataset expected value to compare 1:1
RNA-seq alignment/quantification (STAR/RSEM) STAR→RSEM OUT generic upstream step; no paper-specific expected output value; raw data is the restricted ORIEN cohort

Blockers (why this is a drop)

  1. no_code — No Code Availability statement. Harvested code_url github.com/alxdobin/STAR is a 404 (user/repo does not exist); it is a typo of alexdobin/STAR, the generic STAR RNA-seq aligner — not this paper's analysis pipeline. No authors' analysis code exists to run, and no third-party tool can substitute because (see 2–4) the inputs/expected outputs are not pinnable.
  2. data_restricted — the central quantitative results all derive from the ORIEN/Avatar cohort, which is on-request (not downloadable).
  3. docs_insufficient — the five T-cell signature gene lists are not tabulated (only loose marker genes in §2.6); z-scoring only cited as "per Lee et al." The signature definitions cannot be reconstructed unambiguously.
  4. no_expected_result (public path) — for the only fully-public, accession-backed cohort (GSE165278, n=22) the paper reports no dataset-specific value; the public metric is a pooled average AUROC across 12 cohorts.

Decision

DROP. Primary drop_reason = data_restricted (paper's own data is on-request), compounded by no_code + docs_insufficient + no pinnable public expected result. No «our HPC» compute spent — resolved at the control-plane screening stage. Per the brief, drops are valid and a result must not be fabricated to avoid one.

Figures / tables: Figure 1Figure 3Table

No individual results have been recorded for this entry yet.

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 31/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🔴1. Data identity
🔴2. Endpoint comparability
🟡3. Location of the main deviation
🔴4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Input / endpoint not comparable 1:1
+1 pts
From: Q1 · Data identity 🔴
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +7

This is a clean DROP (data_restricted): every reported number (Fig 1 cross-malignancy p-values, Fig 3 melanoma survival, Table 3 ORIEN AUROC, the responder analysis) derives from the controlled-access ORIEN/Avatar cohort, so no 1:1 input exists. The drop is compounded by no code (the only link is a 404 typo of the generic alexdobin/STAR aligner) and under-specified methods (signature gene lists not tabulated; pooled-only public AUROC with no per-dataset target). The limitation sits on the data-availability/authors' side, not in our methodology, and no fabrication is indicated — nothing was reproduced, so q5/q7 are scored as unestablished (yellow) rather than red per the restricted-data principle. Overall a well-justified, non-critical drop.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

74.7 k
tokens (I/O) · 3.1 M incl. cache
7 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.