Integrating multi-omics data reveals the antitumor role and clinical benefits of gamma-delta T cells in triple-negative breast cancer.
The main results reproduced, with only marginal, non-material deviations.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- Nothing in this column.
- 🟡Could not use the authors’ exact input data
- 🟡Reported values were only indirectly comparable
- 🟡A deviation arose in the data or preprocessing
- 🔴A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The deviation was non-trivial in magnitude
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
PARTIAL, 1:1-where-possible. Described well enough only at a high level: the paper ships NO analysis pipeline (its single code link, junjunlab/GseaVis, is a generic GSEA-plotting package) and never prints its central 20-gene gamma-delta signature, so exact reproduction of its numbers is impossible from what is published. Using the paper's named tools (limma + clusterProfiler GSEA + GseaVis) on the paper's data type (TCGA-BRCA, UCSC Xena), with a signature built from the genes the paper DOES name (TRDC/TRGC1/TRGC2 turned out absent from the gene-level matrix, leaving the cytotoxic genes GNLY/GZMB/NKG7/XCL1/XCL2 -> effectively a cytotoxic/NK signature): 10/11 checked claims reproduce in DIRECTION. Immune-checkpoint correlations are all strongly positive (LAG3 0.80 vs 0.84 and CTLA4 0.83 vs 0.76 within tolerance; CD274 0.58 vs 0.73 lower). 7/7 immune/inflammatory hallmark pathways are significantly enriched in the high-gd group (TNFA 2.23 vs 2.35 and JAK-STAT 2.38 vs 2.54 within tolerance; IFN-alpha/gamma, allograft, inflammatory, complement enriched but NES systematically ~0.5-1.0 lower, plausibly because the paper used the TNBC subset + GSVA + CIBERSORTx abundance). One genuine MISMATCH: OXPHOS, reported down-regulated (NES=-1.71), reproduces as non-significant/slightly positive (NES=+0.60, p=1.0). NOT attempted (out of scope / hard 20%): radiomics model (manual 3D-Slicer segmentation, R=0.94/AUC=0.633), IHC + 61 in-house specimens survival (wet-lab), exact CIBERSORTx gd abundance (unpublished Vd1/Vd2 reference + license), METABRIC treatment-response KM, full scRNA integration to 46,318 cells / Monocle / CellChat, maftools/GISTIC2 mutation enrichment.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 59assessed: 2026-06-14 ⛓ 52de848012c5
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-14
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15no human curator yet
- Last updated
- 2026-09-19
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetWhether gamma-delta (γδ) T cell infiltration plays an antitumor role and holds clinical/prognostic value in breast cancer, particularly triple-negative breast cancer (TNBC), given controversial and cancer-type-dependent evidence for γδT cell function.
- ★ High γδT cell infiltration is associated with favorable prognosis in TNBC but not in HR-positive or HER2-positive breast cancer finding
- ★ γδT cell infiltration is associated with somatic mutations in MUC16, DOCK11, LAMA2, and RYR1 finding
- ★ γδT cells activate immune-related pathways and suppress malignancy-associated pathways in the TNBC microenvironment mechanism
- ★ Intrinsic lineage evolution may drive the direct cytotoxic effects of γδT cells mechanism
- ★ γδT cells interact with antigen-presenting cells (cDCs and macrophages) via ligand-receptor pairs mechanism
- ★ Patients with high γδT abundance benefit more from chemotherapy or radiotherapy alone than from their combination finding
- ★ Patients with high γδT abundance are more likely to benefit from immunotherapy finding
- ★ A DCE-MRI-based radiomic model can estimate γδT cell abundance in TNBC patients resource
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| single-cell RNA-seq | primary TNBC tumor tissue (GSE176078, GSE180286) | none | cell lineage evolution, cell-cell interactions | — |
| single-cell RNA-seq | purified human γδ T cells from 3 donors (GSE128223) | none | Vδ1/Vδ2 subtype identification and marker expression (CD3D, TRDC, TRGC1, TRGC2) | — |
| bulk RNA-seq with digital cytometry (CIBERSORTx) | TCGA-BRCA tumor tissue | none | γδT cell abundance estimation | CIBERSORTx |
| bulk RNA-seq and clinical data analysis | METABRIC dataset | none | γδT abundance and differential gene expression/pathway enrichment | — |
| immunohistochemistry | 61 TNBC surgical specimens | none | γδ+ T cell counts per high-magnification field | TCR γ/δ antibody TCR1061, clone 5A6.E9, Thermo Fisher |
| somatic mutation (SNV) analysis | TCGA-BRCA TNBC tumor DNA | none | differential mutation frequency between high/low γδT groups | maftools |
| copy number variation analysis | TCGA-BRCA tumor DNA | none | CNV landscape differences between high/low γδT groups | GISTIC2 |
| radiomics (DCE-MRI feature extraction) | TCGA-BRCA TNBC patients (n=11) | none | MRI texture/shape features correlated with γδT cell abundance | 1.5T GE whole-body MRI system; 3D Slicer |
- ▲ High γδT cell abundance associated with favorable prognosis in TNBC by bioinformatic analysis; not significant in HR-positive (p=0.39) or HER2-positive (p=0.79) patients
- ▲ IHC confirmed high γδ+ T cell infiltration (n=31) vs low (n=30, median cutoff 26/HPF) associated with favorable prognosis in TNBC p=0.026
- – MUC16, DOCK11, LAMA2, and RYR1 mutations more frequent in high γδT group; DNAH7 more frequent in low γδT group p<0.05
- ▲ GSEA showed immune-activation pathways (defense response, innate immune response, adaptive immune response) upregulated in high γδT group
- ▲ Hallmark pathways including TNF-α, IFN-α, allograft rejection, JAK-STAT, IFN-γ, inflammatory response, and complement signaling upregulated in high γδT group NES 2.35–3.59, adjusted p<0.001
- ▼ Oxidative phosphorylation pathway downregulated in high γδT group NES=-1.71, adjusted p<0.001
- – Amplifications at 1q21.3/1q23.3 more frequent in high γδT group; deletions at 4q34.3/13q12.12 more frequent in low γδT group
- – Clinical characteristics (age, tumor stage, T-stage, race) showed no significant correlation with γδT cell abundance p=0.296, 0.215, 0.274, 0.766
- pvalue p=0.026 (IHC-based survival difference between high and low γδ+ T cell TNBC groups)
- pvalue p<0.05 (Mutation frequency of MUC16, DOCK11, LAMA2, RYR1 in high γδT infiltration group)
- fold_change NES=3.59, adjusted p<0.001 (IFN-γ pathway upregulation in high γδT group)
- fold_change NES=-1.71, adjusted p<0.001 (Oxidative phosphorylation downregulation in high γδT group)
- count 8,107 γδ T cells (Sorted single-cell γδ T cell profiling divided into Vδ1/Vδ2)
- count 46,318 cells from 12 primary TNBC patients (scRNA-seq of TNBC tumor tissue)
- pvalue p=0.39 (Prognostic association of γδT abundance in HR-positive breast cancer (not significant))
- pvalue p=0.79 (Prognostic association of γδT abundance in HER2-positive breast cancer (not significant))
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This multi-omics study integrated scRNA-seq trajectory inference, bulk RNA-seq differential expression and gene set enrichment analysis (GSEA), somatic mutation profiling, digital cytometry (CIBERSORTX), retrospective IHC in a single-institution cohort, and DCE-MRI radiomics. The primary prognostic question was addressed with log-rank tests on γδT cell abundance dichotomised at the median across TCGA-BRCA molecular subtypes and an IHC validation cohort (n = 61). Pathway enrichment, mutation differences, and immune correlates were assessed with limma, Fisher exact tests, and Pearson correlation; a radiomic prediction model was built with elastic net regression in n = 11 TNBC cases.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Log-rank test | Overall survival comparison between high and low γδT cell groups in TCGA-BRCA stratified by molecular subtype (TNBC, HR-positive, HER2-positive) and in IHC validation cohort | IHC cohort: n = 61 (31 high, 30 low); TCGA-BRCA TNBC subset n not explicitly stated | not stated |
| Fisher exact test | Identification of somatic gene mutations with significantly different frequencies between high and low γδT cell infiltration groups in TCGA-BRCA TNBC | High group n = 54; low group n not explicitly stated | not stated |
| limma linear model (differential expression) | Differential gene expression between high γδT (n = 60) and low γδT (n = 61) groups in TCGA-BRCA TNBC bulk RNA-seq | n = 60 (high) vs n = 61 (low) | not stated |
| Gene set enrichment analysis (GSEA) via clusterProfiler | Hallmark and Gene Ontology Biological Process pathway enrichment ranked by fold-change from limma differential expression, high vs low γδT cell groups | n = 60 vs n = 61 | not stated |
| Pearson correlation | Correlation of γδT signature score vs immune checkpoint inhibitor response scores (easier package); radiomic features vs γδT cell abundance (feature selection, p < 0.05); radiomic score vs γδT cell abundance (model evaluation) | null | not stated |
| Elastic net regression (α = 0.5, 10-fold cross-validation) via glmnet | Radiomic linear model to estimate γδT cell abundance from DCE-MRI features | n = 11 TCGA-BRCA TNBC cases with imaging data | not stated |
| ROC curve analysis | Discriminative performance of radiomic model for identifying patients with high vs low γδT cell infiltration | n = 11 | na |
-
γδT cell abundance was dichotomised at the median to define high and low groups for survival and downstream comparisons↳ Could also: Continuous Cox proportional-hazards regression with γδT abundance as a continuous predictor could also be used — Continuous modelling preserves information lost by dichotomisation, avoids sensitivity to the choice of cut-point, and yields a hazard ratio with a confidence interval as a direct effect-size estimate
-
Separate log-rank tests were performed within each molecular subtype and in the IHC cohort without a stated multiplicity correction across the family of tests↳ Could also: A single Cox model including subtype as a covariate and a γδT-by-subtype interaction term could also address the same question — An interaction model provides a formal statistical test of whether the γδT prognostic effect differs by subtype and produces a single inferential family, reducing the accumulation of type I error across multiple comparisons
-
Fisher exact tests were applied gene-by-gene to identify somatic mutations differing between high and low γδT groups across the full mutational landscape↳ Could also: FDR correction (e.g., Benjamini-Hochberg) across all gene-level Fisher tests could also be applied, as is standard in somatic mutation association studies — When hundreds of genes are tested simultaneously, FDR control limits the expected proportion of false discoveries, a routine practice in pan-cancer mutation analyses
-
Pearson correlation was used to filter radiomic features and assess association between radiomic scores and γδT abundance in a dataset of n = 11↳ Could also: Spearman rank correlation could also be used, particularly when variable distributions cannot be verified as normal in very small samples — Spearman correlation is more robust to non-normality and outliers, which are common in small radiomic and deconvolution-derived abundance datasets
-
An elastic net radiomic model was trained and internally validated with 10-fold cross-validation in n = 11 samples↳ Could also: Leave-one-out cross-validation (LOOCV) could also be used as an internal validation strategy when n is very small — With n = 11, each 10-fold held-out test set contains approximately 1 sample; LOOCV trains on n − 1 samples per iteration, maximising data use for fitting while still providing an approximately unbiased estimate of generalisation error
-
limma (originally designed for microarray log-intensities) was used for differential expression analysis of bulk RNA-seq data↳ Could also: DESeq2 or edgeR could also be applied for count-based RNA-seq differential expression — DESeq2 and edgeR model discrete read counts with a negative binomial distribution that is natural for RNA-seq; limma-voom extends limma to count data and is also widely accepted — all three approaches are commonly compared in RNA-seq benchmarking studies
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
Downstream reach in the literature
128 downstream papers · 3 datasetsHow widely the datasets deposited by this paper are reused across the whole literature (Europe PMC), beyond our assessed set. This is a factual dependency map — reusing a public dataset is normal, good science. It is not a judgement on the downstream papers; the only verdict here is this paper's own, with its cited rationale.
- Temporal profiling of the breast tumour microenviron... 2022 · 159 cites
- Cellular architecture of human brain metastases. 2022 · 151 cites
- Identification of the novel exhausted T cell CD8 + m... 2024 · 111 cites
- Molecular mechanisms and therapeutic significance of... 2024 · 91 cites
- Context-dependent activation of STING-interferon sig... 2023 · 76 cites
- BIDCell: Biologically-informed self-supervised learn... 2024 · 63 cites
- Single-cell RNA sequencing reveals cell heterogeneit... 2021 · 128 cites
- Cancer cell plasticity and MHC-II-mediated immune to... 2023 · 53 cites
- High B7-H3 expression with low PD-L1 expression iden... 2024 · 27 cites
- The tumor microenvironment and triple-negative breas... 2022 · 18 cites
- Augmentation of the RNA m6A reader signature is asso... 2022 · 16 cites
- Molecular, metabolic, and functional CD4 T cell para... 2023 · 14 cites
- Single-cell RNA sequencing unveils the shared and th... 2019 · 185 cites
- Multiomics profiling reveals the benefits of gamma-d... 2024 · 31 cites
- Costimulation of γδTCR and TLR7/8 promotes Vδ2 T-cel... 2021 · 25 cites
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-40197136
Paper: Wang G et al. (2025) Integrating multi-omics data reveals the antitumor role and clinical benefits of gamma-delta (γδ) T cells in triple-negative breast cancer (TNBC). BMC Cancer. DOI 10.1186/s12885-025-14029-8. PMCID PMC11974128.
"Code" link in record: https://github.com/junjunlab/GseaVis
(HEAD 89a55f97f3134117a9a31ce98840c485f21bc4c2, public, MIT-ish/NOASSERTION).
Important: GseaVis is a generic GSEA-plotting R package (third-party, by
junjunlab), NOT the authors' own analysis pipeline. The paper's actual analysis
code is not shared. Per brief P16, running a third-party tool on the paper's
own data is a valid reproduction; we therefore use GseaVis (the cited tool) for
the GSEA step, and standard, named tools (Seurat, GSVA, clusterProfiler, limma)
for the rest.
Data: GSE176078 (Wu et al. 2021 BRCA scRNA atlas; 10 TNBC samples used), GSE180286, GSE128223 (scRNA); TCGA-BRCA + METABRIC + TCIA (bulk/clinical/imaging).
Central reproducibility caveat (drives every grade)
Almost every numeric downstream claim (immune-checkpoint correlations, survival
split, GSEA NES, METABRIC treatment response, radiomic score) is computed from a
γδ T-cell signature / abundance that the paper derives as "the top 20 markers
(p<0.05) of γδT cells" — but this 20-gene list is never printed. Only the
identity markers (CD3D, TRDC, TRGC1, TRGC2) and trajectory cytotoxic genes (GNLY,
GZMB, NKG7, XCL1, XCL2) are named. So an exact 1:1 reproduction of the reported
numbers is not possible from the shipped text; we reconstruct a defensible γδ
signature from the named genes and report honestly (expect partial, not exact).
This unpublished-input dependency is itself the key auditable finding.
IN SCOPE — attempted (clearly-specified, self-contained pipeline outputs)
| id | result | reported | pipeline (named in paper) | data |
|---|---|---|---|---|
| C1 | γδ-signature ↔ immune-checkpoint correlation (Pearson R): CD274, CTLA4, LAG3 | R=0.73 / 0.76 / 0.84, all p<0.001 | GSVA score + correlation | TCGA-BRCA (UCSC Xena) |
| C2 | GSEA hallmark NES, high vs low γδ group (TNFα, IFNα, IFNγ, allograft rej., JAK-STAT, inflammatory, complement, OXPHOS) | NES 2.35/3.09/3.59/3.12/2.54/3.11/2.92 / −1.71 | limma DEG → clusterProfiler GSEA → GseaVis plot | TCGA-BRCA |
Both reuse one TCGA-BRCA download + one γδ signature, and C2 exercises the cited GseaVis package directly.
OUT OF SCOPE — not attempted (with reason)
- Radiomics model (R=0.94, AUC=0.633, two-term score): requires manual tumor
segmentation in 3D Slicer on TCIA DCE-MRI for 11 patients (manual/wet-lab-like),
not a reproducible automated pipeline. → out (
non_pipeline-like / manual). - IHC γδ infiltration & survival (p=0.026) on 61 in-house TNBC specimens: wet-lab + private specimens. → out.
- Exact CIBERSORTx γδ abundance: needs the authors' Vδ1/Vδ2 scRNA reference signature matrix + CIBERSORTx account/license; reference not shipped. We use a GSVA signature score as a transparent proxy (noted as a method substitution).
- METABRIC treatment-response KM (chemo p=0.056 / radio p=0.021 / combo p=0.15): depends on the unpublished γδ abundance split; skipped (hard 20%).
- Full scRNA integration to exactly 46,318 cells across 3 datasets + Monocle/ Vector trajectory + CellChat: heavy; the integration/QC choices that fix the exact cell count are under-specified. Optional stretch, not core.
- maftools/GISTIC2 mutation enrichment (MUC16/DOCK11/… ; GISTIC2 "online defaults"): depends on the γδ split + online tool; skipped.
Honest expectation
C1 likely reproduces the direction & rough magnitude (immune checkpoints
co-vary strongly with any T/cytotoxic signature) → partial/within-tol.
C2 likely reproduces the set of enriched immune pathways with NES>0 but not the
exact NES values (different DEG ranking from a different γδ split) → `pa
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
The reproduction confirms the paper's qualitative core — the γδ/cytotoxic signature is strongly positively correlated with immune checkpoints (CTLA4 0.83 vs 0.76, LAG3 0.80 vs 0.84) and 7/7 immune/inflammatory hallmarks are significantly enriched — but exact values are not derivable because the authors never publish their 20-gene signature, ship no pipeline (only the generic GseaVis plotter), and use an unpublished CIBERSORTx reference. The deviations (CD274 0.73→0.58, NES systematically ~0.5-1.0 lower, and one OXPHOS direction flip -1.71→+0.60 n.s.) are best explained by forced methodology substitutions (cytotoxic/NK z-score proxy, whole cohort vs TNBC) driven by authors-side under-specification, not by fabrication. Severity is moderate overall (direction holds for 10/11 claims) and the γδ-specific claim is simply untestable since TRDC/TRGC1/TRGC2 are absent from the Xena gene-level matrix.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at [email protected].
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.