Analysis of Tumor-Infiltrating T-Cell Transcriptomes Reveal a Unique Genetic Signature across Different Types of Cancer.
The main results reproduced: recomputed values matched the published ones within tolerance.
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Described well enough for the COUNT/SET-OPERATION claims, NOT for the upstream pipeline. No analysis code is shipped (github.com/tumourTcells/Cytoscape-file is a README pointing to a 308MB Google-Drive gzip of the final ClueGO Cytoscape session files), so align->TPM->cluster->subset gene-list generation cannot be re-run and is out of scope (also: contradictory marker thresholds in Methods). What IS reproducible is whether the paper's reported numbers are recomputable from the deposited supplement (MDPI S1, retrieved via EuropePMC) + the ClueGO networks, and they overwhelmingly are, EXACTLY: C1 exclusive genes 652/69/557; C3 Table 2 GO/Reactome (all 12 values 1335/950/1388, 270/159/249, 376/271/385, 2059/1746/1966); C5 network clusters 93/92/43/123/83/81 (Table S9 rows == distinct ClueGO GOGroups, also matching the .cluego sessions); C6 validation pathways 218/194; C7d proteomics common pathways 561; and the Figure-4A 'exclusive GO terms' 38/33/72/80/100/38 (Table S10). ONE flagged discrepancy: the Results-text 'exclusive GO terms' values 17/17/12/86/54/25 are NOT supported by any deposited table and contradict the paper's OWN Figure 4A/Table S10 (38/33/72/80/100/38) -- no p-value threshold reproduces them. This is an internal inconsistency / possible error in the publication; flagged for human review, not asserted as fabrication. C2 (Table 1 totals) and C7a-c (proteomics protein counts) are uncheckable because the upstream intermediates were not deposited. Overall: a faithful 1:1 reproduction of every deposited-derivable reported number except the single self-contradictory C4 text claim; partial because the upstream pipeline and a few non-deposited numbers could not be attempted.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-29
- Rubric version
- not recorded
- Assessed by
- —
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetThe paper tests whether tumor-infiltrating CD4+, CD8+ and Treg T-cell subsets exhibit a shared, cancer-type-independent transcriptomic signature that is distinct from healthy tissue-resident T-cells, in order to identify subset-specific pathways in malignancy.
- ★ Common genes shared across five cancer types differ from those found in nonmalignant tissue-resident T-cells for each subset (CD4-T, CD8-T, Treg) finding
- ★ Cytokine signaling, especially the Th2-type cytokine response, is the top overrepresented pathway in Tregs from malignant samples finding
- ★ A bioinformatics pipeline combining scRNA-seq dimensionality reduction/clustering with GO and Reactome pathway analysis can identify exclusive transcriptomic pathways for tumor-infiltrating T-cell subsets method
- ★ Pathway findings from public scRNA-seq datasets were validated against in-house RNA-seq and proteomic data from T-cell subsets cultured under malignant conditions method
- ★ 652, 69 and 557 genes are exclusive to malignant-derived CD4-T, CD8-T and Treg cells, respectively, relative to nonmalignant counterparts finding
- Regulation of macromolecule metabolic process is positively enriched in malignant Tregs but negatively enriched in nonmalignant Tregs finding
- Peroxiredoxin activity, threonine-type peptidase activity and NADH dehydrogenase activity are recurrent molecular functions common to malignant CD4, CD8 and Treg cells finding
- T-cell subsets can be identified from scRNA-seq data using classical marker genes (cd3g, cd8, cd4, foxp3) as a computational alternative to flow cytometry method
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| single-cell RNA-seq | human breast cancer tissue (GSE114727, GSE75688) | malignant vs nonmalignant tissue origin | gene expression (TPM) in CD4-T, CD8-T, Treg | — |
| single-cell RNA-seq | human lung cancer tissue (GSE126030, GSE99254) | malignant vs nonmalignant tissue origin | gene expression (TPM) in CD4-T, CD8-T, Treg | — |
| single-cell RNA-seq | human colorectal cancer tissue (GSE108989) | malignant vs nonmalignant tissue origin | gene expression (TPM) in CD4-T, CD8-T, Treg | — |
| single-cell RNA-seq with PCA/t-SNE dimensionality reduction and clustering | human melanoma tissue (GSE72056, GSE123139) | none (malignant tissue only) | cell-type clustering and identification of T-cell subsets | Scikit-learn PCA |
| single-cell RNA-seq with PCA/t-SNE dimensionality reduction and clustering | human head and neck cancer tissue (GSE103322) | none (malignant tissue only) | cell-type clustering and identification of T-cell subsets | Scikit-learn PCA |
| Gene Ontology (GO) enrichment and network analysis | in silico gene sets derived from CD4-T, CD8-T, Treg (malignant vs nonmalignant) | none | exclusive and common biological process/molecular function/cellular component annotations | Cytoscape with ClueGO plugin |
| Reactome pathway analysis | in silico gene sets derived from CD4-T, CD8-T, Treg (malignant vs nonmalignant) | none | overrepresented reactome pathways (e.g., protein metabolism, RNA metabolism, antigen presentation) | Reactome Pathway Database |
| RNA-seq and proteomics | in-house cultured T-cell subsets | malignant culture environment | validation of pathway signatures identified in silico (e.g., Th2 cytokine signaling) | — |
- – 652 genes (11.74%) exclusive to malignant-derived CD4-T cells versus nonmalignant CD4-T 652 genes / 11.74%
- – 69 genes (1.10%) exclusive to malignant-derived CD8-T cells versus nonmalignant CD8-T 69 genes / 1.10%
- – 557 genes (12%) exclusive to malignant-derived Tregs versus nonmalignant Tregs 557 genes / 12%
- ▲ Th2-type cytokine signaling identified as the top overrepresented pathway in malignant Tregs, consistent across melanoma, colorectal and oral cancer
- – Regulation of macromolecule metabolic process cluster is positive in malignant Tregs but negative in nonmalignant Tregs
- – Exclusive GO biological process terms: 38 malignant vs 33 nonmalignant CD4; 72 malignant vs 80 nonmalignant CD8; 100 malignant vs 38 nonmalignant Treg
- – Total genes identified: 67,917 malignant CD4, 63,038 malignant CD8, 56,827 malignant Treg (5 tissues); 24,436 nonmalignant CD4, 28,341 nonmalignant CD8, 20,877 nonmalignant Treg (2 tissues)
- – 5771 Reactome annotations identified for malignant data versus 6262 for nonmalignant data
- count 652 genes (11.74%) (exclusive genes in malignant CD4-T vs nonmalignant)
- count 69 genes (1.10%) (exclusive genes in malignant CD8-T vs nonmalignant)
- count 557 genes (12%) (exclusive genes in malignant Treg vs nonmalignant)
- count 7490 biological process; 1404 molecular function; 2035 cellular component; 12,033 reactome pathways (total GO/reactome annotations across all analyzed samples)
- count 5771 (malignant) vs 6262 (nonmalignant) (total Reactome pathway annotations by condition)
- count 93/92 malignant/nonmalignant CD4; 43/123 CD8; 83/81 Treg (ClueGO network clusters per subset and condition)
- count 38/33 CD4; 72/80 CD8; 100/38 Treg (malignant/nonmalignant) (exclusive biological process GO terms per subset and condition)
- count 67,917/63,038/56,827 (malignant CD4/CD8/Treg); 24,436/28,341/20,877 (nonmalignant CD4/CD8/Treg) (total gene counts identified per subset after filtering low-expression genes)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This is an in-silico bioinformatics study that reanalyzed publicly available single-cell RNA-seq datasets (GEO) from five tumor types, using PCA and t-SNE for dimensionality reduction/clustering and fixed marker-gene thresholds (cd3g, cd8, cd4, foxp3) to identify CD4-T, CD8-T and Treg cells. Common versus exclusive gene sets between malignant and nonmalignant samples were derived via Venn-diagram-style set comparison, and functional interpretation was performed with Gene Ontology and Reactome enrichment analysis visualized as networks in Cytoscape/ClueGO, with p-values reported (including as -log(p-value) in scatter plots) for annotation/pathway terms. Findings were further compared qualitatively against RNA-seq and proteomic data from the authors' own in-house cultured T-cell experiments. Results are reported primarily as gene/annotation counts, percentages, and pathway enrichment terms rather than as classical group-comparison statistics (e.g., no t-tests, ANOVA, or dispersion measures such as SD/SEM are described in the reported sections).
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| PCA-based component selection (Scikit-learn) followed by t-SNE dimensionality reduction and clustering | Identification of T-cell subpopulations from melanoma (GSE72056) and head and neck cancer (GSE103322) datasets, Figure 2A,B | component/variance analysis on the full single-cell gene expression matrices per dataset; exact cell counts not stated in this excerpt | not stated |
| Marker-gene threshold classification (cd3g, cd8, cd4, foxp3 presence/absence) | Assignment of cells to CD4-T, CD8-T and Treg subsets across all datasets | all cells passing prior filtering/clustering per dataset (Table 1) | not stated |
| Gene set intersection / Venn diagram comparison | Identification of common and exclusive genes across cancer types and between malignant vs nonmalignant samples per T-cell subset, Figure 2C-G, Table S2 | gene counts per dataset/subset as reported in Table 1 | not stated |
| Gene Ontology (GO) enrichment analysis (Gene Ontology Consortium) | Biological process, molecular function and cellular component annotation of malignant vs nonmalignant gene sets, Tables S3-S8 | gene sets identified per subset/condition | not stated |
| Reactome pathway enrichment analysis | Pathway annotation of malignant vs nonmalignant CD4-T, CD8-T and Treg gene sets, Section 2.4 | gene sets identified per subset/condition | not stated |
| ClueGO/Cytoscape network-based term enrichment with associated p-values (test underlying the p-values not specified in this excerpt) | Biological process network clustering and exclusive/common term identification, Figure 3, Figure 4B-C, Table S9 | GO/Reactome term sets per condition | not stated |
-
Cell-subset identity was assigned using fixed marker-gene presence/absence thresholds (cd3g>0, cd8, cd4, foxp3 combinations).↳ Could also: Reference-based or module-score annotation tools (e.g., SingleR, Seurat's AddModuleScore, or scType) — These approaches assign cell identity using continuous scoring or reference correlation rather than hard on/off thresholds, which can also help capture intermediate or transitional expression states across single cells.
-
Genes common or exclusive to malignant vs nonmalignant samples were determined via list intersection (Venn-diagram-style comparison) after a median-TPM expression cutoff.↳ Could also: A formal per-gene differential expression test (e.g., Wilcoxon rank-sum in Seurat/Scanpy, or pseudobulk modeling with DESeq2/edgeR) with FDR correction — A statistical DE test would also quantify the magnitude and significance of expression differences per gene, complementing a presence/absence list-based comparison.
-
GO and Reactome enrichment p-values from ClueGO were reported for many terms without the correction method specified in this excerpt.↳ Could also: Explicitly reporting a Benjamini-Hochberg FDR or the specific ClueGO correction setting (e.g., Bonferroni step-down) — Explicit reporting of the multiple-testing correction would also make clear how the false discovery rate was controlled across the large number of GO/Reactome terms evaluated.
-
Data from five independent GEO datasets/tumor types were merged and compared using a per-dataset upper-median TPM cutoff.↳ Could also: Batch-integration methods such as Harmony or Seurat's canonical correlation analysis (CCA) integration — Integration methods can also reduce study-specific technical variation before combining gene lists across independently generated single-cell datasets.
-
Dimensionality reduction and clustering used PCA followed by t-SNE.↳ Could also: UMAP combined with graph-based Louvain/Leiden clustering (as implemented in Seurat/Scanpy) — This combination is a widely used current default for scRNA-seq that can also preserve both local and global structure during visualization and clustering.
-
Pathways identified from scRNA-seq were compared qualitatively to independent in-house RNA-seq and proteomic data.↳ Could also: A formal overlap-significance test (e.g., hypergeometric/Fisher's exact test) or gene set enrichment analysis (GSEA) between the two data sources — A quantitative concordance test would also provide a statistical measure of how significantly the scRNA-seq-derived pathways align with the independent RNA-seq/proteomic findings.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-36232369
Title: Analysis of Tumor-Infiltrating T-Cell Transcriptomes Reveal a Unique Genetic Signature across Different Types of Cancer. (Vidal et al., Int J Mol Sci 2022; PMC9569723; DOI 10.3390/ijms231911065)
What the paper actually did (pipeline overview)
Re-analysis of public scRNA-seq from 8 tumor datasets (5 cancer types) + 1 validation RNA-seq, plus an in-house proteomics validation. Per-cell-type (CD4-T, CD8-T, Treg) gene lists were derived for malignant vs nonmalignant origin, exclusive genes/pathways were computed, and networks were built with Cytoscape/ClueGO.
Pipeline (as described in Methods):
- QC FastQC → STAR align hg38 → TPM (cutoff = upper median TPM)
- PCA (scikit-learn) + t-SNE + agglomerative clustering for cell-type ID
- Marker-gene filtering into CD4-T / CD8-T / Treg subsets (CD3G/CD8/CD4/FOXP3)
- Per-subset gene lists → Venn/exclusive-set analysis (InteractiVenn) → exclusive genes
- GO (Gene Ontology Consortium, 2-May-2020) + Reactome enrichment
- Cytoscape v3.8.2 + ClueGO v2.5.7 (p<0.001, kappa) → networks, clusters, exclusive terms
- Validation: RNA-seq (GSE171638) + in-house proteomics (timsTOF, MSFragger 3.2, Perseus)
Code / data availability
- Code repo (github.com/tumourTcells/Cytoscape-file): contains ONLY a README. It is NOT analysis code — it just instructs the reader to download pre-built ClueGO network files from a Google Drive link (bit.ly/3vfQFYA → drive.google.com/...1RZoch5aGnOzFWOYkQj92UEyQS6Yl_aDZ). No pipeline source code exists for alignment, TPM, clustering, T-cell filtering, exclusive-gene set ops, or enrichment. (P16 third-party-tool exemption N/A: there is no runnable tool here either, only the final ClueGO session files.)
- Supplement (mdpi.com/.../s1, 5.4 MB zip): contains the DERIVED result tables:
- Table S1: marker genes used for cell typing
- Table S2: total genes identified per subset (CD4/CD8/Treg) — the per-subset gene lists
- Tables S3–S8: GO annotations (malignant/nonmalignant × CD4/CD8/Treg)
- Table S9: ClueGO cluster analysis (all conditions)
- Table S10: exclusive biological pathways
- Table S11: common pathways across RNA-seq + proteomic validation
- Data: 9 GEO series (public, open): GSE114727, GSE75688, GSE126030, GSE99254, GSE108989, GSE72056, GSE123139, GSE103322 (+ GSE171638 validation).
IN SCOPE (pipeline-derived, attempt to reproduce)
Because no pipeline code is shipped, the upstream steps (align→TPM→cluster→subset→gene list) cannot be re-run faithfully (no code, contradictory marker thresholds, 9 datasets). What CAN be reproduced is whether the paper's reported numbers are recomputable from the deposited derived data (supplement) — i.e. the InteractiVenn/ClueGO counting steps, which IS a pipeline step:
| id | reported result | reproduce by |
|---|---|---|
| C1 | exclusive genes: malig CD4=652, CD8=69, Treg=557 (%) | recompute exclusive sets from S2 gene lists (set ops) |
| C2 | Table 1 total genes per subset (67917, 63038, …) | count unique genes per subset in S2 |
| C3 | Table 2 GO/Reactome counts (BP/MF/CC/Reactome × CD4/CD8/Treg) | count annotation rows in S3–S8 |
| C4 | ClueGO exclusive GO terms (17/17/12/86/54/25) | count exclusive terms in S10 |
| C5 | network clusters (93/92/43/123/83/81) | count clusters in S9 + the Google-Drive ClueGO network files |
| C6 | RNA-seq validation common pathways: CD4=218, Treg=194 | count rows in S11 |
| C7 | proteomics: 2558 diff / 4410 quant / 1692 in scRNA / 561 common | count rows in S11 / proteomics table |
OUT OF SCOPE
- Upstream scRNA reprocessing (align→TPM→PCA→cluster→subset gene lists): no code shipped, marker definitions in Methods are internally contradictory (e.g. CD8+ defined with FOXP3>0; Treg defined CD4>0,FOXP3>0 same as some CD4 rule), 8 heterogeneous datasets, undocumented per-dataset PCA/cluster choices. Not faithfully reproducible → documented blocker.
- Wet-lab: T-cell isolation/a
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.