Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Analysis of Tumor-Infiltrating T-Cell Transcriptomes Reveal a Unique Genetic Signature across Different Types of Cancer.

Int J Mol Sci · 2022
87/100 3/4
Why this verdict

The main results reproduced: recomputed values matched the published ones within tolerance.

Reproduced on the brainbox compute brainarbeit.com
How its reproducibility compares
87/100
Reproducibility score
0.7 SD above mean
vs. all fields · 1173 studies
🎯 Scores higher than 72% of all assessed papers rank 301 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough for the COUNT/SET-OPERATION claims, NOT for the upstream pipeline. No analysis code is shipped (github.com/tumourTcells/Cytoscape-file is a README pointing to a 308MB Google-Drive gzip of the final ClueGO Cytoscape session files), so align->TPM->cluster->subset gene-list generation cannot be re-run and is out of scope (also: contradictory marker thresholds in Methods). What IS reproducible is whether the paper's reported numbers are recomputable from the deposited supplement (MDPI S1, retrieved via EuropePMC) + the ClueGO networks, and they overwhelmingly are, EXACTLY: C1 exclusive genes 652/69/557; C3 Table 2 GO/Reactome (all 12 values 1335/950/1388, 270/159/249, 376/271/385, 2059/1746/1966); C5 network clusters 93/92/43/123/83/81 (Table S9 rows == distinct ClueGO GOGroups, also matching the .cluego sessions); C6 validation pathways 218/194; C7d proteomics common pathways 561; and the Figure-4A 'exclusive GO terms' 38/33/72/80/100/38 (Table S10). ONE flagged discrepancy: the Results-text 'exclusive GO terms' values 17/17/12/86/54/25 are NOT supported by any deposited table and contradict the paper's OWN Figure 4A/Table S10 (38/33/72/80/100/38) -- no p-value threshold reproduces them. This is an internal inconsistency / possible error in the publication; flagged for human review, not asserted as fabrication. C2 (Table 1 totals) and C7a-c (proteomics protein counts) are uncheckable because the upstream intermediates were not deposited. Overall: a faithful 1:1 reproduction of every deposited-derivable reported number except the single self-contradictory C4 text claim; partial because the upstream pipeline and a few non-deposited numbers could not be attempted.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-29
Rubric version
not recorded
Assessed by
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The paper tests whether tumor-infiltrating CD4+, CD8+ and Treg T-cell subsets exhibit a shared, cancer-type-independent transcriptomic signature that is distinct from healthy tissue-resident T-cells, in order to identify subset-specific pathways in malignancy.

Core claims
  • Common genes shared across five cancer types differ from those found in nonmalignant tissue-resident T-cells for each subset (CD4-T, CD8-T, Treg) finding
  • Cytokine signaling, especially the Th2-type cytokine response, is the top overrepresented pathway in Tregs from malignant samples finding
  • A bioinformatics pipeline combining scRNA-seq dimensionality reduction/clustering with GO and Reactome pathway analysis can identify exclusive transcriptomic pathways for tumor-infiltrating T-cell subsets method
  • Pathway findings from public scRNA-seq datasets were validated against in-house RNA-seq and proteomic data from T-cell subsets cultured under malignant conditions method
  • 652, 69 and 557 genes are exclusive to malignant-derived CD4-T, CD8-T and Treg cells, respectively, relative to nonmalignant counterparts finding
  • Regulation of macromolecule metabolic process is positively enriched in malignant Tregs but negatively enriched in nonmalignant Tregs finding
  • Peroxiredoxin activity, threonine-type peptidase activity and NADH dehydrogenase activity are recurrent molecular functions common to malignant CD4, CD8 and Treg cells finding
  • T-cell subsets can be identified from scRNA-seq data using classical marker genes (cd3g, cd8, cd4, foxp3) as a computational alternative to flow cytometry method
Experimental setups
Assay System Perturbation Readout Platform
single-cell RNA-seq human breast cancer tissue (GSE114727, GSE75688) malignant vs nonmalignant tissue origin gene expression (TPM) in CD4-T, CD8-T, Treg
single-cell RNA-seq human lung cancer tissue (GSE126030, GSE99254) malignant vs nonmalignant tissue origin gene expression (TPM) in CD4-T, CD8-T, Treg
single-cell RNA-seq human colorectal cancer tissue (GSE108989) malignant vs nonmalignant tissue origin gene expression (TPM) in CD4-T, CD8-T, Treg
single-cell RNA-seq with PCA/t-SNE dimensionality reduction and clustering human melanoma tissue (GSE72056, GSE123139) none (malignant tissue only) cell-type clustering and identification of T-cell subsets Scikit-learn PCA
single-cell RNA-seq with PCA/t-SNE dimensionality reduction and clustering human head and neck cancer tissue (GSE103322) none (malignant tissue only) cell-type clustering and identification of T-cell subsets Scikit-learn PCA
Gene Ontology (GO) enrichment and network analysis in silico gene sets derived from CD4-T, CD8-T, Treg (malignant vs nonmalignant) none exclusive and common biological process/molecular function/cellular component annotations Cytoscape with ClueGO plugin
Reactome pathway analysis in silico gene sets derived from CD4-T, CD8-T, Treg (malignant vs nonmalignant) none overrepresented reactome pathways (e.g., protein metabolism, RNA metabolism, antigen presentation) Reactome Pathway Database
RNA-seq and proteomics in-house cultured T-cell subsets malignant culture environment validation of pathway signatures identified in silico (e.g., Th2 cytokine signaling)
Key results
  • 652 genes (11.74%) exclusive to malignant-derived CD4-T cells versus nonmalignant CD4-T 652 genes / 11.74%
  • 69 genes (1.10%) exclusive to malignant-derived CD8-T cells versus nonmalignant CD8-T 69 genes / 1.10%
  • 557 genes (12%) exclusive to malignant-derived Tregs versus nonmalignant Tregs 557 genes / 12%
  • Th2-type cytokine signaling identified as the top overrepresented pathway in malignant Tregs, consistent across melanoma, colorectal and oral cancer
  • Regulation of macromolecule metabolic process cluster is positive in malignant Tregs but negative in nonmalignant Tregs
  • Exclusive GO biological process terms: 38 malignant vs 33 nonmalignant CD4; 72 malignant vs 80 nonmalignant CD8; 100 malignant vs 38 nonmalignant Treg
  • Total genes identified: 67,917 malignant CD4, 63,038 malignant CD8, 56,827 malignant Treg (5 tissues); 24,436 nonmalignant CD4, 28,341 nonmalignant CD8, 20,877 nonmalignant Treg (2 tissues)
  • 5771 Reactome annotations identified for malignant data versus 6262 for nonmalignant data
Key statistics
  • count 652 genes (11.74%) (exclusive genes in malignant CD4-T vs nonmalignant)
  • count 69 genes (1.10%) (exclusive genes in malignant CD8-T vs nonmalignant)
  • count 557 genes (12%) (exclusive genes in malignant Treg vs nonmalignant)
  • count 7490 biological process; 1404 molecular function; 2035 cellular component; 12,033 reactome pathways (total GO/reactome annotations across all analyzed samples)
  • count 5771 (malignant) vs 6262 (nonmalignant) (total Reactome pathway annotations by condition)
  • count 93/92 malignant/nonmalignant CD4; 43/123 CD8; 83/81 Treg (ClueGO network clusters per subset and condition)
  • count 38/33 CD4; 72/80 CD8; 100/38 Treg (malignant/nonmalignant) (exclusive biological process GO terms per subset and condition)
  • count 67,917/63,038/56,827 (malignant CD4/CD8/Treg); 24,436/28,341/20,877 (nonmalignant CD4/CD8/Treg) (total gene counts identified per subset after filtering low-expression genes)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This is an in-silico bioinformatics study that reanalyzed publicly available single-cell RNA-seq datasets (GEO) from five tumor types, using PCA and t-SNE for dimensionality reduction/clustering and fixed marker-gene thresholds (cd3g, cd8, cd4, foxp3) to identify CD4-T, CD8-T and Treg cells. Common versus exclusive gene sets between malignant and nonmalignant samples were derived via Venn-diagram-style set comparison, and functional interpretation was performed with Gene Ontology and Reactome enrichment analysis visualized as networks in Cytoscape/ClueGO, with p-values reported (including as -log(p-value) in scatter plots) for annotation/pathway terms. Findings were further compared qualitatively against RNA-seq and proteomic data from the authors' own in-house cultured T-cell experiments. Results are reported primarily as gene/annotation counts, percentages, and pathway enrichment terms rather than as classical group-comparison statistics (e.g., no t-tests, ANOVA, or dispersion measures such as SD/SEM are described in the reported sections).

Replicationbiological Sample sizeSample sizes are given as gene/annotation counts per dataset and condition (e.g., Table 1: gene counts by tissue/GEO ID for malignant and nonmalignant CD4-T, CD8-T, Treg cells), reflecting pooled public single-cell datasets from multiple patients per cancer type rather than a pre-specified power calculation. GroupsMalignant vs nonmalignant tumor-infiltrating T-cell subsets (CD4-T, CD8-T, Treg) across five cancer types (breast, lung, colorectal, melanoma, head and neck) Pairingunpaired Randomization/blindingna Dispersionnone Exact p-valuesyes Effect sizesno Confidence intervalsno
Statistical tests used
Test Applied to n Assumptions
PCA-based component selection (Scikit-learn) followed by t-SNE dimensionality reduction and clustering Identification of T-cell subpopulations from melanoma (GSE72056) and head and neck cancer (GSE103322) datasets, Figure 2A,B component/variance analysis on the full single-cell gene expression matrices per dataset; exact cell counts not stated in this excerpt not stated
Marker-gene threshold classification (cd3g, cd8, cd4, foxp3 presence/absence) Assignment of cells to CD4-T, CD8-T and Treg subsets across all datasets all cells passing prior filtering/clustering per dataset (Table 1) not stated
Gene set intersection / Venn diagram comparison Identification of common and exclusive genes across cancer types and between malignant vs nonmalignant samples per T-cell subset, Figure 2C-G, Table S2 gene counts per dataset/subset as reported in Table 1 not stated
Gene Ontology (GO) enrichment analysis (Gene Ontology Consortium) Biological process, molecular function and cellular component annotation of malignant vs nonmalignant gene sets, Tables S3-S8 gene sets identified per subset/condition not stated
Reactome pathway enrichment analysis Pathway annotation of malignant vs nonmalignant CD4-T, CD8-T and Treg gene sets, Section 2.4 gene sets identified per subset/condition not stated
ClueGO/Cytoscape network-based term enrichment with associated p-values (test underlying the p-values not specified in this excerpt) Biological process network clustering and exclusive/common term identification, Figure 3, Figure 4B-C, Table S9 GO/Reactome term sets per condition not stated
Approaches that could also have been used
  • Cell-subset identity was assigned using fixed marker-gene presence/absence thresholds (cd3g>0, cd8, cd4, foxp3 combinations).
    Could also: Reference-based or module-score annotation tools (e.g., SingleR, Seurat's AddModuleScore, or scType) — These approaches assign cell identity using continuous scoring or reference correlation rather than hard on/off thresholds, which can also help capture intermediate or transitional expression states across single cells.
  • Genes common or exclusive to malignant vs nonmalignant samples were determined via list intersection (Venn-diagram-style comparison) after a median-TPM expression cutoff.
    Could also: A formal per-gene differential expression test (e.g., Wilcoxon rank-sum in Seurat/Scanpy, or pseudobulk modeling with DESeq2/edgeR) with FDR correction — A statistical DE test would also quantify the magnitude and significance of expression differences per gene, complementing a presence/absence list-based comparison.
  • GO and Reactome enrichment p-values from ClueGO were reported for many terms without the correction method specified in this excerpt.
    Could also: Explicitly reporting a Benjamini-Hochberg FDR or the specific ClueGO correction setting (e.g., Bonferroni step-down) — Explicit reporting of the multiple-testing correction would also make clear how the false discovery rate was controlled across the large number of GO/Reactome terms evaluated.
  • Data from five independent GEO datasets/tumor types were merged and compared using a per-dataset upper-median TPM cutoff.
    Could also: Batch-integration methods such as Harmony or Seurat's canonical correlation analysis (CCA) integration — Integration methods can also reduce study-specific technical variation before combining gene lists across independently generated single-cell datasets.
  • Dimensionality reduction and clustering used PCA followed by t-SNE.
    Could also: UMAP combined with graph-based Louvain/Leiden clustering (as implemented in Seurat/Scanpy) — This combination is a widely used current default for scRNA-seq that can also preserve both local and global structure during visualization and clustering.
  • Pathways identified from scRNA-seq were compared qualitatively to independent in-house RNA-seq and proteomic data.
    Could also: A formal overlap-significance test (e.g., hypergeometric/Fisher's exact test) or gene set enrichment analysis (GSEA) between the two data sources — A quantitative concordance test would also provide a statistical measure of how significantly the scRNA-seq-derived pathways align with the independent RNA-seq/proteomic findings.
Software: Python / Scikit-learn (PCA) · Cytoscape · ClueGO (Cytoscape plugin) · Gene Ontology Consortium annotation resources · Reactome Pathway Database

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-36232369

Title: Analysis of Tumor-Infiltrating T-Cell Transcriptomes Reveal a Unique Genetic Signature across Different Types of Cancer. (Vidal et al., Int J Mol Sci 2022; PMC9569723; DOI 10.3390/ijms231911065)

What the paper actually did (pipeline overview)

Re-analysis of public scRNA-seq from 8 tumor datasets (5 cancer types) + 1 validation RNA-seq, plus an in-house proteomics validation. Per-cell-type (CD4-T, CD8-T, Treg) gene lists were derived for malignant vs nonmalignant origin, exclusive genes/pathways were computed, and networks were built with Cytoscape/ClueGO.

Pipeline (as described in Methods):

  • QC FastQC → STAR align hg38 → TPM (cutoff = upper median TPM)
  • PCA (scikit-learn) + t-SNE + agglomerative clustering for cell-type ID
  • Marker-gene filtering into CD4-T / CD8-T / Treg subsets (CD3G/CD8/CD4/FOXP3)
  • Per-subset gene lists → Venn/exclusive-set analysis (InteractiVenn) → exclusive genes
  • GO (Gene Ontology Consortium, 2-May-2020) + Reactome enrichment
  • Cytoscape v3.8.2 + ClueGO v2.5.7 (p<0.001, kappa) → networks, clusters, exclusive terms
  • Validation: RNA-seq (GSE171638) + in-house proteomics (timsTOF, MSFragger 3.2, Perseus)

Code / data availability

  • Code repo (github.com/tumourTcells/Cytoscape-file): contains ONLY a README. It is NOT analysis code — it just instructs the reader to download pre-built ClueGO network files from a Google Drive link (bit.ly/3vfQFYA → drive.google.com/...1RZoch5aGnOzFWOYkQj92UEyQS6Yl_aDZ). No pipeline source code exists for alignment, TPM, clustering, T-cell filtering, exclusive-gene set ops, or enrichment. (P16 third-party-tool exemption N/A: there is no runnable tool here either, only the final ClueGO session files.)
  • Supplement (mdpi.com/.../s1, 5.4 MB zip): contains the DERIVED result tables:
    • Table S1: marker genes used for cell typing
    • Table S2: total genes identified per subset (CD4/CD8/Treg) — the per-subset gene lists
    • Tables S3–S8: GO annotations (malignant/nonmalignant × CD4/CD8/Treg)
    • Table S9: ClueGO cluster analysis (all conditions)
    • Table S10: exclusive biological pathways
    • Table S11: common pathways across RNA-seq + proteomic validation
  • Data: 9 GEO series (public, open): GSE114727, GSE75688, GSE126030, GSE99254, GSE108989, GSE72056, GSE123139, GSE103322 (+ GSE171638 validation).

IN SCOPE (pipeline-derived, attempt to reproduce)

Because no pipeline code is shipped, the upstream steps (align→TPM→cluster→subset→gene list) cannot be re-run faithfully (no code, contradictory marker thresholds, 9 datasets). What CAN be reproduced is whether the paper's reported numbers are recomputable from the deposited derived data (supplement) — i.e. the InteractiVenn/ClueGO counting steps, which IS a pipeline step:

id reported result reproduce by
C1 exclusive genes: malig CD4=652, CD8=69, Treg=557 (%) recompute exclusive sets from S2 gene lists (set ops)
C2 Table 1 total genes per subset (67917, 63038, …) count unique genes per subset in S2
C3 Table 2 GO/Reactome counts (BP/MF/CC/Reactome × CD4/CD8/Treg) count annotation rows in S3–S8
C4 ClueGO exclusive GO terms (17/17/12/86/54/25) count exclusive terms in S10
C5 network clusters (93/92/43/123/83/81) count clusters in S9 + the Google-Drive ClueGO network files
C6 RNA-seq validation common pathways: CD4=218, Treg=194 count rows in S11
C7 proteomics: 2558 diff / 4410 quant / 1692 in scRNA / 561 common count rows in S11 / proteomics table

OUT OF SCOPE

  • Upstream scRNA reprocessing (align→TPM→PCA→cluster→subset gene lists): no code shipped, marker definitions in Methods are internally contradictory (e.g. CD8+ defined with FOXP3>0; Treg defined CD4>0,FOXP3>0 same as some CD4 rule), 8 heterogeneous datasets, undocumented per-dataset PCA/cluster choices. Not faithfully reproducible → documented blocker.
  • Wet-lab: T-cell isolation/a
Figures / tables: Table
C1a
Reported
652 exclusive genes (malignant CD4)
Reproduced
652
exact
C1b
Reported
69 exclusive genes (malignant CD8)
Reproduced
69
exact
C1c
Reported
557 exclusive genes (malignant Treg)
Reproduced
557
exact
C2a-f
Reported
Table 1 total genes per subset (67917/63038/56827/24436/28341/20877)
Reproduced
uncheckable
partial
C3a
Reported
GO BiologicalProcess malig CD4/CD8/Treg = 1335/950/1388
Reproduced
1335/950/1388
exact
C3b
Reported
GO MolecularFunction malig CD4/CD8/Treg = 270/159/249
Reproduced
270/159/249
exact
C3c
Reported
GO CellularComponent malig CD4/CD8/Treg = 376/271/385
Reproduced
376/271/385
exact
C3d
Reported
Reactome malig CD4/CD8/Treg = 2059/1746/1966
Reproduced
2059/1746/1966
exact
C4-Fig4A
Reported
exclusive GO terms (Figure 4A): malig/nonmalig CD4=38/33, CD8=72/80, Treg=100/38
Reproduced
38/33, 72/80, 100/38
exact
C4-text
Reported
exclusive GO terms (Results text): 17/17/12/86/54/25
Reproduced
Table S10 yields 38/33/72/80/100/38; no p-value cutoff reproduces 17/12/54
did not match
C5a-f
Reported
network clusters 93/92/43/123/83/81 (malig/nonmalig CD4/CD8/Treg)
Reproduced
93/92/43/123/83/81
exact
C6a
Reported
RNA-seq validation common pathways CD4 = 218
Reproduced
218
exact
C6b
Reported
RNA-seq validation common pathways Treg = 194
Reproduced
194
exact
C7d
Reported
proteomics common pathways = 561
Reproduced
561
exact
C7a-c
Reported
proteomics 2558 differential / 4410 quantifiable / 1692 in scRNA gene list
Reproduced
uncheckable
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

No assessment has been recorded yet.
🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.