Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Single-Cell Analysis Reveals Characterization of Infiltrating T Cells in Moderately Differentiated Colorectal Cancer.

Front Immunol · 2021
L1 45/100 3/4
⚑ Flagged for review — a reproduced result did not match the reported value

Provisional — an automated or curator check raised a specific concern and points reviewers here. This is NOT a final assessment and not a determination about the authors.

Why this verdict

The main result did not reproduce in this reproduction attempt. Where our recomputation produced values that differ from the published ones, those discrepancies are listed below. This is a single automated attempt — not peer review and not a finding of error or misconduct — and differences can also arise from data access, undocumented parameters or the computing environment. The verdict can be contested via “report an error”.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +6
✓ What held up
  • Reported values were directly comparable
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🟡A deviation arose in the data or preprocessing
  • 🔴A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
45/100
Reproducibility score
1.7 SD below mean
vs. all fields · 1173 studies
🎯 Scores higher than 6% of all assessed papers rank 1103 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough to RE-RUN the method, but reported NUMBERS only partially reproduce. The paper is a secondary re-analysis of public GSE108989 with the third-party sscClust tool (P16). PROFILE: GSE108989 is clean/complete (11138 cells, 10805 post-QC, 12 patients) and delivers what it promises -- the discrepancies are in the downstream paper, not the deposit. REPRODUCED: the 12-patient deposit (exact), the 5-moderately-differentiated patient selection (identified them: 3 colon + 2 rectal), the 8-tumor/7-blood cluster STRUCTURE, and all major marker-defined T-cell subtypes (Treg, exhausted CD8-TEX, naive, TEMRA/TEFF) using the paper's own markers. NOT REPRODUCED: the reported cell counts 1632 tumor / 1252 blood (observed 1472 / 1164; reaching the reported N requires including LOW-differentiated patients -> possible inconsistency with the stated 'moderately differentiated' criterion), the 12547-gene count (matches no threshold; identical for two different cell sets), and per-cluster counts. NOT ATTEMPTED: iTALK ligand-receptor totals and limma DEG counts (both sit downstream of clusters that are not exactly reproducible -- no published seed, undefined NMI k-selection, and a manual marker-based merge -- and the colon-vs-rectal DE is patient-level confounded at 3 vs 2 patients); Metascape GO/KEGG + PPI/MCODE are out of scope (online/manual). Faithful reimplementation of sscClust's documented algorithm used because cor.BLAS/ssc.build live in the uninstalled companion sscVis package. All grades provisional -- human audit required; cell-count and gene-count inconsistencies flagged.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 45
    assessed: 2026-06-18 ⛓ 6e226d9b574f
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-18
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-18
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

What are the characteristics of tumor-infiltrating and peripheral blood T cells in moderately differentiated colorectal cancer, and how do tumor-infiltrating T cell populations and gene expression differ between colon cancer and rectal cancer?

Core claims
  • Eight distinct T cell populations are identifiable in CRC tumor tissue and seven in peripheral blood by unsupervised clustering of scRNA-seq data. finding
  • Tumor-Treg (C1) is strongly correlated with Th17 cells (C4) in tumor tissue. finding
  • CD8+ tissue-resident memory T cells (CD8+ TRM) are positively correlated with CD8+ intraepithelial lymphocytes (CD8+ IEL). finding
  • Colon and rectal cancers differ in the composition of tumor-infiltrating T cell populations, with CD8+ IEL found only in rectal cancer and the majority of CD8+ Tex found in colon cancer. finding
  • Ligand-receptor crosstalk including checkpoint pairs (e.g., Treg CD80–Th17 CTLA4, Treg CD274–Th17 PDCD1) and cytokine pair CCL4–CCR8 between CD8+ Tex and Tumor-Treg occurs among tumor-infiltrating T cells. mechanism
  • T cells from colon and rectal cancer tissues show changes in gene expression pattern, with cluster-specific differentially expressed genes (e.g., TNF, CXCR3 up; CXCR6, CCR6 down in colon Tumor-Treg). finding
  • Reanalysis of published scRNA-seq data restricted to moderately differentiated CRC samples can characterize functionally distinct T cell subsets. method
  • A combined pipeline (sscClust, K-means with NMI, iTALK, Limma, Metascape) was used to cluster cells, infer crosstalk, and analyze differential expression. method
Experimental setups
Assay System Perturbation Readout Platform
single-cell RNA sequencing (reanalysis of public data) tumor tissue T cells from 5 moderately differentiated CRC patients none gene expression matrix; T cell cluster identity (1632 cells, 12547 genes)
single-cell RNA sequencing (reanalysis of public data) peripheral blood T cells from 5 moderately differentiated CRC patients none gene expression matrix; T cell cluster identity (1252 cells, 12547 genes)
unsupervised clustering / cell type identification CD4+ and CD8+ T cells (tumor and blood) none number of clusters via NMI index; tSNE visualization; marker gene expression sscClust R package
ligand-receptor interaction (cell-cell crosstalk) analysis eight tumor T cell clusters; seven blood T cell clusters none number of ligand-receptor pairs by category (growth factor, cytokine, checkpoint, other) iTALK R package (2648 ligand-receptor pairs)
correlation analysis T cell clusters (tumor and blood) none Pearson correlation coefficient between cluster average expression Corrplot R package
differential expression analysis tumor-infiltrating T cell clusters, colon cancer vs rectal cancer none (comparison by tumor location) DEGs (adjusted P<0.05, |logFC|>1) Limma R package
functional enrichment and protein-protein interaction analysis DEGs per tumor T cell cluster none GO biological process/KEGG/Reactome terms; PPI network and MCODE modules Metascape; Cytoscape v3.4.0
Key results
  • Eight distinct T cell clusters identified in tumor tissue (Tumor-Treg, CD4+TRM, CD4+TEM, Th17, CD8+TEM, CD8+TEX, CD8+TRM, CD8+IEL)
  • CD8+ Tex (C6) cells predominantly in colon cancer (177 cells) versus rectal cancer (22 cells) 88.94% vs 11.06%
  • CD8+ IEL (C8) cells found exclusively in rectal cancer 54 of 54 cells in rectal cancer
  • 7852 ligand-receptor pairs identified among eight tumor T cell clusters (636 growth factor, 1170 cytokine, 395 checkpoint, 5651 other) 7852 pairs
  • Tumor-Treg (C1) showed 112 DEGs between colon and rectal cancer (43 up, 69 down); Th17 (C4) showed only 9 DEGs 112 and 9 DEGs
  • Tumor-Treg PPI network contained 51 genes and 80 interactions; module 2 (CCL5, CCR6, CXCR3, CXCR6) enriched in chemokine signaling 51 genes, 80 interactions
  • Strong correlation between Tumor-Treg (C1) and Th17 (C4); CD8+TEM (C5) strongly correlated with CD4+TEM, CD8+IEL, CD8+TRM, and CD8+TEX
  • C1 contained 547 infiltrating Treg cells (339 colon, 208 rectal) 547 cells
Key statistics
  • count 1632 T cells from tumor tissue; 12,547 genes (tumor tissue scRNA-seq cells, moderately differentiated patients)
  • count 1252 T cells from peripheral blood; 12,547 genes (peripheral blood scRNA-seq cells, moderately differentiated patients)
  • count 7852 ligand-receptor pairs (among eight tumor T cell clusters)
  • count 4546 ligand-receptor pairs (among seven peripheral blood T cell clusters)
  • count 88.94% (177 cells) colon vs 11.06% (22 cells) rectal (distribution of CD8+ Tex (C6) cells)
  • count 112 DEGs (43 up, 69 down) (C1 Tumor-Treg colon vs rectal cancer)
  • other adjusted P<0.05 and |logFC|>1 (DEG screening threshold (Benjamini & Hochberg))
  • pvalue Log(q-value) -8.4 (C1_MCODE_2 enrichment for chemokine receptors bind chemokines)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This study performed a secondary analysis of publicly available scRNA-seq data, selecting five moderately differentiated CRC patients from a larger 12-patient dataset, and profiled 1632 tumor-infiltrating and 1252 peripheral-blood T cells independently. Unsupervised K-means clustering with NMI-based optimal cluster selection was used to define T cell subtypes, visualised with tSNE. Differential gene expression between colon and rectal cancer samples was tested per T cell cluster using the Limma empirical-Bayes framework with Benjamini-Hochberg correction, and results were primarily reported as DEG counts, cell proportions, and enriched pathway terms.

Replicationbiological Sample sizeFive moderately differentiated CRC patients selected from a 12-patient public dataset (Zhang et al.); 1632 tumor-tissue T cells and 1252 peripheral-blood T cells retained after screening; patient-level n not broken down by colon vs. rectal subgroup GroupsColon cancer vs. rectal cancer (within tumor-tissue and peripheral-blood T cell clusters); tumor tissue vs. peripheral blood analysed as independent datasets Pairingunclear Randomization/blindingnot stated Dispersionnone Exact p-valuesno Effect sizesno Confidence intervalsno Multiplicity correctionBenjamini-Hochberg FDR
Statistical tests used
Test Applied to n Assumptions
K-means clustering with NMI-based cluster number selection Identification of CD4+ and CD8+ T cell subtypes in tumor tissue and peripheral blood (2–20 clusters pre-set, optimal number chosen by maximal NMI) 1632 cells (tumor tissue); 1252 cells (peripheral blood) not stated
Pearson correlation coefficient (Cor function, R) Inter-cluster correlation heatmaps based on average gene expression per cluster (tumor tissue: 8 clusters; peripheral blood: 7 clusters) 8 clusters (tumor) / 7 clusters (peripheral blood); averages across cells per cluster not stated
Limma empirical-Bayes moderated t-statistic (classical Bayesian method) Differential expression analysis of each T cell cluster: colon cancer vs. rectal cancer Varies by cluster (e.g., C1: 339 colon + 208 rectal cells); 5 patients total not stated
Hypergeometric enrichment test (via Metascape, parameters: Min Overlap=3, P cutoff=0.05, Min Enrichment=1.5) GO biological process, KEGG pathway, and Reactome pathway enrichment for DEG lists per cluster; also for MCODE module genes DEG lists per cluster (e.g., C1: 112 DEGs) not stated
Ligand-receptor pair matching (iTALK, top 50% expressed genes per cluster matched to 2648 non-redundant pairs) Cell-cell crosstalk characterisation among 8 tumor-tissue and 7 peripheral-blood T cell clusters null na
MCODE algorithm (Molecular Complex Detection) for PPI module identification PPI networks constructed from DEGs per cluster; modules identified for functional enrichment null na
Approaches that could also have been used
  • Differential expression between colon and rectal cancer was tested using Limma, a framework originally developed for microarray and bulk RNA-seq data, applied directly to single-cell counts.
    Could also: MAST (Model-based Analysis of Single-cell Transcriptomics) or a pseudo-bulk approach (summing counts per patient per cluster, then applying DESeq2 or edgeR) could also have been used. — MAST explicitly models the bimodal, zero-inflated expression distributions common in scRNA-seq; pseudo-bulk methods treat the patient (n=5) rather than the cell as the unit of replication, which better matches the biological independence structure and is increasingly recommended for differential testing in scRNA-seq data.
  • K-means clustering was used for T cell subtype identification, with the cluster number selected by maximal NMI across a pre-set range of 2–20.
    Could also: Graph-based clustering algorithms such as Louvain or Leiden (as implemented in Seurat or Scanpy) could also have been used. — Graph-based methods do not require pre-specifying a cluster number range, are widely adopted in recent scRNA-seq workflows, and can handle non-spherical cluster geometries that K-means assumes to be compact and equally sized.
  • tSNE was used for two-dimensional visualisation of T cell clusters.
    Could also: UMAP (Uniform Manifold Approximation and Projection) could also have been used for low-dimensional embedding. — UMAP has become a common alternative to tSNE in scRNA-seq analyses; it tends to better preserve global structure and runs faster on large cell numbers, while tSNE better separates local neighbourhood structure.
  • Inter-cluster relationships were characterised using Pearson correlation of cluster-average expression profiles.
    Could also: Spearman rank correlation, or trajectory/pseudotime analyses (e.g., Monocle, Slingshot), could also have been used. — Spearman correlation is more robust to outlier genes and skewed expression distributions common in scRNA-seq; trajectory analyses could additionally capture continuous developmental relationships among T cell states rather than pairwise similarities between discrete cluster averages.
  • Differences in T cell subtype composition between colon and rectal cancer were described using raw cell counts and percentages without a formal statistical test.
    Could also: A Fisher's exact test, chi-square test, or a permutation-based test on the proportion of cells per cluster could also have been applied. — A formal test would quantify whether the observed proportion differences (e.g., CD8+ IEL present only in rectal cancer; 88.94% of CD8+ T_EX in colon cancer) exceed what might arise by chance given the small number of patients, complementing the descriptive cell counts.
  • Only the five moderately differentiated patients were selected from the 12-patient source dataset, with differentiation grade treated as an exclusion criterion.
    Could also: All 12 patients could also have been analysed with differentiation grade included as a covariate (e.g., in the Limma model) or using batch-aware integration methods. — Including all patients would increase the effective sample size and statistical power, while a covariate term for differentiation grade would adjust for its potential confounding effect, enabling broader generalisability of the findings.
Software: R / sscClust · R / corrplot · R / iTALK · R / limma · Metascape (online) · Cytoscape 3.4.0

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

scope.md — pmid-33584715

Paper: Yang et al. 2021, Front Immunol. "Single-Cell Analysis Reveals Characterization of Infiltrating T Cells in Moderately Differentiated Colorectal Cancer." (re-analysis)

  • Data: GEO GSE108989 (Zhang et al. 2018 Nature CRC T-cell Smart-seq2 deposit; 12 patients).
  • Code: github.com/Japrin/sscClust (Zhang-lab clustering tool) — THIRD-PARTY tool applied to public data (P16 case: equally valid).

In scope (pipeline-derived)

  1. Patient/cell SELECTION: pick the 5 "moderately differentiated" patients from the 12; count tumor-tissue T cells (reported 1632), peripheral-blood T cells (1252), genes (12547).
  2. CLUSTERING via sscClust: top-1500 HVG (SD) -> Spearman cell-cell corr ("iCor") -> kmeans k=2..20 (nstart=50,iter.max=1000) -> NMI to pick k -> tSNE. CD4 & CD8 clustered SEPARATELY, then MANUALLY merged by Zhang-marker genes. Reported: tumor 8 clusters (4 CD4 + 4 CD8), blood 7 clusters (4 CD4 + 3 CD8); per-cluster cell counts (Table 1A/3A).
  3. DE: limma (adj.P<0.05, |logFC|>1, BH) colon vs rectal per cluster -> DEG counts (Table 1B/3B).
  4. Ligand-receptor: iTALK on top-50%-expressed genes vs 2648 LR pairs -> 7852 (tumor)/4546 (blood).

Out of scope (manual / online / non-pipeline)

  • Metascape GO/KEGG/Reactome enrichment (online tool, manual).
  • PPI (BioGrid/InWeb/OmniPath) + MCODE modules via Metascape + Cytoscape (manual/online).
  • Manual cluster->subtype merging & naming (judgement, marker-based; not algorithmic).
  • Patient differentiation grading (wet-lab/clinical, from data descriptor).

Key reproducibility blockers (recorded honestly)

  • No random seed published; kmeans + tSNE stochastic.
  • "NMI to pick k" underspecified (no reference labels given); manual merge to 4 not specified.
  • Cell counts 1632/1252 do NOT match the deposit counts for the 5 pure-"moderate" patients (observed 1472/1164); reaching 1632/1252 requires including LOW-differentiated patients.
Figures / tables: TableFig.1Fig.6Fig.2Fig.7Fig.3
C1
Reported
12 patients in GSE108989
Reproduced
12 patient GSMs / 11138 cells
exact
C2
Reported
5 moderately-differentiated patients selected
Reproduced
identified P1207,P0123,P0413(colon),P0309,P0411(rectum)
exact
C3
Reported
1632 tumor T cells
Reproduced
1472 (5 moderate patients, GEO deposit)
did not match
C4
Reported
1252 blood T cells
Reproduced
1164
did not match
C5
Reported
12547 genes
Reproduced
no threshold matches; identical value for 2 different cell sets
did not match
C6
Reported
8 tumor clusters (4 CD4 + 4 CD8)
Reproduced
8 recovered at paper's merged k
partial
C7
Reported
7 blood clusters (4 CD4 + 3 CD8)
Reproduced
7 recovered at paper's merged k
partial
C8
Reported
Tumor-Treg subtype (FOXP3/CCR8)
Reproduced
FOXP3/IL2RA/CCR8/LAYN high in tumor CD4 C1,C3
partial
C9
Reported
CD8-TEX exhausted subtype (PDCD1/HAVCR2/LAYN)
Reproduced
clearly recovered (tumor CD8 C2)
partial
C10
Reported
blood naive/TCM/TEMRA/Treg subtypes
Reproduced
all CD4(4/4) + CD8(3/3) recovered by markers
partial
C11
Reported
per-cluster cell counts (Table 1A/3A)
Reproduced
do not match (e.g. paper Treg 547 vs my max CD4 287)
did not match
C12
Reported
ligand-receptor pairs 7852/4546 (iTALK)
Reproduced
not attempted (downstream of non-reproducible clusters)
partial
C13
Reported
DEG counts colon vs rectal per cluster (limma)
Reproduced
not attempted (patient-confounded, downstream of clusters)
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 45/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟢2. Endpoint comparability
🟡3. Location of the main deviation
🔴4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +6

On clean, complete public data (GSE108989) the structure (8 tumor / 7 blood clusters) and all major marker-defined T-cell subtypes (Treg, exhausted CD8-TEX, naive, TEMRA) reproduce qualitatively, so the central biological conclusion holds in limited form. However, the reported cell counts (1632/1252 vs observed 1472/1164) are only reachable by including LOW-differentiated patients — a likely authors'-side inconsistency with the stated 'moderately differentiated' selection — and the 12547-gene count is implausible (identical for two different cell sets, matching no threshold). Per-cluster counts, iTALK LR pairs, and limma DEGs sit downstream of stochastic, seed-less, manually-merged clusters and were not reproducible. Net: a solid re-analysis with moderate, mostly-explainable deviations plus two flagged numeric anomalies on the authors' side — overall yellow.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

325.9 k
tokens (I/O) · 22.5 M incl. cache
34 min
runtime · 0.04 CPU-h
3.1 GB
peak RAM
1
HPC jobs
hummel
machine