Heterogeneity of cancer-associated fibroblasts in head and neck squamous cell carcinoma.
The main results reproduced, with only marginal, non-material deviations.
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
PARTIAL reproduction (fresh re-run after requeue; «job»). The brief's accession GSE139324 is a CD45+ FACS-sorted immune atlas (Cillo 2020) and contains no fibroblasts, so the paper's defining CAF result is NOT derivable from it (flagged as a data/claim mismatch, not fabrication). I therefore applied the paper's named third-party tool (pySCENIC, P16-valid) to the paper's OWN unsorted CAF-containing series GSE173647: scanpy QC + leiden + relative-z lineage assignment yielded a 14,860-cell fibroblast subset, then pyscenic grn/ctx/aucell (hg38 cisTarget v10) recovered 166 active regulons. Overlap with the figure-only Fig 1F TF set is concordant but not 1:1: 5 of 19 hand-read confident TFs (CEBPG, CREB3, FOXM1, SOX4, TCF4; 8 of 23 incl. uncertain), several top-ranked by AUCell (SOX4 #23, FOXM1 #36, CREB3 #49). This independently confirms the earlier run («job», 189 regulons, 6/19): a stable core SOX4/FOXM1/CREB3/CEBPG recovers in BOTH runs, the remainder shuffles -> the non-exact match is GRNBoost2 stochasticity. Three understood reasons for partial vs exact: figure-only low-res reference with no supplementary table; stochastic GRNBoost2; single-series subset vs the paper's 3-series integration. NOT attempted: exact 7-subtype top-5 matching (hard 20%), and all non-pySCENIC tools (LASSO CAFscore, NMF, hdWGCNA, BayesPrism, CellPhoneDB, spatial) which are out of scope. INFRA NOTE: «infra»+/home user quota was exhausted (janitor clean_beegfs du-timeout failing every tick); cleared via a compute-node SLURM reclaim job (2220566) before compute could run. Honest partial.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 50assessed: 2026-06-20 ⛓ a291d5922bd2
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-23
- Rubric version
- not recorded
- Assessed by
- —
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetThe study tests whether cancer-associated fibroblasts (CAFs) in head and neck squamous cell carcinoma (HNSCC) are functionally heterogeneous subpopulations with distinct prognostic significance, immune-suppressive roles, and metabolic activity, particularly focusing on a CKS2+ inflammatory CAF (iCAF) subset.
- ★ Seven distinct CAF subsets exist in HNSCC, identified via integration of scRNA-seq, bulk transcriptomic, and spatial transcriptomic data. finding
- ★ CKS2+ iCAFs are significantly associated with poor prognosis and are localized in close proximity to cancer cells. finding
- ★ CKS2+ iCAFs negatively correlate with cytotoxic CD8+ T cells and NK cells, but positively correlate with exhausted CD8+ T cells. finding
- ★ Patient clusters with high CKS2+ iCAFs (Cluster 3) or high CKS2- iCAFs/CENPF-/MYLPF- myCAFs (Cluster 2) show limited immunotherapeutic response. finding
- ★ Cancer cells show close intercellular communication with CKS2+ iCAFs and CENPF+ myCAFs via specific ligand-receptor pairs. finding
- ★ CKS2+ iCAFs display the highest metabolic activity among CAF subsets, notably in alanine/aspartate/glutamate metabolism, glycolysis/gluconeogenesis, and glycosaminoglycan biosynthesis. finding
- ★ A six-gene-module-based prognostic risk model was constructed from genes correlated with the CKS2+ iCAFs subset using LASSO regression. method
- ★ High infiltration of CKS2+ CAFs, validated by immunohistochemistry in a patient cohort, indicates worse overall survival. finding
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| single-cell RNA sequencing (scRNA-seq) | 24 HNSCC tissue samples (human) | none | cell type identity, cluster-specific transcripts, CAF subset marker expression | Seurat; CellTypist database |
| bulk RNA-seq / TCGA transcriptomic analysis | TCGA-HNSCC cohort, 503 samples with survival data | none | gene expression, survival/prognosis association, CAF subset proportion (deconvolution) | TCGAbiolinks; BayesPrism |
| spatial transcriptomics | HNSCC tissue sections (GSE181300) | none | spatial distribution of CAF subsets relative to epithelial cells | STUtility; Spacexr |
| immunohistochemistry (IHC) | 52 primary HNSCC patient specimens | none | CKS2+ CAF abundance scored and correlated with prognosis | Polink2-plus reagent kit |
| regulon/transcription factor activity analysis | CAF populations from scRNA-seq data | none | top 5 specific transcription factors per CAF subset | pySCENIC |
| co-expression network analysis (WGCNA) | fibroblast expression matrix (CKS2+ iCAFs subset) | none | gene modules correlated with prognosis, module eigengene connectivity (kME) | hdWGCNA |
| cell-cell communication analysis | epithelial cells (malignant/normal) and CAF subsets from scRNA-seq/TCGA | none | ligand-receptor interaction pairs, enrichment score of interactions | CellPhoneDB; Copykat |
| metabolic flux prediction | fibroblast/CAF subsets from scRNA-seq | none | activity of glucose and other metabolic pathways per CAF subset | scFEA; scMetabolism |
| immune cell infiltration estimation (ssGSEA) and immunotherapy response prediction | TCGA-HNSCC bulk samples | none | immune cell infiltration ratio; predicted immunotherapy response | ssGSEA; TIDE; TCIA |
- ▼ Fibroblast and epithelial cell clusters showed significant poor prognosis by Scissor analysis
- – CKS2+ iCAFs and PTN+ myCAFs associated with poor prognosis; CENPF-/MYLPF- myCAFs associated with more favorable prognosis
- – CKS2+ iCAFs mainly distributed in areas of high epithelial cell concentration in spatial transcriptomic sections
- – Negative correlation between CKS2+ iCAFs and cytotoxic CD8+ T cells/NK cells; positive correlation with exhausted CD8+ T cells
- – Cluster 3 (high CKS2+ iCAFs) and Cluster 2 (high CKS2- iCAFs/CENPF-/MYLPF- myCAFs) showed no significant immunotherapeutic response
- – Close interactions confirmed between cancer cells and CKS2+ iCAFs/CENPF+ myCAFs with identified ligand-receptor pair
- ▲ CKS2+ iCAFs showed highest metabolic activity among the seven CAF subsets
- ▼ Patients with high CKS2+ CAF infiltration had poor overall survival (IHC validation)
- count 24 HNSCC samples (GSE139324, GSE173647, GSE173964) (scRNA-seq sample source)
- count 73,660 cells retained after QC (cells after removing >20% mitochondrial genome content)
- count 2,000 variable genes selected (used for PCA in scRNA-seq and CAF subclustering)
- count 503 TCGA samples with survival time (TCGA-HNSCC cohort used for prognostic analysis)
- count 7 CAF subsets identified (CAF heterogeneity classification)
- count 52 primary HNSCC patients (IHC validation cohort)
- other top 25 genes by kME value selected from 6 modules (brown, yellow, midnight blue, pink, red, light cyan) (gene sets for prognostic model construction)
- pvalue P < 0.05 (significance threshold; *P<0.05, **P<0.01, ***P<0.001) (statistical significance criteria used throughout, including CellPhoneDB ligand-receptor retention)
- other z-score enrichment score > 1 (threshold indicating strong cell-cell interaction in bulk_ScrNA analysis)
- other soft threshold power of 5 (used for topological overlap matrix (TOM) calculation in hdWGCNA)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This bioinformatics study integrated public scRNA-seq data (24 HNSCC samples, 73,660 cells), bulk RNA-seq data from TCGA (503 samples), spatial transcriptomic data, and an in-house IHC cohort (52 patients) to characterize CAF heterogeneity. Main statistical procedures included unsupervised clustering, Kaplan-Meier survival analysis with log-rank testing, LASSO-based prognostic model construction, Pearson correlation between deconvolved cell proportions, and ANOVA/chi-square tests for immunotherapy response comparisons. Significance was uniformly set at P<0.05 and reported as threshold categories (*, **, ***).
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Log-rank test | Kaplan-Meier overall survival comparisons between CKS2+ CAF high vs. low IHC groups and TCGA risk groups | 52 (IHC cohort); 503 (TCGA cohort) | not stated |
| Student's t-test | Inter-group differences in gene expression and deconvolved cell proportions | null | not stated |
| Pearson correlation | Correlation between proportions of 7 CAF subpopulations and 3 CD8+ T cell/NK cell subsets in TCGA bulk data | 503 | not stated |
| Chi-square test | Differences in patient immunotherapy response among NMF-defined CAF subtypes | null | not stated |
| One-way ANOVA | Immunotherapy response scores across four NMF clusters | null | not stated |
| LASSO regression | Feature selection and construction of prognostic risk model from top-kME genes of six hdWGCNA modules | 503 split 1:1 train/test | not stated |
| ROC analysis (AUC) | Performance assessment of the LASSO-based prognostic model | null | na |
| CellPhoneDB permutation test (P<0.05 threshold) | Ligand-receptor interaction inference between CAF subsets and epithelial cells | null | not stated |
-
Multiple independent t-tests were used for inter-group comparisons without a stated check of normality or homoscedasticity↳ Could also: Wilcoxon rank-sum (Mann-Whitney U) test could also be applied, particularly for deconvolved proportions which are bounded [0,1] and often non-normally distributed — Non-parametric tests make no distributional assumption; they are commonly preferred when sample distributions are unknown or when values are compositional/proportional
-
P-values for all hypothesis tests were reported only as threshold categories (*, **, ***) rather than exact values↳ Could also: Reporting exact P-values (e.g., P=0.023) is also standard and required by many journals — Exact P-values allow readers to apply their own thresholds and facilitate meta-analysis; threshold categories alone can mask borderline results
-
No multiple-testing correction was applied across the many simultaneous Pearson correlations (7 CAF subsets × 4 immune populations) and across ANOVA/t-test comparisons↳ Could also: Benjamini-Hochberg FDR correction could also be applied across this family of tests — With ~28 simultaneous correlation tests, the expected number of false positives at P<0.05 is ~1.4; FDR control is a standard approach to characterize the false-discovery rate in such exploratory multi-test settings
-
Survival analysis used the log-rank test and Kaplan-Meier curves dichotomizing patients at the median CKS2 IHC score↳ Could also: A Cox proportional hazards model treating the IHC score as a continuous variable (or using multivariable adjustment for known confounders such as stage and HPV status) could also be used — Cox regression retains the full information in a continuous score, avoids the arbitrariness of a median cut-point, and allows adjustment for clinical covariates that may confound the prognostic association
-
Effect sizes (e.g., hazard ratios, Pearson r values, odds ratios) were not reported alongside P-values↳ Could also: Reporting effect size estimates with 95% confidence intervals alongside each P-value is also standard practice — Effect sizes convey clinical or biological magnitude independently of sample size; confidence intervals communicate estimation uncertainty and are requested by CONSORT, STROBE, and most statistical reporting guidelines
-
The LASSO prognostic model was validated in a randomly split 1:1 internal test set from the same TCGA cohort↳ Could also: External validation in an independent cohort, or repeated k-fold cross-validation, could also be used alongside or instead of a single random split — A single random split can have high variance depending on the random seed; k-fold cross-validation or an independent external cohort provides a more stable and generalisable estimate of model performance
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-37320872
Paper: Mou T et al. Heterogeneity of cancer-associated fibroblasts in head and neck squamous cell carcinoma. Transl Oncol 2023. PMID 37320872 / PMC10277597 / DOI 10.1016/j.tranon.2023.101717.
Named code (brief): https://github.com/aertslab/pySCENIC (a third-party tool, NOT the authors' own pipeline repo — the authors ship no own code). Named data (brief): GEO GSE139324.
Datasets the paper actually used (from Methods)
The paper integrates three scRNA-seq series + 1 spatial + TCGA:
| Accession | What it is | Contains CAFs? |
|---|---|---|
| GSE139324 (brief's accession) | Cillo et al. 2020 — CD45+ FACS-sorted immune cells (131,224 cells, 63 samples) | NO — sorted immune only; fibroblasts are CD45‑negative stroma |
| GSE173647 | HNSCC biopsies, 2 pts pre/post cetuximab (unsorted; raw 10x triplets) | YES (stroma present) |
| GSE173964 | 3 primary HNSCC tumors (unsorted; normalized mtx) | YES (stroma present) |
| GSE181300 | Visium spatial (8 sections) | spatial |
| TCGA-HNSC | 503 bulk RNA-seq (survival) | bulk |
Critical auditability note. The brief's accession GSE139324 is CD45+ sorted immune cells and contains essentially no fibroblasts/CAFs. The CAF result that defines this paper therefore cannot be derived from GSE139324; it must come from the unsorted series GSE173647 / GSE173964. We reproduce the pySCENIC step on an unsorted, CAF-containing dataset (GSE173647) — the paper's own data.
In scope (pipeline-derived, attempted)
- pySCENIC TF-regulon analysis of fibroblasts/CAFs (Fig 1F). The named tool
(pySCENIC) run on the paper's actual data. We reconstruct a fibroblast subset
from GSE173647 with a standard Seurat/Scanpy pipeline, then run
pyscenic grn → ctx → aucell(hg38 cisTarget v10 databases) and compare the recovered active regulons to the TF set shown in Figure 1F ("Top 5 active transcription factors of the seven CAF subsets").
In scope but NOT attempted (the hard ~20%, with reason)
- Exact 7 CAF subtypes + per-subtype top‑5 TFs. The subtyping depends on a 10-tool, under-specified workflow (integration of 3 series, CellTypist annotation, unspecified clustering resolution). Reproducing the exact 7 clusters and their per-cluster top‑5 regulons 1:1 is the hard 20% and is not pursued; pySCENIC GRN inference (GRNBoost2) is also stochastic.
Out of scope (not pipeline / not attempted)
- IHC of CKS2 in 52 patients (wet-lab).
- TCGA LASSO "CAFscore" 13-gene prognostic model, NMF immunotherapy clusters, hdWGCNA modules, BayesPrism deconvolution, CellPhoneDB L-R pairs, Copykat, scFEA/scMetabolism, Scissor, spatial (STUtility/Spacexr) — each a separate tool; outside the named-tool (pySCENIC) reproduction target and the 80/20 budget.
Claim-pinning limitation
The pySCENIC result (Fig 1F) is published only as a low-resolution heatmap
image — there is no machine-readable supplementary TF table. The reported
TF names were read by hand from the figure at high zoom and are recorded with a
confidence flag in original/claims.tsv.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.