Clustering and machine learning-based integration identify cancer associated fibroblasts genes' signature in head and neck squamous cell carcinoma.
The main results reproduced, with only marginal, non-material deviations.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- Nothing in this column.
- 🟡Could not use the authors’ exact input data
- 🟡Reported values were only indirectly comparable
- 🟡A deviation arose in the data or preprocessing
- 🟡A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The deviation was non-trivial in magnitude
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
DESCRIBED WELL ENOUGH for the signature core; partial 1:1. Reproduced the ssGSEA 'CafS' (31 CAF genes) and its correlations on the brief-pinned public cohort GSE65858 (n=270 Illumina HumanHT-12; the paper's reported values are on a merged 868-sample combined cohort, so this is an independent-cohort check, not an identical-matrix recompute). All 31 CAF genes present. 3 of 4 in-scope correlation claims reproduce WITHIN TOLERANCE on the independent cohort: EMT r=0.862 vs 0.87, Angiogenesis r=0.819 vs 0.82 (near-exact, delta<=0.008), PDGFB r=0.629 vs 0.61. The CSF1R correlation MISMATCHES: r~0.07 (non-significant) vs reported 0.33 - the weakest reported immune association is the one that fails to reproduce (flagged). ClassDiscovery 2-subtype structure is consistent (imposed k=2 -> clean 114/156 split) but not independently re-derived. NOT ATTEMPTED (hard last 20%, out of scope): the 9-gene Mime/RSF prognostic model (train C-index 0.95, AUC 0.93/0.99/0.97) - trained model weights not deposited, non-deterministic 107-combination ML search, and those numbers are flagged as overfit-pattern (0.95 train vs 0.61 val); and the scRNA-seq GSE164690/irGSEA single-cell branch (separate heavy pipeline). Note: the paper's 'code' link github.com/chuiqin/irGSEA is a THIRD-PARTY scRNA GSEA tool, not the authors' signature/prognostic code, which ships no runnable code - reproduction was done by re-running the described ssGSEA pipeline on the named public data (valid per P16).
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 63assessed: 2026-06-15 ⛓ 77d1bcb676b3
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-15
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetCancer-associated fibroblasts (CAFs) drive an immunosuppressive, invasion-promoting tumor microenvironment in head and neck squamous cell carcinoma, and a CAFs-gene-derived signature/score can stratify patient prognosis and reveal targetable mechanisms.
- ★ Clustering of 31 CAFs genes across 868 HNSCC samples identifies two distinct molecular patterns (C1, C2) with different survival outcomes finding
- ★ A ssGSEA-based CAFs score (CafS) quantifies CAFs gene expression and serves as an independent prognostic factor method
- ★ High CafS is associated with immunosuppression (increased TAM/macrophage infiltration, reduced cytotoxic T cells), worse prognosis, and higher likelihood of HPV-negative status finding
- ★ High CafS/C1 cluster shows enrichment of carcinogenic pathways including angiogenesis, epithelial-mesenchymal transition, and coagulation finding
- ★ MDK and NAMPT ligand-receptor crosstalk between CAFs and other cell clusters may mechanistically drive immune escape mechanism
- ★ A random survival forest (RSF) model, selected from 107 machine learning algorithm combinations across 10 algorithms, most accurately and stably classifies HNSCC patient prognosis method
- Single-cell RNA-seq identifies CAFs (FAP, MMP11, PDGFRA, PDGFRB, ADAMTS2, SFRP2) and endothelial cell (PLVAP, KDR, PTPRB) clusters among 18 total clusters in HNSCC tumors resource
- The 31 CAFs genes are functionally enriched in extracellular matrix organization, collagen-containing ECM, endopeptidase activity, and focal adhesion pathway finding
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| ssGSEA scoring | HNSCC patient tumor samples (TCGA-HNSC, GSE65858, GSE41613, Combined cohort) | none | CafS score quantifying 31 CAFs genes expression | R (v4.1.2) |
| Unsupervised clustering (ClassDiscovery) | Combined HNSCC cohort (TCGA + GSE41613 + GSE65858) | none | CAFs gene expression pattern clusters (C1/C2) | R package ClassDiscovery |
| Immune infiltration estimation (ssGSEA and CIBERSORT) | Combined HNSCC bulk RNA-seq cohort | none | Abundance of 28 (ssGSEA) and 24 (CIBERSORT) immune cell types | CIBERSORT algorithm; R ESTIMATE package for StromalScore |
| GSVA hallmark pathway enrichment | Combined HNSCC bulk RNA-seq cohort | none | Pathway enrichment scores (EMT, coagulation, angiogenesis, hypoxia, etc.) and correlation with CafS | GSVA, MSigDB hallmark gene sets |
| Single-cell RNA sequencing | GSE164690 HNSCC single-cell cohort, 5 patients, CD45-negative cells | none | Cell cluster identity (18 clusters), CAFs/endothelial marker expression, ligand-receptor cellular crosstalk | Seurat, SingleR, CellChat, UMAP |
| Deconvolution of bulk RNA-seq using single-cell reference | TCGA-HNSC bulk cohort with GSE164690 scRNA-seq reference | none | Estimated proportion of CAFs in bulk tumor tissue | MuSiC |
| LASSO and 10-algorithm/107-combination machine learning survival modeling | TCGA-HNSC, GSE41613, GSE65858, Combined cohort; external validation GSE42743 | none | Concordance index (C-index) of prognostic risk models | RSF, Enet, Lasso, Ridge, stepwise Cox, CoxBoost, plsRcox, SuperPC, GBM, survival-SVM |
| Quantitative real-time PCR | HNSCC cell lines CAL-27, FaDu vs. normal nasopharyngeal epithelial cell line NP69 | none (cell line comparison) | mRNA expression of PGAM1 and ENO1 (normalized to β-actin) | cDNA Synthesis Kit, qPCR |
- ▼ C1 cluster showed significantly worse survival than C2 cluster in combined cohort log-rank p=0.018
- ▼ High CafS group had worse survival across Combined, TCGA, GSE65858, GSE41613, and external GSE42743 cohorts log-rank p=0.013, 0.036, 0.00028, 0.049, 0.0072
- ▲ CafS was significantly higher in HPV-negative than HPV-positive patients (TCGA and GSE65858) p=6e-10 (TCGA), p=0.002 (GSE65858)
- ▲ CafS positively correlated with tumor invasion-related pathways (EMT, coagulation, angiogenesis) and metabolism/DNA-damage pathways EMT r=0.87; Coagulation r=0.82; Angiogenesis r=0.82; Hypoxia r=0.53; UV-response-down r=0.60
- ▲ CafS positively correlated with tumor-associated macrophage (TAM) markers CCL2, CLEC7A, CSF1, CSF1R, PDGFB r=0.28 to r=0.61 (PDGFB r=0.61 highest)
- ▲ StromalScore was significantly higher in C1 than C2 cluster p=1.61e-36
- ▲ Random survival forest (RSF) model achieved the highest average C-index across four validation datasets among 107 ML algorithm combinations
- – Single-cell analysis of 10,244 cells from 5 HNSCC patients identified 18 clusters including CAFs and endothelial cells based on canonical markers 10,244 cells; 18 clusters
- correlation r=0.87, p=1.77e-268 (CafS vs. EMT pathway score, combined cohort)
- correlation r=0.61, p=1.28e-88 (CafS vs. PDGFB (TAM marker) expression, combined cohort)
- pvalue p=0.018 (Log-rank survival difference between C1 and C2 clusters, combined cohort)
- pvalue p<2.22e-16 (CafS score difference between C1 and C2 clusters)
- pvalue p=6e-10 (CafS difference between HPV-positive and HPV-negative patients, TCGA cohort)
- pvalue p=1.61e-36 (StromalScore difference between C1 and C2 groups)
- count 10,244 cells; 18 clusters (Single-cell RNA-seq of CD45-negative cells from 5 HNSCC patients (GSE164690))
- count 868 samples (Total HNSCC samples pooled from TCGA-HNSC, GSE65858, and GSE41613 cohorts)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This study applied unsupervised clustering (ClassDiscovery) to 868 HNSCC samples across three public cohorts to derive two CAF molecular subtypes, then built a continuous CAF activity score (CafS) via ssGSEA for survival stratification and immune profiling. Kaplan-Meier log-rank tests were the primary inferential tool for survival comparisons across five cohorts; Pearson correlations linked CafS to immune markers and hallmark pathways. A prognostic model was developed by screening 107 machine-learning algorithm combinations (anchored on LASSO feature selection) and selecting the combination with the highest average concordance index (C-index) across four cohorts, with cell-line RT-qPCR providing experimental confirmation of two candidate genes.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Log-rank test (Kaplan-Meier) | Survival comparison C1 vs C2 clusters (combined cohort p=0.018, GSE42743 p=0.034) and high-CafS vs low-CafS (combined p=0.013, TCGA p=0.036, GSE41613 p=0.00028, GSE65858 p=0.049, GSE42743 p=0.0072) | Combined cohort n=868; individual cohort n not separately stated | not stated |
| Two-sample t-test | CafS score comparison between C1 and C2 clusters (Figure 4A, p<2.22e-16) | n=868 (combined cohort) | not stated |
| Pearson correlation | CafS vs five TAM markers (CCL2, CLEC7A, CSF1, CSF1R, PDGFB) and CafS vs five hallmark pathways (EMT, coagulation, angiogenesis, hypoxia, UV-response-down) | n=868 (combined cohort) | not stated |
| Wilcoxon rank-sum test | Stated as the general method for continuous and ordered categorical variable comparisons (Statistical Analysis section); implied for HPV status, TMB, smoking, and tumor stage comparisons | null | not stated |
| ssGSEA (single-sample gene set enrichment analysis) | Construction of CafS score from 31 CAF genes; quantification of 28 immune cell populations (Bindea et al. set) | n=868 | na |
| CIBERSORT deconvolution algorithm | Differential immune cell infiltration (24 fractions) between C1 and C2 clusters (Figure 5B) | n=868 | na |
| GSVA (gene set variation analysis) | Hallmark signaling pathway enrichment comparison across high- and low-CafS groups (Figure 6A) | n=868 | na |
| Concordance index (C-index) comparison across 107 machine-learning algorithm combinations (RSF, Enet, Lasso, Ridge, stepwise Cox, CoxBoost, plsRcox, SuperPC, GBM, survival-SVM) | Prognostic model selection across four cohorts (GSE41613, GSE65858, TCGA-HNSC, combined); optimal model validated in GSE42743 | n=868 across four cohorts; GSE42743 n not stated | na |
-
Pearson correlation was used to link CafS (an ssGSEA-derived enrichment score) to immune markers and pathway scores↳ Could also: Spearman rank correlation could also be used for the same associations — ssGSEA scores are bounded, non-normally distributed enrichment estimates; Spearman rank correlation requires neither linearity nor normality and is commonly preferred for gene-expression-derived composite scores
-
A two-sample t-test was used to compare CafS between C1 and C2 clusters↳ Could also: A Wilcoxon rank-sum (Mann-Whitney U) test could also be applied for this comparison — ssGSEA scores may not follow a normal distribution; a rank-based test is distribution-free and is a common alternative for enrichment-score group comparisons in computational genomics
-
CafS was dichotomized into high and low groups using a data-derived optimal cut-off to generate the survival groups used in Figures 4B–F↳ Could also: A pre-specified cut-off (e.g., median split) or CafS modeled as a continuous predictor in a Cox proportional hazards regression could also be used — Selecting an optimal cut-off from the same data used for testing can inflate the type I error rate; using the score continuously in Cox regression avoids information loss from dichotomization and reduces optimization bias
-
Multiple log-rank tests were conducted across five cohorts and several grouping variables without stated correction for multiple comparisons↳ Could also: A Bonferroni or Benjamini-Hochberg FDR correction applied across the family of survival comparisons could also be used — Performing many survival tests on related groupings increases the chance of spurious significant findings; adjusting for multiple comparisons is standard in multi-cohort survival studies and would allow readers to assess the overall false discovery rate
-
Immune cell infiltration was compared between high- and low-CafS groups across 28 ssGSEA populations and 24 CIBERSORT fractions with significance denoted by asterisks (*/**/***) without stated multiplicity adjustment↳ Could also: Benjamini-Hochberg FDR correction across all simultaneous immune-fraction comparisons could also be used — With 28–52 simultaneous comparisons, applying an FDR procedure is a standard approach to limit the expected proportion of spurious findings in immune deconvolution analyses
-
The best prognostic model was selected by comparing average C-index across four partially overlapping cohorts (three of which were also used to develop the combined training set)↳ Could also: k-fold cross-validation strictly within a designated training cohort, followed by evaluation in a fully held-out test cohort, could also be used — Selecting among 107 algorithm combinations based on performance in cohorts that overlap with training data conflates model selection with external validation; a clean train/tune/test partition or nested cross-validation more rigorously estimates generalization performance
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-37065499
Paper: Wang Q, Zhao Y, Wang F, Tan G. Clustering and machine learning-based integration identify cancer associated fibroblasts genes' signature in head and neck squamous cell carcinoma. Front Genet 2023. PMID 37065499 · PMC10098459 · DOI 10.3389/fgene.2023.1111816.
Pipeline structure (as described in Methods)
- 31 CAF genes taken from prior literature (Kürten et al. scRNA-seq;
Chakravarthy et al. bulk RNA-seq). Fixed input list — NOT a pipeline output.
List (verbatim from paper, with obvious typos noted):
GAPDH, ENO1, ITGA6, PGK1, TGFBI, ACTN1, FTH1, KDELR2, CD82, SSR3, A2M, PHLDA1, TSC22D1, ISG15, PRSS23, PGAM1, SFRP2, PDGFRB, CEBPB, TNFRSF12A, MMP9, SNAI2(="SNA12"), ADAMTS2, MMP11, MMP12, COL4A6, STEAP1, ITGAX, ADAMTS14, TLL1, COL4A4(31 genes). - CafS = single-sample GSEA (ssGSEA) score of those 31 genes per sample. Tool: ssGSEA (paper does not name the package; GSVA is the canonical R impl).
- ClassDiscovery consensus clustering of the 31-gene matrix → two CAF subtypes C1 / C2 (k=2; no distance/linkage stated).
- Mime-style ML integration: 10 algorithms × 107 combinations → an optimal Random Survival Forest model on 9 genes (ENO1, TSC22D1, ISG15, PGAM1, SFRP2(="SFEP2"), PDGFRB, ITGAX, ADAMTS14, TLL1) + CafS → prognostic risk score.
- Downstream: survival (KM), AUC, immune/pathway correlation, single-cell (GSE164690) analysis using irGSEA (the GitHub link in the brief — a THIRD-PARTY single-cell GSEA tool, NOT the authors' own pipeline code; used only for the single-cell pathway step).
Datasets named
- TCGA-HNSC (training), GSE65858 / GSE41613 / GSE42743 (bulk validation), GSE164690 (scRNA-seq). "Combined cohort" = merged bulk (868 samples).
- Brief pins GSE65858 (Wichmann 2015, public, RNA-seq, n≈270, has survival).
IN SCOPE (clearly-specified, public-data, deterministic — the 80%)
The reproducible core is the ssGSEA CafS score and its correlations/clustering, which need only a public expression matrix + a standard tool (GSVA) + MSigDB Hallmark gene sets. ssGSEA is rank-based → robust to the exact normalization.
| id | reported result | what we reproduce |
|---|---|---|
| C1 | CafS vs EMT correlation r = 0.87 (p=1.77e-268) | ssGSEA CafS vs Hallmark EMT ssGSEA, Pearson r |
| C2 | CafS vs Angiogenesis r = 0.82 (p=3.19e-215) | ssGSEA CafS vs Hallmark Angiogenesis ssGSEA, Pearson r |
| C3 | CafS vs PDGFB r = 0.61; CafS vs CSF1R r = 0.33 | gene-level Pearson against CafS |
| C4 | ClassDiscovery → 2 CAF subtypes C1/C2 | consensus clustering (k=2) of 31-gene matrix; cluster sizes |
We compute on the brief-pinned GSE65858 (and, if feasible, TCGA-HNSC) as a single-cohort stand-in for the "combined cohort". Reported correlations are on the merged 868-sample cohort; we therefore test whether the reported strong correlation is an intrinsic, reproducible property rather than reproducing the exact merge. This is stated honestly, not as 1:1 on the identical matrix.
OUT OF SCOPE / hard last 20% (NOT attempted — with reason)
- 9-gene RSF prognostic model AUC/C-index (train C-index 0.95; val 0.61; AUC 1/3/5y up to 0.93/0.99/0.97). Requires the trained Mime/RSF model (weights NOT deposited), the exact batch-corrected combined cohort, and a non-deterministic 107-combination ML search. Not byte-reproducible; the headline AUCs are also implausibly high (train C-index 0.95) — flagged for the auditor.
- Single-cell GSE164690 / irGSEA (18 clusters, 10,244 cells, CAF proportion KM p=0.0029) — separate heavy scRNA-seq pipeline; peripheral to the signature.
Possible-fabrication / red flags for the human auditor
- Train C-index 0.95 with validation 0.61 is a large generalization gap (classic overfit / optimistic training metric). AUCs of 0.93–0.99 at 1–5 yr are unusually high for HNSCC prognosis. Worth scrutiny.
- The "code" link (irGSEA) is not t
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
The CAF-signature core reproduces well: EMT (r=0.862 vs 0.87), Angiogenesis (r=0.819 vs 0.82) and PDGFB (r=0.629 vs 0.61) all land within tolerance on an independent cohort (GSE65858, n=270), strong evidence the correlations are intrinsic — this is mainly a data-availability limitation (the authors' merged n=868 matrix was never deposited), i.e. our-method/data side, not an authors' defect. The notable failure is CSF1R (reported r=0.33 → 0.07, n.s.), a significance flip on the weakest reported association, plausibly cohort/batch-driven but not derivable here. The prognostic model (C-index 0.95, AUC up to 0.99) is uncheckable and overfit-flagged because no trained model was deposited, and the cited code is a third-party tool. Overall: solid, explainable, partial reproduction — substantive but not fabrication-grade.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.