Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Heterogeneity of cancer-associated fibroblasts in head and neck squamous cell carcinoma.

Transl Oncol · 2023
50/100 3/4
Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
How its reproducibility compares
50/100
Reproducibility score
1.4 SD below mean
vs. all fields · 1173 studies
🎯 Scores higher than 8% of all assessed papers rank 1026 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

PARTIAL reproduction (fresh re-run after requeue; «job»). The brief's accession GSE139324 is a CD45+ FACS-sorted immune atlas (Cillo 2020) and contains no fibroblasts, so the paper's defining CAF result is NOT derivable from it (flagged as a data/claim mismatch, not fabrication). I therefore applied the paper's named third-party tool (pySCENIC, P16-valid) to the paper's OWN unsorted CAF-containing series GSE173647: scanpy QC + leiden + relative-z lineage assignment yielded a 14,860-cell fibroblast subset, then pyscenic grn/ctx/aucell (hg38 cisTarget v10) recovered 166 active regulons. Overlap with the figure-only Fig 1F TF set is concordant but not 1:1: 5 of 19 hand-read confident TFs (CEBPG, CREB3, FOXM1, SOX4, TCF4; 8 of 23 incl. uncertain), several top-ranked by AUCell (SOX4 #23, FOXM1 #36, CREB3 #49). This independently confirms the earlier run («job», 189 regulons, 6/19): a stable core SOX4/FOXM1/CREB3/CEBPG recovers in BOTH runs, the remainder shuffles -> the non-exact match is GRNBoost2 stochasticity. Three understood reasons for partial vs exact: figure-only low-res reference with no supplementary table; stochastic GRNBoost2; single-series subset vs the paper's 3-series integration. NOT attempted: exact 7-subtype top-5 matching (hard 20%), and all non-pySCENIC tools (LASSO CAFscore, NMF, hdWGCNA, BayesPrism, CellPhoneDB, spatial) which are out of scope. INFRA NOTE: «infra»+/home user quota was exhausted (janitor clean_beegfs du-timeout failing every tick); cleared via a compute-node SLURM reclaim job (2220566) before compute could run. Honest partial.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 50
    assessed: 2026-06-20 ⛓ a291d5922bd2
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-23
Rubric version
not recorded
Assessed by
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The study tests whether cancer-associated fibroblasts (CAFs) in head and neck squamous cell carcinoma (HNSCC) are functionally heterogeneous subpopulations with distinct prognostic significance, immune-suppressive roles, and metabolic activity, particularly focusing on a CKS2+ inflammatory CAF (iCAF) subset.

Core claims
  • Seven distinct CAF subsets exist in HNSCC, identified via integration of scRNA-seq, bulk transcriptomic, and spatial transcriptomic data. finding
  • CKS2+ iCAFs are significantly associated with poor prognosis and are localized in close proximity to cancer cells. finding
  • CKS2+ iCAFs negatively correlate with cytotoxic CD8+ T cells and NK cells, but positively correlate with exhausted CD8+ T cells. finding
  • Patient clusters with high CKS2+ iCAFs (Cluster 3) or high CKS2- iCAFs/CENPF-/MYLPF- myCAFs (Cluster 2) show limited immunotherapeutic response. finding
  • Cancer cells show close intercellular communication with CKS2+ iCAFs and CENPF+ myCAFs via specific ligand-receptor pairs. finding
  • CKS2+ iCAFs display the highest metabolic activity among CAF subsets, notably in alanine/aspartate/glutamate metabolism, glycolysis/gluconeogenesis, and glycosaminoglycan biosynthesis. finding
  • A six-gene-module-based prognostic risk model was constructed from genes correlated with the CKS2+ iCAFs subset using LASSO regression. method
  • High infiltration of CKS2+ CAFs, validated by immunohistochemistry in a patient cohort, indicates worse overall survival. finding
Experimental setups
Assay System Perturbation Readout Platform
single-cell RNA sequencing (scRNA-seq) 24 HNSCC tissue samples (human) none cell type identity, cluster-specific transcripts, CAF subset marker expression Seurat; CellTypist database
bulk RNA-seq / TCGA transcriptomic analysis TCGA-HNSCC cohort, 503 samples with survival data none gene expression, survival/prognosis association, CAF subset proportion (deconvolution) TCGAbiolinks; BayesPrism
spatial transcriptomics HNSCC tissue sections (GSE181300) none spatial distribution of CAF subsets relative to epithelial cells STUtility; Spacexr
immunohistochemistry (IHC) 52 primary HNSCC patient specimens none CKS2+ CAF abundance scored and correlated with prognosis Polink2-plus reagent kit
regulon/transcription factor activity analysis CAF populations from scRNA-seq data none top 5 specific transcription factors per CAF subset pySCENIC
co-expression network analysis (WGCNA) fibroblast expression matrix (CKS2+ iCAFs subset) none gene modules correlated with prognosis, module eigengene connectivity (kME) hdWGCNA
cell-cell communication analysis epithelial cells (malignant/normal) and CAF subsets from scRNA-seq/TCGA none ligand-receptor interaction pairs, enrichment score of interactions CellPhoneDB; Copykat
metabolic flux prediction fibroblast/CAF subsets from scRNA-seq none activity of glucose and other metabolic pathways per CAF subset scFEA; scMetabolism
immune cell infiltration estimation (ssGSEA) and immunotherapy response prediction TCGA-HNSCC bulk samples none immune cell infiltration ratio; predicted immunotherapy response ssGSEA; TIDE; TCIA
Key results
  • Fibroblast and epithelial cell clusters showed significant poor prognosis by Scissor analysis
  • CKS2+ iCAFs and PTN+ myCAFs associated with poor prognosis; CENPF-/MYLPF- myCAFs associated with more favorable prognosis
  • CKS2+ iCAFs mainly distributed in areas of high epithelial cell concentration in spatial transcriptomic sections
  • Negative correlation between CKS2+ iCAFs and cytotoxic CD8+ T cells/NK cells; positive correlation with exhausted CD8+ T cells
  • Cluster 3 (high CKS2+ iCAFs) and Cluster 2 (high CKS2- iCAFs/CENPF-/MYLPF- myCAFs) showed no significant immunotherapeutic response
  • Close interactions confirmed between cancer cells and CKS2+ iCAFs/CENPF+ myCAFs with identified ligand-receptor pair
  • CKS2+ iCAFs showed highest metabolic activity among the seven CAF subsets
  • Patients with high CKS2+ CAF infiltration had poor overall survival (IHC validation)
Key statistics
  • count 24 HNSCC samples (GSE139324, GSE173647, GSE173964) (scRNA-seq sample source)
  • count 73,660 cells retained after QC (cells after removing >20% mitochondrial genome content)
  • count 2,000 variable genes selected (used for PCA in scRNA-seq and CAF subclustering)
  • count 503 TCGA samples with survival time (TCGA-HNSCC cohort used for prognostic analysis)
  • count 7 CAF subsets identified (CAF heterogeneity classification)
  • count 52 primary HNSCC patients (IHC validation cohort)
  • other top 25 genes by kME value selected from 6 modules (brown, yellow, midnight blue, pink, red, light cyan) (gene sets for prognostic model construction)
  • pvalue P < 0.05 (significance threshold; *P<0.05, **P<0.01, ***P<0.001) (statistical significance criteria used throughout, including CellPhoneDB ligand-receptor retention)
  • other z-score enrichment score > 1 (threshold indicating strong cell-cell interaction in bulk_ScrNA analysis)
  • other soft threshold power of 5 (used for topological overlap matrix (TOM) calculation in hdWGCNA)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This bioinformatics study integrated public scRNA-seq data (24 HNSCC samples, 73,660 cells), bulk RNA-seq data from TCGA (503 samples), spatial transcriptomic data, and an in-house IHC cohort (52 patients) to characterize CAF heterogeneity. Main statistical procedures included unsupervised clustering, Kaplan-Meier survival analysis with log-rank testing, LASSO-based prognostic model construction, Pearson correlation between deconvolved cell proportions, and ANOVA/chi-square tests for immunotherapy response comparisons. Significance was uniformly set at P<0.05 and reported as threshold categories (*, **, ***).

Replicationbiological Sample size24 scRNA-seq samples (three GEO datasets); 503 TCGA-HNSCC samples with survival data; 52 primary HNSCC patients for IHC; spatial transcriptomics from GSE181300 (sample count not stated) GroupsCAF subsets (7 types); IHC CKS2 high vs. low; TCGA risk score high vs. low; NMF immunotherapy clusters (3-4 groups) Pairingunpaired Randomization/blindingstated Dispersionnone Exact p-valuesno Effect sizesno Confidence intervalsno Multiplicity correctionnone stated
Statistical tests used
Test Applied to n Assumptions
Log-rank test Kaplan-Meier overall survival comparisons between CKS2+ CAF high vs. low IHC groups and TCGA risk groups 52 (IHC cohort); 503 (TCGA cohort) not stated
Student's t-test Inter-group differences in gene expression and deconvolved cell proportions null not stated
Pearson correlation Correlation between proportions of 7 CAF subpopulations and 3 CD8+ T cell/NK cell subsets in TCGA bulk data 503 not stated
Chi-square test Differences in patient immunotherapy response among NMF-defined CAF subtypes null not stated
One-way ANOVA Immunotherapy response scores across four NMF clusters null not stated
LASSO regression Feature selection and construction of prognostic risk model from top-kME genes of six hdWGCNA modules 503 split 1:1 train/test not stated
ROC analysis (AUC) Performance assessment of the LASSO-based prognostic model null na
CellPhoneDB permutation test (P<0.05 threshold) Ligand-receptor interaction inference between CAF subsets and epithelial cells null not stated
Approaches that could also have been used
  • Multiple independent t-tests were used for inter-group comparisons without a stated check of normality or homoscedasticity
    Could also: Wilcoxon rank-sum (Mann-Whitney U) test could also be applied, particularly for deconvolved proportions which are bounded [0,1] and often non-normally distributed — Non-parametric tests make no distributional assumption; they are commonly preferred when sample distributions are unknown or when values are compositional/proportional
  • P-values for all hypothesis tests were reported only as threshold categories (*, **, ***) rather than exact values
    Could also: Reporting exact P-values (e.g., P=0.023) is also standard and required by many journals — Exact P-values allow readers to apply their own thresholds and facilitate meta-analysis; threshold categories alone can mask borderline results
  • No multiple-testing correction was applied across the many simultaneous Pearson correlations (7 CAF subsets × 4 immune populations) and across ANOVA/t-test comparisons
    Could also: Benjamini-Hochberg FDR correction could also be applied across this family of tests — With ~28 simultaneous correlation tests, the expected number of false positives at P<0.05 is ~1.4; FDR control is a standard approach to characterize the false-discovery rate in such exploratory multi-test settings
  • Survival analysis used the log-rank test and Kaplan-Meier curves dichotomizing patients at the median CKS2 IHC score
    Could also: A Cox proportional hazards model treating the IHC score as a continuous variable (or using multivariable adjustment for known confounders such as stage and HPV status) could also be used — Cox regression retains the full information in a continuous score, avoids the arbitrariness of a median cut-point, and allows adjustment for clinical covariates that may confound the prognostic association
  • Effect sizes (e.g., hazard ratios, Pearson r values, odds ratios) were not reported alongside P-values
    Could also: Reporting effect size estimates with 95% confidence intervals alongside each P-value is also standard practice — Effect sizes convey clinical or biological magnitude independently of sample size; confidence intervals communicate estimation uncertainty and are requested by CONSORT, STROBE, and most statistical reporting guidelines
  • The LASSO prognostic model was validated in a randomly split 1:1 internal test set from the same TCGA cohort
    Could also: External validation in an independent cohort, or repeated k-fold cross-validation, could also be used alongside or instead of a single random split — A single random split can have high variance depending on the random seed; k-fold cross-validation or an independent external cohort provides a more stable and generalisable estimate of model performance
Software: R 4.1.3 · Seurat · pySCENIC · BayesPrism · hdWGCNA · NMF (brunet method) · CellPhoneDB · STUtility / Spacexr · scFEA / scMetabolism · Copykat · TCGAbiolinks / ssGSEA

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-37320872

Paper: Mou T et al. Heterogeneity of cancer-associated fibroblasts in head and neck squamous cell carcinoma. Transl Oncol 2023. PMID 37320872 / PMC10277597 / DOI 10.1016/j.tranon.2023.101717.

Named code (brief): https://github.com/aertslab/pySCENIC (a third-party tool, NOT the authors' own pipeline repo — the authors ship no own code). Named data (brief): GEO GSE139324.

Datasets the paper actually used (from Methods)

The paper integrates three scRNA-seq series + 1 spatial + TCGA:

Accession What it is Contains CAFs?
GSE139324 (brief's accession) Cillo et al. 2020 — CD45+ FACS-sorted immune cells (131,224 cells, 63 samples) NO — sorted immune only; fibroblasts are CD45‑negative stroma
GSE173647 HNSCC biopsies, 2 pts pre/post cetuximab (unsorted; raw 10x triplets) YES (stroma present)
GSE173964 3 primary HNSCC tumors (unsorted; normalized mtx) YES (stroma present)
GSE181300 Visium spatial (8 sections) spatial
TCGA-HNSC 503 bulk RNA-seq (survival) bulk

Critical auditability note. The brief's accession GSE139324 is CD45+ sorted immune cells and contains essentially no fibroblasts/CAFs. The CAF result that defines this paper therefore cannot be derived from GSE139324; it must come from the unsorted series GSE173647 / GSE173964. We reproduce the pySCENIC step on an unsorted, CAF-containing dataset (GSE173647) — the paper's own data.

In scope (pipeline-derived, attempted)

  • pySCENIC TF-regulon analysis of fibroblasts/CAFs (Fig 1F). The named tool (pySCENIC) run on the paper's actual data. We reconstruct a fibroblast subset from GSE173647 with a standard Seurat/Scanpy pipeline, then run pyscenic grn → ctx → aucell (hg38 cisTarget v10 databases) and compare the recovered active regulons to the TF set shown in Figure 1F ("Top 5 active transcription factors of the seven CAF subsets").

In scope but NOT attempted (the hard ~20%, with reason)

  • Exact 7 CAF subtypes + per-subtype top‑5 TFs. The subtyping depends on a 10-tool, under-specified workflow (integration of 3 series, CellTypist annotation, unspecified clustering resolution). Reproducing the exact 7 clusters and their per-cluster top‑5 regulons 1:1 is the hard 20% and is not pursued; pySCENIC GRN inference (GRNBoost2) is also stochastic.

Out of scope (not pipeline / not attempted)

  • IHC of CKS2 in 52 patients (wet-lab).
  • TCGA LASSO "CAFscore" 13-gene prognostic model, NMF immunotherapy clusters, hdWGCNA modules, BayesPrism deconvolution, CellPhoneDB L-R pairs, Copykat, scFEA/scMetabolism, Scissor, spatial (STUtility/Spacexr) — each a separate tool; outside the named-tool (pySCENIC) reproduction target and the 80/20 budget.

Claim-pinning limitation

The pySCENIC result (Fig 1F) is published only as a low-resolution heatmap image — there is no machine-readable supplementary TF table. The reported TF names were read by hand from the figure at high zoom and are recorded with a confidence flag in original/claims.tsv.

Figures / tables: Fig 1F
C1
Reported
scRNA-seq atlas uses GSE139324 (brief accession) as a single-cell HNSCC dataset
Reproduced
GSE139324 verified (NCBI series_matrix) as Cillo et al. 2020 CD45+ FACS-sorted IMMUNE atlas (63 GSMs). Real and complete, but contains no CD45-negative stroma, so the paper's CAF compartment cannot derive from it. CAFs reconstructed from the paper's own unsorted series GSE173647 (36,316 raw -> 35,940 post-QC -> 14,860 fibroblast/CAF cells, 21,261 genes).
partial
C4
Reported
pySCENIC top-5 active TF regulons of the 7 CAF subsets (Fig 1F; figure-only heatmap, no machine-readable supplementary TF table; reference TF set hand-read = 19 confident + 4 uncertain)
Reproduced
FRESH RUN («job»): pySCENIC 0.12.1 (grn[GRNBoost2 seed=777] -> ctx[cisTarget hg38 v10 clust] -> aucell) on GSE173647 fibroblasts (14,860 cells) recovered 166 active regulons. Overlap with hand-read Fig 1F set: 5/19 confident (CEBPG, CREB3, FOXM1, SOX4, TCF4); 8/23 including uncertain (+ETS1, RB1, TGIF1). Recovered TFs rank highly by mean AUCell among 166: SOX4 #23, FOXM1 #36, CREB3 #49, CEBPG #52, TCF4 #84. ROBUSTNESS: a prior independent run («job», 189 regulons) gave 6/19 confident; the stable core SOX4/FOXM1/CREB3/CEBPG recovers in BOTH runs while the rest shuffles -> direct evidence the non-1:1 match is driven by GRNBoost2 stochasticity, not error.
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

No assessment has been recorded yet.
🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

809.1 k
tokens (I/O) · 64.3 M incl. cache
316 min
runtime · 13.7 CPU-h
74.6 GB
peak RAM
4 (2 failed)
HPC jobs
hummel
machine