A novel oxidative stress-related gene signature as an indicator of prognosis and immunotherapy responses in HNSCC.
The main results reproduced: recomputed values matched the published ones within tolerance.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Reported values were directly comparable
- ✓No authors-side cause for any deviation
- ✓Reported values are derivable from the shared data
- ✓Any deviation was negligible
- ✓The central claim held under reproduction
- 🟡Could not use the authors’ exact input data
- 🟡A deviation arose in the data or preprocessing
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Described well enough -> 1:1 (approximate). The paper's headline is a 9-gene oxidative-stress prognostic risk score whose exact coefficients are printed in the Results; there is no authors' code repo (the registry pySCENIC pointer covers only the single-cell regulon sub-analysis). We applied the PUBLISHED formula (no LASSO refit) to the paper's own public data on «our HPC» and reproduced all 5 checkable pipeline metrics within tolerance: significant KM separation in BOTH cohorts (TCGA p<1e-7, GSE41613 p=0.045/0.0037), AUCs ~0.65-0.73 vs reported 0.67-0.74, and the multivariate Cox HR (z-scored 2.54 vs reported 2.55, near-exact). The z-scored variant matches reported numbers best, implying the authors standardized expression (unstated). Honest substitutions: GDC Xena hub (paper's FPKM source) now 403s, so TCGA expression came from the legacy Xena HiSeqV2 hub (different normalization) and used 7/9 genes (JCHAIN/IGJ and FDCSP absent from that matrix); GSE41613 used all 9/9, n=97 exact. NOT attempted (skipped 20%): de-novo LASSO gene selection (CV-random, no seed), the scRNA-seq arm (GSE103322 clustering/CellChat/pySCENIC regulons), the immunotherapy arm (PRJEB23709/TIDE), ssGSEA/CIBERSORT/GSVA immune descriptives, and IHC/HPA/GEPIA2 protein validation (wet-lab/manual). No fabrication signal: every reported value reproduces in direction and approximate magnitude.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 85assessed: 2026-06-15 ⛓ c2ac1937afa7
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-15
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusThere may be a correlation between oxidative stress-related genes and the tumor microenvironment in HNSCC; the study tests whether oxidative stress-related genes can be used to define molecular subtypes and build a scoring model predicting HNSCC prognosis and immunotherapy responses.
- ★ A nine-gene oxidative stress-related scoring (OSRS) model (AREG, CES1, CSTA, FDCSP, JCHAIN, IFFO2, PGLYRP4, SPOCK2, SPINK6) predicts overall survival in HNSCC. resource
- ★ AREG and CES1 are prognostic risk factors, while CSTA, FDCSP, JCHAIN, IFFO2, PGLYRP4, SPOCK2 and SPINK6 are protective prognostic factors. finding
- ★ Consensus clustering of 74 oxidative stress-related genes identifies three HNSCC subgroups (Cluster1/2/3) with significantly different prognosis, Cluster1 having the poorest OS. finding
- ★ The OSRS model is validated in external datasets GSE41613, GSE103322 and the PRJEB23709 immunotherapy cohort, and SPINK6 IHC staining of nasopharyngeal cancer samples validates the panel. method
- ★ Oxidative stress prognostic signature subgroups are associated with cellular communication, immune microenvironment composition, transcription factor activation and immunotherapy responses. mechanism
- ★ High OSRS risk score is associated with significantly shorter overall survival in the TCGA-HNSCC cohort. finding
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| bulk RNA-seq (FPKM expression profiling) for consensus clustering, DEG, Cox/LASSO prognostic modeling | TCGA-HNSCC, 494 tumor samples (training cohort) | none | gene expression, overall survival, risk score | — |
| bulk transcriptome expression and survival validation | GSE41613 HNSCC, 97 tumor samples | none | gene expression, overall survival | — |
| single-cell RNA-seq (scRNA-seq) with clustering, trajectory (monocle2), cell communication (CellChat), TF network (SCENIC) | GSE103322 HNSCC, 18 primary tumor samples, 5902 cells (2205 malignant) | none | cell type identity, malignant subgroups, TF regulon activity, cell-cell communication | — |
| immunotherapy response prediction | PRJEB23709 immunotherapy cohort, 77 oxidative stress-related genes | immune checkpoint inhibitor treatment | predictive efficacy of signature for immunotherapy | — |
| immunohistochemistry (IHC) | nasopharyngeal squamous cell carcinoma, 10 FFPE patient samples | none | SPINK6 protein staining / IHC score | rabbit anti-human SPINK6 polyclonal antibody (Cusabio CSB-PA744263LA01HU), Leica Bond System |
| immune cell infiltration / TME estimation (ssGSEA, CIBERSORT/LM22, xCell, ESTIMATE) | TCGA-HNSCC tumor samples | none | relative abundance of immune cell types, immune/stromal/purity scores | — |
| drug sensitivity prediction (oncoPredict calcPhenotype) | TCGA-HNSCC training cohort | drug | IC50 values correlated with OSRS | GDSC and CTRP databases |
- ▼ High-risk OSRS group had significantly shorter overall survival in TCGA-HNSCC log-rank p < 0.001
- – OSRS model ROC AUCs for 1-, 3-, and 5-year survival AUC = 0.694, 0.692, 0.673
- – Consensus clustering identified three subgroups with Cluster1 showing poorer OS n = 197/140/157
- – 22 of the DEGs were significantly associated with OS by univariate Cox regression 22 genes
- – Nine most predictive OS factors selected by LASSO-Cox from the 22 genes 9 genes
- – Significant difference in tumor staging between oxidative stress subgroups p < 0.05
- – SPINK6 IHC scores in nasopharyngeal carcinoma samples validated the gene panel IHC score 160 in 5/10 patients
- pvalue log-rank p < 0.001 (OS difference between high- and low-risk OSRS groups in TCGA)
- other AUC 0.694, 0.692, 0.673 (1-, 3-, 5-year ROC AUC of OSRS model)
- count three subgroups n = 197/140/157 (consensus clustering of TCGA-HNSCC)
- count 216 DEGs (DEGs among three oxidative stress expression patterns)
- count 22 genes (DEGs significantly associated with OS by univariate Cox)
- count 74 of 77 genes expressed (oxidative stress-related genes from Harmonizome expressed in training cohort)
- count 5902 cells; 2205 malignant (single-cell transcriptome after QC in GSE103322)
- mean 65.2 ± 12.917 (age of nasopharyngeal carcinoma IHC patients)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
The study used retrospective bulk and single-cell RNA-seq data from public repositories (TCGA-HNSCC training cohort n=494; GSE41613, PRJEB23709, and GSE103322 validation cohorts) to construct an oxidative stress-related gene scoring model for HNSCC prognosis. Unsupervised consensus clustering identified three molecular subgroups; DEGs between subgroups were filtered by limma, then prognostic genes were selected sequentially via univariate Cox regression (p<0.05) and LASSO-Cox regression with 10-fold cross-validation, yielding a nine-gene risk score. The score was dichotomized at the cohort median and evaluated with Kaplan-Meier/log-rank tests, time-dependent ROC curves, and uni/multivariate Cox models; downstream immune and drug-sensitivity analyses used Wilcoxon rank-sum, Kruskal-Wallis, and Spearman correlation tests.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Log-rank test | OS comparisons between consensus clustering subgroups (Cluster1/2/3) and between high- vs low-OSRS groups in training and validation cohorts | 494 (TCGA training); 97 (GSE41613 validation) | not stated |
| Univariate Cox proportional hazard regression | Screening 216 DEGs for OS association; 22 genes retained at p<0.05 | 494 | not stated |
| LASSO-Cox regression with 10-fold cross-validation | Selection of nine key prognostic genes from 22 univariate-significant DEGs; penalty parameter lambda chosen by cross-validation | 494 | not stated |
| Multivariate Cox proportional hazard regression | Assessing independent prognostic value of OSRS alongside clinical covariates; nomogram construction | 494 | not stated |
| limma moderated t-test (DEG analysis) | Identifying differentially expressed genes between three oxidative stress consensus clusters; |log2FC|≥1, FDR<0.05 | 494 | not stated |
| Wilcoxon rank-sum test | Inter-group comparisons of immune cell infiltration scores (ssGSEA, CIBERSORT, xCell, ESTIMATE) between high- and low-OSRS groups | 494 | not stated |
| Kruskal-Wallis test | Multiple-group comparisons across three consensus clustering subgroups | 494 | not stated |
| Spearman rank correlation | Correlation between OSRS and predicted IC50 values (drug sensitivity) from GDSC/CTRP databases | 494 | not stated |
| Time-dependent ROC / AUC (timeROC) | Discriminative performance of the OSRS model for predicting OS at 1, 3, and 5 years (AUC: 0.694, 0.692, 0.673) | 494 | na |
-
The OSRS risk score was dichotomized at the cohort median into high- and low-risk groups for all survival comparisons↳ Could also: The risk score could also be retained as a continuous predictor in Cox models, or an optimal cutpoint method (e.g., maximally selected rank statistics via the 'maxstat' R package) could be used — Median dichotomization is reproducible and widely understood, but analyzing the score continuously avoids information loss and does not depend on the distribution of the specific cohort; reporting both approaches side-by-side is common in prognostic signature papers
-
Univariate Cox regression (p<0.05) was used as a pre-filter on 216 DEGs before applying LASSO-Cox to the retained 22 genes↳ Could also: LASSO-Cox (or elastic net Cox) could also be applied directly to all DEGs without a prior univariate significance filter — Pre-filtering on univariate p-values can introduce selection bias when predictors are correlated, because marginal significance does not reflect joint importance; penalized regression applied to the full candidate set lets the regularization handle multicollinearity directly
-
The optimal number of consensus clusters (k=3) was selected using the relative change in area under the CDF curve↳ Could also: The gap statistic, average silhouette width, or prediction strength could also be used to guide cluster number selection — Different internal validity indices can favor different values of k; reporting multiple indices together, or showing results for a range of k values, can strengthen confidence in the selected solution
-
Multiple Wilcoxon rank-sum and Spearman correlation tests were performed across many immune cell types and drug compounds without a stated multiplicity correction for this family of tests↳ Could also: Benjamini-Hochberg FDR correction applied across the full family of simultaneous comparisons (e.g., all 28 immune cell types together) would also control the expected false discovery rate — When many tests are performed in parallel on related endpoints, even a modest FDR correction reduces the probability that a nominally significant individual result is a false positive; this is particularly relevant given the large number of immune cell comparisons
-
AUC values for 1-, 3-, and 5-year survival prediction were reported as point estimates only↳ Could also: Time-dependent AUC with 95% bootstrap confidence intervals (directly available within the timeROC package) could also be reported — Confidence intervals around AUC convey the precision of each estimate and allow informal or formal comparison against a null AUC of 0.5, or against alternative prognostic models applied to the same cohort
-
Immune cell infiltration was estimated in parallel using three independent methods (ssGSEA, CIBERSORT/LM22, and xCell)↳ Could also: A single deconvolution method could be designated as the primary endpoint a priori, with the other two serving as pre-specified sensitivity analyses — Designating a primary method before analysis makes the inferential plan explicit; reporting all three as co-equal results is informative for triangulation but benefits from a clear statement of which method was selected a priori and why
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-38157249
Paper: Li Z et al. "A novel oxidative stress-related gene signature as an indicator of prognosis and immunotherapy responses in HNSCC." Aging (Albany NY) 2023. PMID 38157249 · PMCID PMC10781479 · DOI 10.18632/aging.205323.
Full text obtained via Europe PMC fulltext API (PMC10781479) — open access, CC BY. No captcha. (PMC HTML itself is captcha-walled; Europe PMC mirror is clean.)
What kind of paper this is
A bulk + single-cell bioinformatics signature paper built on public data:
- TCGA-HNSCC (training, 494 tumor samples; FPKM expression + OS).
- GSE41613 (bulk validation, 97 oral-SCC tumors, HG-U133 Plus 2.0 + OS).
- GSE103322 (scRNA-seq, 18 tumors / 2205 malignant cells) — single-cell arm.
- PRJEB23709 (immunotherapy cohort) — ICI-response arm.
Methods named in text: limma (DEGs), survival/survminer (KM, log-rank), glmnet (LASSO-Cox), timeROC (AUC), ConsensusClusterPlus, GSVA/clusterProfiler, ssGSEA, CIBERSORT, TIDE, CellChat, pySCENIC (regulon/RSS, single-cell arm).
NOTE on the registry code pointer: the room's
code_urlisgithub.com/aertslab/pySCENIC. pySCENIC IS genuinely used — but only for the single-cell transcription-factor regulon (RSS/CSI) sub-analysis, NOT for the headline prognostic signature. There is no authors' own code repo; the "reproduction" here is the brief-sanctioned P16 case: apply the published signature (with its printed coefficients) to the paper's own public data and check whether the reported survival metrics reproduce.
IN SCOPE (clearly-specified, low-hanging — what we attempt)
The 9-gene Oxidative-Stress Risk Score (OSRS) is printed verbatim with its coefficients (Results, "Prognostic signatures"):
Score = -SPOCK2*0.096 - JCHAIN*0.044 - CSTA*0.0004 - SPINK6*0.106
+ AREG*0.123 - FDCSP*0.025 - IFFO2*0.053 + CES1*0.023
- PGLYRP4*0.058
Because the exact coefficients are given, we DO NOT need to re-derive them by LASSO (which depends on CV-fold randomness — that is the hard last 20%, skipped). We apply the published formula to each cohort, median-split into high/low risk, and reproduce these reported numbers:
| # | Result | Reported | Cohort | Location |
|---|---|---|---|---|
| C1 | KM high-vs-low OS log-rank p | < 0.001 | TCGA-HNSC | Fig 3A |
| C2 | time-dependent AUC 1/3/5-yr | 0.694 / 0.692 / 0.673 | TCGA-HNSC | Fig 3B |
| C3 | multivariate Cox HR for risk score | 2.55 (p<0.001) | TCGA-HNSC | Fig 3G |
| C4 | KM high-vs-low OS (significant) | high-risk shorter OS | GSE41613 | Fig 4A |
| C5 | time-dependent AUC 1/3/5-yr | 0.739 / 0.692 / 0.681 | GSE41613 | Fig 4B |
Pipelines: log2(FPKM)/array intensity -> linear risk score -> survival (KM, log-rank) + timeROC AUC + Cox. All run on «our HPC».
OUT OF SCOPE (not attempted, with reason)
- De-novo LASSO gene selection (which 9 of the 22 univariate genes survive): depends on 10-fold CV randomness + unstated seed -> the hard 20%, skipped.
- scRNA-seq arm (GSE103322: clustering, CellChat, pySCENIC regulons/RSS): heavy, exploratory, no single pinnable number; out of 80/20 budget.
- Immunotherapy arm (PRJEB23709 / TIDE): separate cohort + raw-seq processing.
- ssGSEA / CIBERSORT immune-infiltration boxplots, GSVA, nomogram: descriptive, no single hard reported value cheaply checkable.
- IHC / HPA / GEPIA2 protein validation: wet-lab + manual, non-pipeline.
Honesty note
Exact bit-reproduction is not expected: the paper does not state the expression
transform (raw FPKM vs log2), the median-split tie handling, or the timeROC
cause/weighting. We report our values 1:1 next to the reported ones and grade
provisionally (exact / within-tol / partial / mismatch). The human auditor decides.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
The paper's headline 9-gene oxidative-stress risk score reproduces well: applying the verbatim published coefficients to the paper's own public cohorts recovers significant KM separation in both (TCGA p=3.4e-08, GSE41613 p=0.045/0.0037), AUCs of ~0.65–0.73 vs reported 0.67–0.74, and a multivariate Cox HR of 2.54 vs the reported 2.55. The deviations are minor and sit on our/data-availability side: the GDC Xena FPKM hub now 403s, so TCGA came from a legacy hub with different normalization and only 7/9 genes, and the best-matching z-scored variant reflects an unstated standardization choice. No fabrication signal — every reported value is derivable in direction and approximate magnitude. Overall yellow: the core claim is fully confirmed but the reproduction is approximate rather than a 1:1 bit-match due to explainable host/preprocessing substitutions.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.