A novel oxidative stress-related gene signature as an indicator of prognosis and immunotherapy responses in HNSCC.
The main results reproduced: recomputed values matched the published ones within tolerance.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Reported values were directly comparable
- ✓No authors-side cause for any deviation
- ✓Reported values are derivable from the shared data
- ✓Any deviation was negligible
- ✓The central claim held under reproduction
- 🟡Could not use the authors’ exact input data
- 🟡A deviation arose in the data or preprocessing
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Described well enough -> 1:1 (approximate). The paper's headline is a 9-gene oxidative-stress prognostic risk score whose exact coefficients are printed in the Results; there is no authors' code repo (the registry pySCENIC pointer covers only the single-cell regulon sub-analysis). We applied the PUBLISHED formula (no LASSO refit) to the paper's own public data on «our HPC» and reproduced all 5 checkable pipeline metrics within tolerance: significant KM separation in BOTH cohorts (TCGA p<1e-7, GSE41613 p=0.045/0.0037), AUCs ~0.65-0.73 vs reported 0.67-0.74, and the multivariate Cox HR (z-scored 2.54 vs reported 2.55, near-exact). The z-scored variant matches reported numbers best, implying the authors standardized expression (unstated). Honest substitutions: GDC Xena hub (paper's FPKM source) now 403s, so TCGA expression came from the legacy Xena HiSeqV2 hub (different normalization) and used 7/9 genes (JCHAIN/IGJ and FDCSP absent from that matrix); GSE41613 used all 9/9, n=97 exact. NOT attempted (skipped 20%): de-novo LASSO gene selection (CV-random, no seed), the scRNA-seq arm (GSE103322 clustering/CellChat/pySCENIC regulons), the immunotherapy arm (PRJEB23709/TIDE), ssGSEA/CIBERSORT/GSVA immune descriptives, and IHC/HPA/GEPIA2 protein validation (wet-lab/manual). No fabrication signal: every reported value reproduces in direction and approximate magnitude.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 85assessed: 2026-06-15 ⛓ c2ac1937afa7
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-15
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15no human curator yet
- Last updated
- 2026-09-19
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetThe study tests whether oxidative stress-related gene expression patterns can define molecular subtypes of HNSCC that correlate with the tumor immune microenvironment and can be used to build a prognostic/immunotherapy-response scoring model.
- ★ Consensus clustering of 74 oxidative stress-related genes identifies three distinct HNSCC molecular subgroups (Cluster1-3) with significantly different overall survival and tumor stage distribution finding
- ★ A nine-gene oxidative stress-related risk scoring model (OSRS), built from LASSO-Cox regression on DEGs between the subgroups, predicts overall survival in HNSCC method
- ★ AREG and CES1 are prognostic risk factors, while CSTA, FDCSP, JCHAIN, IFFO2, PGLYRP4, SPOCK2 and SPINK6 are protective prognostic factors in the OSRS model finding
- ★ The OSRS model was validated using independent datasets (GSE41613, GSE103322, PRJEB23709) method
- ★ Immunohistochemical staining of SPINK6 in nasopharyngeal cancer clinical samples supports the gene panel finding
- ★ OSRS subgroups differ in cell-cell communication, immune microenvironment composition, transcription factor activity, and predicted immunotherapy/drug response finding
- 77 oxidative stress-related genes were retrieved from the Harmonizome Biocarta Pathways database, of which 74 were expressed in the training cohort resource
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| bulk RNA-seq (FPKM) + clinical survival data, consensus clustering | TCGA-HNSCC tumor samples (n=494) | none | oxidative stress gene expression subgroups, overall survival | TCGAbiolinks / ConsensusClusterPlus |
| bulk RNA-seq + survival data (validation cohort) | GSE41613 HNSCC tumor samples (n=97) | none | validation of OSRS prognostic signature | GEO |
| single-cell RNA-seq | GSE103322, 18 primary HNSCC tumors, 5902 cells (2205 malignant) | none | malignant cell subgroups, trajectory (pseudotime), cell-cell communication | Harmony batch correction, monocle2, CellChat |
| bulk RNA-seq, immunotherapy cohort | PRJEB23709 (BioProject) | immune checkpoint inhibitor (ICI) treatment | predictive efficacy of OSRS for immunotherapy response | — |
| Immunohistochemistry (IHC) | FFPE nasopharyngeal squamous cell carcinoma clinical samples (n=10) | none | SPINK6 protein expression/IHC score | Leica Bond HRP-conjugated SP system; rabbit anti-human SPINK6 antibody |
| Drug sensitivity prediction (IC50 estimation) | TCGA-HNSCC cohort (in silico) | none | correlation of OSRS with predicted IC50 values | oncoPredict (GDSC, CTRP databases) |
| Transcription factor regulatory network / regulon activity (SCENIC, AUCell) | TCGA-HNSCC / scRNA-seq malignant cells | none | transcription factor activity scores, regulon specificity | RcisTarget, SCENIC, AUCell |
| Tumor immune microenvironment deconvolution (ssGSEA, CIBERSORT, xCell, ESTIMATE) | TCGA-HNSCC tumor samples | none | immune cell infiltration fractions, immune/stromal/purity scores by OSRS group | CIBERSORT/LM22, IOBR (xCell), ESTIMATE |
- ▼ Three oxidative stress-related subgroups identified (Cluster1 n=197, Cluster2 n=140, Cluster3 n=157); Cluster1 had significantly poorer overall survival
- – 216 DEGs identified among the three oxidative stress expression patterns
- – 22 of the DEGs were significantly associated with OS by univariate Cox regression
- – LASSO-Cox regression selected 9 predictive genes forming the OSRS model
- ▼ High-risk OSRS group had significantly shorter overall survival than low-risk group log-rank p<0.001
- – OSRS model accurately predicted 1-, 3-, and 5-year overall survival AUC = 0.694 (1yr), 0.692 (3yr), 0.673 (5yr)
- – SPINK6 IHC staining in nasopharyngeal cancer samples validated the gene panel
- pvalue log-rank p < 0.001 (OS difference between high- and low-OSRS risk groups, TCGA-HNSCC)
- other AUC = 0.694, 0.692, 0.673 (1-, 3-, 5-year OS prediction accuracy of OSRS model)
- count 216 DEGs (DEGs among the three oxidative stress-related clusters)
- count 22 prognostic DEGs (DEGs significantly associated with OS by univariate Cox regression)
- count 9 genes (final predictive genes selected by LASSO regression for OSRS model)
- count three clusters, n=197/140/157 (sample distribution across oxidative stress consensus clusters)
- fold_change |log2FC| ≥ 1, FDR < 0.05 (DEG screening threshold criteria)
- count 74 of 77 genes expressed (oxidative stress-related genes from Harmonizome used in analysis)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
The study used retrospective bulk and single-cell RNA-seq data from public repositories (TCGA-HNSCC training cohort n=494; GSE41613, PRJEB23709, and GSE103322 validation cohorts) to construct an oxidative stress-related gene scoring model for HNSCC prognosis. Unsupervised consensus clustering identified three molecular subgroups; DEGs between subgroups were filtered by limma, then prognostic genes were selected sequentially via univariate Cox regression (p<0.05) and LASSO-Cox regression with 10-fold cross-validation, yielding a nine-gene risk score. The score was dichotomized at the cohort median and evaluated with Kaplan-Meier/log-rank tests, time-dependent ROC curves, and uni/multivariate Cox models; downstream immune and drug-sensitivity analyses used Wilcoxon rank-sum, Kruskal-Wallis, and Spearman correlation tests.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Log-rank test | OS comparisons between consensus clustering subgroups (Cluster1/2/3) and between high- vs low-OSRS groups in training and validation cohorts | 494 (TCGA training); 97 (GSE41613 validation) | not stated |
| Univariate Cox proportional hazard regression | Screening 216 DEGs for OS association; 22 genes retained at p<0.05 | 494 | not stated |
| LASSO-Cox regression with 10-fold cross-validation | Selection of nine key prognostic genes from 22 univariate-significant DEGs; penalty parameter lambda chosen by cross-validation | 494 | not stated |
| Multivariate Cox proportional hazard regression | Assessing independent prognostic value of OSRS alongside clinical covariates; nomogram construction | 494 | not stated |
| limma moderated t-test (DEG analysis) | Identifying differentially expressed genes between three oxidative stress consensus clusters; |log2FC|≥1, FDR<0.05 | 494 | not stated |
| Wilcoxon rank-sum test | Inter-group comparisons of immune cell infiltration scores (ssGSEA, CIBERSORT, xCell, ESTIMATE) between high- and low-OSRS groups | 494 | not stated |
| Kruskal-Wallis test | Multiple-group comparisons across three consensus clustering subgroups | 494 | not stated |
| Spearman rank correlation | Correlation between OSRS and predicted IC50 values (drug sensitivity) from GDSC/CTRP databases | 494 | not stated |
| Time-dependent ROC / AUC (timeROC) | Discriminative performance of the OSRS model for predicting OS at 1, 3, and 5 years (AUC: 0.694, 0.692, 0.673) | 494 | na |
-
The OSRS risk score was dichotomized at the cohort median into high- and low-risk groups for all survival comparisons↳ Could also: The risk score could also be retained as a continuous predictor in Cox models, or an optimal cutpoint method (e.g., maximally selected rank statistics via the 'maxstat' R package) could be used — Median dichotomization is reproducible and widely understood, but analyzing the score continuously avoids information loss and does not depend on the distribution of the specific cohort; reporting both approaches side-by-side is common in prognostic signature papers
-
Univariate Cox regression (p<0.05) was used as a pre-filter on 216 DEGs before applying LASSO-Cox to the retained 22 genes↳ Could also: LASSO-Cox (or elastic net Cox) could also be applied directly to all DEGs without a prior univariate significance filter — Pre-filtering on univariate p-values can introduce selection bias when predictors are correlated, because marginal significance does not reflect joint importance; penalized regression applied to the full candidate set lets the regularization handle multicollinearity directly
-
The optimal number of consensus clusters (k=3) was selected using the relative change in area under the CDF curve↳ Could also: The gap statistic, average silhouette width, or prediction strength could also be used to guide cluster number selection — Different internal validity indices can favor different values of k; reporting multiple indices together, or showing results for a range of k values, can strengthen confidence in the selected solution
-
Multiple Wilcoxon rank-sum and Spearman correlation tests were performed across many immune cell types and drug compounds without a stated multiplicity correction for this family of tests↳ Could also: Benjamini-Hochberg FDR correction applied across the full family of simultaneous comparisons (e.g., all 28 immune cell types together) would also control the expected false discovery rate — When many tests are performed in parallel on related endpoints, even a modest FDR correction reduces the probability that a nominally significant individual result is a false positive; this is particularly relevant given the large number of immune cell comparisons
-
AUC values for 1-, 3-, and 5-year survival prediction were reported as point estimates only↳ Could also: Time-dependent AUC with 95% bootstrap confidence intervals (directly available within the timeROC package) could also be reported — Confidence intervals around AUC convey the precision of each estimate and allow informal or formal comparison against a null AUC of 0.5, or against alternative prognostic models applied to the same cohort
-
Immune cell infiltration was estimated in parallel using three independent methods (ssGSEA, CIBERSORT/LM22, and xCell)↳ Could also: A single deconvolution method could be designated as the primary endpoint a priori, with the other two serving as pre-specified sensitivity analyses — Designating a primary method before analysis makes the inferential plan explicit; reporting all three as co-equal results is informative for triangulation but benefits from a clear statement of which method was selected a priori and why
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-38157249
Paper: Li Z et al. "A novel oxidative stress-related gene signature as an indicator of prognosis and immunotherapy responses in HNSCC." Aging (Albany NY) 2023. PMID 38157249 · PMCID PMC10781479 · DOI 10.18632/aging.205323.
Full text obtained via Europe PMC fulltext API (PMC10781479) — open access, CC BY. No captcha. (PMC HTML itself is captcha-walled; Europe PMC mirror is clean.)
What kind of paper this is
A bulk + single-cell bioinformatics signature paper built on public data:
- TCGA-HNSCC (training, 494 tumor samples; FPKM expression + OS).
- GSE41613 (bulk validation, 97 oral-SCC tumors, HG-U133 Plus 2.0 + OS).
- GSE103322 (scRNA-seq, 18 tumors / 2205 malignant cells) — single-cell arm.
- PRJEB23709 (immunotherapy cohort) — ICI-response arm.
Methods named in text: limma (DEGs), survival/survminer (KM, log-rank), glmnet (LASSO-Cox), timeROC (AUC), ConsensusClusterPlus, GSVA/clusterProfiler, ssGSEA, CIBERSORT, TIDE, CellChat, pySCENIC (regulon/RSS, single-cell arm).
NOTE on the registry code pointer: the room's
code_urlisgithub.com/aertslab/pySCENIC. pySCENIC IS genuinely used — but only for the single-cell transcription-factor regulon (RSS/CSI) sub-analysis, NOT for the headline prognostic signature. There is no authors' own code repo; the "reproduction" here is the brief-sanctioned P16 case: apply the published signature (with its printed coefficients) to the paper's own public data and check whether the reported survival metrics reproduce.
IN SCOPE (clearly-specified, low-hanging — what we attempt)
The 9-gene Oxidative-Stress Risk Score (OSRS) is printed verbatim with its coefficients (Results, "Prognostic signatures"):
Score = -SPOCK2*0.096 - JCHAIN*0.044 - CSTA*0.0004 - SPINK6*0.106
+ AREG*0.123 - FDCSP*0.025 - IFFO2*0.053 + CES1*0.023
- PGLYRP4*0.058
Because the exact coefficients are given, we DO NOT need to re-derive them by LASSO (which depends on CV-fold randomness — that is the hard last 20%, skipped). We apply the published formula to each cohort, median-split into high/low risk, and reproduce these reported numbers:
| # | Result | Reported | Cohort | Location |
|---|---|---|---|---|
| C1 | KM high-vs-low OS log-rank p | < 0.001 | TCGA-HNSC | Fig 3A |
| C2 | time-dependent AUC 1/3/5-yr | 0.694 / 0.692 / 0.673 | TCGA-HNSC | Fig 3B |
| C3 | multivariate Cox HR for risk score | 2.55 (p<0.001) | TCGA-HNSC | Fig 3G |
| C4 | KM high-vs-low OS (significant) | high-risk shorter OS | GSE41613 | Fig 4A |
| C5 | time-dependent AUC 1/3/5-yr | 0.739 / 0.692 / 0.681 | GSE41613 | Fig 4B |
Pipelines: log2(FPKM)/array intensity -> linear risk score -> survival (KM, log-rank) + timeROC AUC + Cox. All run on «our HPC».
OUT OF SCOPE (not attempted, with reason)
- De-novo LASSO gene selection (which 9 of the 22 univariate genes survive): depends on 10-fold CV randomness + unstated seed -> the hard 20%, skipped.
- scRNA-seq arm (GSE103322: clustering, CellChat, pySCENIC regulons/RSS): heavy, exploratory, no single pinnable number; out of 80/20 budget.
- Immunotherapy arm (PRJEB23709 / TIDE): separate cohort + raw-seq processing.
- ssGSEA / CIBERSORT immune-infiltration boxplots, GSVA, nomogram: descriptive, no single hard reported value cheaply checkable.
- IHC / HPA / GEPIA2 protein validation: wet-lab + manual, non-pipeline.
Honesty note
Exact bit-reproduction is not expected: the paper does not state the expression
transform (raw FPKM vs log2), the median-split tie handling, or the timeROC
cause/weighting. We report our values 1:1 next to the reported ones and grade
provisionally (exact / within-tol / partial / mismatch). The human auditor decides.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
The paper's headline 9-gene oxidative-stress risk score reproduces well: applying the verbatim published coefficients to the paper's own public cohorts recovers significant KM separation in both (TCGA p=3.4e-08, GSE41613 p=0.045/0.0037), AUCs of ~0.65–0.73 vs reported 0.67–0.74, and a multivariate Cox HR of 2.54 vs the reported 2.55. The deviations are minor and sit on our/data-availability side: the GDC Xena FPKM hub now 403s, so TCGA came from a legacy hub with different normalization and only 7/9 genes, and the best-matching z-scored variant reflects an unstated standardization choice. No fabrication signal — every reported value is derivable in direction and approximate magnitude. Overall yellow: the core claim is fully confirmed but the reproduction is approximate rather than a 1:1 bit-match due to explainable host/preprocessing substitutions.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at [email protected].
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.