Comprehensive analysis of mitophagy in HPV-related head and neck squamous cell carcinoma.
The main results reproduced, with only marginal, non-material deviations.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- Nothing in this column.
- 🟡Could not use the authors’ exact input data
- 🟡Reported values were only indirectly comparable
- 🟡A deviation arose in the data or preprocessing
- 🟡A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The deviation was non-trivial in magnitude
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Described well enough to attempt; NO own code (the 'code' link github.com/tidyverse/ggplot2 is the generic plotting library), so reproduction = re-running the described R/Bioconductor pipeline on the paper's own public data (GEO GSE65858, TCGA-HNSC). RESULT = partial. A1 EXACT: GSE65858 = 60 HPV16+/196 negative, matching the paper's 60/196 to the sample. A2 WITHIN-TOL/PARTIAL: TCGA-HNSC clinical HPV via TCGAbiolinks gives p16 41+/74-; the paper's 72 HPV- = within-tol (74), its 31 HPV+ lies between p16+ (41) and ISH+ (22), most consistent with the RNA-seq-available subset. B (76 up/321 down TCGA DESeq2 DEGs) NOT completed: the R env built and the TCGA STAR-counts download started, but the run aborted on a one-line script bug (fwrite with=FALSE on a data.frame) before the DE step; not rerun per the operator's finalize-now instruction + 80/20. NOT attempted: LASSO risk model (C), timeROC AUC (D), consensus-clustering/GSVA/CIBERSORTx (E) = the model/immune 'last 20%' (seed/CV/web-server dependent), and qRT-PCR (F) = wet-lab. FLAGS for human auditor: (1) the paper reports IDENTICAL training and validation timeROC AUCs (0.792/0.696/0.898) — possible copy error/fabrication; (2) no analysis code shipped.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 73assessed: 2026-06-15 ⛓ f523a0754b0c
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-15
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusThe authors hypothesize that high-risk HPV infection alters the mitophagy process in head and neck squamous cell carcinoma (HNSCC), and that this contributes to the better prognosis, distinct immune cell infiltration, and tumour development seen in HPV-associated HNSCC compared to non-HPV-associated HNSCC.
- ★ In HPV-associated HNSCC, the mitophagy process affects tumour development, immune cell infiltration and prognosis. finding
- ★ Hub genes NOS2, IL17REL, TMSB15A and TUBB4A show significantly higher expression in HPV-related HNSCC than in non-HPV-related HNSCC. finding
- ★ HPV-related HNSCC has a unique immunological characteristic of high CD8+ T cell expression. finding
- ★ A mitophagy regulatory scoring model (risk score from DMBX1, KRT2, NPPC) predicts patient prognosis/overall survival. method
- ★ Mitophagy-regulating genes SLC25A4, SLC25A5, TP53, TSC2 and USP30 are highly expressed in HPV-positive HNSCC samples. finding
- ★ Higher mitophagy score is associated with reduced infiltration of CD4-positive memory T cells and M2 macrophages. finding
- TP53, VPS13C, VPS13D and AMBRA1 carry mutations in HPV-related HNSCC. finding
- Integration of TCGA/GEO bioinformatics with qRT-PCR experimental validation characterizes the mitophagy signature of HPV-associated HNSCC. resource
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Bulk RNA-seq gene expression profiling / differential expression analysis | Human TCGA-HNSCC tissue samples (72 HPV-negative, 31 HPV-positive) | none (HPV status grouping by p16) | Differentially expressed genes (76 up, 321 down); mitophagy regulatory gene expression | TCGA database (TCGAbiolinks v2.22.4, DESeq2 v1.34.0) |
| Microarray/gene expression profiling (validation cohort) | Human GSE65858 GEO samples (196 HPV-negative/p16-negative, 60 HPV-positive/p16-positive) | none (HPV status grouping) | Gene expression for score model construction/validation | GEO database (GEOquery) |
| qRT-PCR | Human tissue: 6 HPV-negative HNSCC and 6 normal tonsil tissue | none (disease vs normal) | mRNA expression of NOS2, IL17REL, TMSB15A, TUBB4A (normalized to GAPDH, 2^-ΔΔCt) | StepOnePlus PCR system, Thermo Fisher; ChamQ Universal SYBR qPCR Master Mix |
| Somatic mutation analysis (waterfall/lollipop) | 31 HPV-positive TCGA-HNSCC samples | none | Mutation frequency and sites in TP53, VPS13C, VPS13D, AMBRA1 | GDC software / TCGA |
| Immune cell infiltration deconvolution | TCGA + GEO HPV-positive HNSCC training/validation expression matrices | none | Proportion of 22 immune cell types by mitophagy score group | CIBERSORTx with LM22 signature matrix |
| Pathway/gene set enrichment analysis (GO, KEGG, GSEA, GSVA, ssGSEA) | TCGA-HNSCC mitophagy subtype samples (subtype A vs B) | none | Enriched biological processes and pathways | clusterProfiler v4.2.2, GSVA v25, limma v3.50.1, MSigDB c2.cp.kegg |
| Prognostic scoring model (univariate Cox + LASSO regression, Kaplan-Meier, time-dependent ROC) | Combined TCGA + GSE65858 HPV-positive HNSCC | none | Risk score, overall survival, AUC | glmnet v4.1.3, survival v3.2.13, timeROC v0.4, SVA (batch removal) |
| PCA dimensionality reduction / consensus clustering | TCGA-HNSCC tissue samples | none | Sample separation by HPV status; mitophagy subtypes (70 A, 33 B) | rgl v0.107.10, ConsensusClusterPlus v1.62.0 |
- ▼ NOS2, IL17REL, TMSB15A and TUBB4A expression significantly lower in non-HPV-related HNSCC tissue than in normal tonsil tissue by qRT-PCR P < 0.05
- ▲ SLC25A4, SLC25A5, TP53, TSC2 and USP30 highly expressed in HPV-positive HNSCC samples
- – Differential expression analysis between mitophagy subtypes yielded 76 upregulated and 321 downregulated DEGs logFC >1 or <-1, adj P < 0.05
- ▼ CD4-positive memory T cells and M2 macrophages reduced in higher mitophagy-score group
- – TP53 had 4 mutation points; VPS13C, VPS13D and AMBRA1 each had 1 mutation point in HPV-positive samples TP53: 4 sites; others: 1 site each
- – DEGs enriched in muscle system processes, muscle contraction, keratinization, sarcomere, and cardiac muscle contraction / calcium signalling pathways FDR < 0.05
- – Prognostic genes affected KEGG ribosome, focal adhesion, spliceosome, DNA replication and HALLMARK E2F targets, EMT, G2M checkpoint, MYC targets V1, myogenesis FDR < 0.25
- ▲ HPV-related HNSCC characterized by high-level CD8+ T cell expression
- count 72 HPV-negative and 31 HPV-positive TCGA-HNSCC tissue samples (TCGA cohort after removing invalid samples)
- count 196 HPV-negative (p16-) and 60 HPV-positive (p16+) samples (GSE65858 GEO validation cohort)
- count 76 upregulated and 321 downregulated DEGs (391 total) (DESeq2 differential analysis between mitophagy subtypes A and B)
- other risk score = (DMBX1*0.118814319) + (KRT2*-0.000610081) + (NPPC*0.004718129) (LASSO/Cox-derived mitophagy prognostic risk score formula)
- pvalue P < 0.05 (qRT-PCR significance for NOS2, IL17REL, TMSB15A, TUBB4A; CIBERSORTx filter)
- count 12 functional mitophagy regulatory genes (GO:1901524) (gene set used to build mitophagy score)
- count 70 type A and 33 type B samples (TCGA samples regrouped by mitophagy regulatory genes)
- other logFC > 1 / < -1 and adj P < 0.05 (DEG up/down-regulation thresholds)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This study combined retrospective transcriptomic analysis of TCGA (n=103 HNSCC samples) and GEO (GSE65858; n=256 samples) datasets with qRT-PCR validation on 6 paired tissue specimens to characterise mitophagy-related gene differences between HPV-positive and HPV-negative HNSCC. Differential expression was identified via DESeq2 following unsupervised consensus clustering into mitophagy subtypes; pathway enrichment was assessed with GSEA, GSVA, and ClusterProfiler. A three-gene prognostic risk score was constructed by univariate Cox screening followed by LASSO regression, and survival differences between high- and low-score groups were evaluated with Kaplan-Meier log-rank tests and time-dependent ROC AUC.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| DESeq2 Wald test (negative-binomial GLM, adjusted P value) | Differential expression between consensus-cluster mitophagy subtypes A and B in TCGA-HNSCC; logFC > 1 or < -1 and adj P < 0.05 as thresholds | 70 subtype-A vs 33 subtype-B (re-classified from 103 TCGA samples) | not stated |
| Univariate Cox proportional-hazards regression | Screening association between each differential gene's expression and overall survival in combined HPV-positive dataset; P < 0.1 retention threshold | 31 TCGA HPV-positive + 60 GEO HPV-positive = 91 samples combined after batch correction | not stated |
| LASSO-penalized regression (glmnet, family = binomial) for subtype scoring model | Classification scoring model built from DEGs derived from mitophagy subtypes | null | not stated |
| LASSO-penalized Cox regression (glmnet) for prognostic model | Variable selection in training set from combined HPV-positive dataset; produced three-gene risk score (DMBX1, KRT2, NPPC) with Cox coefficients | null | not stated |
| Log-rank test with Kaplan-Meier analysis | Overall survival comparison between high-risk-score and low-risk-score groups in validation set | null | not stated |
| Time-dependent ROC / AUC (timeROC package) | Prognostic discrimination accuracy of the three-gene risk score in training and validation sets | null | na |
| Independent Student's t-test | Comparison of normally distributed continuous variables between two groups; stated as the method for qRT-PCR gene expression comparisons (P < 0.05 threshold) | n=6 per group for qRT-PCR | not stated |
| Mann-Whitney U test (Wilcoxon rank-sum test) | Comparison of non-normally distributed continuous variables between two groups | null | not stated |
| GSEA (gene set enrichment analysis; c2.cp.kegg.v6.2 and HALLMARK gene sets from MSigDB) | Enriched pathways and hallmark gene sets across mitophagy subtypes A and B in TCGA-HNSCC; FDR < 0.25 threshold | 103 TCGA-HNSCC samples re-classified into subtypes | not stated |
| GSVA (gene set variation analysis) with limma differential enrichment | Per-sample pathway enrichment scoring and differential enrichment between subtypes; FDR < 0.05 | 103 TCGA-HNSCC samples | not stated |
| CIBERSORTx deconvolution (LM22 signature matrix, permutation-based P value) | Estimation of 22 immune cell infiltration proportions in training and validation sets; samples with P >= 0.05 excluded | null | not stated |
| ConsensusClusterPlus consensus clustering | Unsupervised subtype identification of HPV-related HNSCC based on 12 mitophagy regulatory gene expression; k=2 selected | 103 TCGA-HNSCC samples | not stated |
-
The qRT-PCR experiment was designed with matched pairs (each HNSCC tissue case paired one-to-one with a normal tonsil tissue case, divided into 6 groups), but the statistical analysis section describes an independent Student's t-test for normally distributed variables↳ Could also: A paired Student's t-test or Wilcoxon signed-rank test on the matched pairs could also be used — When the experimental design explicitly creates matched pairs processed together, a paired test accounts for within-pair correlation introduced by RNA extraction batch or handling; this can increase statistical power by removing nuisance variability attributable to the pairing factor, which is particularly relevant at n=6 per group
-
Univariate Cox regression was used to pre-screen genes (P < 0.1) before applying LASSO Cox regression for prognostic model construction↳ Could also: LASSO-penalized Cox regression applied directly to all candidate differential genes without a prior univariate screening step could also be used — Applying LASSO directly lets the regularization parameter alone govern variable inclusion, which is its intended behavior; a two-stage filter (univariate screen then LASSO) can introduce selection bias because genes narrowly excluded by the univariate threshold never enter the LASSO, potentially omitting predictors whose signal only emerges in a multivariate context
-
Consensus clustering was used to define mitophagy subtypes (A and B), and DESeq2 differential expression was then applied between those same subtypes in the same TCGA dataset↳ Could also: Deriving subtype labels from one dataset (e.g., TCGA) and applying them to an independent held-out dataset (e.g., GEO) for differential expression or survival analysis could also be used — When clustering and downstream hypothesis testing are performed on the same samples, effect estimates and associations can be optimistic because the clustering step has already separated samples by their expression; applying labels to an independent cohort provides an assessment of how well the subtypes and associated gene signatures generalize
-
SVA (surrogate variable analysis) was used to remove batch effects when integrating the TCGA RNA-seq and GEO microarray/RNA-seq datasets↳ Could also: ComBat-seq (designed for raw RNA-seq count data) or limma's removeBatchEffect (for normalized log-expression matrices) could also be applied depending on the data type of each source — ComBat-seq operates on raw counts before normalization and preserves count-data distributional properties appropriate for DESeq2 downstream; the choice among batch-correction methods can affect differential expression results, and a sensitivity comparison across methods is sometimes reported to assess robustness
-
qRT-PCR results were reported using only asterisk notation (P < 0.05) without a numerical dispersion measure or exact P values↳ Could also: Reporting individual data points overlaid on bar or box plots with mean ± SD (or median with IQR) and exact P values could also be used — With n=6 per group, individual data points allow readers to see the full distribution and identify outliers; SD conveys absolute biological variability while exact P values allow independent re-analysis and meta-analysis; many journals' reporting guidelines (e.g., Nature Research, ICMJE) recommend this for small-n experiments
-
The prognostic risk score was evaluated using time-dependent AUC point estimates without accompanying confidence intervals↳ Could also: Bootstrapped confidence intervals for the time-dependent AUC, or a C-index with 95% CI, could also be reported alongside the point estimate — A point AUC estimate alone conveys the central tendency of model discrimination but not its precision; confidence intervals are particularly informative for small validation cohorts (the HPV-positive validation set is a subset of 91 samples) where the AUC estimate may be unstable, and allow readers to assess whether the performance interval overlaps with chance (0.5)
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-37161060
Paper: Yanan L. et al. (2023) Comprehensive analysis of mitophagy in HPV-related head and neck squamous cell carcinoma. Sci Rep 13:7480. PMID 37161060 / PMC10170109 / DOI 10.1038/s41598-023-34698-4.
Code availability — IMPORTANT
The registry code_url is https://github.com/tidyverse/ggplot2. This is not the
authors' analysis code — it is the generic ggplot2 plotting library. The paper ships
no own analysis repository and no analysis scripts. It is a secondary
("dry-lab") bioinformatic re-analysis of public data using named R/Bioconductor packages.
So reproduction = re-running the described pipeline on the paper's own public data,
per HARD RULE 2 (third-party tool on the paper's data is equally valid). There is no
repo to pin; we pin tool versions the paper states instead.
Data used by the paper (all public)
- TCGA-HNSC (RNA-seq counts + clinical HPV status): 31 HPV+ / 72 HPV− (103 with HPV status).
- GEO GSE65858 (Illumina microarray; Wichmann et al. HNSCC): 60 p16+ / 196 p16− (256 of 270).
Pipeline-derived results (candidate claims)
| id | result | tool | in scope? |
|---|---|---|---|
| A1 | GSE65858 sample composition 60 p16+/196 p16− | GEOquery metadata | YES — clean metadata count |
| A2 | TCGA-HNSC HPV split 31+/72− | TCGAbiolinks clinical | YES — clean metadata count |
| B | TCGA HPV+ vs HPV− DEGs: 76 up + 321 down (|logFC|>1, padj<0.05) | DESeq2 v1.34.0 | YES — best-effort pipeline output |
| C | LASSO risk model: riskscore = DMBX1·0.1188 + KRT2·(−0.00061) + NPPC·0.00472 | glmnet | OUT (hard 20%; seed/CV dependent) |
| D | timeROC AUC: 1/3/4-yr = 0.792/0.696/0.898 | timeROC | OUT — also a red flag: train & validation AUCs reported identical |
| E | Consensus clustering k=2, GSVA mitophagy score, CIBERSORTx immune (15/22 cells) | various | OUT (downstream of model; 20%) |
| F | qRT-PCR validation of NOS2/IL17REL/TMSB15A/TUBB4A | wet-lab | OUT — not a pipeline result |
What we attempt
A1, A2 (clean metadata 1:1) and B (best-effort DESeq2). We explicitly do not attempt C/D/E (the prognostic-model "last 20%", seed/CV/CIBERSORTx-dependent) or F (wet-lab). No completeness claim.
Red flags to record
- No analysis code shipped (ggplot2 placeholder).
- Train-set and validation-set timeROC AUCs reported as numerically identical (0.792/0.696/0.898) — possible copy error or fabrication; flagged for the human auditor.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
On the testable input claims the reproduction is strong: GSE65858 60 HPV16+/196 negative matches exactly, and the TCGA HPV- count (74 vs reported 72) is within tolerance. The deviations sit on the input/cohort side — the TCGA HPV+ count (31 vs 41 p16+/22 ISH+) is undefined by the authors and had to be self-redefined — rather than in core statistics we recomputed. The heaviest output claim (76/321 DESeq2 DEGs) was not completed (a one-line script bug aborted the run), and the prognostic model/AUCs were out of scope, so the central conclusion is neither confirmed nor refuted. Two authors'-side concerns are recorded for audit: no analysis code is shipped (the 'code' link is ggplot2) and the identical train/validation timeROC AUCs look fabrication-suspect, though both are unverified here.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.