Identifying Molecular Subtypes and 6-Gene Prognostic Signature Based on Hypoxia for Optimizing Targeted Therapies in Non-Small Cell Lung Cancer.
The main results reproduced, with only marginal, non-material deviations.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Any deviation was negligible
- ✓The central claim held under reproduction
- 🟡Could not use the authors’ exact input data
- 🟡Reported values were only indirectly comparable
- 🟡A deviation arose in the data or preprocessing
- 🔴A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Partial reproduction. The repo (jianwem/pro2022) is the authors' own R script but NOT self-contained: it ships no input data (all read from the first author's local e:/ drive or an unshipped raw_dat/ tree) and leaves 5 core helper functions (Sig_cox, DEGs_limma, estimate_score, plot_RS_model, ggsurvplotKM) undefined, so the TCGA-NSCLC discovery pipeline (consensus clustering -> LASSO -> stepAIC, incl. the exact 6-gene coefficients and Fig 9A) is NOT bit-reproducible. Per the brief we reproduced the clearly-specified public-data validation on the listed accession GEO:GSE31210 (GPL570, 226 LUAD tumors, 35 OS events) on «our HPC» (R 4.5.3 / GEOquery 2.78.0 / survival / timeROC; «job»). Applying the paper's published 6-gene risk formula gives a highly significant high- vs low-risk OS split (log-rank P=8.9e-06, HR=5.82), reproducing the paper's validation claim of P<0.0001 (Fig 9B, C3=within-tol); time-dependent AUC 0.67/0.65/0.74 at 1/3/5 yr. All 6 signature genes map in GSE31210 (C1); the 2-subtype design is confirmed in code (C4). NOT reproduced / out of scope: the exact LASSO/stepAIC coefficient VALUES (C2 partial -- un-rederivable from shipped artifacts, an auditability gap rather than evidence of fabrication; refit signs match 3/6) and the TCGA discovery KM (C5, drop_reason docs_insufficient). Caveat recorded: GSE31210 series-matrix expression is linear-scale, not log2, so absolute risk scores are off the paper's TPM scale though the stratification is robust. No fabrication signal found for the in-scope result. Grades are provisional; a human auditor must re-derive and sign AUDIT.md.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 59assessed: 2026-06-14 ⛓ 58d3ddb11123
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-14
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetThe study tests whether NSCLC patients can be classified into distinct molecular subtypes based on hypoxia-related gene expression, and whether a hypoxia-based prognostic gene signature can predict survival and guide selection of patients for targeted/immune/chemotherapies.
- ★ NSCLC samples can be classified into two molecular subtypes (C1 and C2) based on hypoxia-related gene expression via consensus clustering finding
- ★ C1 subtype shows significantly better overall survival than C2 subtype finding
- ★ Oncogenic pathways (hypoxia, EMT, NOTCH signaling, p53 signaling, TGF-β, WNT-β catenin) are more enriched in C2 subtype finding
- ★ C1 and C2 subtypes show distinct tumor microenvironment/immune cell infiltration profiles finding
- ★ C1 subtype is predicted to have a more favorable response to immunotherapy and chemotherapy than C2 finding
- ★ A 6-gene hypoxia-related prognostic signature classifies NSCLC patients into high-risk and low-risk groups with robust prognostic ability method
- Molecular subtyping is robust and reproducible across TCGA and independent GEO cohorts (validated by SubMap analysis) method
- The six prognostic genes may represent novel targets for understanding hypoxia-associated NSCLC mechanisms and developing targeted therapies resource
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| bulk RNA-seq differential expression (limma) | TCGA-NSCLC (LUAD+LUSC tumor samples) | none (comparison of C1 vs C2 molecular subtype) | differentially expressed genes (up/down-regulated) | TCGA RNA-seq; limma R package |
| microarray gene expression profiling | GEO cohorts (GSE31210, GSE30219, GSE50081, GSE19188, GSE37745, GSE29013) NSCLC patients | none (comparison of C1 vs C2 molecular subtype) | gene expression profiles, differentially expressed genes | GPL570 Affymetrix array; affy rma normalization |
| unsupervised consensus clustering | TCGA-NSCLC dataset and GSE cohorts | none | molecular subtype assignment (C1/C2) based on 101 prognosis-associated hypoxia genes | ConsensusClusterPlus R package v1.48.0 |
| GO/KEGG enrichment and GSEA pathway analysis | TCGA-NSCLC dataset and GSE cohorts | none (C1 vs C2 comparison) | enriched biological pathways and terms | clusterProfiler v3.16.0; MSigDB h.all.v7.3.symbols.gmt |
| immune cell deconvolution (CIBERSORT, MCP-counter, EPIC, ESTIMATE) | TCGA-NSCLC dataset (and GSE cohorts for validation) | none (C1 vs C2 comparison) | immune cell fractions, immune score, stromal score, immune checkpoint expression | CIBERSORT; MCP-counter; EPIC; ESTIMATE algorithms |
| TIDE (Tumor Immune Dysfunction and Exclusion) analysis and SubMap | TCGA-NSCLC dataset, GSE cohorts, and anti-PD-1 treated dataset (GSE78220) | none / in silico prediction of ICB response | TIDE score, proportion of predicted positive/negative immunotherapy responders | TIDE algorithm; GenePattern SubMap |
| in silico drug sensitivity prediction | TCGA-NSCLC dataset | none (predicted response to 7 chemotherapeutic drugs) | estimated IC50 for bexarotene, doxorubicin, embelin, etoposide, gemcitabine, mitomycin C, vinorelbine | pRRophetic R package |
| survival/prognostic model construction (univariate Cox, LASSO, stepAIC, multivariate Cox) | TCGA-NSCLC dataset (with GSE cohorts for gene screening) | none | 6-gene risk score, high-risk/low-risk classification, overall survival | MASS R package (stepAIC); survminer (surv_cutpoint) |
- ▲ C1 subtype showed significantly more favorable overall survival than C2 P=0.00011
- ▲ Hypoxia pathway enrichment score was significantly higher in C2 than C1 in both TCGA-NSCLC and GSE cohorts P<0.0001
- ▲ C1 subtype had a higher proportion of alive patients and milder clinical stages than C2 P=0.0391
- – 610 DEGs identified between subtypes in TCGA-NSCLC dataset (496 up, 114 down in C2); 903 DEGs in GSE cohorts (552 up, 381 down in C2) 610 and 903 DEGs
- – 14 of 22 immune cell types differed significantly between subtypes (CIBERSORT); C2 had higher M0 macrophages and lower CD8 T-cells than C1 P<0.0001
- ▲ C2 subtype had significantly higher immune score and stromal score (ESTIMATE) than C1, indicating greater immune infiltration
- ▲ C1 subtype showed lower TIDE score and higher proportion of positive immunotherapy response than C2 P<0.0001
- ▲ In GSE cohorts, positive immunotherapy response rate was 57% in C1 versus 26% in C2 57% vs 26%
- pvalue P=0.00011 (OS difference between C1 and C2 subtypes in TCGA-NSCLC dataset)
- pvalue P<0.0001 (Hypoxia pathway enrichment difference between C1 and C2 (TCGA and GSE cohorts))
- pvalue P=0.0391 (Distribution of survival status between C1 and C2 subtypes)
- count 45 and 56 hypoxia-related prognostic genes; 101 intersected genes (Genes from univariate Cox regression in TCGA-NSCLC and GSE cohorts used for clustering)
- count 610 DEGs (496 up, 114 down) (DEGs between C2 and C1 subtypes in TCGA-NSCLC dataset, FDR<0.05 and |FC|>1.5)
- count 903 DEGs (552 up, 381 down) (DEGs between C2 and C1 subtypes in GSE cohorts (801 samples))
- pvalue P<0.0001 (Differential expression of 14/22 immune cells (CIBERSORT) and immune checkpoints (33/47) between subtypes)
- other 57% vs 26% (Positive immunotherapy response rate in C1 vs C2 subtypes in GSE cohorts)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This bioinformatics study used unsupervised consensus clustering of hypoxia-related gene expression (pre-filtered by univariate Cox regression) to define two NSCLC molecular subtypes in combined TCGA and GEO cohorts. Between-subtype comparisons used log-rank tests (survival), chi-squared tests (clinical features), Wilcoxon tests (pathway enrichment scores), and Student t-tests (immune checkpoint expression). A 6-gene prognostic signature was built through sequential univariate Cox regression, LASSO penalised regression, and stepAIC optimisation, with survival stratification evaluated by Kaplan-Meier curves and multivariate Cox regression. Results were reported primarily as p-values and significance thresholds, with fold-change criteria applied to differential expression.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Log-rank test | Overall survival comparison between C1 and C2 subtypes in TCGA-NSCLC dataset and GSE cohorts; survival of five pan-cancer immune subtypes; TIDE-defined immunotherapy responders vs non-responders | 801 samples noted for GSE cohorts; TCGA-NSCLC total n not explicitly stated in excerpt | not stated |
| Chi-squared test | Distribution of clinical features (survival status, age, gender, smoking, stage) between C1 and C2; distribution of pan-cancer immune subtypes across molecular subtypes | null | not stated |
| Wilcoxon test (two-sample) | Comparison of hypoxia pathway enrichment scores between C1 and C2 in TCGA-NSCLC dataset and GSE cohorts | null | not stated |
| Student t-test | Expression of 47 immune checkpoints compared between C1 and C2 subtypes | null | not stated |
| Univariate Cox proportional-hazards regression | Screening of hypoxia-related genes associated with prognosis in TCGA-NSCLC and GSE cohorts (p<0.05 threshold); initial prognostic gene filtering | null | not stated |
| LASSO (least absolute shrinkage and selection operator) regression | Reduction of candidate prognostic genes to a smaller feature set | null | not stated |
| StepAIC (stepwise Akaike information criterion) with multivariate Cox regression | Final optimisation and construction of 6-gene prognostic model; univariate and multivariate Cox for independence of risk score | null | not stated |
| Limma moderated t-statistic (differential expression) | Identification of DEGs between C1 and C2 subtypes in TCGA-NSCLC and GSE cohorts (FDR<0.05, |FC|>1.5) | null | not stated |
| Decision curve analysis (DCA) | Evaluation of clinical utility/net benefit of risk score and nomogram | null | na |
-
Student t-test was applied to compare expression of 47 immune checkpoints between subtypes in a single family of tests without explicit multiplicity correction↳ Could also: Apply Benjamini-Hochberg FDR correction across the 47 tests, or use a non-parametric Wilcoxon rank-sum test per checkpoint (consistent with the Wilcoxon test used elsewhere in the paper for enrichment score comparisons) — Controlling the false discovery rate across a family of 47 tests limits expected false positives among the reported significant checkpoints; non-parametric testing avoids the normality assumption for expression values that may be skewed
-
LASSO regression alone was used for feature selection before stepAIC/Cox model building↳ Could also: Elastic net regularisation (alpha between 0 and 1, combining L1 and L2 penalties) could also be used for feature selection — Elastic net can handle groups of correlated predictors—as co-regulated hypoxia pathway genes often are—by distributing coefficients across a correlated group rather than arbitrarily selecting one, which may produce a more stable gene signature
-
The optimal cut-off to dichotomise continuous risk scores into high- and low-risk groups was determined by the surv_cutpoint function (maximally selected rank statistics)↳ Could also: The continuous risk score could also be used directly in a Cox regression without dichotomisation, or a pre-specified median split could be applied — Data-driven cut-point optimisation can inflate the apparent prognostic separation; using the risk score as a continuous predictor or a pre-specified split preserves the full statistical information and avoids this optimism
-
Consensus clustering with k evaluated from 2 to 10 was used to define molecular subtypes, with k=2 selected based on visual inspection of consensus matrices↳ Could also: Formal cluster validity indices such as average silhouette width or gap statistic could also be used to guide the choice of k — Quantitative indices provide a reproducible, criterion-based justification for the number of clusters that complements the visual consensus matrix inspection approach
-
Limma (a method developed for microarray and RNA-seq data) was used for differential expression in the TCGA RNA-seq portion of the data↳ Could also: For count-based RNA-seq data such as TCGA, DESeq2 or edgeR could also be applied, as they model negative-binomial count distributions directly — Count-based methods leverage the raw count structure of RNA-seq data and have well-characterised statistical properties for that data type; limma-voom can bridge this gap but the voom transformation step is not explicitly described
-
Prognostic gene candidates were pre-filtered by univariate Cox regression at p<0.05 before LASSO, without correction for the number of genes tested↳ Could also: A more stringent pre-filtering threshold (e.g., FDR<0.05) or direct application of penalised regression to the full hypoxia gene set without a pre-filter could also be used — Uncorrected univariate pre-filtering of many genes can pass a substantial number of false positives into the LASSO step; a corrected threshold or no pre-filter allows the regularisation itself to perform selection across the full candidate set
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
scope.md — pmid-35509605
Paper: Lin J, Chen S, Xiao L, Wang Z, Lin Y, Xu S (2022) Identifying
Molecular Subtypes and 6-Gene Prognostic Signature Based on Hypoxia for
Optimizing Targeted Therapies in Non-Small Cell Lung Cancer. Int J Gen Med
15:4523–4538. PMID 35509605 / PMC9058021 / DOI 10.2147/ijgm.s352238.
Repo: https://github.com/jianwem/pro2022 (HEAD main, pushed 2022-02-24).
Ships exactly two files: README.md (one line: "# Scripts") and
Hypoxia_LUAD_immune_model.R (2359 lines).
Rule 2 — code provenance (P16)
The repo is the authors' own analysis code (single monolithic R script for this paper's figures/tables), not a third-party tool. Reproduction is therefore of the authors' own pipeline. (Per the operator clarification, a third-party tool applied to the same data would have been equally valid; here the own-code condition holds.)
Reproducibility surface — what the repo actually provides
The script is not self-contained:
- Input data is NOT shipped. Every input is read from the first author's
local Windows drive (
e:/public/...,e:/datas/...) or from an unshippedraw_dat//model_select/tree, e.g.raw_dat/TCGA_LUAD_TPM_genesymbol.txt,raw_dat/Clinical BCR XML.merge.txt,raw_dat/MSigDB7.3_HALLMARK_HYPOXIA.txt,raw_dat/GSE31210/GSE31210_model_dat.txt,model_select/traindat_median_31.txt. - Helper functions are NOT defined. The pipeline calls
Sig_cox(),DEGs_limma(),estimate_score(),plot_RS_model(),ggsurvplotKM(),surv(), and the palettecolor8/color9— none of which are defined in the script or shipped anywhere. They live in the author's private R environment.
Consequence: the script cannot be executed end-to-end as published, and the exact discovery-pipeline numbers (LASSO/stepAIC coefficients, consensus-cluster assignments) cannot be reproduced bit-for-bit from the shipped artifacts. This is recorded honestly; we do not fabricate those values.
Per Rule 2/6 we do not drop the paper for this. Instead we reproduce the clearly-specified, public-data, low-hanging validation result that the paper itself anchors on — the 6-gene signature applied to the listed accession GSE31210 — re-implementing the well-described steps.
Pipeline inventory (what produces the reported numbers)
- TCGA-NSCLC discovery (TCGA-LUAD+LUSC TPM + BCR-XML clinical) → TPM filter
- log2 → univariate Cox on HALLMARK_HYPOXIA genes → ConsensusClusterPlus (km, euclidean, ward.D2, reps=100, pItem=0.8, seed=1984) → 2 subtypes (C1/C2) → ESTIMATE immune score → immune groups → limma DEGs → protective/risk DEG intersection → LASSO Cox (glmnet, set.seed(31)) → stepAIC → 6-gene risk model + coefficients.
- External validation: apply/refit the 6-gene Cox in GEO cohorts GSE31210, GSE30219, GSE50081 (and GSE19188/GSE37745/GSE29013 mentioned) → median-split RiskScore → KM log-rank + time-dependent ROC.
- Downstream (out of scope below): ssGSEA, GO/KEGG, nomogram, TIDE/IMvigor210 immunotherapy-response, drug-sensitivity.
IN SCOPE (public data, tractable on «our HPC») — what we attempt
- C-INSCOPE-1 — GSE31210 6-gene signature survival split. Reproduce the
external-validation result for the listed accession GSE31210 (public GEO,
GPL570, 226 LUAD tumors with overall-survival fields
death+days before death/censor). Pipeline (faithful to script §"GSE31210 数据集", lines ~951–967): buildGSE31210_model_datfrom GEO → fitcoxph(Surv(OS.time,OS) ~ ALDOA+EFNA1+GPC4+HOXB9+PGM2+PLAUR)→ linear predictor as RiskScore → median split High/Low → log-rank P (+ HR, time-dependent AUC). Anchor: paper reports validation P<0.0001 (Fig 9B). - C-INSCOPE-2 — published risk-formula direct application. Apply the
paper's printed coefficients
0.173·ALDOA −0.139·EFNA1 −0.071·GPC4 +0.059·HOXB9 +0.14·PGM2 +0.157·PLAUR(Fig 8 / Results) to GSE31210 expression → median split → log-rank P.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
The in-scope external-validation claim reproduces robustly: applying the paper's published 6-gene risk formula to GEO:GSE31210 yields log-rank P=8.9×10⁻⁶, HR 5.82 [2.41,14.03] in the correct direction, satisfying the paper's P<0.0001 (Fig 9B), so the central prognostic conclusion holds (q6 green, q7 green). The deviations are on the authors'/availability side: the repo is not self-contained (no data, hard-coded e:/ paths, 5 undefined helper functions), so the Fig 8 coefficients and the LASSO/stepAIC derivation of the gene set are not re-derivable (q4 red, q5 yellow). A documented scale caveat (GSE31210 linear vs paper log2-TPM) plausibly explains the EFNA1 sign flip and PGM2/PLAUR≈0 on refit. No fabrication signal for the reproduced result, but the auditability gaps make this a solid-with-deviations rather than 1:1 case (q8 yellow).
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.