Hypoxia-associated genes as predictors of outcomes in gastric cancer: a genomic approach.
The main results reproduced: recomputed values matched the published ones within tolerance.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓No relevant deviation in data/preprocessing
- ✓No authors-side cause for any deviation
- ✓Reported values are derivable from the shared data
- ✓Any deviation was negligible
- 🟡Could not use the authors’ exact input data
- 🟡Reported values were only indirectly comparable
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Described well enough to reproduce the paper's CENTRAL pipeline step 1:1. The paper predicts per-cell hypoxia status on GSE183904 with the third-party tool CHPF (github.com/yihan1221/CHPF, commit c8f8a29); we ran CHPF unmodified on its shipped example data («our HPC» SLURM, R 4.2.3 / Seurat 4.3.0.1 / GSVA 1.46.0 / Python 3.7.12) and compared to the authors' shipped expected outputs. The two DETERMINISTIC outputs reproduce BYTE-FOR-BYTE (identical SHA256): the 536 high-confidence GMM cells (99 hyp + 437 norm) and the 500 Wilcoxon feature genes. The full cell_status.csv is 99.7% concordant (249 vs 248 hypoxic of 1000) — the only divergence is 3 of the 464 LightGBM-predicted 'other' cells, which is the expected behaviour of a tool defect: CHPF.py seeds KFold but leaves Python's global RNG (random.sample) unseeded, so 'other'-cell predictions are non-deterministic by construction. No fabrication concern: every shipped value is derivable from the shipped data+code. NOT ATTEMPTED (the hard ~20%): the paper's GSE183904/TCGA-specific downstream results (H1-H4/N1-N4 neoplastic subpopulations, WGCNA modules, the 5-gene EHF/EIF1AD/GLA/KEAP1/MAGED2 LassoCox OS model) and the qRT-PCR wet-lab validation. Reason: the repo ships only the CHPF tool + a generic example, not the paper's GSE183904 intermediate matrices or the WGCNA/Lasso scripts; that chain depends on unpinned preprocessing/clustering/model-selection parameters. GSE183904 is public and obtainable if a future deeper attempt is warranted. Overall status 'partial': the in-scope tool reproduction succeeded cleanly; the paper's headline downstream model was deliberately not chased.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 96assessed: 2026-06-14 ⛓ 3de8504ac81c
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-14
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusThe study investigates the effects of hypoxia-related genes in stomach adenocarcinoma (STAD) and tests whether a hypoxia-associated gene signature derived from hypoxic tumor cell subpopulations can serve as a prognostic model for patient outcomes.
- ★ Most neoplastic cells, fibroblasts, endothelial cells, and myeloid cells in STAD are in a hypoxic state, while most mast, NK/T, and B cells are non-hypoxic finding
- ★ Neoplastic cells comprise four hypoxic (H1-H4) and four non-hypoxic (N1-N4) subpopulations, with H1 having the highest degree of hypoxia finding
- ★ A prognostic model built from five H1-specific transcription factors (EHF, EIF1AD, GLA, KEAP1, MAGED2) predicts overall survival, with worse OS in high-risk patients resource
- ★ The five H1-specific transcription factors are more highly expressed in gastric cancer cell lines than in a normal gastric epithelial cell line finding
- ★ Hypoxia score is positively correlated with Angiogenesis, Apoptosis, EMT, and Invasion scores in tumor cell subpopulations finding
- H1 and N1 subpopulations have entirely distinct (zero-overlap) key transcription factor sets finding
- CHPF software combined with single-cell transcriptomics and hypoxia gene clusters can classify cells into hypoxic and non-hypoxic groups method
- CD74-CD44 and MIF signaling show high communication activity between immune and neoplastic cells mechanism
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| scRNA-seq analysis (Seurat, Harmony, UMAP/TSNE, Louvain) | 26 primary gastric cancer samples (GSE183904) | none | cell type clustering, marker gene expression, hypoxia classification | — |
| Bulk RNA expression / prognostic modeling (LassoCox, univariate Cox) | TCGA STAD patients (n=368); GSE15460 validation (n=248) | none | overall survival risk score, high/low-risk stratification | — |
| Cellular hypoxia prediction (CHPF) | STAD single-cell transcriptomes | none | hypoxic vs non-hypoxic cell classification | CHPF software (github.com/yihan1221/CHPF) |
| WGCNA and GO-BP/KEGG enrichment | STAD single-cell module genes | none | hypoxia-associated gene modules and pathway enrichment | WGCNA / clusterProfiler |
| Cell-cell communication analysis | STAD immune and neoplastic cells | none | ligand-receptor interactions, signaling pathways | CellChat |
| Single-cell CNV analysis | STAD tumor cells (endothelial cells as reference) | none | CNV scores | InferCNV |
| Single-cell transcription factor / gene regulatory network analysis | H1 and N1 neoplastic subpopulations | none | top 1% key transcription factors | SCENIC / GRNboost2 |
| qRT-PCR | GES-1, AGS, BGC823, MGC803 cell lines (ATCC) | none | mRNA expression of EHF, EIF1AD, GLA, KEAP1, MAGED2 normalized to 18S rRNA | SYBR Premix Ex Taq, CFX96 Real-Time PCR System (Bio-Rad) |
- ▲ Fibroblasts had the highest proportion of hypoxic cells, followed by endothelial cells, myeloid cells, and neoplastic cells
- ▲ H1 subpopulation showed the highest hypoxia score among tumor subpopulations
- ▼ Five H1-specific transcription factor model predicts OS with significantly worse OS in high-risk patients
- ▲ Higher expression of the five transcription factors in gastric cancer cell lines vs normal GES-1 cell line
- ▲ Hypoxia score positively correlated with Angiogenesis, Apoptosis, EMT, and Invasion scores
- – 106 key transcription factors identified in H1 and 85 in N1, with zero overlap 106 (H1), 85 (N1), 0 overlap
- ▲ Proportion of hypoxic cells significantly increased in stage III STAD tissues
- – High CD74-CD44 communication activity between immune and tumor cells; MIF pathway important
- count n=368 (TCGA gastric cancer patients used for modeling)
- count n=248 (GSE15460 GEO validation dataset)
- count 26 primary gastric cancer samples (GSE183904 single-cell dataset)
- count 106 (key transcription factors identified in H1 subpopulation)
- count 85 (key transcription factors identified in N1 subpopulation)
- other |Cor|>0.2 and P<0.05 (threshold for H1-specific TF correlation with Hallmark pathways)
- other 5% (five-year survival rate of 6% in metastatic STAD setting)
- pvalue p<0.05, log2FC>0.25, expression proportion>0.1 (thresholds for differential gene expression between clusters)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This retrospective bioinformatics study of stomach adenocarcinoma (STAD) integrated scRNA-seq data (GSE183904, 26 primary samples) with bulk RNA-seq cohorts from TCGA (n=368, training) and GEO/GSE15460 (n=248, external validation). Single-cell analysis used CHPF-based hypoxia classification, WGCNA gene-module identification, and SCENIC transcription-factor inference to nominate H1-subpopulation-specific candidates; a LassoCox algorithm then reduced these to a five-gene risk score. Risk-group survival differences were assessed by Kaplan-Meier analysis and Cox regression, and model genes were validated in four gastric cancer cell lines versus a normal epithelial line by qRT-PCR. The paper text supplied is truncated before the full prognostic-model results section.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| LassoCox (LASSO-penalized Cox proportional-hazards regression) | Construction of five-gene prognostic risk score for overall survival | 368 (TCGA training cohort) | not stated |
| Univariate Cox proportional-hazards regression | Pre-screening of H1-specific transcription factors for association with overall survival prior to LASSO | 368 (TCGA training cohort) | not stated |
| Kaplan-Meier survival analysis (log-rank test implied; survival package in R) | Overall survival comparison between high-risk and low-risk groups in TCGA and GSE15460 cohorts | 368 (TCGA); 248 (GSE15460) | not stated |
| Wilcoxon rank-sum test or Student's t-test (choice stated to depend on data characteristics) | Comparisons of continuous variables across groups (e.g., pathway signature scores between hypoxic and non-hypoxic subpopulations) | — | not stated |
| Pearson correlation coefficient | Relationship between hypoxia score and pathway scores (Angiogenesis, Apoptosis, EMT, Invasion); correlation between H1-specific transcription factors and Hallmark pathways at |Cor|>0.2, P<0.05 | — | not stated |
| Chi-square test | Comparisons of categorical clinical variables | — | not stated |
| Wilcoxon rank-sum-based differential expression (Seurat FindAllMarkers; p<0.05, log2FC>0.25, expression proportion>0.1) | Differential gene expression between single-cell clusters and cell types in scRNA-seq data | 26 primary gastric cancer samples (GSE183904) | not stated |
-
The median of the continuous LassoCox risk score was used to dichotomize patients into high- and low-risk groups↳ Could also: An optimal-cutpoint method (e.g., maximally selected rank statistics via the maxstat package) or retaining the risk score as a continuous variable in a Cox model could also be used — Median splitting anchors the threshold to the specific cohort distribution and discards within-group variation; optimal-cutpoint methods select the threshold supported by the data, while continuous-score analysis avoids information loss from dichotomization entirely
-
Pearson correlation was used to relate hypoxia scores to pathway activity scores derived from GSVA↳ Could also: Spearman rank correlation could also have been applied to the same data — Spearman correlation requires no assumption of bivariate normality and is less sensitive to outliers, which is often advantageous for enrichment scores that may not follow a Gaussian distribution
-
Differential gene expression between single-cell clusters used a nominal p<0.05 threshold without explicit FDR correction across the many simultaneous gene–cluster tests↳ Could also: Benjamini-Hochberg FDR adjustment on the adjusted_p_val column already computed by FindAllMarkers, or use of a pseudobulk method (e.g., DESeq2 on aggregated counts per sample), could also be applied — Testing thousands of genes across multiple clusters raises the expected number of false positives; FDR control is standard in scRNA-seq differential expression, and pseudobulk approaches additionally account for within-sample correlation among cells
-
The criterion for choosing between the Wilcoxon rank-sum test and Student's t-test for continuous-variable comparisons was described only as 'based on the data's characteristics,' without specifying the decision rule↳ Could also: A pre-specified normality criterion (e.g., Shapiro-Wilk test, or a rule such as always use Wilcoxon for n<30 or visibly skewed distributions) could also be documented explicitly — Transparent pre-specification of the test-selection rule improves reproducibility and removes ambiguity about how the final test was chosen for each comparison
-
GO-BP and KEGG enrichment analyses were conducted for WGCNA modules and cell subpopulations; reporting of adjusted versus nominal p-values is not specified in the available text↳ Could also: Reporting Benjamini-Hochberg-adjusted q-values alongside nominal p-values, as is the default output of clusterProfiler, could also be done — Enrichment analyses test many gene sets simultaneously; q-values convey the expected false-discovery rate among reported terms and are the standard reporting unit in the field
-
Model performance was evaluated with an external validation cohort (GSE15460) after training on TCGA↳ Could also: k-fold cross-validation (e.g., 10-fold) within the TCGA training set, or bootstrap resampling to estimate internal optimism, could also accompany external validation — Internal cross-validation provides a bias-corrected estimate of discriminative performance that is independent of any single train/test split and complements external validation by quantifying in-sample optimism
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-40129974
Paper: Yang S, Jiang Y, Yang Z. Hypoxia-associated genes as predictors of outcomes in gastric cancer: a genomic approach. Front Immunol 2025. DOI 10.3389/fimmu.2025.1553477 · PMCID PMC11931070.
Tool reproduced (third-party, per BRIEF rule 2): CHPF — https://github.com/yihan1221/CHPF (Cellular Hypoxia status Predicting Framework; originally from Zhang et al., Theranostics 2023). The paper states: "Cellular hypoxia status was predicted using the CHPF software" on GSE183904 scRNA-seq. CHPF is the central pipeline step of the paper — every downstream result (hypoxic cell-type calls, H1–H4 subpopulations, the 5-gene LassoCox model) is conditioned on CHPF's per-cell hypoxia/normoxia labels.
Pipeline map of the paper's computational results
| # | Reported result | Pipeline | In scope? |
|---|---|---|---|
| R1 | Per-cell hypoxia/normoxia status (input to everything) | CHPF (GSVA-ssGSEA → mclust GMM → high-confidence cells → LightGBM ensemble) | YES — reproduce the tool 1:1 on its shipped example |
| R2 | "Most neoplastic/fibroblast/endothelial/myeloid cells are hypoxic" on GSE183904 | Seurat preprocessing + CHPF per sample (qualitative) | partial / out of easy scope |
| R3 | 4 hypoxic (H1–H4) + 4 non-hypoxic (N1–N4) neoplastic subpopulations | Seurat re-clustering of neoplastic cells | out of scope (free params) |
| R4 | WGCNA hypoxia module + GO-BP/KEGG enrichment | WGCNA | out of scope (free params) |
| R5 | 5-gene LassoCox model (EHF, EIF1AD, GLA, KEAP1, MAGED2); OS prediction | TCGA-STAD + LassoCox | out of scope (TCGA chain, free params) |
| R6 | qRT-PCR of 5 genes in GES-1/AGS/BGC823/MGC803 | wet lab | out of scope (non-pipeline) |
What we attempt (the clearly-specified, low-hanging 80%)
R1 — CHPF tool functional reproduction on its shipped example data. The repo
ships a complete, self-contained example: input (exmaple/expr.RData,
exmaple/Hypoxia_geneset.gmt) and the authors' own expected outputs
(exmaple/result/cell_status.csv, temp/label_highconfi.csv,
temp/expr_highconfi.csv, 500 feature files). We run CHPF.py unmodified
(R/4.2.3 + Seurat/GSVA/GSEABase/corrplot/mclust + Python 3.7.3 +
lightgbm/sklearn) and compare our output to the shipped expected output.
Determinism analysis (drives the grading)
- Deterministic steps → expect EXACT match: GSVA-ssGSEA scores, mclust GMM
(G=2), high-confidence-cell selection, top-500 Wilcoxon feature genes, and the
high-confidence labels
label_highconfi.csv. - Non-deterministic step → expect high concordance, not exact: the LightGBM
ensemble in
CHPF.pycallsrandom.sample()on Python's global RNG which is never seeded (onlyKFold(random_state=12234)is seeded). So the per-tree balanced subsampling — and hencelabel_others.csvand the "other"-cell part ofcell_status.csv— varies run-to-run by design. We grade these by concordance %, and flag the missing seed as a reproducibility defect.
What we do NOT attempt and why (the hard ~20%)
- R2–R5 end-to-end on GSE183904 + TCGA. Requires re-deriving per-sample CHPF input matrices from raw GEO (Seurat QC/normalization params not fully pinned), re-clustering neoplastic cells into H1–H4/N1–N4 (cluster count, resolution, marker thresholds not pinned), WGCNA (soft-power/module params not pinned), and a TCGA-STAD LassoCox chain (lambda selection, split not pinned). The repo ships only the CHPF tool + a generic example, not the paper's own GSE183904 intermediate matrices or the WGCNA/Lasso scripts. Reproducing the headline 5-gene signature is therefore not a clearly-specified low-hanging output.
- R6 qRT-PCR is wet-lab, non-pipeline.
Data availability (control-plane checks, no compute)
- Code: github.com/yihan1221/CHPF — public, not archived, default branch
main, last push 2025-09-05. Plain git blobs (no LFS needed). - Paper's data: GEO GSE183904 — Public since 2021-10-05, 31 primary gastri
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
The reproduction is a clean 1:1 run of the paper's central third-party tool (CHPF) on its shipped example data: two of three key outputs (536 high-confidence cells = 99 hyp + 437 norm; 500 Wilcoxon feature genes) are SHA256 byte-identical, and the only divergence (249 vs 248 hypoxic of 1000) is a known unseeded-RNG defect in the tool, not a methodological disagreement. No fabrication concern — every tested value is derivable from the shared data+code. However, scope was deliberately limited to the upstream labeling step: the paper's actual headline (the 5-gene LassoCox OS model on GSE183904/TCGA, H1–H4/N1–N4 subpopulations, WGCNA, qRT-PCR) was not attempted, so the central scientific conclusion is only partly supported. Severity of the observed deviation is negligible; the residual uncertainty is on coverage (our method/scope), not on the authors' side.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.