Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Identification of a novel 10 immune-related genes signature as a prognostic biomarker panel for gastric cancer.

Cancer Med · 2021
not yet assessed 2/4
Why this verdict

The main result did not reproduce in this reproduction attempt. Where our recomputation produced values that differ from the published ones, those discrepancies are listed below. This is a single automated attempt — not peer review and not a finding of error or misconduct — and differences can also arise from data access, undocumented parameters or the computing environment. The verdict can be contested via “report an error”.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Main result did not reproduce
Decisive
From: Q7 · Core claim 🔴
Main result did not reproduce
Decisive
From: Q8 · Severity of the miss (overall human judgment) 🔴
✓ What held up
  • Nothing in this column.
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🔴Reported values were only indirectly comparable
  • 🔴A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🔴The central claim did not (fully) hold under reproduction
  • 🔴Overall, the reproduction showed a material discrepancy
Reproduction agent’s raw note

DROP (no_code). Data is available and public (GSE62254), but the publication ships no analysis pipeline code; the single linked GitHub 'Code' URL is the generic vioplot plotting package, not the authors' LASSO/Cox immune-gene-signature pipeline. No pinnable reported value can be reproduced from shipped code, so no claims were graded and no «our HPC» compute was run. Not attempted: reconstructing the 10-gene LASSO-Cox signature, risk score, and survival/nomogram validation from Methods text alone (would be a from-scratch reimplementation, not a reproduction of shipped code; explicitly out of scope per the 80/20 + no-fabrication rules). Verdict is provisional and human-auditable: a reviewer who locates an authors' code supplement not surfaced here could re-screen to 'eligible'.

💻 Code ↗ 🗄 Data: GSE62254

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment
    assessed: 2026-06-15 ⛓ 75563584f32b
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-15
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

Because immune infiltrating cells in the tumor microenvironment correlate with gastric cancer development and progression, can a prognostic signature based on immune-related genes (IRGs) be developed to predict overall survival in gastric cancer patients?

Core claims
  • A 10 immune-related gene signature (BMPR1B, GHR, IL11RA, INHBB, NPR3, OBP2A, PTN, R3HDML, TAC1, TPM2) was constructed via WGCNA combined with LASSO-Cox and predicts overall survival in gastric cancer resource
  • The signature effectively predicts 1-, 3-, and 5-year OS and stratifies patients into high- and low-risk groups with worse prognosis in the high-risk group finding
  • The signature is an independent prognostic factor in the training and two external validation datasets by multivariate Cox regression finding
  • A nomogram combining the signature with clinical information provides strong discrimination (c-index 0.756) for predicting survival resource
  • The risk score correlates with multiple immune infiltrating cell types including CD8 T cells, CD4 memory T cells, NK cells, and macrophages mechanism
  • GSEA revealed significant pathways enriched between risk groups, including TGF-beta and Wnt signaling pathways finding
  • Combining WGCNA and LASSO-Cox on immune-related genes is an effective method to identify candidate prognostic biomarkers method
Experimental setups
Assay System Perturbation Readout Platform
Microarray gene expression profiling (WGCNA + LASSO-Cox prognostic modeling) Gastric cancer patient tumors (training dataset GSE62254, n=300) none Co-expression modules, risk score, overall survival prediction Affymetrix Human Genome U133 Plus 2.0 Array (GPL570)
Microarray gene expression profiling (signature validation) Gastric cancer patient tumors (validation dataset I, GSE15459, n=192) none Risk score, time-dependent ROC, Kaplan-Meier OS GPL570
Microarray gene expression profiling (signature validation) Gastric cancer patient tumors (validation dataset II, GSE84437, n=433) none Risk score, time-dependent ROC, Kaplan-Meier OS GPL6947
RNA-seq differential expression analysis (DESeq2) TCGA-STAD tumor (n=342) vs normal (n=30) none Differentially expressed genes (log2|FC|≥1, p<0.05)
Immune infiltration estimation (ESTIMATE and CIBERSORTx deconvolution) Gastric cancer tumors (GSE62254) none Immune/stromal scores, immune cell type fractions vs risk score CIBERSORTx (100 permutations); ESTIMATE
Gene set enrichment analysis (GSEA) Gastric cancer tumors (GSE62254), high vs low risk score groups none Enriched KEGG pathways (nominal p<0.01, FDR<25%) c2.cp.kegg.v6.2 gene set, 1000 permutations
Key results
  • Signature predicted 1-, 3-, 5-year OS in training dataset (GSE62254) AUC 0.681, 0.741, 0.72
  • Signature predicted 1-, 3-, 5-year OS in validation dataset I (GSE15459) AUC 0.57, 0.619, 0.694
  • Signature predicted 1-, 3-, 5-year OS in validation dataset II (GSE84437) AUC 0.559, 0.624, 0.585
  • High risk score group had significantly worse overall survival in training dataset p<0.0001
  • High risk score group had worse OS in validation datasets I and II GSE15459 p=0.0043; GSE84437 p=0.013
  • Risk score was an independent prognostic factor by multivariate Cox in training dataset HR 2.76 (2.13–3.58), p<0.001
  • Nomogram combining signature and clinical features showed strong discrimination c-index 0.7555
  • Five WGCNA modules correlated with OS; 266 prognostic genes identified in yellow module MEyellow r=-0.23, p=7e-05
Key statistics
  • correlation r=0.9 (scale-free R2) (WGCNA soft-threshold power 3 selected)
  • other c-index 0.7555135 (Nomogram discrimination ability, training dataset)
  • fold_change HR 1.4064 (1.1847–1.67), p=9.75E-05 (INHBB multivariate Cox, strongest individual gene)
  • pvalue MEyellow r=-0.23, p=7e-05 (Module-OS correlation, yellow module)
  • count 87 candidate genes (Intersection of 4383 TCGA-STAD DEGs and 266 survival-related genes)
  • count 1211 overlapping IRGs (GSE62254 and TCGA-STAD intersected with ImmPort IRGs for WGCNA)
  • other RS HR 2.72 (2.15–3.44), p<0.001 (Univariate Cox of risk score, training dataset)
  • other RS HR 1.72 (1.27–2.33), p<0.001 (Multivariate Cox of risk score, validation dataset II)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The study applied a sequential bioinformatics pipeline — WGCNA (Pearson correlation co-expression networks) followed by LASSO-Cox regression — to derive a 10-gene immune-related risk score (RS) from a training cohort of 300 gastric cancer patients (GSE62254), then validated the RS in two independent GEO datasets (GSE15459, n=192; GSE84437, n=433). Predictive performance was quantified via time-dependent ROC curves (AUC at 1, 3, and 5 years) and a bootstrap-validated nomogram (C-index); survival stratification used Kaplan-Meier analysis with log-rank tests, and RS independence was confirmed by univariate and multivariate Cox regression. Immune infiltration was characterized by ESTIMATE scores and CIBERSORTx deconvolution, and pathway enrichment was assessed by GSEA.

Replicationbiological Sample sizeDataset sample sizes stated: training n=300, validation I n=192, validation II n=433, TCGA cancer n=342, TCGA normal n=30; no formal power calculation or sample size justification reported GroupsHigh RS vs. low RS (median cut-off); gastric cancer vs. normal tissue (TCGA DEG analysis) Pairingunpaired Randomization/blindingnot stated Dispersionnone Exact p-valuesyes Effect sizesyes Confidence intervalsyes Multiplicity correctionFDR <25% for GSEA (explicitly stated); DESeq2 applies Benjamini-Hochberg FDR internally by default, though paper cites raw p<0.05 for DEG threshold; no correction stated for univariate Cox gene screening or WGCNA module-trait correlations
Statistical tests used
Test Applied to n Assumptions
Pearson correlation (WGCNA module-eigengene vs. clinical trait) Correlation of nine module eigengenes with OS, sex, death, and age in GSE62254 to select survival-correlated modules 300 not stated
Univariate Cox regression Screening of all IRGs within five survival-correlated WGCNA modules for OS association; 266 genes retained at p<0.05 300 not stated
DESeq2 Wald test Differentially expressed gene analysis between TCGA-STAD cancer and normal samples (|log2FC|≥1, p<0.05) 372 (cancer n=342, normal n=30) not stated
LASSO-Cox regression Feature selection reducing 87 candidate genes to the final 10-gene signature 300 not stated
Multivariate Cox regression Independent prognostic factor assessment of RS alongside gender, age, and stage in training and both validation datasets 300 (training), 192 (validation I), 433 (validation II) not stated
Kaplan-Meier / log-rank test OS comparison between high RS and low RS groups in training and both validation datasets 300, 192, 433 respectively not stated
Time-dependent ROC (tROC / AUC) Predictive accuracy for 1-, 3-, and 5-year OS in training and both validation datasets 300, 192, 433 respectively na
GSEA permutation test (1000 permutations) Pathway enrichment between high RS and low RS groups in GSE62254; threshold: nominal p<0.01 and FDR<25% 300 not stated
CIBERSORTx permutation test (100 permutations) Immune cell-type fraction estimation from GSE62254 bulk expression data 300 not stated
Bootstrap resampling (1000 iterations) for C-index Internal validation of nomogram discrimination ability 300 not stated
Approaches that could also have been used
  • RS was dichotomized at the median to create high and low RS groups for Kaplan-Meier and group-level analyses
    Could also: Continuous RS could be retained as a linear or spline predictor in Cox regression without dichotomization — Treating RS as continuous avoids information loss inherent in median splitting and yields an HR per unit change in RS that is interpretable across the full prognostic range; restricted cubic splines can additionally reveal whether the RS-hazard relationship is linear
  • Univariate Cox regression was applied to all IRGs in WGCNA survival-correlated modules (yielding 266 genes at p<0.05) without a reported multiple-testing correction before LASSO input
    Could also: Benjamini-Hochberg FDR correction could also be applied to the family of univariate Cox p-values at this screening step — With hundreds of simultaneous tests, an FDR adjustment characterizes which associations exceed a pre-specified false-discovery threshold, providing an additional description of the confidence level of genes entering downstream LASSO selection
  • LASSO-Cox was used for feature selection followed by a separate multivariate Cox fit to obtain final coefficients
    Could also: Elastic net Cox regression (combining L1 and L2 penalties) would also perform simultaneous selection and shrinkage in a single model — Elastic net can be more stable than pure LASSO when predictors are correlated — a likely scenario for co-expressed immune-related genes — potentially producing a more reproducible gene panel across independent datasets
  • CIBERSORTx was the sole deconvolution method used to estimate immune cell fractions from bulk expression data
    Could also: TIMER, xCell, EPIC, or MCP-counter would also estimate immune infiltration from microarray or bulk RNA-seq profiles — Comparing estimates across two or more deconvolution algorithms can characterize which immune-infiltration associations are robust to methodological assumptions, since each tool uses different reference matrices and statistical models
  • Predictive accuracy was reported as time-point-specific AUC from time-dependent ROC curves (1, 3, and 5 years)
    Could also: The integrated Brier score or the concordance index (Harrell's C) over the full follow-up would also summarize discriminative and calibration performance — The integrated Brier score simultaneously captures calibration and discrimination across the entire survival curve rather than at fixed horizons, offering a complementary summary of model accuracy
  • Nomogram internal validation used a 1000-resample bootstrap C-index within the single training dataset
    Could also: K-fold cross-validation would also estimate within-dataset generalization error by holding out folds during model fitting — K-fold cross-validation provides a direct estimate of prediction error on unseen data partitions and is sometimes reported alongside bootstrap validation to describe optimism correction from different perspectives
Software: R 3.6.1 · R/affy · R/limma · R/DESeq2 · R/WGCNA · R/survival · R/clusterProfiler · R/glmnet · R/rms · R/timeROC · R/estimate · CIBERSORTx · GSEA · X-tile

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
13
Impact: medium
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

GSE62254 GEO in Results (http://purl.org/orb/Results)
also used by 1 paper:
138981 in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
600985 in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
602413 in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
603717 in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
606132 in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
609522 in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
609925 in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
612557 in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
612898 in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
614597 in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
614757 in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
615005 in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
615996 in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
616194 in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
616771 in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
617222 in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
617528 in Supplementary material (http://purl.obolibrary.org/obo/IAO_0000326)
no other assessed paper uses this yet
GSE15459 GEO in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GSE84437 GEO in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet

Downstream reach in the literature

240 downstream papers · 3 datasets

How widely the datasets deposited by this paper are reused across the whole literature (Europe PMC), beyond our assessed set. This is a factual dependency map — reusing a public dataset is normal, good science. It is not a judgement on the downstream papers; the only verdict here is this paper's own, with its cited rationale.

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

No individual results have been recorded for this entry yet.

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 25/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🔴2. Endpoint comparability
🔴3. Location of the main deviation
🟡4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🔴7. Core claim
🔴8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Main result did not reproduce
Decisive
From: Q7 · Core claim 🔴
Main result did not reproduce
Decisive
From: Q8 · Severity of the miss (overall human judgment) 🔴

This is a no-code DROP: the public ACRG cohort (GSE62254) resolves, but the publication ships no runnable analysis pipeline — the only linked 'Code' URL is the generic third-party vioplot plotting package, not the authors' LASSO/Cox immune-gene-signature workflow. Consequently none of the substantive reported values (10-gene signature, risk-score coefficients, KM/HR survival validation, nomogram) were reproduced or compared against any output. The blocker sits on the authors'/deposit side (no code shipped) combined with our scope decision not to reimplement from Methods prose; there is no evidence of fabrication — the values are plausibly derivable in principle from the shared data, just unverifiable here. Overall red because zero claims could be confirmed.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

18.2 k
tokens (I/O) · 552.5 k incl. cache
13 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.