Interpretable artificial intelligence based on immunoregulation-related genes predicts prognosis and immunotherapy response in lung adenocarcinoma.
The main results reproduced, with only marginal, non-material deviations.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- Nothing in this column.
- 🟡Could not use the authors’ exact input data
- 🟡Reported values were only indirectly comparable
- 🟡A deviation arose in the data or preprocessing
- 🔴A deviation was attributed to the published material
- 🔴Reported values were not (fully) derivable from the shared data
- 🟡The deviation was non-trivial in magnitude
- 🔴The central claim did not (fully) hold under reproduction
- 🔴Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
P16 third-party-style minimal repo: github.com/mikelu1997/consensus @1dfbc60 ships only TWO R snippets (combat.txt = merge 4 GEO sets + sva::ComBat; ConsensusClusterPlus.txt = k=2 subtyping) with NO input data, NO probe->gene mapping, and NO README, covering ~2 of the paper's ~12 pipeline stages. The headline deliverables (15-IRG list aside) have no code: LASSO GREM1/PLAU signature, prognostic HR/AUC, and the XGBoost AUC 0.975 ship nothing. We reproduced the ONE shipped stage by reconstructing inputs from public GEO (GSE10072/32863/40791/68465) on «our HPC»: the merged-cohort NORMAL count is EXACT (226), the tumour count is 653 (the paper prints 654 but its own k=2 split sums to 653 -> off-by-one, flag R2), all 15 IRG gene symbols are present (exact), the DEG count is the right order (570 vs 484, partial), and the consensus k=2 two-subtype structure reproduces with an approximate split (limma removeBatchEffect 460/193 vs reported 438/215; vendored ComBat 384/269). Verdict: PARTIAL -- shipped subtyping stage reproduces in structure and approximately in magnitude; prognostic/XGBoost claims NOT attempted (no code). Fabrication flags: R1 'log2(TPM+1)' is impossible for microarray inputs; R2 654-vs-653 off-by-one; R3 XGBoost AUC 0.975 has label leakage (risk group defined from the same IRGs fed to XGBoost). NOT attempted: WGCNA module count, LASSO signature + HR/AUC, TCGA validation, XGBoost/SHAP, ESTIMATE/TIMER/TIDE, scRNA GSE229353, qRT-PCR. Compute-env note: «our HPC» nodes reach NCBI + bioconda CDN but NOT bioconductor.org / the OSN bioconductor-data-package host, so GEOquery/sva were replaced by direct base-R parsing + a faithful vendored ComBat.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 71assessed: 2026-06-15 ⛓ bd4c35b379f5
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-15
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-16no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusCan immunoregulation-related genes (IRGs) be used to build a reliable, interpretable predictive model of prognosis and immune checkpoint inhibitor (ICI) response in lung adenocarcinoma (LUAD)?
- ★ LUAD samples cluster into IRG-high and IRG-low groups, with the IRG-high group showing significantly better survival and greater immune cell infiltration. finding
- ★ High IRG pattern samples show lower TIDE scores, indicating a better predicted response to ICI treatment. finding
- ★ An IRG index (IRGI) model based on two key genes, GREM1 and PLAU, stratifies patients into high- and low-risk groups with distinct prognosis, mutational profiles, and TME immune infiltration. resource
- ★ An interpretable XGBoost machine learning model built on IRGs improves predictive performance (AUC = 0.975). method
- ★ SHAP analysis shows GREM1 has the greatest impact on the overall model prediction. finding
- IRGs shape the diversity and complexity of TME cell infiltration and modulate the balance between immune activation and suppression. mechanism
- 15 candidate IRGs were identified by intersecting DEGs, WGCNA tumor-related module genes, and ImmPort immune-related genes, and are upregulated in LUAD tumor tissue. finding
- ★ IRGI can serve as a biomarker to predict LUAD prognosis and response to ICIs. finding
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| bulk RNA/mRNA expression profiling (DEG analysis) | LUAD tumor vs normal lung tissue (GEO: GSE10072, GSE32863, GSE40791, GSE68465) | none | differentially expressed genes (logFC, P value) | — |
| Weighted Gene Co-expression Network Analysis (WGCNA) | GEO LUAD transcriptome cohort | none | co-expression modules correlated with immune clusters | — |
| PPI network / hub gene analysis | DEGs from LUAD vs normal | none | hub genes by Degree score | STRING database, Cytoscape CytoHubba |
| unsupervised consensus clustering | GEO LUAD samples | none | IRG-based sample clusters | ConsensusClusterPlus |
| immune microenvironment deconvolution (ESTIMATE, TIMER) | LUAD sample groups | none | ESTIMATE/immune/stromal scores, immune cell infiltration abundance | ESTIMATE, TIMER algorithms |
| immunotherapy response prediction (TIDE) and LASSO-Cox prognostic modeling | LUAD GEO cohort; validation in TCGA, GSE72094, Cho cohort, VanAllen cohort | none | TIDE score, RiskScore (IRGI), overall survival | TIDE online tool, glmnet R package |
| single-cell RNA sequencing analysis | 6 immunotherapy-treated LUAD samples (GSE229353); 15,293 cells | immunotherapy | cell type clusters and marker annotation | Seurat R package, CellMarker 2.0 |
| qRT-PCR | BEAS-2B normal lung epithelial cells and A549 LUAD cells | none | PLAU and GREM1 mRNA expression (2^-ΔΔCT vs β-actin) | Takara SYBR Green kit, Roche LightCycler 480 |
- – XGBoost machine learning model predicts outcome with high accuracy AUC = 0.975
- ▲ IRG-high group has significantly better survival and immune infiltration than IRG-low group
- ▼ High IRG pattern group shows better predicted ICI response (lower TIDE)
- – GREM1 has the greatest SHAP impact on overall prediction
- – 484 DEGs identified between normal and tumor samples (276 down, 208 up) 484 DEGs (276 down, 208 up)
- ▲ Black WGCNA module most strongly correlated with immune clusters r = 0.38
- – 35.0% of LUAD samples harbored a mutation in at least one key IRG, COL11A1 most frequent 35.0%
- ▲ 15 candidate IRGs identified by intersecting DEGs, WGCNA module, and ImmPort genes 15 genes
- other AUC = 0.975 (XGBoost model predictive performance)
- correlation r = 0.38, P = 7e-31 (black module correlation with immune clusters)
- count 484 DEGs (276 downregulated, 208 upregulated) (DEGs between normal and LUAD)
- count 178 genes (core PPI network size)
- count 15,293 cells (scRNA-seq cells retained after QC from 6 LUAD samples)
- other 35.0% (LUAD samples with mutation in at least one key IRG)
- other β = 3 (optimal WGCNA soft-thresholding power)
- count 654 tumor and 226 normal samples (included LUAD tumor and normal tissue samples)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This bioinformatics study integrated four GEO LUAD RNA-seq cohorts (654 tumour, 226 normal samples) after ComBat batch-effect correction, applied WGCNA and limma empirical-Bayes methods to identify immunoregulation-related genes, then used unsupervised consensus clustering (ConsensusClusterPlus, 1,000 iterations) to stratify tumours into two subtypes. A LASSO-Cox regression model (10-fold cross-validation) built on 2 key IRGs (GREM1, PLAU) was validated in external cohorts; an XGBoost classifier was added and interpreted with SHAP; overall survival was assessed with Kaplan-Meier curves and time-dependent ROC analysis, with two-sided p < 0.05 as the significance threshold.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Empirical Bayes moderated t-statistic (limma R package) | Differentially expressed genes between 226 normal and 654 LUAD tumour samples | 880 (226 normal + 654 tumour) | not stated |
| Unsupervised consensus clustering (ConsensusClusterPlus, 1,000 iterations) | Stratification of LUAD tumour samples into IRG expression clusters | 654 tumour samples | na |
| Weighted Gene Co-expression Network Analysis (WGCNA), soft threshold β = 3 | Detection of co-expression modules correlated with immune clusters | 880 | not stated |
| LASSO-Cox regression with 10-fold cross-validation (glmnet R package) | Construction of IRGI prognostic risk signature from 15 candidate IRGs | — | not stated |
| Kaplan-Meier survival analysis (log-rank test not explicitly named) | Overall survival comparison: IRG-high vs IRG-low; IRGI high-risk vs low-risk; external validation cohorts | — | na |
| Time-dependent ROC / AUC (timeROC R package) | Predictive performance of IRGI risk model at 2-, 3-, and 4-year survival endpoints | — | na |
| XGBoost machine learning classifier (stratified 70/30 train-validation split) | IRG-based binary classification; overall AUC = 0.975 reported on validation set | — | na |
| Independent Student's t-test (two-sided) | Between-group comparisons of normally distributed continuous variables | — | not stated |
| Mann-Whitney U test (two-sided) | Between-group comparisons of non-normally distributed continuous variables | — | not stated |
| 2−ΔΔCT comparative CT method | qRT-PCR quantification of PLAU and GREM1 in BEAS-2B (normal) vs A549 (LUAD) cell lines | — | not stated |
-
DEG significance was defined by P < 0.05 combined with a dynamic |logFC| threshold (mean + 2 SD), without an explicit false-discovery-rate correction across the ~20,000 genes tested↳ Could also: Applying Benjamini-Hochberg FDR correction (adjusted P < 0.05 or 0.1) directly on the limma output is standard for genome-wide DEG analyses — FDR control directly addresses the multiple-testing problem inherent in genome-wide screens; it would make the DEG list more conservative and is widely expected by reviewers and replication studies
-
A pure LASSO penalty was used for the Cox prognostic model, selecting 2 IRGs from 15 candidates↳ Could also: Elastic net Cox regression (mixing L1 and L2 penalties with α selected by cross-validation) could also have been applied — LASSO tends to arbitrarily select one predictor from a correlated group; elastic net handles correlated predictors more stably, which is relevant when IRGs share co-expression modules as identified by WGCNA
-
XGBoost predictive performance was summarised as a single AUC point estimate (0.975) on a one-time 30% hold-out set↳ Could also: Bootstrap confidence intervals around the AUC, repeated k-fold cross-validation across multiple random splits, or a calibration plot (reliability diagram) could also accompany the point estimate — A single AUC on one hold-out split carries sampling variability; CIs or repeated cross-validation convey the precision of the estimate and make overfitting easier to detect
-
The optimal number of consensus clusters was selected from the cumulative distribution function curve of ConsensusClusterPlus↳ Could also: Gap statistic, average silhouette width, or the Calinski-Harabasz index could also be computed to guide cluster-number selection; non-negative matrix factorisation (NMF) is another established approach for transcriptomic subtyping — Convergent evidence from multiple cluster-validity indices strengthens confidence in the chosen k, reducing reliance on a single graphical criterion
-
Group comparisons used Student's t-test for normally distributed variables and Mann-Whitney U for non-normal data, but the method used to assess normality before this choice is not described↳ Could also: Stating the normality-screening criterion (e.g., Shapiro-Wilk test at P < 0.05, or a QQ-plot inspection) would clarify how the decision between parametric and non-parametric tests was made; alternatively, non-parametric tests could be used uniformly throughout — Documenting the normality decision rule improves reproducibility and allows readers to evaluate whether the parametric assumption was appropriately checked
-
Cell-line validation of GREM1 and PLAU expression by qRT-PCR was reported without stating the number of biological replicates, technical replicates, or a dispersion measure↳ Could also: Reporting the number of independent experiments (biological replicates), technical replicates per experiment, and a dispersion statistic (SD or SEM) with degrees of freedom follows MIQE guidelines for qRT-PCR — Without replicate counts and dispersion, the statistical confidence of the qRT-PCR result cannot be independently assessed; MIQE compliance is increasingly required for reproducibility
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
Assessed papers, coloured by verdict. Click a node to open it.
- No assessed neighbours yet — the network grows as more papers are assessed.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
This is a non-reproducible-as-shipped paper: the repo ships only 2 of ~12 stages with no input data, so every headline deliverable (GREM1/PLAU LASSO signature, prognostic HR 0.49/AUC 0.794, TCGA HR 1.52, XGBoost AUC 0.975) has no code and could not be reproduced — the gap is on the authors' side. The single shipped stage (consensus k=2 subtyping) reproduces in structure and roughly in magnitude (split 460/193 vs 438/215; DEG 570 vs 484) after reconstructing inputs from public GEO, with deviations attributable to undocumented preprocessing. Three self-consistency red-flags make the central conclusion unsafe: an impossible log2(TPM+1) normalization for microarray inputs, a 654-vs-653 off-by-one, and label leakage rendering the flagship AUC 0.975 circular. Overall red — substantive, with fabrication-suspect signals on q5.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.