Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Interpretable artificial intelligence based on immunoregulation-related genes predicts prognosis and immunotherapy response in lung adenocarcinoma.

Front Bioinform · 2025
L1 71/100 3/4
Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Main result did not reproduce
Decisive
From: Q5 · Derivability / plausibility 🔴
Main result did not reproduce
Decisive
From: Q7 · Core claim 🔴
Main result did not reproduce
Decisive
From: Q8 · Severity of the miss (overall human judgment) 🔴
✓ What held up
  • Nothing in this column.
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🟡Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🔴A deviation was attributed to the published material
  • 🔴Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🔴The central claim did not (fully) hold under reproduction
  • 🔴Overall, the reproduction showed a material discrepancy
How its reproducibility compares
71/100
Reproducibility score
at the mean
vs. all fields · 1173 studies
🎯 Scores higher than 38% of all assessed papers rank 694 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

P16 third-party-style minimal repo: github.com/mikelu1997/consensus @1dfbc60 ships only TWO R snippets (combat.txt = merge 4 GEO sets + sva::ComBat; ConsensusClusterPlus.txt = k=2 subtyping) with NO input data, NO probe->gene mapping, and NO README, covering ~2 of the paper's ~12 pipeline stages. The headline deliverables (15-IRG list aside) have no code: LASSO GREM1/PLAU signature, prognostic HR/AUC, and the XGBoost AUC 0.975 ship nothing. We reproduced the ONE shipped stage by reconstructing inputs from public GEO (GSE10072/32863/40791/68465) on «our HPC»: the merged-cohort NORMAL count is EXACT (226), the tumour count is 653 (the paper prints 654 but its own k=2 split sums to 653 -> off-by-one, flag R2), all 15 IRG gene symbols are present (exact), the DEG count is the right order (570 vs 484, partial), and the consensus k=2 two-subtype structure reproduces with an approximate split (limma removeBatchEffect 460/193 vs reported 438/215; vendored ComBat 384/269). Verdict: PARTIAL -- shipped subtyping stage reproduces in structure and approximately in magnitude; prognostic/XGBoost claims NOT attempted (no code). Fabrication flags: R1 'log2(TPM+1)' is impossible for microarray inputs; R2 654-vs-653 off-by-one; R3 XGBoost AUC 0.975 has label leakage (risk group defined from the same IRGs fed to XGBoost). NOT attempted: WGCNA module count, LASSO signature + HR/AUC, TCGA validation, XGBoost/SHAP, ESTIMATE/TIMER/TIDE, scRNA GSE229353, qRT-PCR. Compute-env note: «our HPC» nodes reach NCBI + bioconda CDN but NOT bioconductor.org / the OSN bioconductor-data-package host, so GEOquery/sva were replaced by direct base-R parsing + a faithful vendored ComBat.

💻 Code ↗ 🗄 Data: GSE10072

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 71
    assessed: 2026-06-15 ⛓ bd4c35b379f5
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-15
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-16
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

Can immunoregulation-related genes (IRGs) be used to build a reliable, interpretable predictive model of prognosis and immune checkpoint inhibitor (ICI) response in lung adenocarcinoma (LUAD)?

Core claims
  • LUAD samples cluster into IRG-high and IRG-low groups, with the IRG-high group showing significantly better survival and greater immune cell infiltration. finding
  • High IRG pattern samples show lower TIDE scores, indicating a better predicted response to ICI treatment. finding
  • An IRG index (IRGI) model based on two key genes, GREM1 and PLAU, stratifies patients into high- and low-risk groups with distinct prognosis, mutational profiles, and TME immune infiltration. resource
  • An interpretable XGBoost machine learning model built on IRGs improves predictive performance (AUC = 0.975). method
  • SHAP analysis shows GREM1 has the greatest impact on the overall model prediction. finding
  • IRGs shape the diversity and complexity of TME cell infiltration and modulate the balance between immune activation and suppression. mechanism
  • 15 candidate IRGs were identified by intersecting DEGs, WGCNA tumor-related module genes, and ImmPort immune-related genes, and are upregulated in LUAD tumor tissue. finding
  • IRGI can serve as a biomarker to predict LUAD prognosis and response to ICIs. finding
Experimental setups
Assay System Perturbation Readout Platform
bulk RNA/mRNA expression profiling (DEG analysis) LUAD tumor vs normal lung tissue (GEO: GSE10072, GSE32863, GSE40791, GSE68465) none differentially expressed genes (logFC, P value)
Weighted Gene Co-expression Network Analysis (WGCNA) GEO LUAD transcriptome cohort none co-expression modules correlated with immune clusters
PPI network / hub gene analysis DEGs from LUAD vs normal none hub genes by Degree score STRING database, Cytoscape CytoHubba
unsupervised consensus clustering GEO LUAD samples none IRG-based sample clusters ConsensusClusterPlus
immune microenvironment deconvolution (ESTIMATE, TIMER) LUAD sample groups none ESTIMATE/immune/stromal scores, immune cell infiltration abundance ESTIMATE, TIMER algorithms
immunotherapy response prediction (TIDE) and LASSO-Cox prognostic modeling LUAD GEO cohort; validation in TCGA, GSE72094, Cho cohort, VanAllen cohort none TIDE score, RiskScore (IRGI), overall survival TIDE online tool, glmnet R package
single-cell RNA sequencing analysis 6 immunotherapy-treated LUAD samples (GSE229353); 15,293 cells immunotherapy cell type clusters and marker annotation Seurat R package, CellMarker 2.0
qRT-PCR BEAS-2B normal lung epithelial cells and A549 LUAD cells none PLAU and GREM1 mRNA expression (2^-ΔΔCT vs β-actin) Takara SYBR Green kit, Roche LightCycler 480
Key results
  • XGBoost machine learning model predicts outcome with high accuracy AUC = 0.975
  • IRG-high group has significantly better survival and immune infiltration than IRG-low group
  • High IRG pattern group shows better predicted ICI response (lower TIDE)
  • GREM1 has the greatest SHAP impact on overall prediction
  • 484 DEGs identified between normal and tumor samples (276 down, 208 up) 484 DEGs (276 down, 208 up)
  • Black WGCNA module most strongly correlated with immune clusters r = 0.38
  • 35.0% of LUAD samples harbored a mutation in at least one key IRG, COL11A1 most frequent 35.0%
  • 15 candidate IRGs identified by intersecting DEGs, WGCNA module, and ImmPort genes 15 genes
Key statistics
  • other AUC = 0.975 (XGBoost model predictive performance)
  • correlation r = 0.38, P = 7e-31 (black module correlation with immune clusters)
  • count 484 DEGs (276 downregulated, 208 upregulated) (DEGs between normal and LUAD)
  • count 178 genes (core PPI network size)
  • count 15,293 cells (scRNA-seq cells retained after QC from 6 LUAD samples)
  • other 35.0% (LUAD samples with mutation in at least one key IRG)
  • other β = 3 (optimal WGCNA soft-thresholding power)
  • count 654 tumor and 226 normal samples (included LUAD tumor and normal tissue samples)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This bioinformatics study integrated four GEO LUAD RNA-seq cohorts (654 tumour, 226 normal samples) after ComBat batch-effect correction, applied WGCNA and limma empirical-Bayes methods to identify immunoregulation-related genes, then used unsupervised consensus clustering (ConsensusClusterPlus, 1,000 iterations) to stratify tumours into two subtypes. A LASSO-Cox regression model (10-fold cross-validation) built on 2 key IRGs (GREM1, PLAU) was validated in external cohorts; an XGBoost classifier was added and interpreted with SHAP; overall survival was assessed with Kaplan-Meier curves and time-dependent ROC analysis, with two-sided p < 0.05 as the significance threshold.

Replicationmixed Sample sizeTotal counts from GEO cohorts stated (654 tumour, 226 normal across four datasets); exclusion of samples with missing survival or OS < 1 month noted; no formal power calculation reported; external validation in TCGA, GSE72094, Cho, and VanAllen cohorts; scRNA-seq: 15,293 cells retained after QC from 6 samples GroupsNormal lung vs LUAD tumour; IRG-high vs IRG-low cluster; IRGI high-risk vs low-risk; BEAS-2B vs A549 cell lines (qRT-PCR) Pairingunpaired Randomization/blindingnot stated Dispersionnone Exact p-valuesyes Effect sizesyes Confidence intervalsno Multiplicity correctionnone stated
Statistical tests used
Test Applied to n Assumptions
Empirical Bayes moderated t-statistic (limma R package) Differentially expressed genes between 226 normal and 654 LUAD tumour samples 880 (226 normal + 654 tumour) not stated
Unsupervised consensus clustering (ConsensusClusterPlus, 1,000 iterations) Stratification of LUAD tumour samples into IRG expression clusters 654 tumour samples na
Weighted Gene Co-expression Network Analysis (WGCNA), soft threshold β = 3 Detection of co-expression modules correlated with immune clusters 880 not stated
LASSO-Cox regression with 10-fold cross-validation (glmnet R package) Construction of IRGI prognostic risk signature from 15 candidate IRGs not stated
Kaplan-Meier survival analysis (log-rank test not explicitly named) Overall survival comparison: IRG-high vs IRG-low; IRGI high-risk vs low-risk; external validation cohorts na
Time-dependent ROC / AUC (timeROC R package) Predictive performance of IRGI risk model at 2-, 3-, and 4-year survival endpoints na
XGBoost machine learning classifier (stratified 70/30 train-validation split) IRG-based binary classification; overall AUC = 0.975 reported on validation set na
Independent Student's t-test (two-sided) Between-group comparisons of normally distributed continuous variables not stated
Mann-Whitney U test (two-sided) Between-group comparisons of non-normally distributed continuous variables not stated
2−ΔΔCT comparative CT method qRT-PCR quantification of PLAU and GREM1 in BEAS-2B (normal) vs A549 (LUAD) cell lines not stated
Approaches that could also have been used
  • DEG significance was defined by P < 0.05 combined with a dynamic |logFC| threshold (mean + 2 SD), without an explicit false-discovery-rate correction across the ~20,000 genes tested
    Could also: Applying Benjamini-Hochberg FDR correction (adjusted P < 0.05 or 0.1) directly on the limma output is standard for genome-wide DEG analyses — FDR control directly addresses the multiple-testing problem inherent in genome-wide screens; it would make the DEG list more conservative and is widely expected by reviewers and replication studies
  • A pure LASSO penalty was used for the Cox prognostic model, selecting 2 IRGs from 15 candidates
    Could also: Elastic net Cox regression (mixing L1 and L2 penalties with α selected by cross-validation) could also have been applied — LASSO tends to arbitrarily select one predictor from a correlated group; elastic net handles correlated predictors more stably, which is relevant when IRGs share co-expression modules as identified by WGCNA
  • XGBoost predictive performance was summarised as a single AUC point estimate (0.975) on a one-time 30% hold-out set
    Could also: Bootstrap confidence intervals around the AUC, repeated k-fold cross-validation across multiple random splits, or a calibration plot (reliability diagram) could also accompany the point estimate — A single AUC on one hold-out split carries sampling variability; CIs or repeated cross-validation convey the precision of the estimate and make overfitting easier to detect
  • The optimal number of consensus clusters was selected from the cumulative distribution function curve of ConsensusClusterPlus
    Could also: Gap statistic, average silhouette width, or the Calinski-Harabasz index could also be computed to guide cluster-number selection; non-negative matrix factorisation (NMF) is another established approach for transcriptomic subtyping — Convergent evidence from multiple cluster-validity indices strengthens confidence in the chosen k, reducing reliance on a single graphical criterion
  • Group comparisons used Student's t-test for normally distributed variables and Mann-Whitney U for non-normal data, but the method used to assess normality before this choice is not described
    Could also: Stating the normality-screening criterion (e.g., Shapiro-Wilk test at P < 0.05, or a QQ-plot inspection) would clarify how the decision between parametric and non-parametric tests was made; alternatively, non-parametric tests could be used uniformly throughout — Documenting the normality decision rule improves reproducibility and allows readers to evaluate whether the parametric assumption was appropriately checked
  • Cell-line validation of GREM1 and PLAU expression by qRT-PCR was reported without stating the number of biological replicates, technical replicates, or a dispersion measure
    Could also: Reporting the number of independent experiments (biological replicates), technical replicates per experiment, and a dispersion statistic (SD or SEM) with degrees of freedom follows MIQE guidelines for qRT-PCR — Without replicate counts and dispersion, the statistical confidence of the qRT-PCR result cannot be independently assessed; MIQE compliance is increasingly required for reproducibility
Software: R 4.2.1 · limma · glmnet (LASSO-Cox) · ConsensusClusterPlus · sva / ComBat (batch correction) · survival / survminer · timeROC · Seurat (scRNA-seq) · XGBoost (R package) + shap · ESTIMATE algorithm · TIMER algorithm · TIDE (online tool, tide.dfci.harvard.edu) · STRING / Cytoscape / CytoHubba plugin

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Authors · 6
Citations
2
Impact: low
Foundation confidence
Built on 1 assessed reference(s) · mean reproducibility 50/100
partly built on non-reproducible work
Topics

Assessed papers, coloured by verdict. Click a node to open it.

Built on (assessed references) (1)
Cited by (assessed papers) (0)
  • No assessed neighbours yet — the network grows as more papers are assessed.

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Figures / tables: table
C1
Reported
654 tumour + 226 normal = 880
Reproduced
653 tumour + 226 normal = 879
within tolerance
C2
Reported
484 DEGs (208 up / 276 down)
Reproduced
570 DEGs (215 up / 355 down)
partial
C4
Reported
15 immunoregulation-related genes (named)
Reproduced
15/15 present in merged data
exact
C5
Reported
ConsensusClusterPlus k=2: IRG-high 438 / IRG-low 215 (653)
Reproduced
k=2 reproduced; split removeBatchEffect 460/193, vendored ComBat 384/269 (653)
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 71/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟡2. Endpoint comparability
🟡3. Location of the main deviation
🔴4. Cause of the deviation
🔴5. Derivability / plausibility
🟡6. Severity of the deviation
🔴7. Core claim
🔴8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Main result did not reproduce
Decisive
From: Q5 · Derivability / plausibility 🔴
Main result did not reproduce
Decisive
From: Q7 · Core claim 🔴
Main result did not reproduce
Decisive
From: Q8 · Severity of the miss (overall human judgment) 🔴

This is a non-reproducible-as-shipped paper: the repo ships only 2 of ~12 stages with no input data, so every headline deliverable (GREM1/PLAU LASSO signature, prognostic HR 0.49/AUC 0.794, TCGA HR 1.52, XGBoost AUC 0.975) has no code and could not be reproduced — the gap is on the authors' side. The single shipped stage (consensus k=2 subtyping) reproduces in structure and roughly in magnitude (split 460/193 vs 438/215; DEG 570 vs 484) after reconstructing inputs from public GEO, with deviations attributable to undocumented preprocessing. Three self-consistency red-flags make the central conclusion unsafe: an impossible log2(TPM+1) normalization for microarray inputs, a 654-vs-653 off-by-one, and label leakage rendering the flagship AUC 0.975 circular. Overall red — substantive, with fabrication-suspect signals on q5.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

275.6 k
tokens (I/O) · 18.4 M incl. cache
51 min
runtime · 0.2 CPU-h
2.6 GB
peak RAM
5 (4 failed)
HPC jobs
hummel
machine