Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Machine learning-based identification of biomarkers and drugs in immunologically cold and hot pancreatic adenocarcinomas.

J Transl Med · 2024
L1 42/100 3/4
Why this verdict

The main result did not reproduce in this reproduction attempt. Where our recomputation produced values that differ from the published ones, those discrepancies are listed below. This is a single automated attempt — not peer review and not a finding of error or misconduct — and differences can also arise from data access, undocumented parameters or the computing environment. The verdict can be contested via “report an error”.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Main result did not reproduce
Decisive
From: Q5 · Derivability / plausibility 🔴
Main result did not reproduce
Decisive
From: Q8 · Severity of the miss (overall human judgment) 🔴
✓ What held up
  • Nothing in this column.
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🟡Reported values were only indirectly comparable
  • 🔴A deviation arose in the data or preprocessing
  • 🔴A deviation was attributed to the published material
  • 🔴Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🔴Overall, the reproduction showed a material discrepancy
How its reproducibility compares
42/100
Reproducibility score
1.8 SD below mean
vs. all fields · 1173 studies
🎯 Scores higher than 4% of all assessed papers rank 1120 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

NOT described well enough for full 1:1. The authors' repo (github.com/sangmm12/ML_Hot-cold @ 2699024) is code-only: per-figure R plotting scripts with hardcoded D:/R and «path» paths reading intermediate files (hot/cold cluster labels, merged TCGA+ICGC count matrix, HR-prefiltered gene lists, risk scores, model objects) that are NOT shipped; no README, no env, no data, no entry point. The discovery cohort is TCGA-PAAD+ICGC (not the registry's GSE85916, which is 1 of 4 microarray validation sets). So none of the headline numbers (4620 DEGs, training AUC 0.979-0.986, validation AUCs incl. GSE85916 0.647, oncoPredict drugs) are derivable from shipped artifacts. Following brief rule 2, I instead ran the paper's OWN named tools (ESTIMATE + limma + ConsensusClusterPlus) on the paper's primary PUBLIC cohort TCGA-PAAD (UCSC Xena, 178 tumours, «our HPC» «job»). PARTIAL result: the foundational immune-hot/cold dichotomy reproduces qualitatively - a 2-cluster split has significantly higher immune infiltration in the hot subtype (ImmuneScore 761.5 vs 403.6, p=0.005; corroborated by leukocyte markers p=0.003). But the stronger claim (hot higher on all 3 ESTIMATE scores) only partially holds (Stromal inverted, ESTIMATE n.s.), and DEG counts differ (2872 vs 4620, ratio inverted) - both explained by cohort (TCGA-only vs ComBat-merged) and unshipped labels. NOT ATTEMPTED (hard 20% / not derivable): the 11-algorithm/137-combination ML model and all its AUCs, WGCNA modules, single-cell layer, oncoPredict drugs. Fabrication concern = possible: headline numbers not regenerable from shipped code+data; a lncRNA/pseudogene-dominated signature is 'validated' on a microarray platform that cannot measure most of its features.

💻 Code ↗ 🗄 Data: GSE85916

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 42
    assessed: 2026-06-14 ⛓ 0fbfdac2ec61
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-14
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

Pancreatic adenocarcinomas (PAADs) often have an immunosuppressive 'cold' tumor microenvironment associated with immune checkpoint blockade resistance; the study tests whether immune-based classification combined with machine learning can identify gene signatures, biomarkers, prognostic models, and drug candidates to distinguish and convert cold versus hot PAAD tumors.

Core claims
  • PAAD patients can be stratified into immunologically hot and cold subtypes with distinct immune infiltration and survival outcomes. finding
  • A novel immune-related gene signature (DPIRGs: Downregulated in hot tumors, Prognostic, and Immune-Related Genes) was constructed via Cox regression and WGCNA. resource
  • A consensus machine-learning prognostic model combining Survival Random Forest and partial least squares regression Cox (plsRcox) using DPIRGs gave optimal PAAD prognostic performance. method
  • Biomarkers and potential therapeutic targets including PLEC, TRPV1, and ITGB4 were identified by machine learning and validated for cell-type-specific expression. finding
  • Drug candidates including thalidomide, SB-431542, and bleomycin A2 were identified by their ability to favorably modulate DPIRG expression to turn cold tumors hot. finding
  • DPIRG molecular mechanisms were characterized through genetic/epigenetic alterations, immune infiltration, pathway enrichment, and miRNA regulation analyses. mechanism
  • Immune-related and ligand-receptor signaling pathways are enriched in genes upregulated in hot tumors, whereas metabolism, epidermis development, cytoskeleton, and chromatin pathways are enriched in downregulated genes. finding
Experimental setups
Assay System Perturbation Readout Platform
bulk RNA-seq expression and clinical analysis human PAAD patient cohorts (TCGA-PAAD, ICGC-PAAD-AU, ICGC-PAAD-CA, GSE85916, GSE28735, GSE62452, GSE78229) none gene expression, immune cell fractions, survival/prognosis
single-cell RNA sequencing (scRNA-seq) analysis human PAAD tissue (8 datasets: CRA001160, GSE111672, GSE141017, GSE148673, GSE154778, GSE158356, GSE162708, GSE165399) none cell type-specific biomarker expression
immunohistochemistry / proteomic validation human PAAD and normal tissue none protein expression of biomarkers Human Protein Atlas (HPA)
immune cell deconvolution (CIBERSORT) human PAAD bulk RNA expression data none fractions of 22 immune cell types IOBR package v0.99.9 / CIBERSORT
consensus clustering and tumor purity scoring human PAAD patients none hot/cold immune cluster assignment, StromaScore/ImmuneScore/EstimateScore ConsensusClusterPlus, ESTIMATE R package v1.0.13
differential gene expression and miRNA correlation analysis human PAAD hot vs cold tumors none DEGs and correlated miRNAs limma voom; TCGA miRNA data
machine learning prognostic modeling human PAAD cohorts (TCGA+ICGC training, GSE validation) none C-index, risk score, survival prediction 11 ML algorithms incl. Survival RF, plsRcox, LOOCV
molecular docking and drug sensitivity prediction target proteins from selected genes / PAAD expression data drug (small-molecule compounds) binding patterns, predicted IC50, drug-gene correlation DOCK v6.10, UCSF Chimera, ZINC15, oncoPredict v0.2 (GDSC)
Key results
  • Optimal number of immune consensus clusters was 2, with a significant survival difference between clusters. k=2 clusters
  • 2055 genes upregulated and 2565 genes downregulated in hot versus cold tumors. 2055 up / 2565 down
  • 82 UPIRGs and 96 DPIRGs identified as prognostic immune-related genes (178 total). 82 UPIRGs / 96 DPIRGs / 178 total
  • Survival RF + plsRcox model on DPIRGs achieved the best C-index and was selected as the optimal (mixed) prognostic model from 137 ML combinations. selected from 137 combinations
  • Turquoise module showed strong correlation between gene significance and module membership (cold-related). r=0.81
  • Pink and black module GS-MM correlations were 0.46 and 0.23 respectively. pink=0.46, black=0.23
  • Intersection of DEGs and WGCNA gave 118 overlapping upregulated/HRG genes and 375 overlapping downregulated/CRG genes. 118 and 375 genes
  • Hot tumors had greater infiltration of naïve B cells, CD8+ T cells, M1/M2 macrophages, activated NK cells and others, while cold tumors had more Tregs, resting NK cells, M0 macrophages, and activated mast cells. p<0.001
Key statistics
  • correlation 0.81 (GS vs MM correlation in turquoise WGCNA module)
  • correlation 0.46 (GS vs MM correlation in pink WGCNA module)
  • correlation 0.23 (GS vs MM correlation in black WGCNA module)
  • count 2055 upregulated, 2565 downregulated genes (DEGs in hot vs cold tumors (adj.P<0.01, |logFC|>1))
  • count 137 (ML algorithm combinations evaluated for prognostic model)
  • pvalue p<0.001 (immune cell infiltration differences between hot and cold tumors)
  • count 96 DPIRGs / 82 UPIRGs (prognostic immune-related genes (Cox P<0.05))
  • other 11% (5-year survival rate of pancreatic adenocarcinoma)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This computational bioinformatics study used multi-cohort public bulk RNA-seq data (TCGA, ICGC, GEO) to classify PAAD patients into immunologically hot and cold subtypes via consensus clustering of CIBERSORT-derived immune cell fractions, then identified prognostic immune-related gene signatures (DPIRGs) through WGCNA and univariate Cox regression. A consensus machine learning framework evaluated 137 algorithm combinations under LOOCV, selecting Survival Random Forest plus plsRcox as the optimal prognostic model assessed by C-index and AUC. Differential expression was analyzed with limma-voom, functional enrichment with clusterProfiler (GO/KEGG) and GSVA, and drug–target interactions were explored via molecular docking.

Replicationunclear Sample sizeMultiple public cohorts used (TCGA-PAAD, ICGC-PAAD-AU, ICGC-PAAD-CA, four GEO bulk datasets, eight scRNA-seq datasets); exact per-cohort patient counts not stated in the provided text GroupsImmunologically hot vs cold PAAD tumor subtypes; high-risk vs low-risk prognostic groups Pairingunpaired Randomization/blindingnot stated Dispersionnone Exact p-valuesno Effect sizesyes Confidence intervalsno Multiplicity correctionFDR adjustment (Benjamini-Hochberg implied by 'adjusted P'); adj.P < 0.01 for DEGs; adj.P < 0.05 for GO/KEGG enrichment
Statistical tests used
Test Applied to n Assumptions
limma-voom differential expression DEG analysis between hot and cold tumor clusters (combined TCGA+ICGC cohort); threshold adj.P < 0.01, |logFC| > 1 not stated
Consensus clustering (ConsensusClusterPlus, 1000 iterations, CDF curve for k selection) Immune subtype identification (hot vs cold) from 22 CIBERSORT immune cell fractions not stated
Univariate Cox proportional hazards regression Prognostic gene selection from WGCNA–DEG intersect genes; threshold P < 0.05 not stated
Machine learning ensemble (Survival RF + plsRcox selected as optimal from 137 combinations) with leave-one-out cross-validation (LOOCV) Prognostic model construction and validation; TCGA+ICGC training, GSE cohorts validation; performance by C-index and AUC not stated
Pearson correlation coefficient Gene–miRNA correlations (threshold r > 0.7); WGCNA module–immune cell correlations; GS–MM within modules not stated
Wilcoxon rank-sum test or t-test (selection per comparison not specified) Comparisons of continuous variables between patient subgroups (e.g., immune scores, immune cell fractions) not stated
Chi-squared test Comparisons of categorical variables between patient subgroups not stated
Gene Set Variation Analysis (GSVA, default parameters) Hallmark pathway differences between hot and cold tumor samples not stated
GO and KEGG enrichment (clusterProfiler); threshold adj.P < 0.05 Functional annotation of upregulated and downregulated DEGs in hot vs cold tumors not stated
Approaches that could also have been used
  • Pearson correlation was used for gene–miRNA associations and WGCNA module–immune cell correlations
    Could also: Spearman rank correlation could also be used for these associations — Spearman correlation does not assume linearity or normality, which may be more appropriate for RNA-seq expression values that can be skewed or contain outliers; reporting the choice explicitly would also aid reproducibility
  • Univariate Cox regression was used to pre-filter genes before ML model construction
    Could also: A penalized Cox regression (LASSO or Elastic Net) could also perform variable selection and model fitting simultaneously, without a separate univariate screening step — Penalized regression handles correlated predictors jointly; marginal univariate screening in high-dimensional data can inflate the number of apparently significant genes when predictors are correlated, whereas penalized approaches shrink and select in one step
  • The optimal ML combination was selected by comparing C-index across 137 combinations using the same LOOCV loop used to estimate performance
    Could also: A nested cross-validation design could also be used, with an outer loop for performance estimation and an inner loop for model/algorithm selection — When model selection and performance estimation share the same cross-validation folds, reported performance can be optimistically biased; nested CV separates these steps and yields a less biased performance estimate
  • CIBERSORT (via IOBR) was the sole method used to estimate immune cell fractions for subtype clustering
    Could also: Alternative or complementary deconvolution methods such as TIMER2, EPIC, or xCell could also be applied, or results compared across methods — Different deconvolution algorithms use distinct reference signatures and assumptions about cell-type mixtures; concordance across methods can support the robustness of subtype assignments
  • Model performance (C-index, AUC) was reported as point estimates without confidence intervals
    Could also: Bootstrapped or cross-validation-derived confidence intervals for C-index and time-dependent AUC could also be reported — Confidence intervals convey the precision of performance estimates, which is particularly informative when external validation cohorts differ in size, and allow readers to assess whether differences between models are likely to reflect true differences
  • The paper states that either the Wilcoxon rank-sum test or the t-test was used for continuous variable comparisons, without specifying which was applied to each individual comparison
    Could also: Pre-specifying the selection criterion (e.g., a normality test or a consistent rule such as always using the non-parametric test) and reporting it per comparison is also standard practice — Transparent test selection allows readers to assess the appropriateness of the inference for each specific comparison and supports reproducibility
Software: R 4.0.5 · limma (voom algorithm) · IOBR (wrapping CIBERSORT) 0.99.9 · ConsensusClusterPlus · ESTIMATE 1.0.13 · clusterProfiler · GSVA · WGCNA · survival 3.2-7 · oncoPredict 0.2 · sva (ComBat batch correction) · DOCK 6.10 · UCSF Chimera 1.15 · LigPlus 2019

Result convergence & founder nodes

Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
22
Impact: medium
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-39152432

Paper: Ge et al. 2024, J Transl Med 22:769. "Machine learning-based identification of biomarkers and drugs in immunologically cold and hot pancreatic adenocarcinomas." DOI 10.1186/s12967-024-05590-0 · PMCID PMC11328457. Code: https://github.com/sangmm12/ML_Hot-cold (commit 2699024, main, pushed 2024-01-18; 47 KB; code-only, no license, no README).

Data the paper used

  • Training / discovery cohort: TCGA-PAAD + ICGC-PAAD-AU + ICGC-PAAD-CA, merged into one matrix ("3data") with the ComBat batch-correction algorithm. (NOT the registry accession.)
  • Validation cohorts (microarray GEO): GSE85916, GSE28735, GSE62452, GSE78229. GSE85916 is the registry's data_accession but is only 1 of 4 validation sets.
  • Single-cell: 8 scRNA-seq datasets (CRA001160, GSE111672, …) — out of scope.

Pipeline map (what produces each reported result)

Fig Pipeline Reported result
1 C–H CIBERSORT (22 cells) + ConsensusClusterPlus → 2 clusters; ESTIMATE Immune/Stroma/ESTIMATEScore; voom-limma DEG; clusterProfiler GO/KEGG; GSVA hallmark k=2 optimal; hot vs cold differ in immune scores; 2055 up / 2565 down DEGs (4620 total)
2 WGCNA on merged TPM hot/cold co-expression modules
3,7,8 11 ML algorithms (ANN, RSF, LASSO, Enet, XGBoost, plsRcox, SVM…), 137 combinations → RSF + plsRcox prognostic model training time-ROC AUC 0.979 / 0.983 / 0.986 (1/2/3 yr)
7/9 model riskscore validation GSE28735 0.735 · GSE62452 0.710 · GSE78229 0.789 · GSE85916 0.647 · combined 0.760
9,10 oncoPredict (v0.2) IC50 thalidomide, SB-431542, bleomycin A2, dasatinib, …

Hub genes named: top-5 AL591135.1, AL158201.1, AHNAK2, AK3P5, CEP295NL; 10-gene hot/cold panel HMGA1P2, C1orf195, ITGB4, AC087257.2, AC011611.5, AC034105.3, AC110373.1, AL449212.1, AL354733.3, DCST1 (AUC 0.715–0.728).

In scope (attempted)

  • C1/C2 — the foundational immune hot/cold dichotomy on the paper's primary public cohort (TCGA-PAAD, UCSC-Xena): reproduce with named third-party tools that immune-infiltration consensus clustering yields 2 subtypes (k=2) that separate by ImmuneScore (hot = high). This is the load-bearing step every downstream result keys on. Per brief rule 2, running the named tools on the paper's own data is a valid reproduction even though the exact labels aren't shipped.
  • C3 — DEG up/down counts between the two immune clusters (voom-limma, the paper's method) vs reported 2055/2565.

Out of scope / NOT reproducible from shipped artifacts (why)

  1. All headline ML AUCs (0.979/0.983/0.986; validation 0.647–0.789). The RSF+plsRcox model's gene signature and coefficients are not shipped (risk.csv, model objects, exp_*_HRp.csv HR-prefiltered gene lists are all absent). The riskscore cannot be recomputed → not derivable. docs_insufficient.
  2. Exact 4620-DEG number on the merged cohort. Requires the merged 3-dataset ComBat matrix and the authors' exact hot/cold labels (clusterresult_hc.csv), neither shipped; cluster label/orientation is arbitrary without them.
  3. WGCNA modules, GO/KEGG, GSVA, single-cell, oncoPredict drugs — all depend on the same unshipped intermediates (8 scRNA sets, 137-combo ML framework). 80/20: skipped.

Honesty flags (carry into AUDIT)

  • Repo = figure-plotting scripts with absolute Windows (D:/R/…) + Linux («path») paths to intermediate CSV/TXT files that are not in the repo. No README, no env, no entry point, no data, no labels, no model.
  • The prognostic signature is dominated by lncRNAs / pseudogenes (AL591135.1, AK3P5, HMGA1P2, ACTBP7, FTH1P4/12, KRT8P33, KRT18P7, ANXA2P1, AC0……). Yet it is "validated" on microarray GSE85916 — most of these features are not present on the array platform, making the reported GSE85916 AUC 0.647 hard to derive.
  • Training AUC 0.979–0.986 with a 137-combina
Figures / tables: Fig.1AFig.1Fig.7
C1
Reported
immune-infiltration consensus clustering, optimal k=2
Reproduced
2-way split obtainable on TCGA-PAAD (107 hot / 71 cold); PAC stability did not uniquely favour k=2
partial
C2
Reported
hot > cold on ImmuneScore/StromalScore/ESTIMATEScore
Reproduced
ImmuneScore hot 761.5 vs cold 403.6 (Wilcoxon p=0.005) + leukocyte-marker proxy p=0.003 reproduced; StromalScore inverted; ESTIMATEScore n.s.
partial
C3
Reported
2055 up / 2565 down DEGs (4620) hot vs cold
Reproduced
1679 up / 1193 down (2872) at padj<0.05 |lfc|>0.5; 445/299 strict; same order of magnitude, up/down ratio inverted
did not match
C4
Reported
RSF+plsRcox training time-ROC AUC 0.979/0.983/0.986
Reproduced
not attempted - model signature/coefficients not shipped
partial
C5
Reported
validation AUC GSE85916 0.647
Reproduced
not attempted - model not shipped; lncRNA/pseudogene signature on microarray flagged
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 42/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟡2. Endpoint comparability
🔴3. Location of the main deviation
🔴4. Cause of the deviation
🔴5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🔴8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Main result did not reproduce
Decisive
From: Q5 · Derivability / plausibility 🔴
Main result did not reproduce
Decisive
From: Q8 · Severity of the miss (overall human judgment) 🔴

The foundational immune-hot/cold dichotomy reproduces qualitatively on TCGA-PAAD (ImmuneScore hot 761.5 vs cold 403.6, p=0.005), but the paper's actual deliverable — the RSF+plsRcox prognostic model and its AUCs (0.979–0.986 training, 0.647 on GSE85916) — cannot be reproduced at all because the repo ships only figure-plotting scripts with hardcoded paths to never-deposited data, labels, gene lists and model objects. The verifiable deviations are on the authors'/data-availability side (non-derivable headline numbers) compounded by directional inversions (StromalScore, DEG up/down ratio) attributable to the unshipped merged cohort. Combined with overfitting-prone AUCs and a lncRNA/pseudogene signature 'validated' on a microarray that cannot measure it, this is fabrication-suspect / critical on q5 and q8, with the core conclusion only partially standing.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

168.9 k
tokens (I/O) · 9.4 M incl. cache
16 min
runtime · 0 CPU-h
0.4 GB
peak RAM
1
HPC jobs
hummel
machine