Proneural-mesenchymal antagonism dominates the patterns of phenotypic heterogeneity in glioblastoma.
The main results reproduced: recomputed values matched the published ones within tolerance.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Same input data as the authors
- ✓Reported values were directly comparable
- ✓No relevant deviation in data/preprocessing
- ✓No authors-side cause for any deviation
- ✓Reported values are derivable from the shared data
- ✓Any deviation was negligible
- ✓The central claim held under reproduction
- ✓Overall, the reproduction was clean
- Every checked point held up.
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Described well enough to reproduce: YES. The paper's central computational result (Fig 2A: proneural-mesenchymal antagonism = negative Spearman correlation of MES vs NPC/PN signature scores in GBM scRNA-seq) was reproduced from GEO GSE131928 using the authors' own repo (Harshavardhan-BV/GBM_4states @ 0218b51; the registry 'GSEApy' code_url is a text-mining artifact, GSEApy being the cited scoring tool). Compute ran on «our HPC» (SLURM «job», COMPLETED 5:52, node n096; env+data pre-staged on «infra» by front1, SLURM=compute-only). Two layered reproductions, fresh re-run matching the prior 2026-06-16 run bit-for-bit (newer libs gseapy 1.3.0/scipy 1.18.0). Repro A: recomputing the Spearman correlation matrix from the authors' DEPOSITED per-cell AUCell scores reproduces the published Fig 2A matrix to MACHINE PRECISION (max abs diff ~1e-16 across both 16x16 matrices, Smart-seq2 + 10X) -> the figure is fully backed by deposited data, no fabrication indicated. Repro B: a fully INDEPENDENT end-to-end re-derivation from raw GEO TPM using GSEApy ssGSEA confirms the headline antagonism in sign for Smart-seq2 (both axes; NefMES-NefNPC=-0.53, VerMES-VerPN=-0.87, concordances positive, AC-OPC positive; 91/120 signs agree) and for the Verhaak axis in 10X (-0.31; 95/120 signs agree). The one independent-path exception is the Neftel MES-NPC pair in 10X, weakly POSITIVE under ssGSEA (+0.19) vs negative under the authors' AUCell (-0.42) -- a method/preprocessing-sensitivity caveat (authors deliberately used AUCell, not ssGSEA, for single-cell data; 16201 cells reconstructed from GEO vs 15476 used by authors, repo ships no preprocess/ scripts), NOT a fabrication signal. Overall: 1:1 reproduction of the published figure (exact) plus independent confirmation of the central biological claim (robust in Smart-seq2, partial in 10X). NOT ATTEMPTED (out of scope, in scope.md): meta-analysis over ~90 bulk/~30 sc datasets (Figs 4/5), TCGA survival/Cox HRs (Fig 3), GRN/RACIPE + PCA/J-metric (Fig 6/S*), 10X AUCell-from-raw. This re-run also adds the qc_room health block that the prior (requeued) ROOM_RESULT lacked.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 100assessed: 2026-06-16 ⛓ d41e1ccf4f1d
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-22
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-16no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetThe paper tests whether the four previously proposed GBM molecular/cellular subtypes (Proneural/Neural/Classical/Mesenchymal from TCGA bulk data, and NPC-like/OPC-like/AC-like/MES-like from single-cell data) are truly mutually exclusive and antagonistic states, or whether they instead reflect a lower-dimensional axis of heterogeneity.
- ★ The four proposed GBM molecular subtypes (Proneural, Neural, Classical, Mesenchymal / NPC-like, OPC-like, AC-like, MES-like) are not mutually exclusive or independent of one another finding
- ★ Proneural (or NPC-like)-Mesenchymal is the dominant antagonistic axis of GBM phenotypic heterogeneity, with the other two subtypes lying along this spectrum finding
- ★ The PN/NPC-MES antagonism observed at the gene signature score level is also present at the level of individual gene-gene correlations finding
- ★ Meta-analysis of over 100 bulk and single-cell transcriptomic datasets confirms PN/NPC-MES as the most consistently anti-correlated subtype pair across contexts finding
- ★ PN/NPC states associate with higher cell-cycle activity and lower glycolysis and immune evasion (PD-L1), whereas MES states show the opposite functional profile finding
- ★ Distinct transcription factor sets (RUNX1, FOSL2, BHLHE40 for MES; TCF4, TCF12, MYT1 for PN) show coordinated 'teams-like' correlation behavior mirroring the PN-MES antagonism finding
- ★ Transcription factors associated with the proneural state correlate with better patient prognosis, while those associated with the mesenchymal state correlate with worse prognosis finding
- A J-metric was defined to quantify within-gene-set positive correlation versus between-gene-set negative correlation as a measure of antagonism method
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| gene set enrichment scoring (ssGSEA/AUCell) | scRNA-seq GBM tumor samples (GSE131928) | none | subtype signature enrichment scores and pairwise correlations | ssGSEA/AUCell |
| gene set enrichment scoring (ssGSEA) | TCGA-GBM bulk RNA-seq patient cohort | none | subtype signature enrichment scores and pairwise correlations | ssGSEA |
| pairwise gene expression correlation analysis | GSE131928 (scRNA-seq) and TCGA-GBM (bulk RNA-seq) | none | Spearman correlation among individual signature genes | — |
| Principal component analysis (PCA) | TCGA-GBM and GSE131928 gene expression (Neftel and Verhaak signature genes) | none | PC1 loadings by gene set | — |
| meta-analysis of bulk RNA-seq datasets (80 datasets, 90 consolidated samples) | patient tumor samples, mouse models, and cell lines | none | correlation coefficients between subtype scores and with PD-L1, glycolysis, and cell cycle signatures | — |
| meta-analysis of scRNA-seq datasets (26 datasets, 30 consolidated samples) | GBM tumor samples | none | correlation coefficients between subtype scores and functional signatures | — |
| transcription factor correlation analysis | GSE131928 scRNA-seq and TCGA-GBM bulk RNA-seq | none | correlation of TF expression with NPC/PN and MES signature scores | — |
| survival analysis (Kaplan-Meier, Cox proportional hazards) | TCGA-GBM patient clinical data | high vs. low TF expression | progression-free interval and overall survival, hazard ratios | — |
- ▼ Neftel NPC-like and MES-like signature scores are negatively correlated in both GSE131928 and TCGA-GBM datasets
- ▼ Verhaak PN and MES signature scores form the most antagonistic pair among Verhaak subtypes
- ▼ In bulk RNA-seq meta-analysis, Neftel MES-NPC pair was negatively correlated in 33 of 38 significant samples 86.8%
- ▼ In bulk RNA-seq meta-analysis, Verhaak MES-PN pair was negatively correlated in 45 of 48 significant samples 93.8%
- ▼ In scRNA-seq meta-analysis, Verhaak MES-PN pair was negatively correlated in all significant samples 23/23 (100%)
- – MES signatures positively correlate with PD-L1 in most datasets while NPC/PN signatures negatively correlate with PD-L1 up to 97% (NPC) and 91.7% (MES)
- – MES signatures positively correlate with glycolysis in most datasets while NPC/PN signatures negatively correlate with glycolysis up to 98.4% (MES) and 69% (PN)
- – TCF12 and CXXC4 high expression associate with better overall survival; RUNX1 and BHLHE40 high expression associate with worse overall survival HR 0.626-1.63
- correlation 33/38 (86.8%) negatively correlated (bulk RNA-seq meta-analysis, Neftel MES-like vs NPC-like)
- correlation 45/48 (93.8%) negatively correlated (bulk RNA-seq meta-analysis, Verhaak MES vs PN)
- correlation 23/23 (100%) negatively correlated (scRNA-seq meta-analysis, Verhaak MES vs PN)
- other HR = 0.626, p < 0.001 (TCF12 expression and overall survival (TCGA-GBM))
- other HR = 0.684, p < 0.05 (CXXC4 expression and overall survival (TCGA-GBM))
- other HR = 1.42, p < 0.05 (RUNX1 expression and overall survival (TCGA-GBM))
- other HR = 1.63, p < 0.001 (BHLHE40 expression and overall survival (TCGA-GBM))
- count 80 bulk RNA-seq datasets (90 consolidated samples); 26 scRNA-seq datasets (30 consolidated samples) (meta-analysis dataset scope)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This computational study re-analyzed publicly available bulk and single-cell RNA-seq GBM transcriptomic datasets using gene-set enrichment scoring (ssGSEA and AUCell) followed by pairwise Spearman correlations to assess mutual antagonism among four proposed GBM subtypes. A custom J-metric and PCA on individual gene expression levels were used to quantify and visualize the dominant axis of gene-expression variance. A meta-analysis across over 100 datasets catalogued the proportion of datasets showing significant negative versus positive pairwise correlations, and survival analyses using Kaplan-Meier curves and hazard ratios were applied to TCGA-GBM patient data to assess clinical relevance of identified transcription factors.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Spearman correlation (pairwise, on ssGSEA and AUCell enrichment scores) | Pairwise correlations among four subtype signature scores in GSE131928 and TCGA-GBM (Figures 2A, 6A, 6B) | 28-patient scRNA-seq cohort (GSE131928); TCGA-GBM cohort size not stated in excerpt | not stated |
| Spearman correlation (pairwise, on individual gene expression levels) | Gene-level pairwise correlation within and across Neftel and Verhaak signature gene sets in TCGA-GBM and GSE131928 (Figure 3A) | not stated | not stated |
| J-metric (custom within/cross-set correlation summary statistic, Equation 2) | Quantifying extent of within-gene-set positive correlation and cross-gene-set negative correlation for all pairwise subtype combinations (Figure 3B) | not stated | not stated |
| Principal component analysis (PCA) | Gene expression levels within each signature to identify dominant variance axis; PC1 loadings inspected for subtype enrichment (Figure 3C) | not stated | not stated |
| Spearman correlation + proportion counting (meta-analysis) | Correlation between subtype scores in 80 bulk RNA-seq datasets (90 consolidated samples) and 26 scRNA-seq datasets (30 consolidated samples); significance threshold r < -0.3 or r > 0.3 with p < 0.05 (Figures 4A, 4B, 5A, 5B) | 80 bulk RNA-seq datasets (90 consolidated samples); 26 scRNA-seq datasets (30 consolidated samples) | not stated |
| Kaplan-Meier survival analysis with hazard ratios (implied Cox proportional hazards) | Overall survival and progression-free interval in TCGA-GBM patients stratified by high vs. low expression of NPC/PN- and MES-associated transcription factors (Figures 6C, 6D) | TCGA-GBM patient cohort; exact n not stated in excerpt | not stated |
-
Pairwise Spearman correlations were computed across a large number of gene pairs, subtype pairs, and datasets with p < 0.05 as the significance threshold, without a stated multiple-testing correction↳ Could also: A false discovery rate (FDR) correction (e.g., Benjamini-Hochberg) applied across all simultaneous correlation tests would also be standard for large-scale correlation screens — With thousands of gene-pair or subtype-pair tests across many datasets, the expected number of false positives under an uncorrected p < 0.05 threshold can be substantial; FDR control conveys the same findings while quantifying the proportion of discoveries that may be chance associations
-
The meta-analysis across 80 bulk and 26 scRNA-seq datasets used proportion counting (number of datasets with r < -0.3 and p < 0.05 vs. r > 0.3 and p < 0.05) to summarize cross-dataset trends↳ Could also: A formal random-effects meta-analysis pooling Fisher-z-transformed correlation coefficients across datasets would also be a standard approach for this design — Pooling effect sizes with inverse-variance weighting produces a single summary estimate with a confidence interval and heterogeneity statistic (I²), enabling quantitative comparison of effect magnitude across dataset types (bulk vs. single-cell, in vitro vs. in vivo) rather than binary vote-counting
-
Gene-set enrichment was quantified using ssGSEA and AUCell scores, then Spearman-correlated across samples↳ Could also: GSVA (Gene Set Variation Analysis) or weighted scoring approaches (e.g., Seurat module scores, UCell) could also be used for per-sample gene-set enrichment estimation — Different enrichment methods make distinct assumptions about gene-set score distributions; comparing results across two or more methods (as the paper partly does with ssGSEA and AUCell) can demonstrate robustness of correlation patterns to scoring choice
-
The dominant axis of transcriptomic variance was assessed by inspecting PC1 loadings from PCA restricted to signature genes↳ Could also: Non-negative matrix factorization (NMF) or independent component analysis (ICA) could also be applied to identify latent transcriptomic programs without assuming orthogonality — PCA finds orthogonal linear components maximizing variance, which may not align with biologically interpretable axes; NMF enforces non-negativity and has been used specifically to recover GBM cell-state programs in the literature this paper builds on
-
Survival analyses reported hazard ratios (implied from Cox regression) and Kaplan-Meier curves with p-values thresholded at p < 0.05, with high vs. low expression dichotomization for each TF↳ Could also: Reporting 95% confidence intervals for each hazard ratio alongside the point estimate is also standard practice in survival reporting, and using a continuous expression variable in Cox regression avoids the information loss of dichotomization — Confidence intervals convey estimation uncertainty and clinical effect range; dichotomizing a continuous expression variable at an arbitrary threshold (e.g., median) can inflate Type I error and reduce power compared to modeling expression as continuous
-
The custom J-metric was introduced to summarize within-set positive and cross-set negative correlations into a single antagonism score↳ Could also: Established co-expression network measures (e.g., module preservation statistics from WGCNA, or silhouette scores on correlation-based clusters) could also quantify the separation between gene-set expression profiles — Well-documented metrics with known distributional properties allow comparison with external studies and enable inference about statistical significance of the antagonism score through permutation or bootstrap procedures
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
Assessed papers, coloured by verdict. Click a node to open it.
- No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-38433919
Paper: Harshavardhan BV & Mohit Kumar Jolly (2024). Proneural-mesenchymal antagonism dominates the patterns of phenotypic heterogeneity in glioblastoma. iScience 27(4):109184. DOI 10.1016/j.isci.2024.109184. PMID 38433919.
Code & data found
- Authors' own analysis repo (P16 — own code):
https://github.com/Harshavardhan-BV/GBM_4states(default branchmain, latest commit0218b517d0d7d8c7d59bf9da55f016435868cad7, 2024-07-09; no LICENSE file). Stated in the paper's "Data and code availability": "The scripts used for analysis and the processed data are available at: https://github.com/Harshavardhan-BV/GBM_4states". The registry'scode_url(GSEApy) is a text-mining artifact — GSEApy is the tool the authors call; the real repo is GBM_4states. The repo SHIPS the derived output matrices (Output/<GSE>/...), the signature definitions (Signatures/GBM_states.gmt), and the analysis scripts (GSEA/*.py,GSEA/AUCell.R). It does NOT ship thepreprocess/scripts referenced in the README (the repo's top-level tree has nopreprocess/dir), nor the rawData/or intermediateData_generated/. - Data: GEO
GSE131928(Neftel et al. 2019 GBM scRNA-seq). Two samples:GSM3828672= Smart-seq2 GBM IDHwt —..._processed_TPM.tsv.gz(gene×cell)GSM3828673= 10X GBM IDHwt —..._processed_TPM.tsv.gz(gene×cell)
Pipeline (from Methods + repo)
Per-cell signature scoring → pairwise Spearman correlation across cells.
- Scoring: ssGSEA (
gseapy.ssgsea,sample_norm_method='rank') for bulk; AUCell (R/AUCellfrom SCENIC) for single-cell. Signatures = the 16 sets inSignatures/GBM_states.gmt: Neftel meta-modules NefMES/NefAC/NefOPC/NefNPC, Verhaak subtypes VerPN/VerNL/VerCL/VerMES, plus OXPHOS/GLYC/FAO/G2M/cell-cycle/ PD-L1 etc. - Correlation:
scipy.stats.spearmanr/pandas.DataFrame.corr(method='spearman')on the per-cell score matrix (GSEA/Corr_GSEA.py). - Figure 2A = clustermap of that correlation matrix (
GSEA/hmap_GSEA.py), shipped asOutput/GSE131928/GSEA/GSM382867{2,3}-corr.tsv.
IN SCOPE (pipeline-derived; attempted)
- Fig 2A correlation matrix, exact recomputation (Repro A). Recompute the
Spearman correlation matrix from the authors' deposited per-cell AUCell scores
(
Output/GSE131928/AUCell/GSM*-AUCell.csv) and compare element-wise to the shippedOutput/GSE131928/GSEA/GSM*-corr.tsv. Verifies the published figure is reproducible from the deposited intermediate data (anti-fabrication check). - Central antagonism claim, independent end-to-end (Repro B). Download
GSE131928 TPM matrices from GEO, run GSEApy ssGSEA (the cited tool) with
GBM_states.gmt, compute Spearman correlations across cells, and test the paper's central claims on the resulting matrix:- NefMES–NefNPC < 0 and VerMES–VerPN < 0 (antagonism, Fig 2A);
- NefMES–VerMES > 0 and NefNPC–VerPN > 0 (cross-signature concordance);
- NefAC–NefOPC ≥ 0 (not antagonistic);
- sign agreement vs the authors' shipped AUCell-based matrix across all 120 off-diagonal signature pairs. This is a true independent reproduction with a different scoring method, so it also tests the paper's robustness-to-method claim.
OUT OF SCOPE (not attempted, and why)
- Meta-analysis over ~90 bulk / ~30 sc datasets (Figs 4/5, the "86.8%/93.8%/…"
percentages): requires downloading and preprocessing ~120 GEO/TCGA/CGGA/CCLE
datasets with un-shipped, case-by-case
preprocess/scripts. Tractable but very large; deferred beyond the core. (The single-dataset GSE131928 result is the Fig 2A anchor of the same claim.) - Survival / Cox-regression HRs (TCGA-GBM), GRN/RACIPE, PCA, J-metric, Figs 3/6:
separate pipelines (R
survival, network modeling) on other data; not the central scoring-correlation result. Not attempted here. - Wet-lab / manual / external-DB content: out of scope by constructi
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
Exact, clean reproduction. Repro A regenerates the entire Fig 2A Spearman matrix from the authors' deposited per-cell AUCell scores to machine precision (max|Δ|~1e-16 across both 16x16 matrices), so all eight reported claims are fully derivable from shared GEO data — no fabrication signal. Repro B, a genuinely independent ssGSEA re-derivation from raw TPM, confirms the central proneural-mesenchymal antagonism in sign for Smart-seq2 (both axes) and the Verhaak 10X axis; the lone exception (Neftel MES-NPC in 10X, +0.19 vs -0.42) sits on our methodology side (ssGSEA vs the authors' AUCell, plus un-shipped preprocessing). Overall a 1:1 reproduction with the central conclusion robustly upheld.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.