Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Integrated multiomic analysis reveals disulfidptosis subtypes in glioblastoma: implications for immunotherapy, targeted therapy, and chemotherapy.

Front Immunol · 2024
L1 73/100 3/4
Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +4
✓ What held up
  • Reported values were directly comparable
  • Reported values are derivable from the shared data
  • Any deviation was negligible
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
73/100
Reproducibility score
at the mean
vs. all fields · 1173 studies
🎯 Scores higher than 41% of all assessed papers rank 664 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

INTERIM safety-net version. C1 (scRNA-seq QC retained-cell count on GSE182109 via Scanpy+Scrublet, the third-party tool named in the code link) reproduced within-tol on «our HPC» «job» (COMPLETED): 219,652 cells over the 40 GBM samples / 16 patients (patient count matches paper exactly after excluding 4 LGG samples) and 240,168 over all 44 GSM; paper reports 227,584, bracketed by these and within 3.5% on the patient-matched cohort. The apparent red flag (paper retains MORE cells than the original GSE182109 paper's 201,986) is explained by the paper's permissive QC thresholds, NOT fabrication. Downstream C2/C4 best-effort attempt in progress.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 73
    assessed: 2026-06-15 ⛓ a5fad0ca92bf
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-23
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

Disulfidptosis-related gene expression patterns can be used to molecularly classify glioblastoma (GBM) patients into prognostically distinct subtypes that differ in tumor immune microenvironment and response to immunotherapy, targeted therapy, and chemotherapy.

Core claims
  • Consensus clustering on 32 disulfidptosis-associated genes stratifies GBM patients into two subtypes, DRGcluster A and B, with distinct survival outcomes. finding
  • DRGcluster A patients have improved overall survival compared to DRGcluster B across CGGA, TCGA, and Tiantan cohorts. finding
  • DRGcluster subtypes differ in tumor immune microenvironment composition and predicted response to immunotherapy. finding
  • DRGcluster B is associated with higher IDH mutation rate, 1p19q codeletion, and MGMT promoter methylation in CGGA cohort. finding
  • An 8-gene LASSO-derived High/Low-Risk Disulfidptosis Predictor stratifies patients into risk groups with prognostic and drug-response relevance. method
  • Risk groups defined by the predictor show significant differences in IC50 values for chemotherapy and targeted therapy drugs (pRRophetic algorithm). finding
  • DRGcluster B shows higher occurrence of chromosome 7 gain and chromosome 10 loss, consistent between TCGA and Tiantan data. finding
  • scRNA-seq data from 16 GBM patients (GSE182109) were used to characterize macrophage/microglia gene signature activity. method
Experimental setups
Assay System Perturbation Readout Platform
bulk RNA-seq (transcriptome sequencing) 26 fresh frozen GBM tumor specimens, Tiantan Hospital none gene expression (FPKM) for disulfidptosis gene classification Illumina HiSeq; STAR alignment; featureCounts
public RNA-seq / clinical data analysis TCGA GBM cohort (n=135) and CGGA GBM cohort (n=314) none gene expression, survival, clinical/molecular features UCSC Xena; CGGA database
consensus clustering (k-means, unsupervised) GBM patients (CGGA, TCGA, Tiantan) none classification into DRGcluster A/B subtypes based on 32 disulfidptosis genes R (1000 iterations, 80% resampling)
somatic mutation and copy number alteration analysis TCGA GBM patients (mutation n=390; CNA n=628) none tumor mutational burden (TMB), CNA burden, mutation frequency/type in disulfidptosis genes maftools, GenVisR, GISTIC 2.0, RCircos
immune deconvolution / tumor microenvironment scoring GBM samples (TCGA/CGGA/Tiantan expression data) none Immune/Stromal/ESTIMATE scores, tumor purity, 22 immune cell fractions, ssGSEA enrichment of 29 immune markers ESTIMATE, CIBERSORT, ssGSEA
differential gene expression and pathway enrichment analysis GBM samples, DRGcluster A vs B none DEGs (FDR<0.01, |FC|>1.5), GO/KEGG pathway enrichment, GSVA pathway scores limma, WebGestaltR, GSVA (R)
drug sensitivity prediction GBM samples (expression profiles) vs cell line training set in silico drug response modeling predicted IC50 for chemotherapy/targeted drugs pRRophetic package, 10-fold cross-validation
single-cell RNA-seq 42 samples from 16 GBM patients (GSE182109) none cell clustering (UMAP), macrophage/microglia gene signature scores Scanpy v1.8.2, Scrublet, BBKNN
Key results
  • Optimal cluster number determined as k=2, consistent across TCGA, CGGA, and Tiantan datasets
  • DRGcluster A patients showed improved overall survival vs DRGcluster B
  • Significant gender distribution difference between DRGclusters in CGGA cohort p=0.031
  • DRGcluster B showed higher tendency toward IDH mutation in CGGA and TCGA cohorts p<0.001 (CGGA); p=0.025 (TCGA)
  • 1p19q codeletion significantly enriched in DRGcluster B (CGGA) p<0.001
  • MGMT promoter methylation differed significantly between DRGclusters (CGGA) p<0.01
  • 8-gene predictor validated for survival stratification in CGGA and TCGA test sets
  • 227,584 single-cell transcriptomes retained after QC from 42 samples/16 patients n=227,584 cells
Key statistics
  • count 314 CGGA GBM patients enrolled (training set for predictor)
  • count 135 TCGA GBM patients enrolled (test set for predictor)
  • count 26 Tiantan GBM patients (in-house RNA-seq cohort)
  • count 390 GBM patients for somatic mutation analysis (TMB/mutation analysis)
  • count 628 GBM patients for CNA analysis (GISTIC 2.0 CNA burden analysis)
  • count 227,584 single-cell transcriptomes retained after QC (scRNA-seq from GSE182109, 16 patients/42 samples)
  • pvalue gender difference p=0.031 (CGGA DRGcluster A vs B)
  • fold_change DEG threshold |FC|>1.5, FDR<0.01 (limma differential expression criteria between DRGcluster A and B)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This retrospective multi-cohort transcriptomic study classified GBM patients from TCGA (n=135), CGGA (n=314), and an institutional Tiantan cohort (n=26) into two disulfidptosis-based subtypes (DRGcluster A/B) via unsupervised consensus k-means clustering on 32 genes. Between-subtype survival was evaluated with Kaplan-Meier curves and two-sided log-rank tests. Differential gene expression was quantified using limma with Benjamini-Hochberg FDR correction, and an 8-gene prognostic predictor was constructed with LASSO regression in the CGGA training set and validated in the TCGA test set.

Replicationbiological Sample sizeCGGA training set n=314, TCGA test set n=135, Tiantan institutional cohort n=26; no formal power calculation or sample size justification stated GroupsDRGcluster A vs DRGcluster B; high-risk vs low-risk predictor groups Pairingunpaired Randomization/blindingnot stated Dispersionnone Effect sizesyes Confidence intervalsno Multiplicity correctionBenjamini-Hochberg FDR
Statistical tests used
Test Applied to n Assumptions
Two-sided log-rank test Overall survival comparison between DRGcluster A and B across CGGA, TCGA, and Tiantan cohorts CGGA n=314, TCGA n=135, Tiantan n=26 not stated
limma moderated t-test (empirical Bayes) Differential gene expression between DRGcluster A and B; GSVA KEGG pathway score comparisons between subtypes CGGA n=314 (training set); TCGA n=135 (test set) not stated
LASSO regression (glmnet) Feature selection for 8-gene High/Low-Risk Disulfidptosis Predictor construction CGGA n=314 not stated
Independent Student's t-test Inter-group comparisons of normally distributed continuous variables (as stated in Statistical Analysis section) not stated
Mann-Whitney U test Comparisons involving non-normally distributed or categorical variables not stated
Chi-square test Categorical data comparisons including clinical feature distributions in Table 1 not stated
Kruskal-Wallis test Multi-group comparisons (e.g., disulfidptosis gene expression across WHO grade II, III, IV groups) not stated
Pearson correlation Relationships between normally distributed continuous variables not stated
Spearman correlation Relationships between non-normally distributed continuous variables not stated
10-fold cross-validation with ridge regression (pRRophetic) Predictive accuracy evaluation for IC50 estimates of chemotherapy and targeted therapy drugs not stated
Approaches that could also have been used
  • Survival differences between DRGcluster A and B were assessed with unadjusted Kaplan-Meier log-rank tests, while IDH status, MGMT methylation, and 1p19q codeletion differed significantly between clusters
    Could also: A multivariable Cox proportional hazards model adjusting for IDH status, MGMT methylation, age, and treatment could also be used — Adjusting for covariates that differ between clusters would allow estimation of the independent prognostic contribution of disulfidptosis subtype beyond its correlation with established GBM molecular markers
  • Multiple clinical feature comparisons between DRGcluster A and B in Table 1 (age, gender, IDH, MGMT, 1p19q, chemotherapy, radiotherapy) were tested with individual unadjusted p-values
    Could also: A Bonferroni or Benjamini-Hochberg FDR correction applied across the family of Table 1 tests could also be used — Applying a multiplicity correction to the full set of clinical variable comparisons would be consistent with the FDR adjustment already used for the omics-level analyses in the same paper
  • Unsupervised subtype discovery used consensus k-means clustering on 32 disulfidptosis genes with optimal k selected by CDF area change and PAC algorithm
    Could also: Non-negative matrix factorization (NMF) or consensus hierarchical clustering with Ward linkage would also be established unsupervised subtyping approaches — NMF is widely used for expression-based cancer subtyping and yields metagene-based archetypes; hierarchical clustering with alternative linkage methods can provide a complementary view of cluster structure and stability across different algorithmic assumptions
  • The 8-gene LASSO predictor was trained on CGGA (n=314) and validated in TCGA (n=135) without a bootstrap or repeated cross-validation stability assessment of the variable selection step
    Could also: Bootstrap stability selection or repeated nested cross-validation applied over the LASSO regularization path could also be used to assess gene-selection robustness — These approaches quantify how consistently each gene is selected across resampled datasets, distinguishing robustly selected genes from those that may vary with a single training partition
  • Drug sensitivity (IC50) was predicted using the pRRophetic ridge regression algorithm trained on GDSC cell line pharmacogenomic data without reporting prediction uncertainty
    Could also: Elastic net, random forest, or gradient boosting trained on the same GDSC data, accompanied by prediction intervals or cross-validated R² values, could also be reported — Alternative learners may capture non-linear gene-drug relationships; reporting prediction uncertainty would allow readers to judge whether IC50 differences between groups exceed the estimation error
  • Continuous outcomes such as immune deconvolution scores and GSVA enrichment scores were compared between groups without reporting a dispersion measure
    Could also: Reporting SD, IQR, or 95% confidence intervals alongside group means or medians is also standard practice — Dispersion measures allow readers to judge effect magnitude relative to within-group variability, which is particularly informative for immune scores that can be heterogeneous across GBM samples
Software: R/limma R 4.3.1 and R 4.1.3 (limma version not stated) · R/glmnet (LASSO regression) · R/GSVA · R/WebGestaltR · R/pRRophetic · R/maftools · R/GenVisR · R/RCircos · Python/Scanpy 1.8.2 · SPSS 27.0 · GISTIC 2.0 · STAR (RNA-seq alignment) · featureCounts · Scrublet (doublet detection) · BBKNN (batch correction)

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
8
Impact: low
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

C1
Reported
227,584 single-cell transcriptomes retained after QC (42 samples / 16 GBM patients, GSE182109, Scanpy+Scrublet)
Reproduced
219,652 retained over 40 GBM samples / 16 patients (LGG excluded); 240,168 over all 44 GSM
within tolerance
C1b
Reported
42 samples from 16 GBM patients used for scRNA-seq
Reproduced
GEO ships 44 GSM / 18 patients; dropping the 2 LGG (low-grade glioma) patients gives exactly 16 GBM patients but 40 samples (not 42)
partial
C1c
Reported
paper retains 227,584 cells vs the original GSE182109 publication's 201,986 (anomaly: more cells from fewer samples)
Reproduced
independent re-run with the paper's stated thresholds yields 240,168 (all 44) / 219,652 (40 GBM) - the higher count is explained by permissive QC, not fabrication
within tolerance

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 73/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟢2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟢5. Derivability / plausibility
🟢6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +4

The targeted claim — 227,584 retained single-cell transcriptomes after Scanpy+Scrublet QC on GSE182109 — reproduces cleanly to 219,652 (GBM-only, 40 samples, 16/16 patients matched) and 240,168 (all 44), bracketing the reported value to within -3.49% using the same tool, thresholds and SHA256-verified data. The only deviation is on the input/sample-definition side: the paper's exact 42-sample subset is underspecified (we get 40 after excluding 2 LGG patients) — a mix of authors' underspecification and our self-made GBM-only cohort cut, not a computation defect. Critically, the apparent fabrication red flag (more cells than the original GSE182109's 201,986) is fully exonerated by the paper's permissive QC. Scope is limited, though: the paper's true central conclusions (disulfidptosis subtypes, DEGs, survival) were not attempted, so overall quality is solid-but-partial rather than a full 1:1.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

191 k
tokens (I/O) · 11.5 M incl. cache
47 min
runtime · 0.24 CPU-h
6.5 GB
peak RAM
1
HPC jobs
hummel
machine