Expression Patterns of Immune Genes Reveal Heterogeneous Subtypes of High-Risk Neuroblastoma.
The main results reproduced, with only marginal, non-material deviations.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Reported values are derivable from the shared data
- ✓The central claim held under reproduction
- 🟡Could not use the authors’ exact input data
- 🟡Reported values were only indirectly comparable
- 🟡A deviation arose in the data or preprocessing
- 🟡A deviation was attributed to the published material
- 🟡The deviation was non-trivial in magnitude
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
REPRODUCED. The paper's central pipeline-derived result reproduces 1:1 in BOTH independent cohorts: using the authors' own Supplementary Table S2 subtype labels joined to public OS data, the UHR-vs-HR survival hazard ratios match (GSE49710 HR 1.73->1.74; TARGET 1.68->1.69) and the log-rank p-values match to the printed precision (0.00758 and 0.0021). Gene-list claims are exact (283 immune / 39 DE / 4 named upregulated). As a bonus harder result, the spectral co-clustering that produces the labels was independently re-derived with sklearn on RNA-seq data and agrees with the authors' UHR/HR calls at 95.5% (kappa 0.907) -> no evidence of fabrication. PARTIAL: the GAL multivariate-Cox prognostic claims (C5 GSE49710, C6 TARGET) did not fully reproduce - GAL is prognostic in the correct direction univariately (p=0.0019) but the exact HR and post-adjustment independence depend on an underspecified expression normalization, and C6 also needs TARGET MYCN/expression not openly available. No authors' analysis code exists; reproduction used third-party tools (lifelines, sklearn) on the paper's own data per BRIEF rule P16. NOT attempted: STRING/KEGG enrichment (S5/S6), Venn diagram, pan-cancer gene-selection step.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 73assessed: 2026-06-20 ⛓ 4c722bde8f2b
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-20
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-20no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetHigh-risk neuroblastoma (HR-NB) is a heterogeneous disease treated uniformly, so the paper tests whether expression patterns of immune-related genes can stratify HR-NB into clinically and biologically distinct subtypes with different survival outcomes.
- ★ Unsupervised biclustering of 283 NB-specific immune genes stratifies HR-NB into two reproducible subtypes: ultra-high-risk (UHR-NB) and (revised) HR-NB. finding
- ★ UHR-NB has significantly worse overall survival than HR-NB in multiple independent cohorts. finding
- ★ 39 NB-specific immune genes are differentially expressed between UHR-NB and HR-NB (4 up, 35 down). finding
- ★ The four UHR-NB-upregulated genes (ADAM22, GAL, KLHL13, TWIST1) are consistently upregulated in MYCN-amplified neuroblastoma across five independent cohorts. finding
- ★ GAL is an independent predictor of overall survival beyond MYCN amplification and age. finding
- TWIST1 expression is strongly associated with tumor stage and genetic alterations (1p deletion, 11q deletion, 17q gain). finding
- ★ Downregulated immune genes in UHR-NB are enriched in 43 KEGG immune pathways, whereas upregulated genes are enriched only in the p53 signaling pathway. finding
- ★ UHR-NB is a TP53-related subtype characterized by reduced immune activity. mechanism
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| bulk RNA-seq differential expression (ANOVA, Student's t-test) | pediatric pan-cancer (ALL, AML, NB, WT) | none | identification of 283 NB-specific immune-related genes | — |
| spectral co-clustering / biclustering | TARGET (n=217) and GSE49710 (n=176) HR-NB cohorts | none | sample and gene subtype clusters | — |
| Kaplan-Meier survival analysis | TARGET and GSE49710 HR-NB patients | none | overall survival by subtype | — |
| Cox proportional hazards regression | TARGET and GSE49710 HR-NB patients | none | prognostic value of gene expression for OS, with/without adjustment for age and MYCN | — |
| differential expression (two-sided t-test) | 5 cohorts: GSE85047, GSE45547, GSE120572, GSE73517, GSE19274 | MYCN amplification status | expression of ADAM22, GAL, KLHL13, TWIST1 | — |
| gene expression validation | mouse neuroblastoma cell lines (9464D MYCN-amplified; NXS2, Neuro2A non-amplified) | MYCN amplification (genetic) | expression of ADAM22, GAL, KLHL13, TWIST1 | — |
| one-way ANOVA | GSE45547 and GSE120572 cohorts | tumor stage / genetic alteration (1p/11q/17q) | TWIST1, ADAM22, KLHL13 expression differences | — |
| GSEA/KEGG pathway enrichment and PPI network analysis | TARGET and GSE49710 (4723 immune-related genes from InnateDB) | none | enriched pathways and protein-protein interactions among UHR-NB up/downregulated genes | STRING |
- ▼ UHR-NB (subtype III) had significantly worse OS than HR-NB in TARGET and GSE49710 HR=1.68 (TARGET), HR=1.73 (GSE49710)
- ▼ Two-year survival rate was much lower in UHR-NB than HR-NB in both cohorts 48.8% vs 79% (TARGET); 43.2% vs 75% (GSE49710)
- – 39 NB-specific immune genes differentiated UHR-NB from HR-NB >1.5-fold, p<1e-3; 4 up, 35 down
- ▲ ADAM22, GAL, KLHL13, TWIST1 upregulated in MYCN-amplified NB across all 5 validation cohorts
- ▲ GAL and TWIST1 significant unadjusted OS predictors; GAL remained significant after adjusting for age/MYCN in both cohorts GAL HR=2.005 (TARGET), 2.37 (GSE49710); TWIST1 HR=1.85 (TARGET), 1.81 (GSE49710)
- – TWIST1 upregulated with 1p deletion, downregulated with 11q deletion and 17q gain p=2.4e-11; p=2.8e-5; p=8.8e-7
- – 311 downregulated immune genes enriched in 43 KEGG pathways; 26 upregulated genes enriched only in p53 signaling pathway FDR=0.03 for p53 pathway
- – 26 commonly upregulated and 311 commonly downregulated immune-related genes identified in UHR-NB >1.5-fold, p<1e-4
- pvalue p=0.0021 (UHR-NB vs HR-NB OS difference, TARGET cohort)
- pvalue p=0.0076 (UHR-NB vs HR-NB OS difference, GSE49710 cohort)
- fold_change HR=1.68, 95% CI 1.17-2.4 (hazard ratio UHR-NB vs HR-NB, TARGET)
- fold_change HR=1.73, 95% CI 1.12-2.67 (hazard ratio UHR-NB vs HR-NB, GSE49710)
- count 283 NB-specific immune genes; 39 differentiated (4 up, 35 down) (gene signature sizes)
- pvalue p=2.35e-11 (TWIST1 vs tumor stage, GSE45547)
- pvalue p=5e-5, HR=2.005 (95% CI 1.43-2.81) (GAL Cox regression OS predictor, TARGET)
- other FDR=0.03 (p53 signaling pathway enrichment in 26 upregulated genes)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
The study reanalyzed public pediatric-cancer gene-expression cohorts (TARGET, GSE49710, and several additional GEO datasets) using ANOVA, Student's t-tests, unsupervised spectral co-clustering, Kaplan-Meier/log-rank survival analysis, Cox proportional hazards regression, and KEGG/GSEA pathway enrichment to define and characterize an ultra-high-risk neuroblastoma subtype. Gene selection steps used fold-change plus raw p-value thresholds explicitly described as applied without multiplicity correction, while pathway enrichment results were reported with an FDR value. Results were reported primarily as p-values, hazard ratios with 95% confidence intervals, and fold changes across multiple independent validation cohorts.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| ANOVA and Student's t-test | comparing expression of 4723 immune genes across four pediatric tumor types (ALL, AML, NB, WT) to define 283 NB-specific immune genes | — | not stated |
| Spectral co-clustering (unsupervised biclustering) | stratifying HR-NB patients in TARGET and GSE49710 into subtypes based on 283 NB-specific immune genes | 217 HR-NB (TARGET), 176 HR-NB (GSE49710) | not stated |
| Kaplan-Meier estimation with log-rank-type p-values | overall survival comparison between UHR-NB and HR-NB subtypes in TARGET and GSE49710 | 217 (TARGET), 176 (GSE49710) | not stated |
| Cox proportional hazards regression | survival association of UHR-NB subtype and of GAL/TWIST1 expression, unadjusted and adjusted for age and MYCN amplification | 217 (TARGET), 176 (GSE49710) | not stated |
| Two-sided Student's t-test | 39 differentiated gene identification (UHR-NB vs HR-NB), MYCN amplified vs non-amplified comparisons (5 cohorts), telomere maintenance status comparison | cohort-specific counts stated (e.g., 55/222, 93/550, 83/310, 33/72, 22/78) | not stated |
| One-way ANOVA | gene expression differences across tumor stages and across genetic alteration categories (1p deletion, 11q deletion, 17q gain) | — | not stated |
-
Initial screening of 283 NB-specific immune genes and 39 UHR-NB-differentiated genes used raw p-value thresholds explicitly without multiplicity correction across thousands of genes.↳ Could also: A false discovery rate procedure such as Benjamini-Hochberg (as implemented in tools like DESeq2 or limma) could also be applied — This would formally control the expected proportion of false positives when testing thousands of genes simultaneously, which is a standard practice in high-throughput expression screening.
-
Group differences in gene expression were assessed with Student's t-tests and one-way ANOVA.↳ Could also: Non-parametric equivalents such as the Mann-Whitney U test or Kruskal-Wallis test could also be used — These do not require an assumption of normally distributed expression values, which can be useful when group sizes are small or distributions are skewed.
-
One-way ANOVA was used to compare expression across multiple tumor-stage or genetic-alteration categories.↳ Could also: ANOVA with a post-hoc pairwise correction (e.g., Tukey HSD) could also be reported alongside the omnibus test — This would show which specific stage/alteration group pairs differ while maintaining control of the family-wise error rate across all pairwise comparisons.
-
Subtypes were identified using spectral co-clustering applied separately to two cohorts.↳ Could also: Consensus clustering or non-negative matrix factorization (NMF) could also be used for subtype discovery — These approaches provide a built-in measure of cluster stability across resampling, which can complement cross-cohort replication as evidence of robust subtype structure.
-
Gene set overlap between cohorts was quantified with a Jaccard index and a probability calculation for random overlap.↳ Could also: A formal hypergeometric (or Fisher's exact) enrichment test could also be used to assess overlap significance — This is a widely used standard for testing whether two gene lists share more members than expected by chance, and it directly yields a p-value for the overlap.
-
Cox proportional hazards models were used to assess prognostic value of gene expression, without discussion of the proportional-hazards assumption.↳ Could also: A proportional-hazards assumption check (e.g., Schoenfeld residuals) or a flexible parametric/time-varying-coefficient survival model could also be used — This would provide additional confirmation that the hazard ratio is constant over the follow-up period, which is an assumption underlying the Cox model's interpretation.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — PMID 32629858
Paper: Liu et al., "Expression Patterns of Immune Genes Reveal Heterogeneous Subtypes of High-Risk Neuroblastoma." Cancers (Basel) 2020;12(7):1739. PMC7408437.
Central thesis (pipeline-derived)
Unsupervised spectral co-clustering of 283 NB-specific immune genes splits high-risk neuroblastoma (HR-NB) into subtypes; subtype III = UHR-NB (ultra-high-risk), subtypes I+II = HR-NB. UHR-NB has significantly worse overall survival in two independent cohorts (TARGET, SEQC/GSE49710).
In-scope (pipeline-derived) results attempted
| ID | Result | Pipeline |
|---|---|---|
| C1 | GSE49710 UHR vs HR OS: HR=1.73 (1.12–2.67), log-rank p=0.00758 | Cox PH + KM/log-rank (survminer→lifelines) |
| C2 | TARGET UHR vs HR OS: HR=1.68 (1.17–2.4), p=0.0021 | Cox PH + KM/log-rank |
| C3 | 2-year OS: UHR 43.2/48.8% vs HR 75/79% (GSE49710/TARGET) | KM curve read-off |
| C4 | 283 NB-specific immune genes; 39 common DE (4 up: ADAM22,GAL,KLHL13,TWIST1; 35 down) | ANOVA/t-test gene selection |
| C5 | GAL independent prognosis GSE49710: β=0.650, HR=1.91 (1.18–3.10), p=0.008 | multivariate Cox (GAL+MYCN+age) |
| C6 | GAL independent prognosis TARGET: β=0.258, HR=1.29 (1.10–1.52), p=0.0016 | multivariate Cox |
| CL | Subtype labels themselves (spectral co-clustering reproducibility) | sklearn SpectralCoclustering re-derivation (bonus) |
Out of scope
- Wet-lab / experimental validation: none in this paper (pure bioinformatics).
- STRING/KEGG enrichment (S5/S6) and Venn: descriptive, not attempted.
- Pan-cancer 716-tumor comparison for gene selection: descriptive support.
Code availability finding
The paper's data-availability statement promises code "will be posted on GitHub"
(future tense, no link). The BRIEF-pinned github.com/kassambara/survminer is the
generic third-party survival R package, NOT the authors' analysis code. Per BRIEF
rule P16, applying a third-party tool (survminer/lifelines + sklearn) to the
paper's own data is a valid reproduction path — which is exactly what we did.
Key data sources used
- GSE49710 (Agilent 4x44K microarray, GPL16876): 498 SEQC-NB patients, full clinical annotation incl. high-risk flag, MYCN, age. (Series-matrix expression values are empty; expression only in RAW.tar.)
- GSE62564 (RNA-seq log2RPM, 498 SEQC-NB): provides OS/EFS time+event and gene expression (43,828 RefSeq rows). Same 498 patients as GSE49710 — joined by patient number. This is the "RNA-seq, 19,320 genes" cohort the paper calls GSE49710.
- Supplementary Table S2 (from the paper): per-sample UHR/HR subtype labels for both cohorts — the authors' own assignments (used to reproduce survival 1:1).
- TARGET-NBL clinical via cBioPortal
nbl_target_2018_pubREST API: OS for 217 HR samples, matched to S2 by TARGET-30-XXXX ID.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
Strong reproduction. The paper's central conclusion — that spectral co-clustering of 283 immune genes splits high-risk neuroblastoma into a worse-OS UHR subtype in two independent cohorts — reproduces essentially 1:1: HRs match within rounding and log-rank p-values are identical to printed precision (0.00758, 0.0021), gene-list counts are exact, and the subtype labels are independently re-derivable at 95.5% agreement (κ=0.907), so there is no fabrication signal. The deviations are confined to the secondary GAL multivariate-Cox claims (C5/C6) and a ~7pt KM read-off offset (C3b). These sit on the input/preprocessing side — an underspecified expression normalization (microarray vs RNA-seq scale) plus open-access data gaps (TARGET MYCN/expression) — a mix of authors' underspecification and our self-chosen data substitution, not a core-logic failure. Overall solid with explainable, well-documented deviations.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.