Mutually exclusive teams-like patterns of gene regulation characterize phenotypic heterogeneity along the noradrenergic-mesenchymal axis in neuroblastoma.
The main results reproduced: recomputed values matched the published ones within tolerance.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Reported values are derivable from the shared data
- ✓Any deviation was negligible
- ✓The central claim held under reproduction
- 🟡Could not use the authors’ exact input data
- 🟡Reported values were only indirectly comparable
- 🟡A deviation arose in the data or preprocessing
- 🟡A deviation was attributed to the published material
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Described well enough for the CENTRAL claim. The repo (github.com/manascripts/teamgenepatterns, renamed from Manas-Sehgal/NB_heterogeneity, commit f9842d3) ships the authors' NOR/MES signature + method scripts (PCA/GSEA/swap/linear-fit) but the gene-expression matrices are placeholder stubs, so data was fetched from GEO (GSE9169 n=86, GSE17714 n=22, GPL570) and pre-processed (probe->gene via GPL570.annot, mean-collapse, log2). 1:1 reproduction of the paper's title/Fig-2 thesis: NOR-NOR and MES-MES Spearman positive, NOR-MES negative on BOTH datasets; sample-level NOR-vs-MES scores strongly anti-correlated (-0.96/-0.85); shipped PCA+K-means(k=2) gives 2 phenotype-separated clusters. NOT attempted (the hard ~20%): Fig 3f J-metric r=-0.33 (J-metric code not shipped), Fig 6a/6b multi-dataset percentages (42-75 datasets), exact Fig 6e single-cell -0.58 (scRNA+MAGIC+AUCell), GSEA/ssGSEA. Those crisp numbers remain unverified (not refuted). No fabrication indicators. All compute on «our HPC».
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 87assessed: 2026-06-14 ⛓ 5532febe4d0a
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-14
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetThe paper tests whether the noradrenergic (NOR) and mesenchymal (MES) gene expression programs that define phenotypic heterogeneity in neuroblastoma are organized as mutually antagonistic, mutually exclusive "teams" of genes, analogous to teams-like network behavior seen in EMT and other cancer plasticity axes.
- ★ NOR-specific and MES-specific gene expression patterns are largely mutually exclusive, exhibiting a teams-like behavior across multiple bulk NB transcriptomic datasets finding
- ★ This NOR-MES antagonism is associated with metabolic reprogramming, including glycolysis and fatty acid oxidation (FAO) programs enriched with the mesenchymal state finding
- ★ The NOR-MES antagonism is associated with immunotherapy targets PD-L1 and GD2 finding
- ★ GD2, a cell-surface marker abundant in NB tumors, is associated with the NOR phenotype at both bulk and single-cell levels finding
- ★ Teams-like mutually exclusive patterns are exclusive to NOR/MES-specific gene lists and are not observed among housekeeping genes finding
- ★ K-means clustering (K=2) on bulk transcriptomic datasets segregates samples along PC1 into two clusters that map onto NOR and MES phenotypes via GSEA enrichment method
- J-metric (based on pairwise Spearman correlations) quantifies the strength of teams-like behavior and is negatively correlated with the correlation between NOR and MES signature ssGSEA scores across datasets method
- Replacing NOR/MES genes one at a time with housekeeping genes causes a linear decrease in PC1 variance explained, validating the low-dimensional teams structure finding
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| bulk RNA-seq/microarray gene correlation analysis | NB cell lines and primary tumor samples (GSE9169, GSE17714, GSE28019, GSE64000, GSE66586, GSE78061, GSE19274, GSE44537, GSE73292) | none | pairwise Spearman correlation between 14 NOR-specific and 12 MES-specific genes | — |
| PCA and K-means clustering (K=2) | bulk NB transcriptomic datasets (same GSE cohorts) | none | sample segregation along PC1; cluster identity | Scikit-learn (Python) |
| single-sample gene set enrichment analysis (ssGSEA) / GSEA | bulk NB transcriptomic datasets | none | normalized enrichment scores (NES) for 359-gene NOR and 485-gene MES signatures, Hallmark EMT, FAO, glycolysis, PD-L1, cell cycle gene sets | GSEAPY |
| single-cell RNA-seq (AUCell scoring, MAGIC-imputed) | NB single-cell datasets | none | per-cell AUCell enrichment scores for NOR/MES and other signatures | Rmagic; AUCell |
| PCA-based gene-swapping/control analysis | housekeeping genes vs. NOR/MES gene lists in bulk NB datasets | in silico replacement of NOR/MES genes with housekeeping genes (0 to 26 swaps) | percentage of PC1 variance explained | prcomp (R) |
| J-metric calculation from correlation matrices | 75 bulk NB transcriptomic datasets (meta-analysis) | none | sum of within-team minus between-team Spearman correlations (only p<.05 correlations included) | — |
- – 14 NOR-specific genes correlate positively with each other and negatively with 12 MES-specific genes across 9 bulk datasets
- ▲ PC1 variance explained by the wild-type 26-gene NOR/MES set is consistently higher than that explained by 1000 random 26-gene housekeeping combinations across datasets
- ▼ J-metric and correlation coefficient between NOR and MES signature ssGSEA scores are negatively correlated across 42 datasets with significant NOR-MES correlation r = -0.33, p < .05
- ▼ PC1 variance explained decreases linearly as more NOR/MES genes are swapped for housekeeping genes
- – K-means-derived clusters (K=2) show one cluster enriched for the NOR gene signature and the other for the MES gene signature via GSEA
- correlation r = -0.33, p < .05 (correlation between J-metric and NOR-vs-MES ssGSEA score correlation coefficient across bulk datasets)
- count 14 NOR-specific and 12 MES-specific genes (core gene signature list used for teams/correlation analysis)
- count 369 NOR-specific and 485 MES-specific genes (extended gene signatures from a previous study used for ssGSEA/AUCell)
- count 75 NB transcriptomic datasets (datasets used for meta-analysis (Table S3))
- count 1000 random gene ensembles (control combinations of housekeeping genes and of randomly sampled NOR/MES-extended genes for PCA comparison)
- pvalue p < .05 (significance threshold applied to pairwise gene correlations included in Figure 2 and J-metric calculation)
- count n = 86 (sample size of GSE9169 dataset)
- count n = 100 (sample size of GSE19274 patient-sample dataset)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
The paper analyzes multiple publicly available neuroblastoma bulk and single-cell RNA-seq/microarray datasets to characterize mutually exclusive 'teams-like' gene expression patterns between noradrenergic (NOR) and mesenchymal (MES) cell states. Core methods are pairwise Spearman correlation of NOR/MES gene sets, PCA with K-means clustering (K=2) for sample stratification, and GSEA/ssGSEA for phenotype enrichment scoring. A custom J-metric derived from summed Spearman correlations quantifies phenotypic antagonism, and permutation-based null distributions (1000 random gene ensembles) benchmark the specificity of the observed PC1 variance. Results are reported primarily as correlation coefficients, normalized enrichment scores, J-metric values, and percentage variance explained by PC1, with significance denoted at p<.05.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Spearman's rank correlation | Pairwise gene-gene associations between 14 NOR- and 12 MES-specific genes across all bulk datasets; J-metric calculation; correlation between NOR/MES ssGSEA scores across 75 datasets (r=−0.33, p<.05, based on 42 datasets with significant NOR-MES correlation) | Varies by dataset: GSE9169 n=86, GSE17714 n=22, GSE28019 n=24, GSE64000 n=8, GSE66586 n=10, GSE78061 n=29, GSE19274 n=100, GSE44537 n=6, GSE73292 n=6; 42 datasets for J-metric vs. NOR-MES correlation analysis | not stated |
| Principal component analysis (PCA) | Dimensionality reduction of bulk transcriptomic datasets to segregate NOR and MES phenotypes along PC1; repeated across nine datasets and in gene-swap permutation analysis using prcomp (R) and Scikit-learn (Python) | Varies by dataset (see above) | not stated |
| K-means clustering (K=2) | Sample grouping in each bulk dataset based on PC1 segregation to define NOR and MES clusters for downstream GSEA | Varies by dataset (see above) | not stated |
| Gene Set Enrichment Analysis (GSEA) via GSEAPY | Enrichment of 485-gene MES and 359-gene NOR signatures in K-means-defined clusters across five representative datasets (GSE9169, GSE28019, GSE64000, GSE66586, GSE78061) | Varies by dataset (see above) | not stated |
| Single-sample GSEA (ssGSEA); AUCell for scRNA-seq | Sample-wise normalized enrichment scores for NOR, MES, EMT, FAO, glycolysis, PD-L1, and cell-cycle pathways across 75 bulk datasets; AUCell applied analogously to imputed scRNA-seq matrices | 75 bulk datasets (Table S3); scRNA-seq sample sizes not stated in available text | not stated |
| Permutation test (empirical null via 1000 random gene-set ensembles) | Benchmarking PC1 variance of the 26-gene NOR/MES list against 1000 random 26-gene housekeeping combinations and 1000 random NOR/MES-drawn combinations across datasets; also 100 combinations per swap count in gene-swap analysis | 1000 permutations per dataset; representative results shown for GSE28019, GSE9169, GSE66586 | not stated |
| Linear regression (ordinary least-squares fit) | Fit to mean PC1 variance as a function of the number of NOR/MES genes swapped with housekeeping genes; R-squared (goodness of fit) and Pearson r reported | 100 combinations per swap count; three representative datasets | not stated |
-
Multiple pairwise Spearman correlations (up to 325 pairs per dataset) were tested with significance judged at p<.05, without a stated multiple-testing correction procedure↳ Could also: Benjamini-Hochberg FDR correction or Bonferroni adjustment could also be applied to the family of simultaneous pairwise tests — With hundreds of concurrent tests per dataset repeated across many datasets, the nominal p<.05 threshold yields an elevated expected false-discovery rate; FDR correction would explicitly bound the proportion of spurious discoveries and make the significance claims more interpretable
-
K-means clustering with K fixed at 2 was used to define NOR and MES sample groups in each dataset↳ Could also: Consensus clustering, hierarchical clustering (Ward linkage), or silhouette-score-based selection of optimal K could also partition samples without pre-specifying the number of clusters — K-means requires K to be specified in advance; methods that estimate optimal K from data could independently validate the binary grouping assumption or reveal whether intermediate cell states produce a third cluster in some datasets
-
A permutation null distribution was constructed visually (histograms) to benchmark PC1 variance specificity of the NOR/MES gene set against 1000 random housekeeping-gene combinations↳ Could also: A formal empirical p-value (fraction of permutations exceeding the observed PC1 variance) could also be computed and reported numerically — A single permutation p-value per dataset would quantify how extreme the observed PC1 variance is relative to the null in a directly comparable, reproducible summary statistic, complementing the visual histogram comparison
-
ssGSEA was used to score individual samples for NOR, MES, and pathway enrichment in bulk transcriptomic data↳ Could also: GSVA (gene set variation analysis) or VISION could also provide single-sample pathway activity estimates — Different single-sample scoring algorithms make different distributional assumptions and can have varying sensitivity to gene-set size; comparing ssGSEA results to GSVA would allow assessment of whether the enrichment patterns are robust to the choice of scoring method
-
scRNA-seq data was imputed using Rmagic to reduce sparsity before downstream correlation and enrichment analyses↳ Could also: Alternative imputation methods (e.g., SAVER, ALRA) or analyses on non-imputed data with dropout-aware models could also be applied — Imputation methods vary in how aggressively they smooth expression values, and some can inflate pairwise correlations; sensitivity analyses comparing results with and without imputation, or across imputation tools, would clarify how much conclusions depend on this preprocessing choice
-
Spearman's rank correlation was chosen throughout to measure pairwise gene-expression associations↳ Could also: Mutual information or partial correlation (controlling for confounders) could also quantify gene-gene associations — Spearman captures monotone relationships; mutual information additionally captures non-linear dependencies and requires no assumptions about the form of the relationship, while partial correlation could disambiguate direct from indirect associations mediated by a third gene
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
Assessed papers, coloured by verdict. Click a node to open it.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Reproduction scope — pmid-38230570
Paper: Sehgal et al. 2024, Cancer Biol Ther — "Mutually exclusive teams-like patterns of gene regulation characterize phenotypic heterogeneity along the noradrenergic–mesenchymal axis in neuroblastoma." DOI 10.1080/15384047.2024.2301802.
Code: https://github.com/Manas-Sehgal/NB_heterogeneity → moved/renamed to
https://github.com/manascripts/teamgenepatterns (commit f9842d3, 2025-09-10).
Authors' own code. R 70% / Python 30%. Ships the method scripts (PCA, GSEA,
gene-swap, linear-fit) + the NOR/MES gene signature (signature.gmt) + housekeeping
gene list, but the gene-expression matrices are placeholder stubs (60-byte
"add your matrix here" files) — data must be fetched from GEO and pre-processed by
the user (README points to the lab's EMT_score_calculation microarray pipeline).
Data: publicly accessible bulk + single-cell RNA-seq/microarray from NCBI-GEO. ~75 bulk transcriptomic datasets in Table S3. Primary accession in registry: GSE9169 (n=86, GPL570 Affymetrix U133 Plus 2.0, neuroblastoma SH-SY5Y).
Central thesis (the title)
NOR ("noradrenergic"/adrenergic) and MES (mesenchymal) gene programs behave as two mutually-antagonistic "teams": genes within a team are positively correlated, genes across teams are negatively correlated — a "teams-like" correlation structure, robust across many neuroblastoma datasets.
In scope (pipeline-derived, low-hanging — what we attempt)
| # | Result (paper) | Pipeline | Our target |
|---|---|---|---|
| A | Fig 2 — NOR genes mostly positively correlated with each other, negatively with MES genes (9 datasets) | Spearman correlation of signature genes | mean within-NOR / within-MES / cross NOR–MES Spearman on GSE9169 + GSE17714 |
| B | Mutual exclusivity at sample level (Fig 5/6 antagonism; sc anchor Fig 6e R=−0.58, p<1e-15) | per-sample NOR vs MES score, correlate | Spearman(NOR-score, MES-score) across samples (bulk) |
| C | Fig 3a–e — PCA + K-means (K=2) splits samples into 2 phenotype clusters | shipped PCA.py (StandardScaler→PCA→KMeans k=2) |
PC1/PC2 variance %, 2-cluster split, cluster↔phenotype |
Pipelines named: PCA/K-means via scikit-learn (PCA.py); Spearman correlation
(NumPy/SciPy); signature = repo signature.gmt (NOR ≈ 370 genes, MES ≈ 430 genes).
Out of scope (the hard ~20% — NOT attempted, by design)
- Fig 3f J-metric vs correlation r=−0.33 over 42 datasets — the J-metric ("teams strength") code is not shipped; would require re-implementing the lab's J-metric and harvesting/pre-processing 42 datasets. Skipped (not specified well enough in repo).
- Fig 6a/6b cross-dataset percentages (85.7% / 97.29% / 92.3% / 75.9% over 42–75 datasets) — requires harvesting + uniformly pre-processing 75 datasets + PD-L1/FAO signatures. Volume work, the "completeness" tail. Skipped.
- Fig 6e single-cell IMR-575 R=−0.58 — requires scRNA-seq + MAGIC imputation + AUCell. Different modality; we instead reproduce the bulk antagonism (Claim B) and compare directionally.
- GSEA / ssGSEA (
GSEA.py, Fig 5) — secondary; phenotype labels depend on a hard-coded PC1 cutoff (=10) tuned per dataset. Skipped for the first pass. - Wet-lab / schematic (Fig 1) — not computational.
Honesty note
The repo ships methods, not a turnkey "regenerate Figure X number" harness (data files are stubs; figure-level numbers aggregate dozens of datasets). We therefore reproduce the central, clearly-specified, low-effort computational claim — the teams-like NOR/MES correlation structure and the 2-cluster PCA — on the paper's own primary datasets, and report 1:1 directional/quantitative agreement. We do not claim to regenerate the multi-dataset summary percentages.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
The central thesis — mutually-exclusive NOR/MES 'teams' (positive within-group, negative cross-group correlation) — reproduces cleanly and consistently on the paper's own primary datasets (GSE9169, GSE17714) using the authors' own signature, with strong sample-level antagonism (-0.96/-0.85). Deviations are explainable and sit on the input/methodology side: the repo ships placeholder-stub matrices so data was self-fetched and self-preprocessed, and the single B1 endpoint compared (sc Fig 6e -0.58) is a different modality than our bulk score. The remaining ~20% of crisp numbers (Fig 3f, 6a/6b, 6e, GSEA) were not attempted (code not shipped / large multi-dataset & single-cell pipelines) and remain unverified, not refuted. No fabrication indicators — overall a solid reproduction with explainable, non-1:1 deviations.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.