M6A-mediated molecular patterns and tumor microenvironment infiltration characterization in nasopharyngeal carcinoma.
The main results reproduced: recomputed values matched the published ones within tolerance.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Reported values were directly comparable
- ✓Reported values are derivable from the shared data
- ✓The central claim held under reproduction
- 🟡Could not use the authors’ exact input data
- 🟡A deviation arose in the data or preprocessing
- 🟡A deviation was attributed to the published material
- 🟡The deviation was non-trivial in magnitude
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Described well enough to reproduce. Authors' own repo ships the processed GSE102349 data (counts, clinical-PFS, published cluster labels). On «our HPC» (edgeR 4.0.16, ConsensusClusterPlus 1.66.0) three clearly-specified pipeline steps were re-run. STRONGEST 1:1: consensus clustering reproduces the published 64/24 m6A subtyping almost exactly (63/25, K=2, 98.9% concordance, 1 of 88 samples relabeled). edgeR DEG count 1558 is bracketed by our filtered (1406) and unfiltered (1772) runs with matching up/down direction; exact integer is edgeR-version/prefilter dependent, and all 4 downstream LASSO hub genes reappear as DEGs. Univariate Cox calls fewer prognostic m6A genes (4 vs 10) but 3 of 4 overlap the paper's set and the misses are just over p=0.05 - explained by using log2(CPM+1) as a proxy for the authors' TPM (TPM matrix NOT shipped and not exactly regenerable: per-gene lengths are lost when counts are collapsed to symbols). NOT ATTEMPTED: ssGSEA/CIBERSORT/ESTIMATE/GSVA/pRRophetic/TIDE/SubMap/GSEA/TF-prediction (Figs 2-5,9-10; mostly directional not single-number), the full LASSO signature (needs non-regenerable TPM + two RNG seeds), and the wet-lab 255-sample qPCR/IHC validation (Fig 8; out of scope, manual). No fabrication indicator: every checked number is derivable from the shipped data within margins explained by the TPM-proxy and edgeR version/prefilter ambiguity.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 76assessed: 2026-06-15 ⛓ 51aa23dabd02
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-15
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusWhat are the m6A-related molecular patterns in nasopharyngeal carcinoma (NPC) and how does m6A modification regulate tumor microenvironment (TME) cell infiltration, therapeutic response, and prognosis?
- ★ NPC samples can be divided into two distinct m6A-related molecular subclasses based on prognostic m6A regulators finding
- ★ Cluster 1 is characterized by immune-related and metabolism pathway activation with better response to anti-PD1/anti-CTLA4 and chemotherapy, while cluster 2 shows stromal activation, low HLA/immune checkpoint expression, and worse therapeutic response finding
- ★ Cluster 2 patients have poorer progression-free survival (PFS) than cluster 1 patients finding
- ★ An m6A-related prognostic signature/risk model and nomogram were constructed to predict PFS in NPC patients resource
- ★ REEP2, TMSB15A, DSEL, and ID4 are upregulated in NPC tumor samples versus non-tumor tissue finding
- ★ High expression of REEP2 and TMSB15A is associated with poor survival in NPC patients finding
- REEP2, TMSB15A, DSEL, and ID4 interact with m6A regulators mechanism
- ★ m6A modification plays an important role in regulating TME heterogeneity and complexity in NPC mechanism
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| bulk mRNA RNA-seq (public dataset analysis) | NPC patient tumors (GSE102349, 113 patients; 88 with PFS) | none | mRNA expression, consensus clustering, PFS | Illumina HiSeq 2000 |
| whole genome expression microarray (public dataset analysis) | NPC primary tumor tissues (18) and non-cancerous nasopharyngeal tissues (18) (GSE53819) | none | mRNA expression, subclass validation | Agilent-014850 Whole Human Genome Microarray 4x44K G4112F |
| immune cell infiltration estimation (ssGSEA, CIBERSORT, ESTIMATE) | NPC tumors (GSE102349) | none | immune/stromal scores, relative abundance of 22/24 immune cell types | — |
| immunotherapy response prediction (TIDE, SubMap) | NPC m6A-related subclasses (GSE102349) | none | predicted response to anti-PD-1/anti-CTLA4, HLA and immune checkpoint expression | — |
| chemotherapy response prediction (pRRophetic) | NPC m6A-related subclasses (GSE102349) | none | IC50 of antitumor drugs | GDSC database |
| immunohistochemistry (IHC) | 255 primary NPC samples and adjacent normal nasopharyngeal tissue (Kunming Medical University) | none | protein expression of REEP2, TMSB15A, DSEL, ID4 and immune cell markers (CD20, CD27, CD68, CD4, CD45RO, CD8) | antibodies: DSEL NBP2-56370 Novus, ID4 LS-B9923 Lsbio, REEP2 bs-15198R bioss, TMSB15A sc-271649 Santa Cruz |
| LASSO/Cox prognostic risk model and nomogram construction | NPC patients (GSE102349, 7:3 training/test split) | none | risk score, PFS, time-dependent ROC/AUC | — |
| transcription factor prediction network analysis | REEP2, TMSB15A, DSEL, ID4 hub genes | none | TF-hub gene interaction network | NetworkAnalyst/JASPAR, Cytoscape 3.9.1 |
- – 88 NPC patients clustered into cluster 1 (64 patients) and cluster 2 (24 patients) 64 vs 24
- – 10 of 21 m6A regulators (IGF2BP1, ALKBH5, YTHDF2, ELAVL1, WTAP, LRPPRC, HNRNPA2B1, CBLL1, RBM15B, YTHDF1) significantly associated with PFS
- ▲ All 10 prognostic m6A regulators upregulated in cluster 2 versus cluster 1
- ▼ Cluster 2 patients show poorer PFS than cluster 1 patients (Kaplan-Meier)
- – 1558 DEGs identified between the two m6A-related subclasses 1558 genes
- ▲ REEP2, TMSB15A, DSEL, ID4 upregulated in NPC versus non-tumor samples
- ▼ High REEP2 and high TMSB15A expression associated with poor survival
- count 129,079 new cases and 73,987 deaths in 2018 (global NPC incidence/mortality)
- count 113,659 (85.5%) new cases in Asia in 2020 (NPC geographic distribution)
- count 1558 DEGs (DEGs between two m6A subclasses)
- count n = 255 (primary NPC clinical samples for IHC)
- count 88 (NPC patients with PFS analyzed from GSE102349)
- fold_change |log2 FC| > 1, FDR < 0.05 (DEG identification threshold (edgeR))
- pvalue P < .05 (univariate Cox significance threshold for m6A regulators)
- correlation correlation coefficient > 0.3 (Pearson correlation threshold between prognostic factors and m6A regulators)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This bioinformatics study applied consensus clustering to 88 RNA-seq NPC samples (GSE102349) using 10 prognostic m6A regulators selected by univariate Cox regression, identifying two molecular subclasses. Tumor microenvironment infiltration differences between subclasses were assessed with Wilcoxon rank-sum or Kruskal-Wallis tests on ssGSEA and ESTIMATE scores. A LASSO Cox prognostic signature was built on a 7:3 internal train/test split and evaluated via Kaplan-Meier/log-rank tests and time-dependent ROC; protein-level findings were confirmed in 255 clinical IHC samples with chi-square tests for categorical comparisons.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Univariate Cox proportional hazards regression | Screening 21 m6A regulators for association with PFS to select features for clustering and prognostic modeling | 88 NPC patients (GSE102349) | not stated |
| Consensus clustering (consensusClusterPlus) | Unsupervised subtype identification from expression profiles of 10 prognostic m6A regulators; k selected by CDF plateau | 88 NPC patients (GSE102349) | na |
| Wilcoxon rank-sum test | Comparing GSVA pathway scores, ESTIMATE immune/stromal scores, and CIBERSORT immune cell proportions between the two m6A subclasses | 88 NPC patients | not stated |
| Kruskal-Wallis test | Comparing GSVA pathway/TME scores between subclasses (reported as alternative to Wilcoxon rank-sum depending on context) | 88 NPC patients | not stated |
| edgeR (negative binomial GLM) | Differential expression between two m6A subclasses; cutoffs |log2FC|>1 and FDR<0.05 | 88 NPC patients (GSE102349, RNA-seq) | not stated |
| LASSO Cox regression with ten-fold cross-validation | Feature selection and construction of m6A-related prognostic risk signature from DEGs | Training set (~62 patients, 7:3 split of 88) | not stated |
| Kaplan-Meier survival analysis with log-rank test | PFS comparison between m6A subclasses and between high-risk vs low-risk groups in training, test, and whole sets; also IHC-based survival for REEP2/TMSB15A | 88 NPC patients (GEO); 255 for IHC cohort (n for survival analysis not explicitly restated) | not stated |
| Time-dependent ROC / AUC | Evaluating predictive accuracy of the prognostic risk model across training, test, and whole sets | 88 NPC patients | na |
| Chi-square test | Testing significance of differences in categorical clinical variables | 255 primary NPC clinical samples | not stated |
| Pearson correlation | Correlation between prognostic signature genes and m6A regulators; threshold r>0.3 and p<0.05 | 88 NPC patients | not stated |
| Spearman correlation | Correlation between expression of REEP2, TMSB15A, DSEL, ID4 and other genes for GSEA group classification | 88 NPC patients | not stated |
| SubMap algorithm with Bonferroni correction | Predicting and cross-dataset validation of anti-PD1 and anti-CTLA4 immunotherapy response between m6A subclasses | 88 NPC patients (GSE102349); 18 NPC (GSE53819) | na |
-
Univariate Cox regression screened 21 m6A regulators at p<0.05 without correction for simultaneous testing↳ Could also: Apply FDR (Benjamini-Hochberg) correction to the 21 univariate Cox p-values before selecting regulators for downstream clustering — Testing 21 hypotheses simultaneously raises the expected number of false discoveries; FDR adjustment is the standard in multi-gene survival screening and is already employed at the DEG step of this same study, so applying it here would make the filtering criteria consistent throughout the pipeline
-
Prognostic signature performance was evaluated using a single internal 7:3 random train/test split of the 88-patient cohort↳ Could also: Repeated k-fold cross-validation or bootstrap-based internal validation could also be used to assess prognostic model performance — With n=88 a single random split allocates only ~26 patients to the test set, making performance metrics sensitive to which patients happen to fall in that partition; repeated cross-validation or bootstrap uses all available patients for both training and evaluation, yielding more stable and less optimistic AUC and concordance estimates
-
Consensus clustering was used to identify the two molecular subclasses, with k chosen by the CDF plateau criterion↳ Could also: Non-negative matrix factorization (NMF) is another widely used approach for identifying molecular subtypes from expression data — NMF and consensus clustering optimize different objective functions and can yield subtypes with different boundaries; showing concordance—or quantifying discordance—between methods is a common sensitivity check and can help establish whether the two-subtype solution is robust to algorithmic choice
-
Pearson correlation was used for gene-regulator associations while Spearman correlation was used separately for gene-gene associations in the GSEA step↳ Could also: Spearman correlation could be used consistently for both correlation analyses throughout the paper — Spearman correlation makes no assumption of bivariate normality and is robust to outliers and non-linear monotone relationships; applying it uniformly would reduce assumption sensitivity and methodological inconsistency between analysis steps
-
Continuous clinical variables (age, tumor size) were summarized as mean ± SD↳ Could also: Median with interquartile range (IQR) could also summarize continuous clinical variables — Clinical measurements such as age and tumor size are often right-skewed; median and IQR describe the distribution more robustly when normality is not verified, and are the format recommended by many clinical reporting guidelines (e.g., CONSORT, STROBE)
-
edgeR was used for differential expression between the two RNA-seq subclasses↳ Could also: DESeq2 is another standard tool for RNA-seq differential expression analysis — DESeq2 and edgeR use different normalization strategies and dispersion estimation models; comparing DEG lists from both tools and reporting the overlap is a common practice for identifying the most reproducible findings, since genes called by both methods tend to be more robust
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-38532632
Paper: Wang et al. 2024, M6A-mediated molecular patterns and tumor microenvironment
infiltration characterization in nasopharyngeal carcinoma, Cancer Biol Ther.
Repo: https://github.com/DrWangfeng/nasopharyngeal-carcinoma (authors' OWN R code; P16 N/A).
Data: GSE102349 (113 NPC patients; 88 with PFS), GSE53819 (18 tumor + 18 normal).
The repo SHIPS the processed data: GSE102349_counts_symbol.txt, ..._clinical_PFS.txt,
..._Cluster_res.txt (published labels), GSE53819_NPC_exprSet.txt, etc.
In scope (pipeline-derived, attempted)
| id | result | method (pipeline) | reported | location |
|---|---|---|---|---|
| C1 | # prognostic m6A regulators | univariate Cox (survival pkg), p<0.05 | 10 of 21 | Results, Fig 1g / Table S1 |
| C2 | m6A consensus clusters | ConsensusClusterPlus (hc, euclidean, seed99, maxK6) | 2 clusters: 64 / 24 | Fig 1c |
| C3 | DEGs between clusters | edgeR, |log2FC|>1 & FDR<0.05 | 1558 (928 up, 630 down) | Fig 6a,b / Table S12 |
Attempted-but-fragile (reported, not primary)
- LASSO 4 hub genes (REEP2, TMSB15A, DSEL, ID4): chain = TPM → edgeR DEG → univ-Cox (caret seed 359) → cv.glmnet (seed 9). Depends on TPM which is not exactly regenerable (gene lengths lost on symbol collapse). We check only whether the 4 hub genes appear in our DEG set, not the full LASSO. Not graded as a primary claim.
Out of scope (not attempted)
- Wet-lab / clinical validation cohort (255 NPC qPCR/IHC, Fig 8) — manual, not pipeline.
- ssGSEA/CIBERSORT/ESTIMATE/GSVA/pRRophetic/TIDE/SubMap/GSEA/TF-prediction (Figs 2–5, 9–10): in principle pipeline-derived but mostly directional ("higher in cluster1") rather than a single checkable number; skipped under 80/20.
- t-SNE plots (visual, non-numeric).
Key honest caveat
Authors' C1/C2 use TPM (GSE102349_TPM_symbol.txt, NOT shipped; count2TPM.R needs
per-gene lengths from featurecounts, also not shipped). We substitute log2(CPM+1)
from the shipped symbol-level counts as a monotone proxy. C3 (edgeR) uses raw counts
directly and needs no TPM → cleanest data point. The edgeR DEG script itself
(GSE102349_edgeR_diff.txt) is also not shipped; we implement it per the README spec.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
This is a solid partial reproduction of the m6A-NPC study. The headline molecular subtyping reproduces almost 1:1 (consensus clustering 63/25 vs reported 64/24, 98.9% concordance, 1/88 relabeled), DEGs are bracketed (1406–1772 vs reported 1558, same direction), and all four LASSO hub genes reappear as DEGs. The one real deviation — 4 vs 10 prognostic m6A genes (C1) — sits on the input/normalization side: it is driven by a log2(CPM+1) proxy that was necessary because the authors did not deposit the (non-regenerable) TPM matrix, not by any computational error or fabrication. Net: deviations are moderate, explainable, and shared between our method choice and an authors'-side data-availability gap; the central conclusion holds.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.