Bioinformatics Analysis of the Characteristics and Correlation of m6A Methylation in Breast Cancer Progression.
The main result did not reproduce in this reproduction attempt. Where our recomputation produced values that differ from the published ones, those discrepancies are listed below. This is a single automated attempt — not peer review and not a finding of error or misconduct — and differences can also arise from data access, undocumented parameters or the computing environment. The verdict can be contested via “report an error”.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- Nothing in this column.
- 🟡Could not use the authors’ exact input data
- 🟡Reported values were only indirectly comparable
- 🔴A deviation arose in the data or preprocessing
- 🔴A deviation was attributed to the published material
- 🔴Reported values were not (fully) derivable from the shared data
- 🟡The deviation was non-trivial in magnitude
- 🔴The central claim did not (fully) hold under reproduction
- 🔴Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Paper is RETRACTED (Hindawi paper-mill, Contrast Media Mol Imaging 2023;2023:9874020). Described well enough only in parts. Two clearly-specified pipeline outputs reproduced on the paper's own public data: (C1) TCGA-BRCA top-10 most-mutated genes via maftools matched 9/10 (only rank-10 HMCN1 vs SYNE1 differs, MAF-version noise; all canonical BRCA drivers recovered) => essentially reproduced/within-tol; (C2) the brief's code link TIDEpy ran cleanly with default params on the full GSE96058 matrix (3409 samples) => pipeline step reproducible but the paper prints no TIDE number, so partial (P16 third-party-tool reproduction). NOT attempted (the hard ~20%): the paper's central '4 m6A subtypes' are defined by sign of undefined 'Set1'/'Set2' z-scores - the classification is not derivable from the text, so all subtype-conditioned results (CIBERSORT 11/22, subtype survival, stemness ranking, exon-skip splicing) cannot be reproduced without guessing. Fabrication/quality flag (C3): the m6A 'writer' gene list is wrong in the paper itself (METTL15 listed twice and is not an m6A writer; canonical METTL3/METTL14 absent), a strong paper-mill signal consistent with the retraction. All heavy compute on «our HPC»/«infra»; only small results + pointers on «host».
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 48assessed: 2026-06-15 ⛓ 0afa19cd01fd
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-15
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusThe study tests whether breast cancer can be stratified into m6A methylation subtypes based on m6A-related gene expression, and whether these subtypes differ systematically in survival, prognosis, mutation, copy number variation, immune infiltration, drug sensitivity, and tumor stemness, thereby clarifying the role of m6A methylation in breast cancer progression.
- ★ Breast cancer samples can be divided into 4 m6A subtypes (quiescent, m6A methylation, protein-binding, mixed) based on expression of m6A-related genes, with consistent proportions across datasets finding
- ★ The four m6A subtypes show significantly different survival rates, with the mixed subtype having higher survival and the protein-binding subtype having significantly lower survival finding
- ★ TP53 mutation and copy number loss are most pronounced in the protein-binding subtype among tumor driver genes finding
- ★ RAD54B copy number is increased in the protein-binding subtype while other DNA damage repair genes show copy number deletion across subtypes finding
- ★ The four m6A subtypes differ significantly in macrophages M0 and resting mast cells infiltration, which correlate with patient prognosis finding
- ★ The m6A methylated subtype has the highest tumor stemness index (mRNAsi), indicating high m6A reader expression drives breast cancer progression finding
- ★ m6A gene expression contributes to breast cancer response to anthracyclines, with the quiescent/resting group having the fewest drug-sensitive samples finding
- Number of exon-skip (ES) alternative splicing events differs among the four m6A subtypes while other splicing types do not finding
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Bulk RNA-seq transcriptome analysis / m6A subtype classification (Z-score of Set1/Set2 gene expression) | TCGA-BRCA breast cancer tissue (human) | none | Expression levels of 21 m6A-related genes; subtype assignment | TCGA-BRCA dataset (1,217 samples; 1,104 tumor) |
| Bulk gene expression / m6A subtype classification | GSE96058 breast cancer samples (human) | none | m6A-related gene expression; subtype assignment | GEO GSE96058 dataset (3,409 samples) |
| Gene expression subtyping and anthracycline sensitivity analysis | GSE25066 breast cancer samples (human) | anthracycline chemotherapy (clinical) | Number of drug-sensitive vs nonsensitive samples per subtype | GEO GSE25066 dataset (508 samples) |
| Survival/prognosis analysis | TCGA-BRCA breast cancer (human, including triple-negative) | none | Survival rate by subtype; clinicopathological correlation | R survival package; Circos |
| Alternative splicing analysis | TCGA-BRCA breast cancer (human) | none | PSI of 7 splicing types (AP, ES, RD, AT, ME, AD, AA) | TCGA SpliceSeq |
| Mutation and copy number variation analysis | TCGA-BRCA breast cancer (human) | none | Mutation/CNV of top tumor driver genes and DNA damage repair genes (GO:0006281) | R Maftools; MSigDB V7.1 |
| Immune cell infiltration deconvolution | TCGA-BRCA and GSE96058 breast cancer (human) | none | Proportion of 22 immune cell types; prognostic association | CIBERSORT V1.01; R survival coxph |
| Immune efficacy and tumor stemness analysis | TCGA-BRCA and GSE96058 breast cancer (human) | none | TIDE immune efficacy score; mRNAsi stemness index | TIDEpy (default parameters); mRNAsi from PMID 29625051 |
- – Survival differs significantly among the 4 m6A subtypes; mixed subtype has higher survival than the other three
- ▼ Protein-binding subtype has significantly lower survival; m6A methylated vs protein-bound subgroups differ significantly in prognosis (including TNBC)
- ▲ TP53 has the most mutations and copy number loss in the protein-binding subtype
- ▲ RAD54B copy number increased in protein-bound subtype; other 9 DNA damage repair genes show copy number deletion
- – 11 of 22 immune cell types show significantly different infiltration among the 4 subgroups; macrophages M0 and resting mast cells prognostically relevant
- ▼ NK cells resting and macrophages M0 associated with survival in all samples (HR<1); dendritic cells resting and mast cells resting in TNBC (HR<1) HR<1
- ▲ Stemness index (mRNAsi) ranked low to high: protein-bonding, mixed, quiescent, m6A methylated (highest)
- ▼ Quiescent/resting group has the fewest anthracycline drug-sensitive samples among the 4 subtypes
- count 1,217 samples including 1,104 sequenced tumor tissue samples (TCGA-BRCA dataset size)
- count 3409 sequenced samples including 136 biological replicates (GSE96058 dataset size)
- count 508 samples (GSE25066 dataset size)
- other 0.1–0.4% of total adenosine residues (abundance of m6A methylation in eukaryotic cells)
- count 22 immune cell types analyzed, 11 significantly different among 4 subgroups (CIBERSORT immune infiltration)
- fold_change increased by 100 times (IGF2BP1 mRNA expression in fibrolamellar HCC (cited literature))
- pvalue p < 0.05 (threshold for statistical significance (ANOVA/t-test))
- count 21 m6A methylation-related genes (8 writer, 2 eraser, 11 reader) (m6A gene set used for subtyping)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This retrospective bioinformatics study classified breast cancer samples from three public datasets (TCGA-BRCA, GSE96058, GSE25066) into four m6A methylation subtypes via a Z-score median split of two pre-defined gene sets. Survival analyses, univariate and multivariate Cox proportional hazards models, ANOVA, and t-tests were applied to compare subtypes across prognosis, immune cell infiltration, genomic features, drug sensitivity, and tumor stemness. Results were reported with a p < 0.05 significance threshold, with hazard ratios provided for Cox analyses.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Kaplan-Meier survival analysis with implied log-rank test (survival R package) | Overall survival comparison across the 4 m6A subtypes in all samples and TNBC subgroup; also per immune cell type (Figures 2a, 2b, 3b, 3c, 4f, 4g) | TCGA-BRCA 1,104 tumor samples; GSE96058 3,409 samples; GSE25066 508 samples — exact per-analysis n not stated | not stated |
| Univariate Cox proportional hazards regression (coxph) | Association of each of 22 immune cell infiltration proportions with patient survival (Figures 4b, 4c) | TCGA-BRCA and GSE96058 combined samples; exact n not stated | not stated |
| Multivariate Cox proportional hazards regression (risk score model) | Risk score derived from immune cell infiltration proportions weighted by beta coefficients (Figures 4d, 4e) | not stated | not stated |
| One-way ANOVA | Differences in proportions of 22 immune cell types across the 4 m6A subtypes; stated as general method for subtype comparisons in Section 2.8 (Figure 4a) | not stated | not stated |
| t-test (specific type not further specified) | Differences in mRNAsi tumor stemness index among the 4 m6A subtypes (Figure 6) | TCGA-BRCA samples; exact n not stated | not stated |
| Descriptive count comparison (no formal inferential test named) | Differences in number of anthracycline-sensitive vs. non-sensitive samples across 4 subtypes in GSE25066 (Figure 5b) | GSE25066: 508 samples; exact subgroup n not stated | na |
-
Samples were classified into 4 subtypes by a 2×2 median Z-score split of two pre-defined gene sets (Set1: writers+readers; Set2: erasers), fixing both the number of groups and their boundaries a priori↳ Could also: Unsupervised consensus clustering (e.g., NMF or k-means with silhouette width or gap statistic to select the optimal k) applied to the full m6A gene expression matrix — Consensus clustering derives subgroup structure empirically from the data, provides quantitative stability metrics for the chosen k, and does not require pre-specifying the number or direction of group boundaries — allowing subtypes that may not align to the median threshold to emerge
-
A t-test was used to compare the mRNAsi stemness index across 4 subtype groups simultaneously↳ Could also: One-way ANOVA followed by a post-hoc pairwise test (e.g., Tukey HSD or Games-Howell if variances differ) to compare all pairs — ANOVA with post-hoc correction is the standard framework for a continuous outcome across more than two groups; it tests the global null hypothesis first and controls the family-wise error rate across all pairwise comparisons, which multiple independent t-tests do not
-
ANOVA was applied separately to each of the 22 immune cell types to identify those differing across m6A subtypes↳ Could also: Apply Benjamini-Hochberg FDR correction across the 22 ANOVA p-values, or use MANOVA treating the 22 proportions jointly — Testing 22 outcomes independently at p < 0.05 increases the expected number of false positives; FDR correction quantifies and limits this inflation, and MANOVA additionally accounts for correlations among cell-type proportions (which CIBERSORT outputs are constrained to sum to 1)
-
Univariate Cox regression was run separately for each of the 22 immune cell types to pre-select prognostic variables, which were then entered into a multivariate model↳ Could also: LASSO-penalized Cox regression (e.g., R package glmnet) applied to all 22 immune cell proportions simultaneously for data-driven variable selection — Univariate pre-filtering followed by multivariate modeling can introduce optimistic bias in coefficient estimates; penalized regression performs selection and estimation in a single regularized step and handles correlated predictors more stably
-
Drug sensitivity differences across subtypes were presented as raw counts of sensitive vs. non-sensitive samples per subtype without a named inferential test↳ Could also: Chi-square test of independence or Fisher's exact test (for sparse cells) comparing the distribution of drug sensitivity across the 4 subtypes, with Cramér's V as an effect size — A formal test of independence provides a p-value and effect size alongside the descriptive counts, allowing readers to assess whether the count differences exceed what would be expected by chance given the subtype sample sizes
-
Survival results were reported with p < 0.05 as the threshold and hazard ratios without confidence intervals↳ Could also: Report 95% confidence intervals for each HR, restricted mean survival time (RMST) differences at a pre-specified horizon, and/or absolute survival rate differences at clinically relevant time points — Confidence intervals communicate estimation uncertainty and clinical magnitude; RMST differences are directly interpretable in time units and do not require the proportional-hazards assumption to hold over the entire follow-up period
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
Downstream reach in the literature
191 downstream papers · 2 datasetsHow widely the datasets deposited by this paper are reused across the whole literature (Europe PMC), beyond our assessed set. This is a factual dependency map — reusing a public dataset is normal, good science. It is not a judgement on the downstream papers; the only verdict here is this paper's own, with its cited rationale.
- Survival analysis across the entire transcriptome id... 2021 · 751 cites
- Epithelial-mesenchymal transition spectrum quantific... 2014 · 604 cites
- A genomic predictor of response and survival followi... 2011 · 543 cites
- Differential response to neoadjuvant chemotherapy am... 2013 · 532 cites
- Inhibition of fatty acid oxidation as a therapy for... 2016 · 438 cites
- Aerobic glycolysis tunes YAP/TAZ transcriptional act... 2015 · 358 cites
- Clinical Value of RNA Sequencing-Based Classifiers f... 2018 · 165 cites
- bc-GenExMiner 4.5: new mining module computes breast... 2021 · 142 cites
- Abundance of Regulatory T Cell (Treg) as a Predictiv... 2020 · 95 cites
- Bulk and single-cell transcriptome profiling reveal... 2021 · 89 cites
- A lncRNA prognostic signature associated with immune... 2020 · 86 cites
- Prognosis and Dissection of Immunosuppressive Microe... 2022 · 73 cites
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-35655723
Paper: Zhao P. et al. "Bioinformatics Analysis of the Characteristics and Correlation of m6A Methylation in Breast Cancer Progression." Contrast Media Mol Imaging 2022;2022:4416439. PMID 35655723 / PMC9148239.
⚠️ RETRACTED. Retraction notice: Contrast Media Mol Imaging 2023 Sep 27;2023:9874020. The journal (Hindawi) retracted large batches of paper-mill bioinformatics articles. This RU is therefore also of interest as a fabrication-quantification datapoint, not only as a reproduction.
Datasets named in the paper
- TCGA-BRCA (1217 samples; 1104 sequenced tumor tissue) — public (GDC).
- GSE96058 (SCAN-B; 3409 sequenced incl. 136 biological replicates) — public (GEO).
- GSE25066 (508 samples) — public (GEO).
Tools named in the paper
SVA/ComBat (batch), RUVSeq, survival (KM + uni/multivariate Cox), TCGA SpliceSeq
(alternative splicing), maftools (mutation/CNV), CIBERSORT (22 immune cell
fractions), TIDE / TIDEpy (immunotherapy-response surrogate, the brief's code
link github.com/liulab-dfci/TIDEpy), mRNAsi stemness index (from PMID 29625051).
In scope (attempted — pipeline-derived, clearly specified, low-hanging)
- TCGA-BRCA top mutated genes via maftools. Paper Results report a "top-10
tumour driver gene" list (TTN, TP53, PIK3CA, CDH1, GATA3, KMT2C, MAP3K1, MUC16,
RYR2, HMCN1). This is deterministic given a standard TCGA-BRCA MAF + maftools
getGeneSummary. → clean 1:1 comparison. («job») - TIDEpy on the paper's data (GSE96058). Runs the exact named code link
(
liulab-dfci/TIDEpy, default parameters) on the paper's own GEO matrix — P16: a third-party tool on the paper's data is an equally valid reproduction. The paper reports no specific TIDE numeric value, so this demonstrates the pipeline step is runnable/reproducible and yields a cohort TIDE distribution; graded partial (tool reproduces; no printed number to match 1:1). («job»)
Out of scope / NOT attempted (the hard ~20%) — and why
- The 4 m6A subtypes ("quiescent / m6A-methylation / protein-binding / mixed" via Set1≤0/>0, Set2≤0/>0 z-score stratification). The paper never defines what Set1 and Set2 are (which gene group → which score), so the central classification is under-specified / not derivable from the text. Everything downstream depends on it.
- CIBERSORT "11 of 22 immune cell types differ among subtypes", survival differences among subtypes (KM/Cox), stemness index ranking (protein-binding<mixed<quiescent<m6A), alternative-splicing (exon-skip) differences — all conditioned on the undefined 4-subtype labels → not reproducible without guessing the classification. Skipped per 80/20.
- Wet-lab / external: none material here (whole paper is in-silico).
Fabrication / quality flags to verify against the original
- The m6A regulator gene list as text-mined is garbled: e.g. "METTL15" (a rRNA methyltransferase, not an m6A writer) listed twice ("METTL15, METL15"), "RBM3" (should be RBM15), "HRNPA2B1" (should be HNRNPA2B1), and the canonical writers METTL3/METTL14 appear absent. Verify whether this is OCR noise or a real error in the paper (recorded in AUDIT.md).
- All heavy compute on «our HPC»/«infra»; «host» holds only small results + pointers.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
Two clearly-specified peripheral pipeline steps reproduce on the paper's own public data — C1 TCGA-BRCA top-10 mutated genes via maftools (9/10, rank-10 HMCN1 vs SYNE1 is MAF-version noise) and C2 TIDEpy running cleanly on GSE96058 (no printed number to match). But the paper's central thesis — four m6A subtypes defined by the sign of undefined 'Set1'/'Set2' z-scores — is not derivable from the text, so all subtype-conditioned results (CIBERSORT, survival, stemness, splicing) are un-reproducible. The fault is squarely on the authors' side: the m6A writer list is garbled in the paper itself (duplicate METTL15, missing METTL3/METTL14), a paper-mill signature consistent with the retraction. Overall red: reproducible at the edges, non-derivable and fabrication-suspect at the core.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.