Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Bioinformatics Analysis of the Characteristics and Correlation of m6A Methylation in Breast Cancer Progression.

Contrast Media Mol Imaging · 2022
L1 48/100 PQI 83
Why this verdict

The main result did not reproduce in this reproduction attempt. Where our recomputation produced values that differ from the published ones, those discrepancies are listed below. This is a single automated attempt — not peer review and not a finding of error or misconduct — and differences can also arise from data access, undocumented parameters or the computing environment. The verdict can be contested via “report an error”.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Main result did not reproduce
Decisive
From: Q5 · Derivability / plausibility 🔴
Main result did not reproduce
Decisive
From: Q7 · Core claim 🔴
Main result did not reproduce
Decisive
From: Q8 · Severity of the miss (overall human judgment) 🔴
✓ What held up
  • Nothing in this column.
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🟡Reported values were only indirectly comparable
  • 🔴A deviation arose in the data or preprocessing
  • 🔴A deviation was attributed to the published material
  • 🔴Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🔴The central claim did not (fully) hold under reproduction
  • 🔴Overall, the reproduction showed a material discrepancy
How its reproducibility compares
48/100
Reproducibility score
1.5 SD below mean
vs. all fields · 1173 studies
🎯 Scores higher than 6% of all assessed papers rank 1092 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Paper is RETRACTED (Hindawi paper-mill, Contrast Media Mol Imaging 2023;2023:9874020). Described well enough only in parts. Two clearly-specified pipeline outputs reproduced on the paper's own public data: (C1) TCGA-BRCA top-10 most-mutated genes via maftools matched 9/10 (only rank-10 HMCN1 vs SYNE1 differs, MAF-version noise; all canonical BRCA drivers recovered) => essentially reproduced/within-tol; (C2) the brief's code link TIDEpy ran cleanly with default params on the full GSE96058 matrix (3409 samples) => pipeline step reproducible but the paper prints no TIDE number, so partial (P16 third-party-tool reproduction). NOT attempted (the hard ~20%): the paper's central '4 m6A subtypes' are defined by sign of undefined 'Set1'/'Set2' z-scores - the classification is not derivable from the text, so all subtype-conditioned results (CIBERSORT 11/22, subtype survival, stemness ranking, exon-skip splicing) cannot be reproduced without guessing. Fabrication/quality flag (C3): the m6A 'writer' gene list is wrong in the paper itself (METTL15 listed twice and is not an m6A writer; canonical METTL3/METTL14 absent), a strong paper-mill signal consistent with the retraction. All heavy compute on «our HPC»/«infra»; only small results + pointers on «host».

💻 Code ↗ 🗄 Data: GSE96058

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 48
    assessed: 2026-06-15 ⛓ 0afa19cd01fd
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-15
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

The study tests whether breast cancer can be stratified into m6A methylation subtypes based on m6A-related gene expression, and whether these subtypes differ systematically in survival, prognosis, mutation, copy number variation, immune infiltration, drug sensitivity, and tumor stemness, thereby clarifying the role of m6A methylation in breast cancer progression.

Core claims
  • Breast cancer samples can be divided into 4 m6A subtypes (quiescent, m6A methylation, protein-binding, mixed) based on expression of m6A-related genes, with consistent proportions across datasets finding
  • The four m6A subtypes show significantly different survival rates, with the mixed subtype having higher survival and the protein-binding subtype having significantly lower survival finding
  • TP53 mutation and copy number loss are most pronounced in the protein-binding subtype among tumor driver genes finding
  • RAD54B copy number is increased in the protein-binding subtype while other DNA damage repair genes show copy number deletion across subtypes finding
  • The four m6A subtypes differ significantly in macrophages M0 and resting mast cells infiltration, which correlate with patient prognosis finding
  • The m6A methylated subtype has the highest tumor stemness index (mRNAsi), indicating high m6A reader expression drives breast cancer progression finding
  • m6A gene expression contributes to breast cancer response to anthracyclines, with the quiescent/resting group having the fewest drug-sensitive samples finding
  • Number of exon-skip (ES) alternative splicing events differs among the four m6A subtypes while other splicing types do not finding
Experimental setups
Assay System Perturbation Readout Platform
Bulk RNA-seq transcriptome analysis / m6A subtype classification (Z-score of Set1/Set2 gene expression) TCGA-BRCA breast cancer tissue (human) none Expression levels of 21 m6A-related genes; subtype assignment TCGA-BRCA dataset (1,217 samples; 1,104 tumor)
Bulk gene expression / m6A subtype classification GSE96058 breast cancer samples (human) none m6A-related gene expression; subtype assignment GEO GSE96058 dataset (3,409 samples)
Gene expression subtyping and anthracycline sensitivity analysis GSE25066 breast cancer samples (human) anthracycline chemotherapy (clinical) Number of drug-sensitive vs nonsensitive samples per subtype GEO GSE25066 dataset (508 samples)
Survival/prognosis analysis TCGA-BRCA breast cancer (human, including triple-negative) none Survival rate by subtype; clinicopathological correlation R survival package; Circos
Alternative splicing analysis TCGA-BRCA breast cancer (human) none PSI of 7 splicing types (AP, ES, RD, AT, ME, AD, AA) TCGA SpliceSeq
Mutation and copy number variation analysis TCGA-BRCA breast cancer (human) none Mutation/CNV of top tumor driver genes and DNA damage repair genes (GO:0006281) R Maftools; MSigDB V7.1
Immune cell infiltration deconvolution TCGA-BRCA and GSE96058 breast cancer (human) none Proportion of 22 immune cell types; prognostic association CIBERSORT V1.01; R survival coxph
Immune efficacy and tumor stemness analysis TCGA-BRCA and GSE96058 breast cancer (human) none TIDE immune efficacy score; mRNAsi stemness index TIDEpy (default parameters); mRNAsi from PMID 29625051
Key results
  • Survival differs significantly among the 4 m6A subtypes; mixed subtype has higher survival than the other three
  • Protein-binding subtype has significantly lower survival; m6A methylated vs protein-bound subgroups differ significantly in prognosis (including TNBC)
  • TP53 has the most mutations and copy number loss in the protein-binding subtype
  • RAD54B copy number increased in protein-bound subtype; other 9 DNA damage repair genes show copy number deletion
  • 11 of 22 immune cell types show significantly different infiltration among the 4 subgroups; macrophages M0 and resting mast cells prognostically relevant
  • NK cells resting and macrophages M0 associated with survival in all samples (HR<1); dendritic cells resting and mast cells resting in TNBC (HR<1) HR<1
  • Stemness index (mRNAsi) ranked low to high: protein-bonding, mixed, quiescent, m6A methylated (highest)
  • Quiescent/resting group has the fewest anthracycline drug-sensitive samples among the 4 subtypes
Key statistics
  • count 1,217 samples including 1,104 sequenced tumor tissue samples (TCGA-BRCA dataset size)
  • count 3409 sequenced samples including 136 biological replicates (GSE96058 dataset size)
  • count 508 samples (GSE25066 dataset size)
  • other 0.1–0.4% of total adenosine residues (abundance of m6A methylation in eukaryotic cells)
  • count 22 immune cell types analyzed, 11 significantly different among 4 subgroups (CIBERSORT immune infiltration)
  • fold_change increased by 100 times (IGF2BP1 mRNA expression in fibrolamellar HCC (cited literature))
  • pvalue p < 0.05 (threshold for statistical significance (ANOVA/t-test))
  • count 21 m6A methylation-related genes (8 writer, 2 eraser, 11 reader) (m6A gene set used for subtyping)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This retrospective bioinformatics study classified breast cancer samples from three public datasets (TCGA-BRCA, GSE96058, GSE25066) into four m6A methylation subtypes via a Z-score median split of two pre-defined gene sets. Survival analyses, univariate and multivariate Cox proportional hazards models, ANOVA, and t-tests were applied to compare subtypes across prognosis, immune cell infiltration, genomic features, drug sensitivity, and tumor stemness. Results were reported with a p < 0.05 significance threshold, with hazard ratios provided for Cox analyses.

Replicationbiological Sample sizeDataset-level sample sizes stated (TCGA-BRCA: 1,217 total / 1,104 tumor; GSE96058: 3,409; GSE25066: 508); no formal sample-size justification or power calculation reported Groups4 m6A subtypes: quiescent, m6A methylation, protein-binding, mixed Pairingunpaired Randomization/blindingnot stated Dispersionnone Exact p-valuesno Effect sizesyes Confidence intervalsno Multiplicity correctionnone stated
Statistical tests used
Test Applied to n Assumptions
Kaplan-Meier survival analysis with implied log-rank test (survival R package) Overall survival comparison across the 4 m6A subtypes in all samples and TNBC subgroup; also per immune cell type (Figures 2a, 2b, 3b, 3c, 4f, 4g) TCGA-BRCA 1,104 tumor samples; GSE96058 3,409 samples; GSE25066 508 samples — exact per-analysis n not stated not stated
Univariate Cox proportional hazards regression (coxph) Association of each of 22 immune cell infiltration proportions with patient survival (Figures 4b, 4c) TCGA-BRCA and GSE96058 combined samples; exact n not stated not stated
Multivariate Cox proportional hazards regression (risk score model) Risk score derived from immune cell infiltration proportions weighted by beta coefficients (Figures 4d, 4e) not stated not stated
One-way ANOVA Differences in proportions of 22 immune cell types across the 4 m6A subtypes; stated as general method for subtype comparisons in Section 2.8 (Figure 4a) not stated not stated
t-test (specific type not further specified) Differences in mRNAsi tumor stemness index among the 4 m6A subtypes (Figure 6) TCGA-BRCA samples; exact n not stated not stated
Descriptive count comparison (no formal inferential test named) Differences in number of anthracycline-sensitive vs. non-sensitive samples across 4 subtypes in GSE25066 (Figure 5b) GSE25066: 508 samples; exact subgroup n not stated na
Approaches that could also have been used
  • Samples were classified into 4 subtypes by a 2×2 median Z-score split of two pre-defined gene sets (Set1: writers+readers; Set2: erasers), fixing both the number of groups and their boundaries a priori
    Could also: Unsupervised consensus clustering (e.g., NMF or k-means with silhouette width or gap statistic to select the optimal k) applied to the full m6A gene expression matrix — Consensus clustering derives subgroup structure empirically from the data, provides quantitative stability metrics for the chosen k, and does not require pre-specifying the number or direction of group boundaries — allowing subtypes that may not align to the median threshold to emerge
  • A t-test was used to compare the mRNAsi stemness index across 4 subtype groups simultaneously
    Could also: One-way ANOVA followed by a post-hoc pairwise test (e.g., Tukey HSD or Games-Howell if variances differ) to compare all pairs — ANOVA with post-hoc correction is the standard framework for a continuous outcome across more than two groups; it tests the global null hypothesis first and controls the family-wise error rate across all pairwise comparisons, which multiple independent t-tests do not
  • ANOVA was applied separately to each of the 22 immune cell types to identify those differing across m6A subtypes
    Could also: Apply Benjamini-Hochberg FDR correction across the 22 ANOVA p-values, or use MANOVA treating the 22 proportions jointly — Testing 22 outcomes independently at p < 0.05 increases the expected number of false positives; FDR correction quantifies and limits this inflation, and MANOVA additionally accounts for correlations among cell-type proportions (which CIBERSORT outputs are constrained to sum to 1)
  • Univariate Cox regression was run separately for each of the 22 immune cell types to pre-select prognostic variables, which were then entered into a multivariate model
    Could also: LASSO-penalized Cox regression (e.g., R package glmnet) applied to all 22 immune cell proportions simultaneously for data-driven variable selection — Univariate pre-filtering followed by multivariate modeling can introduce optimistic bias in coefficient estimates; penalized regression performs selection and estimation in a single regularized step and handles correlated predictors more stably
  • Drug sensitivity differences across subtypes were presented as raw counts of sensitive vs. non-sensitive samples per subtype without a named inferential test
    Could also: Chi-square test of independence or Fisher's exact test (for sparse cells) comparing the distribution of drug sensitivity across the 4 subtypes, with Cramér's V as an effect size — A formal test of independence provides a p-value and effect size alongside the descriptive counts, allowing readers to assess whether the count differences exceed what would be expected by chance given the subtype sample sizes
  • Survival results were reported with p < 0.05 as the threshold and hazard ratios without confidence intervals
    Could also: Report 95% confidence intervals for each HR, restricted mean survival time (RMST) differences at a pre-specified horizon, and/or absolute survival rate differences at clinically relevant time points — Confidence intervals communicate estimation uncertainty and clinical magnitude; RMST differences are directly interpretable in time units and do not require the proportional-hazards assumption to hold over the entire follow-up period
Software: R 4.0 · ComBat (SVA package) 3.36.0 · RUVSeq 1.22.0 · survival (R package) 3.2-3 · Maftools (R package) · CIBERSORT 1.01 · TIDE (TIDEpy) · MSigDB 7.1 · Circos

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
1
Impact: low
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

GSE96058 GEO in Methods (http://purl.org/orb/Methods)
also used by 2 papers:
GSE25066 GEO in Methods (http://purl.org/orb/Methods)
also used by 1 paper:

Downstream reach in the literature

191 downstream papers · 2 datasets

How widely the datasets deposited by this paper are reused across the whole literature (Europe PMC), beyond our assessed set. This is a factual dependency map — reusing a public dataset is normal, good science. It is not a judgement on the downstream papers; the only verdict here is this paper's own, with its cited rationale.

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-35655723

Paper: Zhao P. et al. "Bioinformatics Analysis of the Characteristics and Correlation of m6A Methylation in Breast Cancer Progression." Contrast Media Mol Imaging 2022;2022:4416439. PMID 35655723 / PMC9148239.

⚠️ RETRACTED. Retraction notice: Contrast Media Mol Imaging 2023 Sep 27;2023:9874020. The journal (Hindawi) retracted large batches of paper-mill bioinformatics articles. This RU is therefore also of interest as a fabrication-quantification datapoint, not only as a reproduction.

Datasets named in the paper

  • TCGA-BRCA (1217 samples; 1104 sequenced tumor tissue) — public (GDC).
  • GSE96058 (SCAN-B; 3409 sequenced incl. 136 biological replicates) — public (GEO).
  • GSE25066 (508 samples) — public (GEO).

Tools named in the paper

SVA/ComBat (batch), RUVSeq, survival (KM + uni/multivariate Cox), TCGA SpliceSeq (alternative splicing), maftools (mutation/CNV), CIBERSORT (22 immune cell fractions), TIDE / TIDEpy (immunotherapy-response surrogate, the brief's code link github.com/liulab-dfci/TIDEpy), mRNAsi stemness index (from PMID 29625051).

In scope (attempted — pipeline-derived, clearly specified, low-hanging)

  1. TCGA-BRCA top mutated genes via maftools. Paper Results report a "top-10 tumour driver gene" list (TTN, TP53, PIK3CA, CDH1, GATA3, KMT2C, MAP3K1, MUC16, RYR2, HMCN1). This is deterministic given a standard TCGA-BRCA MAF + maftools getGeneSummary. → clean 1:1 comparison. («job»)
  2. TIDEpy on the paper's data (GSE96058). Runs the exact named code link (liulab-dfci/TIDEpy, default parameters) on the paper's own GEO matrix — P16: a third-party tool on the paper's data is an equally valid reproduction. The paper reports no specific TIDE numeric value, so this demonstrates the pipeline step is runnable/reproducible and yields a cohort TIDE distribution; graded partial (tool reproduces; no printed number to match 1:1). («job»)

Out of scope / NOT attempted (the hard ~20%) — and why

  • The 4 m6A subtypes ("quiescent / m6A-methylation / protein-binding / mixed" via Set1≤0/>0, Set2≤0/>0 z-score stratification). The paper never defines what Set1 and Set2 are (which gene group → which score), so the central classification is under-specified / not derivable from the text. Everything downstream depends on it.
  • CIBERSORT "11 of 22 immune cell types differ among subtypes", survival differences among subtypes (KM/Cox), stemness index ranking (protein-binding<mixed<quiescent<m6A), alternative-splicing (exon-skip) differences — all conditioned on the undefined 4-subtype labels → not reproducible without guessing the classification. Skipped per 80/20.
  • Wet-lab / external: none material here (whole paper is in-silico).

Fabrication / quality flags to verify against the original

  • The m6A regulator gene list as text-mined is garbled: e.g. "METTL15" (a rRNA methyltransferase, not an m6A writer) listed twice ("METTL15, METL15"), "RBM3" (should be RBM15), "HRNPA2B1" (should be HNRNPA2B1), and the canonical writers METTL3/METTL14 appear absent. Verify whether this is OCR noise or a real error in the paper (recorded in AUDIT.md).
  • All heavy compute on «our HPC»/«infra»; «host» holds only small results + pointers.
C1
Reported
TTN, TP53, PIK3CA, CDH1, GATA3, KMT2C, MAP3K1, MUC16, RYR2, HMCN1 (top-10 most-mutated TCGA-BRCA genes, maftools)
Reproduced
PIK3CA, TP53, TTN, MUC16, CDH1, GATA3, MAP3K1, KMT2C, SYNE1, RYR2 (maftools 2.26.0, MC3 BRCA n=1026, by total mutations)
within tolerance
C2
Reported
TIDE (TIDEpy, default params) on GSE96058 / TCGA-BRCA as immune-efficacy indicator; no numeric value printed
Reproduced
TIDEpy 1.3.9 on all 3409 GSE96058 samples (-c Other --force_normalize): Responder 23.6%, TIDE mean 0.351 / median 0.366
partial
C3
Reported
m6A writers: METTL15, METL15, RBM3, RBM15B, WTAP, CBLL1, KIAA1429, ZC3H13
Reproduced
verified garbled in paper text (METTL15 doubled & not an m6A writer; RBM3 should be RBM15; METTL3/METTL14 absent; HRNPA2B1 should be HNRNPA2B1)
did not match

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 48/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟡2. Endpoint comparability
🔴3. Location of the main deviation
🔴4. Cause of the deviation
🔴5. Derivability / plausibility
🟡6. Severity of the deviation
🔴7. Core claim
🔴8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Main result did not reproduce
Decisive
From: Q5 · Derivability / plausibility 🔴
Main result did not reproduce
Decisive
From: Q7 · Core claim 🔴
Main result did not reproduce
Decisive
From: Q8 · Severity of the miss (overall human judgment) 🔴

Two clearly-specified peripheral pipeline steps reproduce on the paper's own public data — C1 TCGA-BRCA top-10 mutated genes via maftools (9/10, rank-10 HMCN1 vs SYNE1 is MAF-version noise) and C2 TIDEpy running cleanly on GSE96058 (no printed number to match). But the paper's central thesis — four m6A subtypes defined by the sign of undefined 'Set1'/'Set2' z-scores — is not derivable from the text, so all subtype-conditioned results (CIBERSORT, survival, stemness, splicing) are un-reproducible. The fault is squarely on the authors' side: the m6A writer list is garbled in the paper itself (duplicate METTL15, missing METTL3/METTL14), a paper-mill signature consistent with the retraction. Overall red: reproducible at the edges, non-derivable and fabrication-suspect at the core.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

143.8 k
tokens (I/O) · 7.1 M incl. cache
16 min
runtime · 0.07 CPU-h
5.2 GB
peak RAM
3 (1 failed)
HPC jobs
hummel
machine