Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Comprehensive analysis of a novel RNA modifications-related model in the prognostic characterization, immune landscape and drug therapy of bladder cancer.

Front Genet · 2023
L1 71/100 3/4
Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +7
✓ What held up
  • Nothing in this column.
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🟡Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🔴A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
71/100
Reproducibility score
at the mean
vs. all fields · 1173 studies
🎯 Scores higher than 38% of all assessed papers rank 694 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough to reproduce the MODEL but NOT, 1:1, its headline validation. The registry code link (github.com/ZhengXia/DaPars) is a text-mining false positive (a cited APA tool, not the paper's analysis code); the paper ships no analysis repo, so per P16 we re-ran the described pipeline on the paper's own public data using the published 14-gene LASSO-Cox model (Suppl. Table S10, recovered in full; text betas match exactly). Applying it on «our HPC»: (1) GSE13507 (a meta-GEO training component, microarray) -> RMS significantly stratifies OS, HR=1.70 (p=0.029), correct direction -> the association is real and our pipeline is correct. (2) TCGA-BLCA, the paper's headline independent validation -> reported HR=1.53 (p=0.006) but we obtain only HR=1.18-1.33 (p=0.06-0.28) with continuous C-index ~0.51 (non-discriminative): direction reproduces, the significant effect does NOT. The works-on-microarray-training / fails-on-RNA-seq-validation pattern is the classic signature of cross-platform transfer loss / over-fit validation; combined with the paper's under-specified TCGA normalization and platform-dependent fixed threshold (median 3.344), this is a partial reproduction with the headline TCGA number flagged for human audit (not asserted as fabrication). NOT attempted (hard ~20%): de-novo LASSO re-derivation (non-deterministic), full 8-set meta-GEO sva assembly (HR=3.00), nomogram, immune/GSVA/TISIDB, GDSC drug, IMvigor210 immunotherapy. All raw/intermediate data kept on «infra»; only small results + pointers on «host».

💻 Code ↗ 🗄 Data: GSE13507

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 71
    assessed: 2026-06-15 ⛓ 22e9f848054d
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-15
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

The study tests whether a prognostic model built from the 'writer' enzymes of five adenine-related RNA modifications (m6A, m6Am, m1A, APA, and A-to-I editing) can characterize the clinical outcome, immune landscape, and therapeutic efficacy of bladder cancer.

Core claims
  • Two distinct RNA modification patterns exist among BCa samples with radically different clinical outcomes and biological characteristics. finding
  • An RNA modification 'writers' score (RMS) model of 14 phenotype-associated prognostic DEGs predicts unfavorable BCa prognosis in both training and validation cohorts. resource
  • RMS-high tumors are enriched for immunosuppressive cell infiltration and activation of EMT, angiogenesis, and IL-6/JAK/STAT3 signaling. finding
  • The RMS model stratifies responsiveness to chemotherapeutic agents and antibody-drug conjugates between RMS-low and -high groups. finding
  • Combining the RMS model with TMB, TNB and PD-L1 improves discrimination of immunotherapy responders from non-responders. finding
  • WGCNA identifies hub genes (e.g., KIAA1429/VIRMA) associated with RNA modifications, validated in human BCa specimens. method
  • Unsupervised clustering of 34 RNA modification 'writers' plus LASSO regression constructs the RMS signature. method
Experimental setups
Assay System Perturbation Readout Platform
Bulk transcriptomics / unsupervised clustering 1,410 BCa patients (meta-GEO of 8 GEO datasets) none RNA modification cluster assignment from 34 'writer' gene expression profiles GEO platforms GPL6102, GPL6947, GPL6244, GPL570
RNA-seq based LASSO model construction/validation meta-GEO training cohort and TCGA-BLCA validation cohort none RMS risk score and prognosis (OS)
Somatic mutation / CNV analysis TCGA-BLCA (412 mutation, 409 CNV, 411 RNA-seq samples) none mutation landscape, copy number variation VarScan2; maftools R package
WGCNA co-expression network analysis 411 TCGA BCa patients (5,657 prognosis-associated genes) none hub genes correlated with clinical traits WGCNA R package
Immunohistochemistry (TMA) 84 paired human BCa and adjacent non-neoplastic bladder tissues none KIAA1429 protein staining score (intensity x proportion) anti-KIAA1429 antibody (25712-1-AP, Proteintech)
RT-qPCR human BCa tissues none KIAA1429 mRNA expression (GAPDH-normalized) Roche LightCycler 480 II; SYBR Green; HiScript II Q RT SuperMix
Drug sensitivity prediction BCa cohorts chemotherapeutic agents / antibody-drug conjugates predicted therapeutic response by RMS group Genomics of Drug Sensitivity in Cancer (GDSC) database
Immune cell infiltration / GSEA BCa TME (meta-GEO) none immune cell infiltration and pathway enrichment by RMS group
Key results
  • RMS positively correlated with unsatisfactory outcome in meta-GEO training cohort HR = 3.00, 95% CI = 2.19–4.12
  • RMS associated with poor outcome in TCGA-BLCA validation cohort HR = 1.53, 95% CI = 1.13–2.09
  • Nomogram showed high prognostic prediction accuracy C-index = 0.785
  • Combining RMS with TMB, TNB and PD-L1 distinguished immunotherapy responders from non-responders AUC = 0.828
  • Immunosuppressive cell infiltration and EMT, angiogenesis, IL-6/JAK/STAT3 signaling enriched in RMS-high group
  • Two distinct RNA modification patterns identified with varying clinical outcomes
Key statistics
  • other HR = 3.00, 95% CI = 2.19–4.12 (RMS prognostic association in meta-GEO training cohort)
  • other HR = 1.53, 95% CI = 1.13–2.09 (RMS prognostic association in TCGA-BLCA validation cohort)
  • other C-index = 0.785 (nomogram prognostic prediction accuracy)
  • other AUC = 0.828 (RMS + TMB/TNB/PD-L1 immunotherapy response discrimination)
  • count 1,410 (BCa samples in meta-GEO cohort)
  • count 34 (RNA modification 'writers' used for clustering)
  • count 14 (RNA modification phenotype-associated prognostic DEGs in RMS model)
  • count 84 (paired BCa samples in tissue microarray)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This retrospective bioinformatics study merged eight GEO bladder cancer datasets into a meta-GEO training cohort (n = 1,410) and used TCGA-BLCA (n = 411) as a validation cohort. Unsupervised clustering of 34 RNA modification writer gene-expression profiles identified two molecular subtypes, and LASSO regression built a 14-gene RNA modifications-related score (RMS) from subtype-associated DEGs. Cox proportional hazards models quantified the prognostic value of RMS, ROC curves and a nomogram assessed predictive accuracy, and WGCNA identified hub genes that were subsequently validated in 84 paired tumor/adjacent-normal tissue microarray samples by IHC and RT-qPCR.

Replicationunclear Sample sizeMeta-GEO training cohort n = 1,410 (8 GEO datasets combined); TCGA-BLCA validation cohort n = 411 RNA-seq / 412 mutation / 409 CNV; IHC/RT-qPCR experimental validation in 84 paired tumor and adjacent-normal BCa samples; no formal power calculation stated GroupsRMS-high vs RMS-low; RNA modification cluster 1 vs cluster 2; immunotherapy responders vs non-responders Pairingmixed Randomization/blindingstated Dispersionnone Exact p-valuesno Effect sizesyes Confidence intervalsyes Multiplicity correctionnone stated
Statistical tests used
Test Applied to n Assumptions
Unsupervised clustering (specific algorithm not stated) Classification of 1,410 BCa samples into two RNA modification patterns based on 34 writer gene-expression profiles 1410 not stated
LASSO (Least Absolute Shrinkage and Selection Operator) regression Feature selection and construction of the 14-gene RMS prognostic signature from RNA modification phenotype-associated DEGs 1410 not stated
Cox proportional hazards model Prognostic evaluation of RMS in meta-GEO training cohort (HR = 3.00, 95% CI 2.19–4.12) and TCGA-BLCA validation cohort (HR = 1.53, 95% CI 1.13–2.09) 1410 (training); 411 (validation) not stated
ROC curve analysis / AUC Prediction of immunotherapy response (AUC = 0.828 for RMS combined with TMB, TNB, PD-L1) and nomogram performance assessment not stated
Concordance index (C-index) Evaluation of nomogram prognostic prediction accuracy (C-index = 0.785) na
Pearson's correlation coefficient WGCNA: association between module eigengenes and clinical traits (tumor stage, histological grade, survival status); module membership (MM) and gene significance (GS) calculations 411 not stated
GSEA (Gene Set Enrichment Analysis) Biological pathway and immune characteristic analysis between RMS-high and RMS-low groups not stated
Decision curve analysis (DCA) Assessment of clinical net benefit / utility of the nomogram na
Approaches that could also have been used
  • Unsupervised clustering was used to identify two RNA modification patterns, but the specific clustering algorithm is not named
    Could also: Consensus clustering (e.g., R/ConsensusClusterPlus) or non-negative matrix factorization (NMF) could also be applied, with explicit reporting of the algorithm, the range of k tested, and stability metrics such as the cophenetic correlation coefficient or silhouette width — Reporting the algorithm and cluster-stability metrics would allow readers to assess whether the two-cluster solution is well-separated and whether additional cluster numbers were considered, which is a standard expectation in molecular subtyping studies
  • LASSO regression was used to select 14 prognostic genes from a larger DEG set
    Could also: Elastic net regression or a random survival forest could also be used for variable selection in a survival context — Elastic net combines LASSO sparsity with ridge grouping and can perform better when predictors are correlated, as co-expressed genes often are; random survival forest makes fewer distributional assumptions and can capture non-linear effects
  • Pearson's correlation was used throughout WGCNA for module-trait associations and gene-level correlations
    Could also: Spearman's rank correlation could also be used for the same purposes — Spearman's correlation is more robust to non-normality and outliers in gene expression data, particularly relevant for TPM-transformed RNA-seq values whose distributions are often right-skewed
  • Batch effects across eight GEO platforms were corrected using the ComBat algorithm
    Could also: limma's removeBatchEffect, surrogate variable analysis (SVA), or PEER factors could also be applied for cross-platform harmonization — SVA estimates latent factors without requiring fully known batch labels and can capture unmeasured sources of systematic variation; reporting post-correction PCA or other diagnostics helps confirm adequate batch removal
  • DEGs between the two RNA modification clusters were identified (method and threshold not specified) without an explicitly stated multiplicity correction
    Could also: Applying a Benjamini-Hochberg FDR correction across all tested genes — and reporting the FDR threshold alongside the number of genes passing it — is a standard approach when screening thousands of genes simultaneously — Stating the FDR threshold and the resulting DEG count allows readers to assess the likely false-discovery burden and to compare the findings with other studies using similar thresholds
  • IHC semi-quantitative scores from 84 paired tumor and adjacent-normal samples were used to validate hub gene expression, but the statistical test applied to the paired comparison is not stated
    Could also: A Wilcoxon signed-rank test (for paired ordinal scores) or a paired t-test on log-transformed continuous H-scores would both be standard for paired IHC data — Paired designs gain power over unpaired analyses by accounting for within-patient variability; naming the test and reporting its p-value and an effect size (e.g., median difference with IQR) supports transparency and reproducibility
Software: R/maftools · R/GEOquery · R/sva (ComBat batch correction) · R/WGCNA

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
5
Impact: low
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

GSE13507 GEO in Methods (http://purl.org/orb/Methods)
also used by 1 paper:
GSE31684 GEO in Methods (http://purl.org/orb/Methods)
also used by 1 paper:
GSE48075 GEO in Methods (http://purl.org/orb/Methods)
also used by 1 paper:
10.5281/zenodo.546110 DOI in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GPL6102 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE104922 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE128959 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE32548 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE83586 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet

Downstream reach in the literature

336 downstream papers · 8 datasets

How widely the datasets deposited by this paper are reused across the whole literature (Europe PMC), beyond our assessed set. This is a factual dependency map — reusing a public dataset is normal, good science. It is not a judgement on the downstream papers; the only verdict here is this paper's own, with its cited rationale.

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Figures / tables: Table
A_signature
Reported
RMS = 0.2062*IFNLR1 + 0.1822*PCDHB11 - 0.1428*TIMM21 + ... + 0.0057*CRELD1 + 0.0028*FOXG1 (14-gene LASSO Cox)
Reproduced
All 14 genes + betas recovered from Suppl. Table S10; the 5 betas printed in the main text match S10 exactly
exact
B_tcga_HR
Reported
TCGA-BLCA validation: OS HR=1.53 (1.13-2.09), p=0.006
Reproduced
HR=1.18 (literal) / 1.33 (z-scored), p=0.28 / 0.057; n=401, 177 events; continuous C-index 0.517
partial
D_gse13507_HR
Reported
meta-GEO training HR=3.00 (2.19-4.12), p=1.06e-11 (GSE13507 is one of its 8 components)
Reproduced
GSE13507 alone: HR=1.70 (1.06-2.75), p=0.029, C-index 0.573; n=165 primary tumors, 69 events
within tolerance
E_nomogram
Reported
Nomogram C-index 0.785; 1/3/5-yr AUC 0.821/0.825/0.806
Reproduced
not attempted (out of 80/20 core; RMS-alone TCGA C-index ~0.51, not comparable)
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 71/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟡2. Endpoint comparability
🟡3. Location of the main deviation
🔴4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +7

The 14-gene RMS signature (Suppl. Table S10) reproduces exactly and stratifies OS significantly in the GSE13507 training component (HR=1.70, p=0.029, correct direction), confirming the pipeline is correct and the meta-GEO association is real. However, the paper's headline independent TCGA-BLCA validation (HR=1.53, p=0.006) is not reproducible from public RNA-seq with the published coefficients — we obtain HR=1.18-1.33 (p=0.28/0.057), C-index ~0.51. The deviation is on the input/authors' side (cross-platform microarray->RNA-seq transfer plus under-specified TCGA normalization and a non-transferable fixed threshold), and the value is not cleanly derivable from the shared data — but the surviving direction and the absence of any 'too-perfect' pattern argue against fabrication. Overall a partial reproduction with the headline TCGA number flagged for human audit.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

212.5 k
tokens (I/O) · 13.3 M incl. cache
24 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.