Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Association between Arsenic Level, Gene Expression in Asian Population, and In Vitro Carcinogenic Bladder Tumor.

Oxid Med Cell Longev · 2022
not yet assessed 2/4
Why this verdict

Part of the results reproduced; minor but material deviations remained.

Reproduced on the brainbox compute brainarbeit.com
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
What did not (or only partly)
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
Reproduction agent’s raw note

PROVISIONAL/IN-PROGRESS. Target = GSE57711 (Data1) of a multi-dataset arsenic/gene-expression meta-analysis. Brief's code attribution (arrayQC_Module) is a MISMATCH: that repo is an Agilent QC tool used only for GSE110852, while GSE57711 is Affymetrix Gene 1.0 ST processed with 'the Affymetrix package in R'. Reproducing via platform-correct oligo/RMA + t-test/ANOVA DE per P16. In-scope: DE gene counts (476/532/232) + 3 FC>=2 sex genes. Out of scope: IPA pathways (proprietary), Random Forest/PLS-DA (stochastic), bladder AUROC (other datasets). «our HPC» tunnel intermittently down at start; offline scope/profiling done first.

💻 Code ↗ 🗄 Data: GSE57711

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment
    assessed: 2026-06-19 ⛓ e47e89d8a948
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-19
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-19
no human curator yet
Last updated
2026-07-31

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

The study tests whether differentially expressed genes derived from arsenic-exposed human populations and arsenic-treated cancer cell lines can yield a molecular signature that stratifies and predicts the risk of developing urothelial (bladder) cancer following arsenic exposure.

Core claims
  • A unique set of 147 genes is associated with arsenic exposure and linked to molecular mechanisms of cancer. finding
  • A three-probe/gene logistic regression model (NKIRAS2, AKTIP, HLA-DQA1) predicts bladder cancer risk, with highest ability for recurrent bladder tumors. resource
  • Integrin-linked kinase (ILK) signaling is one of the most significant pathways shared between both arsenic-exposed populations and is linked to protective response against oxidative damage. mechanism
  • MAPK is among the most active networks linking arsenic-associated genes and ATO-treated myeloma cell lines, indicative of oxidative/nitrosative damage associated with bladder cancer. mechanism
  • Arsenic exposure is mainly associated with organismal injury, immunological/inflammatory/gastrointestinal disease, and increased cancer rates, acting via ROS generation and oxidative stress. finding
  • A three-step integrative analysis (human exposure populations, ATO-treated myeloma cell lines, bladder cancer biopsy datasets) was used to define and validate the predictive signature. method
  • Cancer was the most significant disease and lipid metabolism the most significant molecular/cellular function associated with arsenic-differentially-expressed genes. finding
  • Sex- and arsenic-exposure-associated genes such as XIST, MALAT1, USP9Y, DDX3X, KDM6A, and ZFX were identified as most discriminating. finding
Experimental setups
Assay System Perturbation Readout Platform
Blood cell gene expression microarray (re-analysis of GSE57711, Data1) Human Bangladesh population (29 individuals; low vs high arsenic exposure) Arsenic exposure (low 50-200 µg/L vs high 232-1000 µg/L) Differentially expressed genes by arsenic level and sex Affymetrix (Affymetrix R package)
Blood transcriptome gene expression microarray (re-analysis of GSE110852, Data2) Human Pakistan population (57 individuals; low/medium/high arsenic exposure) Arsenic exposure (urinary: low 0-50, medium 51-100, high >101 µg/g creatinine) Differentially expressed genes by arsenic level and sex Agilent (in-house arrayQC pipeline, BiGCAT-UM)
Gene expression profiling of ATO-treated cancer cell lines (re-analysis of GSE14519) Four multiple myeloma cell lines: U266, MM1S, KMS11, 8226S Arsenic trioxide (ATO) exposure for 6 hr, 28 hr, 48 hr Genes up-/downregulated in response to ATO
Bladder cancer biopsy gene expression microarray (re-analysis of GSE13507) Human bladder tissue (165 primary tumors, 23 recurrent nonmuscle-invasive, 58 surrounding mucosa, 10 normal mucosa) none (tumor vs normal/recurrent comparison) Prognostic gene expression for bladder cancer risk prediction
Bladder cancer biopsy gene expression microarray (re-analysis of GSE3167) Human bladder tissue (28 superficial tumors, 13 muscle-invasive carcinomas, 9 normal) none (carcinoma stage comparison) Validation of gene-based bladder cancer prediction
Key results
  • Three-gene (NKIRAS2, AKTIP, HLA-DQA1) risk model predicts recurrent bladder tumors on training data AUC 0.94 (95% CI 0.744-0.995)
  • Three-gene risk model performance on validation data AUC 0.75 (95% CI 0.343-0.933)
  • 147 genes identified as associated with arsenic exposure and linked to cancer mechanisms 147 genes
  • ANOVA with post hoc test identified differentially expressed probes in Data1 476 probes (476 unique genes)
  • ANOVA with post hoc test identified differentially expressed probes in Data2 529 probes (439 unique genes)
  • PLS-DA showed clear separation between high-arsenic-exposed females and low-arsenic-exposed males in Data1
  • Pearson correlation among Data1 samples shows variability with no clear separation by sex/exposure r = 0.92 to 1
  • Pearson correlation among Data2 samples shows variability with no clear separation by sex/exposure r = 0.75 to 1
Key statistics
  • correlation AUC 0.94 (95% CI: 0.744-0.995) (Training data, recurrent bladder tumor prediction model)
  • correlation AUC 0.75 (95% CI: 0.343-0.933) (Validation data, bladder tumor prediction model)
  • count 147 genes (Arsenic-exposure-associated genes linked to cancer mechanisms)
  • count 476 probes / 476 unique genes (Differentially expressed in Data1 (p<0.05) by ANOVA)
  • count 529 probes / 439 unique genes (Differentially expressed in Data2 (p<0.05) by ANOVA)
  • correlation 0.92 to 1 (Pearson correlation coefficient range, Data1 samples)
  • correlation 0.75 to 1 (Pearson correlation coefficient range, Data2 samples)
  • pvalue alpha (per-gene) 0.00043; 80% power, 2-fold change, SD 0.6 (Power analysis minimum samples for GSE57711)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The paper integrated four publicly available GEO microarray datasets in a three-step analysis: (1) differential gene expression in two arsenic-exposed Asian human cohorts (Bangladesh n=29, Pakistan n=57) using t-tests and one-way ANOVA with Tukey HSD, supplemented by PLS-DA, Random Forest, Pearson correlation, and hierarchical clustering for exploratory profiling; (2) overlap analysis with four ATO-treated myeloma cell line datasets and IPA pathway enrichment; and (3) logistic regression with 10-fold cross-validation and Monte Carlo cross-validation to build and validate a bladder cancer risk prediction model, evaluated by AUROC with 95% CI. Gene filtering at the differential expression stage used a nominal p<0.05 threshold without FDR correction, explicitly to retain genes for downstream evaluation.

Replicationbiological Sample sizePower analysis performed for GSE57711: minimum 14 samples required for 80% power (alpha per-gene 0.00043, fold change 2, SD 0.6, ~11,626 genes, 5 acceptable false positives); sample sizes stated for all four GEO datasets GroupsLow vs. high arsenic exposure (Data1); low vs. medium vs. high arsenic exposure (Data2); male vs. female; ATO-treated vs. control myeloma cell lines; primary/recurrent bladder cancer vs. normal bladder tissue Pairingunpaired Randomization/blindingnot stated Dispersionnone Exact p-valuesno Effect sizesno Confidence intervalsyes Multiplicity correctionNone applied at the gene-filtering stage (t-tests across ~11,600 genes); Tukey HSD applied only for ANOVA post hoc pairwise contrasts
Statistical tests used
Test Applied to n Assumptions
Two-sample t-test Pairwise gene-level comparisons between arsenic exposure categories (low/high; low/medium; medium/high) and sex (male/female) in Data1 and Data2 Data1: n=29 (low n=15, high n=14); Data2: n=57 (low n=18, medium n=19, high n=20) not stated
One-way ANOVA with post hoc Tukey HSD Comparing gene expression across arsenic exposure levels combined with sex effect in Data1 and Data2; yielded 476 probes (Data1) and 529 probes (Data2) at p<0.05 Data1: n=29; Data2: n=57 not stated
Partial Least Squares Discriminant Analysis (PLS-DA) Global gene expression profile visualization and group separation by sex and arsenic exposure in Data1 and Data2 Data1: n=29; Data2: n=57 na
Random Forest (RF) classification Gene importance ranking for distinguishing sex and arsenic exposure categories in Data1 and Data2 Data1: n=29; Data2: n=57 na
Pearson correlation with hierarchical clustering Sample-to-sample correlation heat maps and dendrograms for Data1 and Data2 Data1: n=29; Data2: n=57 not stated
Univariate logistic regression with AUROC, 10-fold cross-validation, and Monte Carlo cross-validation (MCCV) Bladder cancer risk prediction model; trained on GSE13507 (165 primary + 23 recurrent + 58 surrounding + 10 normal samples), validated on GSE3167 (28 superficial + 13 muscle-invasive + 9 normal samples) Training set GSE13507: 256 samples total; validation set GSE3167: 50 samples total not stated
Approaches that could also have been used
  • Gene-level filtering used a nominal p<0.05 threshold without FDR correction across ~11,600 simultaneously tested genes, explicitly to retain candidates for downstream steps
    Could also: Apply Benjamini-Hochberg FDR correction and use a relaxed q-value threshold (e.g., q<0.20 or q<0.30) if a stricter cutoff would remove too many genes — At α=0.05 with ~11,600 genes, approximately 580 false positives are expected by chance; reporting an FDR threshold alongside the nominal threshold would quantify the expected false-discovery rate and help readers calibrate confidence in the gene list while still permitting liberal discovery
  • Multiple pairwise t-tests were run across all exposure category pairs (low/high, low/medium, medium/high) for each gene independently, in addition to ANOVA
    Could also: Use ANOVA as the single omnibus test for all group comparisons, then extract pairwise contrasts exclusively through the Tukey HSD post hoc step — Running t-tests and ANOVA in parallel on the same data introduces redundancy; a unified ANOVA-plus-post-hoc framework consolidates the type-I error control into one procedure and avoids counting the same comparison twice
  • Prediction model performance was summarized primarily by AUROC with 95% CI; no calibration or threshold-specific metrics were reported
    Could also: Report sensitivity, specificity, positive and negative predictive values at a chosen operating threshold, plus a calibration metric such as the Brier score or a calibration plot — AUROC captures overall rank discrimination but does not reflect how well predicted probabilities match observed outcomes or how the model performs at the specific threshold a clinician would use; complementary metrics give a fuller picture of clinical utility
  • A single fixed training dataset (GSE13507) and a single fixed validation dataset (GSE3167) were used, with feature selection (147-gene set) derived from upstream analyses on the same data
    Could also: Nested cross-validation — with feature selection inside the inner loop — or an entirely independent external cohort not touched during any feature selection step — When the same dataset informs both feature selection and model evaluation, optimistic bias in performance estimates can result; separating these steps through nested CV or held-out external validation yields less biased AUC estimates
  • Dispersion of gene expression values (e.g., SD, SEM, IQR) was not reported for any individual gene comparisons
    Could also: Report mean ± SD or median with IQR for key differentially expressed genes, or present box/violin plots for the top candidate genes — Dispersion alongside the central tendency shows whether group differences are large relative to within-group variability, which helps readers assess biological magnitude independent of the p-value and sample size
  • The paper used one-way ANOVA to assess the combined effect of arsenic exposure level and sex, treating exposure-by-sex combinations as a single grouping factor
    Could also: Two-way ANOVA with arsenic exposure and sex as separate factors and an interaction term — A two-way ANOVA with an interaction term explicitly tests whether the effect of arsenic exposure differs between males and females (and vice versa), which is directly relevant given the paper's observation that high-exposure females showed the most distinct expression profiles
Software: R/Bioconductor · IPA (Ingenuity Pathway Analysis) 2020 · R/randomForest · R/e1071 · R/pvclust · R/ggplot2

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-35039759

Paper: Singhal et al. 2022, Oxid Med Cell Longev 2022:3459855. "Association between Arsenic Level, Gene Expression in Asian Population, and In Vitro Carcinogenic Bladder Tumor." PMID 35039759 / PMC8760535.

Datasets the paper relies on

Label Accession Platform N Role
Data1 GSE57711 GPL16522 Affymetrix Human Gene 1.0 ST 29 arsenic-exposed PBMC (Bangladesh) — THIS RU's target
Data2 GSE110852 (arrayQC_Module / Agilent in-house QC) 57 arsenic-exposed, second cohort
Train GSE13507 Illumina bladder-cancer logistic model training
Valid GSE3167 Affymetrix bladder-cancer logistic model validation
(cited) GSE14519 myeloma cell-line comparison

IMPORTANT — code attribution correction

The brief lists Code = github.com/BiGCAT-UM/arrayQC_Module. Reading the paper, that repo is the "in-house QC pipeline" used ONLY for GSE110852 (Data2) — and it is an Agilent FES / GenePix spotted-array limma-wrapper. GSE57711 (Data1) is Affymetrix Gene 1.0 ST and the paper states it was processed with "the Affymetrix package in R." So arrayQC_Module is NOT applicable to this RU's dataset (platform mismatch). Per brief rule P16 we reproduce by running the correct established pipeline (oligo/RMA + differential expression) on GSE57711, which is the platform-appropriate equivalent of the paper's "Affymetrix package in R."

In scope (pipeline-derived, GSE57711)

The reproducible computational outputs for Data1, with their exact paper statements:

id Result Method (paper) Threshold
C1 476 probes = 476 unique gene symbols differentially expressed one-way ANOVA + Tukey HSD post-hoc p<0.05, no FDR
C2 532 genes sex-differentiated (male vs female) t-test p≤0.05
C3 3 of the 532 with fold-change ≥2: PRKY (down), TMSB4Y (down), KI67 (up) t-test + FC p≤0.05, FC≥2
C4 232 genes arsenic-differentiated (low vs high exposure) t-test p<0.05
C5 (profile) GSE57711 = 29 samples (16 M / 13 F; 15 low / 14 high) GEO deposit

Pipeline: download GSE57711 CEL files → RMA via oligo + pd.hugene.1.0.st.v1 → map transcript clusters to gene symbols (hugene10sttranscriptcluster.db) → log + (paper: pareto) scaling → per-comparison t-test / ANOVA at p<0.05, FC≥2 filter.

Out of scope (not pipeline-reproducible / proprietary / stochastic)

  • IPA (Ingenuity Pathway Analysis v2020) — proprietary, licensed; the 145/36/180/50 pathway counts and "6 common pathways" cannot be reproduced.
  • Random Forest top-15 / PLS-DA / hierarchical top-30 — stochastic, seed undocumented.
  • Bladder-cancer logistic-regression AUROC (GSE13507/GSE3167) — different datasets, not GSE57711; the model coefficients are reported but the train/test split & MCCV are underspecified. (May attempt as a stretch if time permits — not the RU's target data.)
  • Myeloma cell-line overlaps (147 genes) — GSE14519 + manual curation, out of scope.
  • Cross-cohort overlaps with Data2 (7 common genes etc.) — needs GSE110852 too.

Honesty flags

  • Probe→gene mapping method NOT described in paper → counts sensitive to annotation choice.
  • "pareto scaling" + "log" + no-FDR t-test = MetaboAnalyst-style; exact group construction for the ANOVA (476) is ambiguous (4 categories vs 2×2). Expect band, not byte-exact.
  • The 3 FC≥2 sex genes (PRKY/TMSB4Y/KI67) are the cleanest checkable result (Y-linked + proliferation marker = robust biology).
Figures / tables: Fig 3Fig 4Table
C1
Reported
476 genes (ANOVA+Tukey p<0.05)
Reproduced
C2
Reported
532 genes (sex t-test p<=0.05)
Reproduced
C3
Reported
3 FC>=2 genes: PRKY down, TMSB4Y down, KI67 up
Reproduced
C4
Reported
232 genes (arsenic t-test p<0.05)
Reproduced
C5
Reported
29 samples (16M/13F; 15 low/14 high)
Reproduced

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 63/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

51.8 k
tokens (I/O) · 2.2 M incl. cache
12 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.