Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Stromal androgen signaling acts as tumor niches to drive prostatic basal epithelial progenitor-initiated oncogenesis.

Nat Commun · 2022
L1 63/100 PQI 88
⚑ Flagged for review — a reproduced result did not match the reported value

Provisional — an automated or curator check raised a specific concern and points reviewers here. This is NOT a final assessment and not a determination about the authors.

Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
✓ What held up
  • The central claim held under reproduction
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🟡Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🔴A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
63/100
Reproducibility score
0.6 SD below mean
vs. all fields · 1173 studies
🎯 Scores higher than 24% of all assessed papers rank 875 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough for a faithful PARTIAL 1:1 reproduction of the pipeline-derived scRNA numbers, working from the paper's own deposited data via a third-party-style direct analysis (P16). BRIEF ACCESSION WRONG: GSE197780 is a different paper (PMID 35754340); correct = GSE174471 (repo README + GEO). Deposited data is a single LOG-NORMALIZED combined matrix (19,541 genes x 16,383 cells); raw 10X/UMI counts not deposited, so the raw QC step and full two-sample Seurat integration could not be re-run from scratch. Split by barcode suffix (.1=Ctrl, plain=ARKO). EXACT matches: Ctrl cell count 10,810; Ctrl mean genes/cell 2,950; ARKO mean genes/cell 1,342. MISMATCH: ARKO cell count deposited 5,573 vs reported 6,848 (1,275 fewer) -- flagged as a discrepancy needing human review (likely cells dropped after the quoted count or a reporting error; Ctrl + both genes/cell values reproduce exactly so the matrix is otherwise faithful). HEADLINE FINDING CONFIRMED: IGFBP3 robustly up in AR-deficient stroma -- whole-sample log2FC +2.54 and, after marker-gating to fibroblasts to remove the composition confound, log2FC +1.94 (36%->75% expressing, p=1.5e-64). Supporting: Igf1 up, Wnt targets (Ccnd1/Cd44/Ctnnb1) down, Ar down, Sp1 down in ARKO -- all consistent with the paper's IGFBP3-IGF1-Wnt mechanism. NOT ATTEMPTED: avg UMI/cell (no raw counts), full integration/cluster annotation (Fig 3 8-cluster / 10-epithelial; raw data not deposited + Matrix/SeuratObject ABI clash; hard 20%), GSEA, pseudotime, all wet-lab figs.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 63
    assessed: 2026-06-15 ⛓ c1e28d0c6c18
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-15
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

The study tests whether stromal androgen receptor (AR) signaling in sonic-hedgehog-responsive Gli1-lineage cells acts as a tumor niche to drive prostatic basal epithelial progenitor-initiated oncogenesis and tumor development.

Core claims
  • Loss of AR in stromal Gli1-lineage cells diminishes prostate epithelial oncogenesis and tumor development. finding
  • AR-deficient Gli1-lineage stroma shows robustly increased IGFBP3 expression via attenuation of AR suppression on Sp1-regulated transcription. mechanism
  • Increased IGFBP3 inhibits IGF1-induced Wnt/β-catenin activation in adjacent basal epithelial cells, repressing oncogenic growth. mechanism
  • Stromal AR deletion reduces atypical Myc+ basal (progenitor/tumor-initiating) cells in PIN lesions. finding
  • Epithelial organoids from stromal AR-deficient mice can regain IGF1-induced oncogenic growth. finding
  • Co-targeting reciprocal stromal/epithelial AR and IGF1 signaling is implicated as a therapeutic strategy for advanced prostate cancer. resource
  • Loss of human prostate tumor basal cell signatures is found in basal cells of stromal AR-deficient mice. finding
Experimental setups
Assay System Perturbation Readout Platform
In vivo tissue recombination + kidney capsule transplantation UGM from Ar^L/Y:Gli1^CreER/+ (ARKO) or Gli1^CreER/+ control + UGE from Ctnnb1^L(ex3)/+:PB^Cre4 (or Pten^L/L:PB^Cre4) embryos, grown in NOD/SCID host mice Stromal Ar KO (tamoxifen-induced) with oncogenic epithelium (stabilized β-catenin or Pten loss) Graft weight/size, pathology (mPIN grade), β-catenin/AR IHC, Ki67+ percentage
Genetic mouse model pathology/IHC Hi-Myc:Ar^L/Y:Gli1^CreER/+ vs Hi-Myc:Gli1^CreER/+ mice (2- and 6-month-old) Stromal Ar deletion with epithelial hMyc transgene Prostate weight/body weight ratio, mPIN/adenocarcinoma grade, Myc/AR/Ki67 IHC
Single-cell RNA sequencing (scRNA-seq) Prostate tissue of 3-month-old Hi-Myc:R26^mTmG/+:Gli1^CreER/+ (HiMyc) and Hi-Myc:R26^mTmG/+:Ar^L/Y:Gli1^CreER/+ (HiMyc-ARKO) mice Stromal Ar KO Cell cluster identities, gene expression (hMycTg, Krt8, Krt14, Ar, mGFP), % hMycTg+ basal/luminal cells, DEGs
Gene set enrichment analysis (GSEA) hMycTg+ basal cells from HiMyc vs HiMyc-ARKO scRNA-seq Stromal Ar KO Enriched signaling pathways (IGF1R, Wnt)
qRT-PCR Sorted prostatic basal epithelial cells from HiMyc vs HiMyc-ARKO mice Stromal Ar KO Expression of hMycTg, Igf1r, Hras, Grb2, Irs2, Akt1, Ctnnb1, Cd44, Tcf7l2, Ccnd1
Immunohistochemistry / co-immunofluorescence Prostate tissue sections of HiMyc and HiMyc-ARKO mice Stromal Ar KO Myc, AR, CK8, CK14, IGF1R, pIGF1R, CyclinD1 staining and colocalization
Prostate organoid assay Organoids derived from mouse basal epithelia of stromal AR-deficient mice IGF1 stimulation ± IGFBP3 incubation IGF1-induced oncogenic organoid growth
Key results
  • ARKO UGM grafts were significantly lighter than control UGM grafts p=0.012
  • Fewer Ki67+ proliferating cells in ARKO UGM grafts p=0.001
  • Reduced prostate weight/body weight ratio in Hi-Myc:Ar^L/Y:Gli1^CreER/+ vs control at 2 and 6 months p=3.97×10^-8 (2mo); p=0.003 (6mo)
  • Lower percentage of hMycTg+ basal cells in HiMyc-ARKO vs HiMyc p<0.01
  • Increased Igf1r downstream genes (Hras, Irs2, Akt1) and Wnt targets (Ctnnb1, Ccnd1, Cd44, Tcf7l2) in hMycTg+ basal cells of HiMyc vs HiMyc-ARKO
  • Positive correlation between Igf1r and downstream/Wnt target genes in HiMyc basal cells, absent in HiMyc-ARKO Spearman r=0.2-0.4
  • Impaired oncogenic growth/pathology with Pten^L/L UGE recombined with AR-deficient UGM
  • IGFBP3 co-incubation with IGF1 attenuates IGF1-induced oncogenic organoid growth
Key statistics
  • correlation r=0.4 (p=1.42×10^-63) (Spearman correlation Igf1r vs Ctnnb1 in HiMyc basal cells)
  • correlation r=0.3 (p=7.15×10^-35) (Spearman correlation Igf1r vs Irs2 in HiMyc basal cells)
  • correlation r=0.3 (p=5.14×10^-42) (Spearman correlation Igf1r vs Tcf7l2 in HiMyc basal cells)
  • pvalue p=3.97×10^-8 (Prostate/body weight ratio reduction in 2-month-old Hi-Myc:Ar^L/Y:Gli1^CreER/+ mice)
  • pvalue p=0.012 (ARKO UGM graft weight lower than control UGM grafts)
  • pvalue p=0.001 (Reduced Ki67+ cell percentage in ARKO UGM grafts)
  • count 10,810 cells (HiMyc); 6,849 cells (HiMyc-ARKO) (scRNA-seq cell numbers per genotype)
  • count 8 cell subsets; 10 epithelial clusters (3 basal, 6 luminal, 1 OE) (scRNA-seq cluster identification)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The paper used two-sided Student's t-tests as the primary inferential method for comparing continuous quantitative outcomes (xenograft weights, prostate/body weight ratios, Ki67+ cell percentages) between two genotype groups with n = 5–8 biological replicates per group. Single-cell RNA-seq data from two pooled mouse samples were processed via UMAP-based clustering and GSEA on differentially expressed genes, and Spearman correlations were reported for seven co-expression pairs within a basal epithelial cell subset. Results were summarized as mean ± SD with exact p-values stated to be available in a Source Data file.

Replicationbiological Sample sizeBiological replicate counts stated per experiment (n = 5 or n = 8 mice per group); no formal power analysis described in the provided text GroupsAR-deficient Gli1-lineage stromal cells (ARKO) vs. Gli1-lineage wild-type controls, in a Hi-Myc prostate oncogenesis background Pairingunpaired Randomization/blindingnot stated DispersionSD Exact p-valuesyes Effect sizesno Confidence intervalsno Multiplicity correctionnone stated
Statistical tests used
Test Applied to n Assumptions
Two-sided Student's t-test Xenograft recombinant weights and Ki67+ cell percentages (Fig. 1b) n = 5 biological replicates per group not stated
Two-sided Student's t-test Prostate weight/body weight ratios at 2 and 6 months (Fig. 2c) n = 5 (2-month) and n = 8 (6-month) biological replicates per group not stated
Proportion comparison (test not named; p < 0.01 stated) Percentage of hMycTg+ cells per total basal or luminal cells by genotype (Fig. 3f) null not stated
Gene Set Enrichment Analysis (GSEA) Differentially expressed genes in hMycTg+ basal cells between HiMyc and HiMyc-ARKO scRNA-seq samples (Fig. 4a) 10,810 cells (HiMyc) and 6,849 cells (HiMyc-ARKO); 2 biological samples total not stated
Spearman correlation Co-expression of Igf1r with Hras, Irs2, Akt1, Ctnnb1, Ccnd1, Cd44, and Tcf7l2 in hMycTg+ basal epithelial cells (Fig. 4c) individual cells from scRNA-seq; exact cell-level n for this subset not stated not stated
Approaches that could also have been used
  • Multiple pairwise two-sided Student's t-tests were used across several figures without a stated family-wise error correction
    Could also: One-way ANOVA followed by a post-hoc correction (e.g., Tukey HSD or Bonferroni) could also be applied when several related outcomes are tested within the same experiment — ANOVA with post-hoc adjustment formally accounts for the inflation of type I error across a family of comparisons, which is standard practice when multiple endpoints are derived from the same experimental cohort
  • Seven Spearman correlations between Igf1r and co-expression partners were each reported with individual p-values, without multiple-testing adjustment
    Could also: Benjamini-Hochberg FDR correction could also be applied across the set of simultaneous correlation tests — Adjusting for the number of parallel correlation tests reduces the expected proportion of false discoveries among the reported associations
  • Continuous outcomes were analyzed with Student's t-tests in groups as small as n = 5 per group
    Could also: A non-parametric rank-based test (Mann-Whitney U / Wilcoxon rank-sum) could also be used for these small-n comparisons — With n = 5 per group, normality cannot be reliably verified; a rank-based test makes no distributional assumption and is frequently preferred for small biological replicate sets in preclinical mouse studies
  • Results are plotted as mean ± SD bar graphs
    Could also: Individual data points (dot plots or strip charts) overlaid on summary statistics could also be shown, as recommended by many journals for small-n studies — Displaying all individual values alongside mean ± SD makes the full distribution visible to readers and is increasingly requested by journals to improve transparency when n is small
  • scRNA-seq differential expression was computed at the single-cell level to generate gene lists for GSEA, but the specific statistical model (e.g., Wilcoxon, MAST, or DESeq2) was not stated in the provided text
    Could also: Pseudo-bulk approaches (e.g., DESeq2 or edgeR applied to per-animal aggregated counts) could also be used when biological replicates exist — Pseudo-bulk methods aggregate cells per biological replicate before testing, which better accounts for within-animal cell-level correlations and is generally considered to provide better-calibrated type I error control than cell-level tests
  • No formal sample size justification or power calculation is reported in the provided text
    Could also: A prospective power analysis based on effect sizes from pilot data or prior publications could also accompany the study design description — Reporting a power calculation allows readers to contextualize whether the study was sized to detect the magnitude of differences observed, and is expected by many funding agencies and journals for in vivo mouse studies
Software: UMAP (uniform manifold approximation and projection) for scRNA-seq visualization · scRNA-seq processing pipeline (specific package, e.g., Seurat or Cell Ranger, not named in provided text)

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
15
Impact: medium
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

GSE197780 GEO in Discussion (http://purl.org/orb/Discussion)
no other assessed paper uses this yet

Downstream reach in the literature

6 downstream papers · 1 datasets

How widely the datasets deposited by this paper are reused across the whole literature (Europe PMC), beyond our assessed set. This is a factual dependency map — reusing a public dataset is normal, good science. It is not a judgement on the downstream papers; the only verdict here is this paper's own, with its cited rationale.

This paper is currently under reproducibility review (see the verdict above). The map below shows where the data in question has propagated — so reuse can be traced, not so the downstream work is presumed affected.

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-36323713

Paper: Hiroto et al., Stromal androgen signaling acts as tumor niches to drive prostatic basal epithelial progenitor-initiated oncogenesis. Nat Commun 2022. PMID 36323713 · PMCID PMC9630272 · DOI 10.1038/s41467-022-34282-w.

Code: https://github.com/wk-kim/HiMYC-ARKO-Gli1_Stromal_AR_Prostate_Tumorigenesis (single script Hiroto_et_al_NatComm_2022_Rcode.R, Seurat scRNA-seq pipeline).

Data accession correction: Brief lists GSE197780 — that is a different paper (enzalutamide/FOXA1, PMID 35754340). The repo README and GEO confirm the correct accession is GSE174471 (PMID 36323713 match; samples GSM5315063 HiMyc-ARKO, GSM5315064 HiMyc). Deposited data = one series-level file GSE174471_HiMyc_HiMyc-ARKO_scRNAseq_Counts.txt.gz (247 MB).

Deposited data characteristics (inspected)

  • Combined matrix: 19,541 genes × 16,383 cells, tab-delimited, genes as rows.
  • Values are log-normalized (Seurat NormalizeData), NOT raw UMI counts.
  • Cell barcodes: 10,810 carry .1 suffix (= HiMyc Ctrl, matches reported 10,810 exactly); 5,573 plain barcodes (= HiMyc-ARKO). .1 stems do not overlap plain barcodes → genuine sample label, not make.unique dedup.
  • Custom transgene features present: MYC-transgene, EGFP, plus Ar, Igfbp3.

IN SCOPE (pipeline-derived, attempted)

Pipeline = Seurat (v3.2.1 in paper) on 10X scRNA-seq.

  • S1 Cell counts after QC filtering per sample (Results: 10,810 Ctrl / 6,848 ARKO). Caveat: deposited matrix is post-QC normalized, so we verify the cell counts present, not re-run the raw QC (raw CellRanger matrices not deposited).
  • S2 Average genes/cell per sample (2,950 Ctrl / 1,342 ARKO) — nonzero genes per cell is preserved under normalization, so this is recomputable.
  • S3 IGFBP3 up in AR-deficient stroma — headline molecular finding; testable as IGFBP3 differential Ctrl vs ARKO (direction + significance) on deposited matrix.
  • S4 (optional / hard 20%) number of cell clusters: 8 major cell subsets (Fig 3b); 10 epithelial subclusters = 3 basal + 6 luminal + 1 OE (Fig 3d). Requires the two-sample integration workflow on raw counts — only approximable from the single deposited normalized matrix.

OUT OF SCOPE (not attempted)

  • All wet-lab / in vivo: tissue recombination xenografts, IHC (Ki67, Myc, β-catenin), mouse genetics, organoid assays, IGFBP3+IGF1 incubation (Figs 1, 2, 4–8 wet-lab).
  • Avg UMI/cell (15,736 Ctrl / 5,041 ARKO): needs raw counts — not deposited.
  • GSEA pathway analysis, Spearman IGF/Wnt correlation matrices, monocle/slingshot pseudotime: depend on the full integrated object / cell-type labels not deposited.

Reproduction approach

Combined normalized matrix → split Ctrl(.1) / ARKO(plain) → per-group cell count, mean genes/cell, and IGFBP3/Ar/EGFP differential (Mann-Whitney). Compute on «our HPC» SLURM («infra»), pull back small JSON. No raw-count re-QC and no full integration re-run (data not deposited for those).

Figures / tables: FigsFig 3bFig 3d
C1
Reported
10,810 (Ctrl cells after QC)
Reproduced
10810
exact
C2
Reported
2,950 (Ctrl genes/cell)
Reproduced
2949.7
exact
C3
Reported
6,848 (ARKO cells after QC)
Reproduced
5573
did not match
C4
Reported
1,342 (ARKO genes/cell)
Reproduced
1342.7
exact
C5
Reported
15,736 / 5,041 UMI/cell
Reproduced
not-reproducible (raw counts not deposited)
did not match
C6
Reported
IGFBP3 robustly increased in AR-deficient (ARKO) stroma
Reproduced
up in ARKO: whole-sample log2FC +2.54 (p~0); fibroblast-gated log2FC +1.94, 36%->75% expressing (p=1.5e-64)
within tolerance
C7
Reported
8 major cell subsets
Reproduced
not-completed (optional)
partial
C8
Reported
10 epithelial subclusters
Reproduced
not-attempted (hard 20%)
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 63/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟡2. Endpoint comparability
🟡3. Location of the main deviation
🔴4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟢7. Core claim
🟡8. Severity of the miss (overall human judgment)
🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

195 k
tokens (I/O) · 17.2 M incl. cache
22 min
runtime · 0.01 CPU-h
2.5 GB
peak RAM
2
HPC jobs
hummel
machine