Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Extensive androgen receptor enhancer heterogeneity in primary prostate cancers underlies transcriptional diversity and metastatic potential.

Nat Commun · 2022
L1 70/100 3/4
Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +4
✓ What held up
  • Same input data as the authors
  • Reported values were directly comparable
  • No relevant deviation in data/preprocessing
  • No authors-side cause for any deviation
  • Any deviation was negligible
What did not (or only partly)
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
70/100
Reproducibility score
at the mean
vs. all fields · 1173 studies
🎯 Scores higher than 37% of all assessed papers rank 732 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

DESCRIBED WELL ENOUGH + 1:1 reproducible for the low-hanging outputs. The arbshet repo (commit ca64a97) ships its own input data inside the repo for several R-Markdown notebooks, so those analyses reproduce WITHOUT downloading GSE217319 (valid per P16). PRIMARY result C1: the authors' Cistrome-GO regulatory-potential notebook hard-codes the gene counts above RP>0.05 as plot annotations (geom_vline 1026 for case met-ARBS, 524 for control); recomputing these from the shipped supplementary table 211115_cistrome_supp.txt via the authors' R read path on «our HPC» (R 4.5.3, SLURM «job») gives EXACTLY 1026 and 524 -> exact 1:1, code internally consistent with shipped data, no fabrication signal. Secondary descriptive reproductions also matched: GSEA Hallmark gene-set counts per group (top5p=5, un=10, mcase=10, mctrl=2) and seqPos distinct TF-motif counts (case=61, control=41). NOT ATTEMPTED (80/20 drop): the genome-wide GLM headline (2026 ARBS regulating 1901 genes) because its notebook needs non-shipped inputs (AR27ac_overlap.txt, 210616_ARBSloop_genes.txt HiChIP loop->gene map, and per-patient AR peaklist .bed files = full GSE217319 + upstream peak processing); also the CNA/dependency/STARR-seq/mutation analyses (external or wet-lab data, out of scope). CAVEAT for the human auditor: C1's 1026/524 were pinned to the authors' published CODE annotation, not located verbatim in the paper main text; confirm against Fig 5 / GO-RP methods.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 70
    assessed: 2026-06-15 ⛓ 35b4e5f7ded9
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-15
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

This study investigates the extent and biological/clinical consequences of inter-tumor heterogeneity in androgen receptor (AR) chromatin binding (enhancer usage) across primary prostate cancers, testing whether such epigenetic heterogeneity drives transcriptional diversity and metastatic potential.

Core claims
  • AR enhancer/chromatin binding usage is highly heterogeneous between primary prostate tumors, with <5% of all AR binding sites shared by half of tumors analyzed. finding
  • AR enhancer heterogeneity is patient-intrinsic rather than tumor-intrinsic, as primary tumor and normal prostate epithelium show strikingly similar heterogeneous ARBS rankings. finding
  • Somatic mutations (SNVs) and germline cQTL SNPs converge on/enrich in commonly-shared AR sites (SH-ARBS) in primary and metastatic tissues. finding
  • Less-frequently shared AR sites (PS/UN-ARBS) associate strongly with AR-driven enhancer activity and gene expression and can drive oncogenic processes. finding
  • Heterogeneous AR enhancer usage in primary disease distinguishes patient outcome and is informative for risk of biochemical relapse. finding
  • Ranked ARBS exhibit hierarchical enhancer activity (inducible, constitutive, inactive) measured via massively parallel reporter assays. finding
  • Enhancer-specific copy number alterations at heterogeneous ARBS interact with promoters of essential PCa genes and drive transcriptional output in metastatic disease. mechanism
  • Extensive QC analyses (NSC, RSC, FRiP, GC%, MSPC, read-depth correlations) confirm observed AR heterogeneity is biological, not technical artifact. method
Experimental setups
Assay System Perturbation Readout Platform
AR ChIP-seq 88 primary prostate cancer patient tumor tissues (mean tumor cell %>80%) none AR binding sites (ARBS) / peaks ranked by prevalence across tumors
AR ChIP-seq 15 normal prostate epithelium tissues none ARBS ranking (n=27,850/27,500)
AR ChIP-seq AR+ PCa cell lines (LNCaP, VCaP, 22Rv1, LNCaP BR bicalutamide-resistant, 42D ENZR enzalutamide-resistant, LSHAR) and controls THP-1 monocytic, MDA-MB-453 breast cancer treatment sensitive/resistant; AR-transduced normal prostate (LSHAR) ARBS overlap with tumor ranked ARBS
RNA-seq primary PCa tissues / metastatic PCa patient cohort none / copy-number altered vs neutral gene expression; log2 fold expression change over copy-number neutral samples
Massively parallel reporter assay (STARR-seq) LNCaP cells vehicle (EtOH) vs DHT enhancer activity classification (inducible/constitutive/inactive) of ARBS
Luciferase reporter assay cell line (subset of STARR-seq library regions) hormone (DHT) vs vehicle enhancer activity status and hormonal dependency validation
H3K27ac Hi-ChIP / 3D-genome (ChIA-PET integration) PCa tissue / VCaP none enhancer-promoter chromatin interactions/loops
Somatic mutation / SNV & CNA / structural variation analysis 200 primary PCa tumors and 101 metastatic PCa tumors none SNV, CNA gains/losses, SVs at ARBS
Key results
  • Fewer than 5% of all AR binding sites are shared by half of the tumors analyzed, indicating high enhancer heterogeneity <5%
  • Primary tumor and normal epithelium ranked ARBS follow strikingly similar distributions, supporting patient-intrinsic heterogeneity
  • Commonly shared ARBS (SH-ARBS) enriched in PCa cell lines and AR-transduced LSHAR, with further shift toward SH-ARBS enrichment upon therapy resistance; overlap absent in THP-1 and MDA-MB-453 controls
  • Germline cQTL SNPs and somatic SNVs in primary and metastatic PCa are enriched in primary SH-ARBS
  • Only 52 of 764 unique risk SNPs overlapped primary ranked ARBS, with no enrichment at particular ARBS 52/764
  • Most-commonly shared ARBS enriched for hormone-dependent (inducible) enhancer activity (n=286) relative to constitutive (n=463) or inactive (n=2467) sites
  • 23 of 101 metastatic PCa patients had copy number gains exclusively at the AR enhancer locus (PS/UN-ARBS) 23/101
  • 25 essential PCa genes interact with CNA-affected ARBS in cell line dependency (DepMap) analysis 25 genes
Key statistics
  • count 69,330 AR binding sites (total ARBS universe across 88 primary prostate tumors)
  • count 7394 peaks per tumor (average) (mean AR ChIP-seq peaks per tumor, FRiP >1.5)
  • count 27,850 (also reported 27,500) (ARBS identified in normal prostate epithelium)
  • pvalue p<0.0001 (****) (hypergeometric test enrichment of SH-ARBS in cell lines; t-test vs LSHAR)
  • count 7422 ARBS (expanded to 20,790 via machine learning; +2495 additional) (ARBS tested for enhancer potential by STARR-seq)
  • count 764 risk SNPs; 4454 cQTL SNPs (q<0.05) (rSNPs and germline allelic imbalance SNPs analyzed at ARBS)
  • count 278,209 primary SNVs; 1,048,576 metastatic SNVs (single nucleotide variations in primary (n=200) and metastatic (n=101) PCa)
  • count 1201 SH-ARBS (number of shared ARBS where SNVs/cQTLs enriched)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The study characterized androgen receptor (AR) chromatin binding heterogeneity across 88 primary prostate tumor tissues via ChIP-seq and integrated multiple external genomic datasets (mutation, enhancer activity, copy number, 3D-genome) to assess functional and clinical consequences of inter-tumor ARBS heterogeneity. Primary statistical comparisons employed two-tailed Student's t-tests for continuous distributions across ARBS categories and cell lines, hypergeometric tests for category enrichment, and two-tailed Fisher's exact tests for mutation rate comparisons. Results were reported as threshold p-values (p < 0.05, p < 0.01, p < 0.0001) with distributions visualized as boxplots using median and interquartile range.

Replicationbiological Sample size88 primary prostate tumor tissues (mean tumor cell percentage >80%) and 15 normal prostate epithelium for ChIP-seq; 200 primary PCa and 101 metastatic PCa for somatic mutation data drawn from external published cohorts; no formal power calculation stated GroupsSH-ARBS vs. PS-ARBS vs. UN-ARBS; tumor vs. normal epithelium; PCa cell lines vs. patient tumors; primary vs. metastatic somatic mutation enrichment Pairingunpaired Randomization/blindingnot stated DispersionIQR Exact p-valuesno Effect sizesno Confidence intervalsno Multiplicity correctionMultiple testing correction applied within MSPC peak-calling QC (specific method not stated); q < 0.05 threshold cited for external cQTL dataset; no correction stated for the primary t-tests, hypergeometric tests, or Fisher's exact tests
Statistical tests used
Test Applied to n Assumptions
two-tailed Student's t-test Comparison of ARBS presence in cell lines across ranked ARBS (Fig. 1c, means compared to LSHAR cells); comparison of STARR-seq enhancer activity distributions across ARBS rank categories (Fig. 2b) 69,330 ARBS across 88 tumors (Fig. 1c); STARR-seq subsets: 286 inducible, 463 constitutive, 2467 inactive ARBS (Fig. 2b) not stated
hypergeometric test Enrichment of SH-ARBS in PCa cell lines (Fig. 1c); enrichment of SH-ARBS at super-enhancer genomic locations (Fig. 2e) SH-ARBS n = 1201; total ARBS n = 69,330 not stated
two-tailed Fisher's exact test Observed vs. expected background primary or metastatic somatic mutation rate across ARBS rank categories (Fig. 2d); applied on untransformed values Primary PCa SNV data from n = 200 patients; metastatic PCa SNV data from n = 101 patients (external cohorts) not stated
multi-sample peak calling with multiple testing correction (MSPC) Quality control validation confirming heterogeneous AR binding sites as true positives across 88 ChIP-seq samples 88 primary PCa ChIP-seq samples not stated
Approaches that could also have been used
  • Multiple independent two-tailed t-tests were used to compare ARBS category distributions across several cell lines and conditions without a stated correction for multiple comparisons
    Could also: One-way ANOVA followed by a post-hoc test (e.g., Tukey HSD) for normally distributed data, or Kruskal-Wallis with Dunn's post-hoc for non-normal distributions, with a single family-wise correction — An omnibus test with post-hoc correction explicitly controls the family-wise error rate when comparing more than two groups simultaneously, which is an additional safeguard when many pairwise comparisons are conducted across figures
  • Student's t-test was applied to compare AR binding site distributions without a stated assessment of normality
    Could also: Mann-Whitney U (Wilcoxon rank-sum) test as a non-parametric alternative — ChIP-seq signal intensities and peak-prevalence distributions are often right-skewed; a rank-based test does not require normality and can be more appropriate when distributional assumptions are not formally verified
  • P-values were reported as threshold categories (p < 0.05, p < 0.01, p < 0.0001) rather than exact values
    Could also: Reporting exact p-values (e.g., p = 0.0023) for each comparison — Exact p-values allow readers to apply their own thresholds, facilitate downstream meta-analyses, and provide more complete evidence summaries; they are increasingly recommended by statistical reporting guidelines
  • Enrichment of ARBS categories in genomic feature sets (super-enhancers, mutation hotspots) was assessed with hypergeometric and Fisher's exact tests assuming a fixed genomic background
    Could also: Permutation-based enrichment testing or bootstrap resampling to define the null distribution empirically — Permutation tests make fewer assumptions about the background null distribution and can better account for the non-uniform genomic distribution of ChIP-seq peaks and local mutation rates along chromosomes
  • Boxplots summarized dispersion using the interquartile range across ARBS categories
    Could also: Violin plots or strip plots overlaying individual data points, or supplementing IQR with 95% confidence intervals for group medians or means — Showing the full distributional shape alongside summary statistics — particularly when group sizes vary considerably across ARBS categories — allows readers to assess skewness, modality, and the density of observations more completely
  • No formal effect-size estimates (e.g., Cohen's d, odds ratios, fold-enrichment with uncertainty) are reported alongside the significance tests
    Could also: Reporting effect sizes with 95% confidence intervals for key comparisons — Effect-size estimates with confidence intervals convey the magnitude and precision of differences independently of sample size, complementing p-values and aiding interpretation of biological relevance
Software: GIGGLE (TF ChIP-seq database overlap tool) · MSPC (Multi-Sample Peak Calling) · DepMap (cancer gene dependency repository)

Result convergence & founder nodes

Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
23
Impact: medium
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

rs710886 RefSNP in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-36450752

Paper: Kneppers et al. 2022, Extensive androgen receptor enhancer heterogeneity in primary prostate cancers underlies transcriptional diversity and metastatic potential. Nat Commun 13:7367. DOI 10.1038/s41467-022-35135-2. Code: https://github.com/jknp/arbshet @ commit ca64a974f1312ee77243992d5cc190ae2589b819 (2022-09-05). Data accession: GEO GSE217319 (AR / H3K27ac ChIP-seq, RNA-seq; built on prior cohort Stelloo et al. 2018, GSE120738).

The repo is a set of R-Markdown analysis notebooks. Crucially, several notebooks ship their own input data inside the repo (.txt/.csv), so those analyses are reproducible from the repo alone, without downloading GSE217319. Per brief rule P16, running the authors' own shipped pipeline on the authors' own shipped data is a valid 1:1 reproduction.

In scope (self-contained → attempted)

# Notebook Shipped input(s) Reproducible output Why in scope
C1 plot/GO_RP/CistromeGO_RPplot.Rmd 211115_cistrome_supp.txt (55,993 rows) # genes with Cistrome-GO regulatory-potential (RP) score > 0.05, case vs control ARBS. The code hard-codes these counts as plot annotations: geom_vline(xintercept = 1026) (case) and geom_vline(xintercept = 524) (control). Input fully shipped; the annotated counts are exactly recomputable from the shipped column RPscore + type. Clean 1:1 self-consistency / reproducibility check.
C2 plot/GSEA/bubble_plot.Rmd 210908_hallmarks_combi_RP0.05.txt (27 rows) # MSigDB Hallmark gene sets per comparison group (top5p, un, mcase, mctrl) and # passing the displayed significance cutoff (vline at −log10(FDRq·p) = 2). Input fully shipped; descriptive counts recomputable. Secondary.
C3 plot/seqpos/seqpos.Rmd 210909_seqpos8_combi.txt (131 rows) # distinct enriched TF motifs (seqPos) for case vs control met-ARBS (wordcloud inputs). Input fully shipped; descriptive counts recomputable. Secondary.

Out of scope (the hard ~20% — NOT attempted, with reason)

Notebook Reported result Why dropped
GLM/ARBS_GLM_genomewide.Rmd "2,026 ARBS regulating 1,901 unique genes" significant ARBS↔expression associations Inputs not fully shipped: needs AR27ac_overlap.txt, 210616_ARBSloop_genes.txt (H3K27ac HiChIP loop→gene map) and the per-patient AR peaklist .bed files (= GSE217319). Requires full GEO download + upstream peak processing. Heavy + under-specified → 80/20 drop.
CNA/mCNA_ARBSintegration.Rmd ARBS × copy-number integration Needs external per-patient copy-number .bed files (/mSV/Quigley/CN/*) not in repo (mCRPC WGS cohort, controlled).
plot/dep/Dependency.Rmd, plot/tal_line.Rmd, functions/Liftover.Rmd dependency / TAL line / liftover panels Mix of external DepMap/Achilles inputs and figure cosmetics; not single pinnable pipeline numbers.
Wet-lab: STARR-seq enhancer activity (286/463/2467), mutation enrichment Experimental / external pipelines, not in this repo. Out of scope by definition.

Primary claim to grade

C1 — case RP>0.05 = 1026, control RP>0.05 = 524 — recomputed from the shipped supplementary table via the authors' R read path on «our HPC».

Figures / tables: Fig 5Fig 4
C1a
Reported
1026
Reproduced
1026
exact
C1b
Reported
524
Reproduced
524
exact
C2
Reported
top5p=5,un=10,mcase=10,mctrl=2
Reproduced
top5p=5,un=10,mcase=10,mctrl=2
partial
C3
Reported
case=61,control=41
Reproduced
case=61,control=41
partial
GLM
Reported
2026 ARBS / 1901 genes
Reproduced
not-attempted
partial

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 70/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟢2. Endpoint comparability
🟢3. Location of the main deviation
🟢4. Cause of the deviation
🟡5. Derivability / plausibility
🟢6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +4

The reproducible low-hanging outputs are 1:1 exact with no fabrication signal: RP>0.05 recomputes to exactly 1026 (case) / 524 (control) from the shipped supplementary table, and the GSEA (top5p=5/un=10/mcase=10/mctrl=2) and seqPos (61/41) descriptive counts also match. The limitation is on the data-availability / our-scope side, not the authors': the central genome-wide GLM headline (2026 ARBS / 1901 genes) was deliberately not attempted because its inputs are not shipped in the repo (require full GSE217319 + peak/loop processing). C1 is also a self-consistency match against the authors' code geom_vline annotation rather than a number pinned in the paper text. Overall a solid partial reproduction — exact where checked, but the central claim remains untested.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

104.5 k
tokens (I/O) · 6.8 M incl. cache
12 min
runtime · 0.01 CPU-h
1.9 GB
peak RAM
1
HPC jobs
hummel
machine