Epigenetic loss of heterogeneity from low to high grade localized prostate tumours.
The main results reproduced: recomputed values matched the published ones within tolerance.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Same input data as the authors
- ✓Reported values are derivable from the shared data
- ✓Any deviation was negligible
- ✓The central claim held under reproduction
- 🟡Reported values were only indirectly comparable
- 🟡A deviation arose in the data or preprocessing
- 🟡A deviation was attributed to the published material
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Described well enough to reproduce the RESULT (not the authors' script). The repo (AlexChitsazan/ProstateTumorATACCode) is two research R-Markdown notebooks with hard-coded local paths, un-shipped .RDS intermediates and deprecated snapATAC v1 -> not directly re-runnable. But GEO GSE171559 ships the processed binned 5kb matrix + metaData, so the pipeline output was reproduced from deposited data with snapATAC2 2.9.0 (the maintained Python successor; P16-valid third-party tool) on «our HPC» SLURM. 1:1 on cohort integrity: 18 samples (exact), 5 Gleason scores (exact), 14,242 cells vs reported 14,251 clustered (within-tol, diff 9). Clustering: recovered 16 clusters (cluster COUNT exact, resolution-tuned; identity moderate-overlap ARI 0.20 / NMI 0.44 vs deposited labels - different engine than the paper's LDA-Gibbs, so not cell-for-cell identical). HEADLINE finding (the paper's title) reproduced clearly and with large effect on two independent metrics: low-grade (Gleason-3) cells mix across many patients (entropy 2.43 bits, per-patient silhouette ~0) while high-grade (Gleason-4) tumours are patient-specific (entropy 0.14 bits, silhouette +0.23) = 'epigenetic loss of heterogeneity from low to high grade'. No fabrication signal. NOT attempted (80/20): raw-read->.snap preprocessing, exact 30-topic LDA Gibbs model + topic-GO, MACS2 differential-accessibility + TF-motif enrichment (FOXA1/HOXB13/CDX2), Cicero co-accessibility, and the wet-lab cyclic-IF NRXN1/NLGN1 imaging (non-pipeline).
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 79assessed: 2026-06-15 ⛓ 8678ac80d63c
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-15
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusWhether single-cell chromatin accessibility profiling can identify molecular and cellular markers distinguishing low-grade (primary Gleason pattern 3) from high-grade (primary Gleason pattern 4) localized prostate tumours and reveal the heterogeneity underlying disease progression.
- ★ Shared chromatin accessibility features among low-grade (Gleason pattern 3) prostate cancer cells are lost in high-grade (Gleason pattern 4) tumours. finding
- ★ Despite loss of shared chromatin features, high-grade tumours are enriched for FOXA1, HOXB13 and CDX2 transcription factor binding sites, indicating a shared trans-regulatory programme. finding
- ★ Two neuronal adhesion molecule genes, NRXN1 and NLGN1 (plus CDH9), are highly accessible in high-grade prostate tumours. finding
- ★ NRXN1 and NLGN1 are expressed in epithelial, endothelial, immune and neuronal cells in prostate cancer as shown by cyclic immunofluorescence. finding
- ★ Single-cell ATAC-seq (combinatorial indexing / sci-ATAC-seq) on flash-frozen prostate tumours captures unbiased chromatin accessibility landscapes including immune and stromal cell types. method
- ★ Gleason pattern 3 tumours share chromatin accessibility constraints and form a single cluster without patient-specific clustering, whereas Gleason pattern 4 tumours form distinct outer clusters. finding
- Cicero co-accessibility analysis reveals an increase in predicted cis-regulatory interactions around the NRXN1 locus in Gleason pattern 4 tumours. mechanism
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| single-cell ATAC-seq (sci-ATAC-seq, combinatorial indexing) | flash-frozen primary human prostate tumours from 18 radical prostatectomy patients | none (Gleason pattern 3 vs 4 comparison) | chromatin accessibility / accessible chromatin peaks per single cell | 96-well plate combinatorial indexing with transposase |
| bulk ATAC-seq (reference comparison) | prostate adenocarcinoma (PRAD) TCGA datasets | none | peak distribution across functional genomic elements | — |
| cyclic immunofluorescence (multiplex imaging) | FFPE prostate tumour tissue sections from patient cohort | none | NRXN1 and NLGN1 protein expression across epithelial, endothelial, immune and neuronal cells | — |
| H&E histopathology | FFPE prostate tumour tissue sections | none | Gleason grade / tumour morphology | — |
- – 14,424 single cells with high-quality sci-ATAC-seq reads recovered from 18 primary prostate cancer samples 14,424 cells
- – 125,569 peaks called from aggregated Gleason pattern 3 and ≥4 tumours 125,569 peaks
- ▲ Higher chromatin accessibility to SCHLAP1 lncRNA locus in Gleason pattern ≥4 vs pattern 3 tumours
- ▲ Genomic regions associated with neuronal adhesion genes NRXN1, NLGN1 and CDH9 are more accessible in Gleason pattern 4 vs 3 tumours; 15 peaks linked to these three genes 15 peaks
- ▲ Increased number of Cicero predicted cis-regulatory interactions around NRXN1 locus in Gleason pattern 4 tumours despite fewer pattern 4 cells
- ▲ Higher accessibility to MYC promoter region in Gleason pattern 4 tumours
- – cisTopic identified 16 cell clusters spanning 30 topics; clusters 7 (stromal), 12 (lymphoid), 14 (myeloid) attracted cells from all samples and were removed from downstream epithelial analysis 16 clusters / 30 topics
- – Gleason pattern 4 cells (5,334) fewer than Gleason pattern 3 cells (7,383) yet show greater NRXN1 co-accessibility 5,334 vs 7,383 cells
- count 14,424 single-cells (high-quality sci-ATAC-seq cells from 18 primary prostate tumours)
- count 125,569 peaks (peaks called from aggregated Gleason pattern 3 and ≥4 tumours)
- count 16 clusters; 30 Topics (cisTopic clustering from 14,251 cells)
- count 5,334 cells (Gleason pattern 4 cells analyzed)
- count 7,383 cells (Gleason pattern 3 cells analyzed)
- count 15 peaks linked to three neuronal adhesion genes (differentially accessible regions linked to NRXN1, NLGN1, CDH9)
- count 6 low-risk, 9 intermediate-risk, 2 high-risk patients (cohort risk stratification by CAPRA-S/MSKCC nomogram (note: sums to 17 as written))
- count 70% Gleason pattern 3 / 30% Gleason pattern 4 cells (one intermediate-risk Gleason 3+4 tumour composition)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
The paper applied combinatorial-indexing single-cell ATAC-seq (sci-ATAC-seq) to 14,424 cells from 18 primary prostate tumours, using Latent Dirichlet Allocation (LDA)-based topic modelling (snapATAC/cisTopic) and UMAP for dimensionality reduction and cell clustering into 16 clusters across 30 topics. Differential chromatin accessibility between Gleason pattern 3 and pattern 4 cells was assessed using snapATAC's differential accessibility functions. Cluster quality was evaluated with silhouette analysis, putative cis-regulatory interactions were inferred with Cicero co-accessibility scoring, and genomic region and GO term enrichment were performed with GREAT.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| snapATAC differential accessibility (underlying statistical test not specified in text) | Gleason pattern 3 vs. Gleason pattern 4 cells across all accessible chromatin regions | 7,383 Gleason pattern 3 cells vs. 5,334 Gleason pattern 4 cells | not stated |
| Latent Dirichlet Allocation (LDA) topic modelling | Cell clustering into 16 clusters across 30 topics from all single cells | 14,424 cells from 18 tumours (14,251 stated in Fig. 2 caption) | na |
| GREAT genomic region enrichment / GO term enrichment (binomial and hypergeometric tests performed internally by GREAT) | GO annotation of accessible regions enriched in Gleason pattern 4 tumours; immune and stromal cluster annotation | — | not stated |
| Silhouette analysis (silhouette score) | Assessment of cluster coherence per patient sample (Gleason 3+3 vs. 4+4) and aggregated Gleason pattern groups | Per-sample cell counts listed in Fig. 3f; range 86–3,757 cells per sample across 18 samples | na |
| Cicero co-accessibility scoring (Pearson correlation-based; qualitative link count comparison, noted as 'not quantitative' in text) | Putative cis-regulatory interactions around NRXN1, NLGN1, and CDH9 loci in Gleason pattern 3 vs. 4 | 5,334 Gleason pattern 4 cells; 7,383 Gleason pattern 3 cells | not stated |
-
Differential chromatin accessibility between Gleason pattern groups was tested with snapATAC's built-in function, whose underlying statistical model is not detailed in the text↳ Could also: Pseudobulk approaches — aggregating per-cell counts to per-patient counts then applying DESeq2 or edgeR Wald/likelihood-ratio tests — could also be used; alternatively, ArchR or Signac/Seurat offer explicitly named tests (Wilcoxon rank-sum, logistic regression, negative binomial) — Pseudobulk methods account for the non-independence of cells nested within patients, a recognised consideration in single-cell differential analyses; naming the underlying test also aids reproducibility and cross-study comparison
-
Putative cis-regulatory interactions were compared between Gleason patterns by qualitatively counting Cicero links, acknowledged in the text as 'not quantitative'↳ Could also: A formal statistical comparison of co-accessibility scores between groups — e.g. Wilcoxon rank-sum test on per-link scores or a permutation test — could also be applied — A test statistic and p-value would allow readers to assess the strength of evidence for differential regulatory connectivity independently of the unequal group sizes (fewer pattern 4 cells), which the authors themselves flag as a limitation
-
GO term enrichment was performed through GREAT, which applies its own internal binomial and hypergeometric tests↳ Could also: Direct hypergeometric or Fisher's exact tests on peaks overlapping curated gene sets (e.g. via clusterProfiler, fgsea, or g:Profiler) with an explicit FDR correction (e.g. Benjamini-Hochberg) could also be applied — Reporting the specific enrichment test, the multiple-testing correction method, and adjusted p-values makes the evidence directly comparable to enrichment analyses in other studies and allows readers to judge significance thresholds independently
-
Results across 125,569 peaks and multiple GO categories are described with significance language, but no multiple-testing correction method is stated↳ Could also: Benjamini-Hochberg FDR control or Bonferroni correction applied to the full family of differential accessibility tests could also be explicitly reported — With a peak universe of this size, even a modest false-positive rate per test translates to many spurious hits; stating the correction method and the FDR threshold used lets readers calibrate the expected number of false discoveries
-
Cluster quality was assessed using silhouette scores displayed as box plots per sample↳ Could also: Adjusted Rand index against known Gleason grade labels, Davies-Bouldin index, or bootstrap-based cluster stability metrics (e.g. clusterboot) could also characterise cluster reproducibility — Complementary metrics capture different aspects of cluster quality — silhouette scores measure intra-vs-inter-cluster cohesion, while label-aware metrics and stability analyses assess biological concordance and robustness to data subsampling
-
Key findings (differential accessibility, co-accessibility link counts) are reported without numeric effect sizes or confidence intervals↳ Could also: Log2 fold-changes in accessibility with 95% confidence intervals, or standardised effect sizes (e.g. Cohen's d on topic scores), could also accompany significance statements — Effect sizes and confidence intervals convey the magnitude and precision of observed differences independently of sample size, enabling meta-analytic reuse and helping readers distinguish statistical from biological significance
Result convergence & founder nodes
Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.
-
Chromatin accessibility at the MYC promoter region is higher in Gleason pattern 4 vs pattern 3 prostate tumours.ATAC-seq human prostate tumor up 2021×1papers★ This paper is the founder (earliest)
-
Chromatin accessibility at the NRXN1, NLGN1, and CDH9 neuronal adhesion gene loci is increased in Gleason pattern 4 vs pattern 3 prostate tumours.ATAC-seq human prostate tumor up 2021×1papers★ This paper is the founder (earliest)
-
Chromatin accessibility at the SCHLAP1 lncRNA locus is higher in Gleason pattern 4 vs pattern 3 prostate tumours.ATAC-seq human prostate tumor up 2021×1papers★ This paper is the founder (earliest)
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
Assessed papers, coloured by verdict. Click a node to open it.
- No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-34911933
Paper: Eksi et al. 2021, Nat Commun 12:7424. "Epigenetic loss of heterogeneity
from low to high grade localized prostate tumours." DOI 10.1038/s41467-021-27615-8.
Code: https://github.com/AlexChitsazan/ProstateTumorATACCode (HEAD be706d5, pushed 2021-11-03; no license, no README content).
Data: GEO GSE171559 (super-series) — single-cell (sci-)ATAC-seq.
What the paper is / pipeline
sci-ATAC-seq (combinatorial indexing) of 14,424 single cells from 18 fresh-frozen
localized prostate tumours. Reads → snapATAC (v1) preprocessing → 5 kb genome
bins (bmat) → binarize → LDA dimensionality reduction (cisTopic-style, "runLDA",
30 topics) → UMAP → Louvain clustering → 16 cell clusters. Clusters annotated as
epithelial Gleason-3 / Gleason-4 / immune / stromal. Headline finding: low-grade
(Gleason pattern 3) cells from different patients mix into one shared cluster,
whereas high-grade (Gleason pattern 4) tumours form distinct patient-specific outer
clusters → "loss of (inter-tumour shared) chromatin heterogeneity" from G3→G4.
Repo reality check
The repo is two research R-Markdown notebooks (snapATAC.Rmd, 2769 lines;
snapATAC_ImmuneStroma.Rmd, 548 lines), not a runnable pipeline:
- Hard-coded local paths (
«path»,«path» Sync/...). - Inputs are per-sample
.snapfiles (built off-repo from raw reads with snaptools) and pre-computed.RDSintermediates that are NOT shipped (ImmuneStromaOnly.RDS,selected.ModelImmuneOnly.RDS,AllSamplesDecember.lda.sp). The expensiverunLDAcall is commented out and replaced byreadRDS(...). snapATACv1 is deprecated/unmaintained and hard to install today. → The authors' code is not directly re-runnable (docs_insufficient at the script level). BUT GEO ships the processed matrices, so the result is reproducible from the deposited data with the maintained successor tool (P16: a third-party tool on the paper's own data is equally valid).
GEO GSE171559 deposited (processed) files — the reproduction inputs
| file | size (gz) | content |
|---|---|---|
…barcodes.txt.gz |
57 KB | 14,242 cell barcodes |
…binned.features.txt.gz |
2.9 MB | 5 kb genome bins (the bmat features) |
…binned.values.mtx.gz |
87 MB | binned cell×bin matrix (bmat) ← clustering input |
…metaData.txt.gz |
522 KB | per-cell annotations (14,242 cells, 14 cols) |
…peak.features.txt.gz / …peak.values.mtx.gz |
2.9 / 37 MB | MACS2 peak matrix (DAR analysis) |
metaData columns: barcode, TN, UM, PP, UQ, CM, MTRatio, PromoterRatio, Sample, GleasonScore, SampleNamePaper, EnrichedCluster, SoftGleasonScore, IncreasedEnrichedResolution. Note: the deposited table does NOT carry the
numeric 16-cluster id; it carries broad labels EnrichedCluster ∈
{G3, G4, Immune, Stroma} and a finer IncreasedEnrichedResolution (8 levels).
In scope (pipeline-derived, attempted)
- R1 cohort size — samples: 18 tumours/patients. (direct: count unique
Sample) - R2 cohort size — cells: 14,424 recovered / 14,251 clustered. (direct: barcode/row count)
- R3 Gleason categories: 5 scores {3+3, 3+4, 4+3, 4+4, 4+5}. (direct: count unique
GleasonScore) - R4 clustering: 16 clusters from the binarized 5 kb bmat. (rerun: spectral/LSI +
Leiden on the deposited bmat with snapATAC2 = maintained snapATAC successor; report
recovered cluster count + ARI vs deposited
EnrichedCluster/IncreasedEnrichedResolution). - R5 loss of heterogeneity (headline): G3 cells cross-patient mix (one dominant
shared cluster, high patient-entropy) while G4 cells form distinct patient-specific
clusters (low patient-entropy). (rerun: per-cluster patient-mixing entropy G3 vs G4;
- Fig-3f-style per-patient pseudobulk silhouette G3 vs G4.)
Out of scope (not attempted, why)
- Raw reads →
.snap(snaptools) preprocessing: needs SRA fastqs + barcode pipeline; the deposited bmat already encodes this step. (80/20: skip) - Exact LDA 30-to
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
Reproduction is solid and fabrication-free: from the authors' own deposited GSE171559 matrices we recover the cohort exactly (18 samples, 5 Gleason scores), the 16-cluster scale, and — most importantly — the paper's headline loss of heterogeneity with a large effect on two independent metrics (entropy G3=2.43 vs G4=0.14 bits; silhouette -0.01 vs +0.23). The only deviations are on our side / methodological: a different clustering engine (snapATAC2 vs the paper's LDA-Gibbs, hence moderate ARI 0.20/NMI 0.44), a resolution-tuned cluster count, and a benign 9-cell (0.06%) deposition gap. Severity is negligible and the core claim holds fully, so overall yellow only because it is not a 1:1 engine-identical rerun and several downstream analyses were intentionally out of scope.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.