A urine extracellular vesicle lncRNA classifier for high-grade prostate cancer and increased risk of progression: A multi-center study.
Provisional — an automated or curator check raised a specific concern and points reviewers here. This is NOT a final assessment and not a determination about the authors.
The main results reproduced, with only marginal, non-material deviations.
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
DESCRIBED WELL ENOUGH? Partially. The classifier is fully specified (explicit LASSO->logistic formula with coefficients and 3 named lncRNAs), but most of its inputs are NOT public. CODE: the paper states verbatim 'This paper does not report original code'; the harvested github.com/mdbrown/rmda is a generic third-party decision-curve-analysis package (P16), not the analysis pipeline. DATA: the central pipeline-derived results (discovery DE of 1,681 lncRNAs; training/val1/val2 AUCs 0.756/0.776/0.761; DCA; cutpoint metrics) rest on a urine RT-qPCR multi-center cohort (n=350/232/251) that was never deposited, and on a lncRNA RNA-seq matrix that was never deposited. KEY AUDITABLE FINDING (possible data-availability flag): GSE147761, the only accession cited for 'RNA sequencing data reported in this paper', actually contains a circRNA matrix (the authors' 2021 sister paper's data) with NONE of this paper's classifier lncRNAs -> C1 is not reproducible because the wrong RNA species was deposited. WHAT WAS REPRODUCED (fully public inputs, end-to-end on «our HPC»): C2 = the TCGA validation AUC. Applying the published fixed-weight formula to TCGA-PRAD STAR-Counts FPKM (GDC API; UCSC Xena bulk files 302-redirect to an S3 bucket that «infra» egress blocks) over all 554 samples with Gleason-grade labels gives AUC 0.695 (log2, CI 0.625-0.760) / 0.672 (raw, CI 0.602-0.737). The reported 0.733 falls INSIDE our log2 bootstrap CI; the negative class n=98 matches the paper exactly and high-grade n=456 vs 453. The point estimate is ~0.04-0.06 lower, consistent with the unstated FPKM-vs-qPCR normalization (coefficients were fit on 2^-dCt). C3 = each marker independently carries high-grade signal (univariate AUC 0.633-0.709), strongest marker = largest coefficient. VERDICT: partial reproduction with one within-tolerance validation point (C2), one confirmed supporting result (C3), and one strong negative data-availability finding (C1). NOT ATTEMPTED: qPCR-cohort training/val AUCs, DCA net-benefit, clinical-utility cutpoint (inputs not deposited); discovery DE (lncRNA matrix not deposited).
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 65assessed: 2026-06-16 ⛓ fc59f4c3af6c
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-16
- Rubric version
- not recorded
- Assessed by
- —
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetThe study tests whether a urine extracellular vesicle lncRNA classifier can detect high-grade prostate cancer (grade group ≥2) and predict risk of progression during active surveillance, with performance superior to PSA-derived tools like PCA3, mpMRI, and standard risk calculators.
- ★ A 3-lncRNA urine extracellular vesicle classifier (Clnc: AC015987.1, CTD-2589M5.4, RP11-363E6.3) detects high-grade PCa with higher accuracy than PCA3, mpMRI, PCPT-RC 2.0, and ERSPC-RC finding
- ★ Clnc at diagnosis is an independent predictor of overall active surveillance progression after adjustment for clinicopathological factors finding
- ★ Six urine extracellular vesicle lncRNAs (ENSG00000224746, ENSG00000233255, ENSG00000254027, ENSG00000225489, ENSG00000255007 upregulated; ENSG00000245025 downregulated) discriminate high-grade PCa from benign/healthy tissue finding
- ★ Urine extracellular vesicle lncRNA levels correlate with matched cancer tissue lncRNA levels finding
- LASSO-based logistic regression was used to select 3 of 6 candidate lncRNAs and build the Clnc diagnostic model method
- Combining Clnc with ERSPC-RC or PCPT-RC improves predictive performance over the risk calculators alone finding
- Random urine samples show high consistency with first-morning urine for Clnc detection finding
- ENSG00000255007 expression is exclusively increased in PCa tumor tissue relative to 33 other cancer/tissue types in TCGA finding
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| RNA sequencing | urinary extracellular vesicles, human patients | none (case-control: high-grade PCa vs benign prostatic hyperplasia) | differential lncRNA expression | — |
| lncRNA expression analysis (public dataset) | TCGA PCa tumor tissue vs normal adjacent tissue; GEO healthy blood leukocytes | none | differentially expressed lncRNAs | — |
| qRT-PCR | paired human high-grade PCa tissue, NAT, and matched urine extracellular vesicles (n=30 each) | none | expression levels of 6 candidate lncRNAs | — |
| qRT-PCR + logistic regression/LASSO modeling | urine extracellular vesicles from training cohort (n=350: 139 high-grade PCa, 211 controls) | none | Clnc classifier score, AUC | — |
| qRT-PCR classifier validation | urine extracellular vesicles, validation cohort 1 (n=232) and validation cohort 2 (n=251) | none | AUC, NPV, biopsy avoidance rate | — |
| lncRNA expression analysis (TCGA) | TCGA PCa tissue (n=499) vs healthy/GG1 tissue | none | AUC of Clnc vs PCA3 | — |
| Spearman correlation, decision curve analysis, calibration analysis | matched urine EV and tumor tissue, multiple cohorts | none | correlation coefficients, net benefit, calibration agreement | — |
| Cox proportional hazards regression, cumulative incidence analysis | prospective active surveillance cohort (n=182) | risk stratification by Clnc cut point (0.3657) | hazard ratio and cumulative incidence of AS progression (overall, PSA, clinical, histopathologic) | — |
- – Clnc AUC in training cohort 0.756 (95% CI, 0.706-0.806)
- – Clnc AUC in validation cohort 1 0.776 (95% CI, 0.713-0.838)
- – Clnc AUC in validation cohort 2 0.761 (95% CI, 0.699-0.822)
- ▲ Clnc AUC in TCGA cohort exceeds PCA3 AUC Clnc 0.733 vs PCA3 0.655, p=0.0046
- – Clnc at 0.3657 cutoff avoids unnecessary biopsies with negative predictive value NPV 70.69%-80.15%; avoided 41.83%-48.28% of biopsies across cohorts
- ▲ Clnc at diagnosis independently predicts overall AS progression in multivariable Cox model HR 2.10 (95% CI, 1.16-3.81), p=0.0146
- ▲ 2-year cumulative incidence of overall AS progression higher in high-risk vs low-risk Clnc group 38.12% (95% CI 24.67-51.57) vs 19.42% (95% CI 11.27-27.57)
- – Five of six candidate lncRNAs upregulated and one downregulated in high-grade PCa tissue/urine vs NAT
- other AUC 0.756 (95% CI, 0.706-0.806) (Clnc performance, training cohort)
- other AUC 0.776 (95% CI, 0.713-0.838) (Clnc performance, validation cohort 1)
- other AUC 0.761 (95% CI, 0.699-0.822) (Clnc performance, validation cohort 2)
- other AUC 0.733 (0.699-0.774) (Clnc performance, TCGA cohort)
- other HR 2.10 (95% CI, 1.16-3.81), p=0.0146 (Clnc as independent predictor of overall AS progression)
- pvalue p < 0.0001 (Median Clnc differs between negative/GG1 and ≥GG2 PCa in training and validation cohorts)
- count NPV 70.69%; avoided 46.86% of all biopsies or 77.73% of unnecessary biopsies (Clnc cut point 0.3657 biopsy avoidance, training cohort)
- other 2-year cumulative incidence: 19.42% (low-risk) vs 38.12% (high-risk) (Overall AS progression by Clnc risk group)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
The paper used differential expression analysis of RNA-seq data to discover candidate lncRNAs, then applied LASSO-penalized logistic regression in a training cohort (n=350) to build a 3-lncRNA classifier (C_lnc) for high-grade prostate cancer detection. Classifier performance was evaluated by AUC comparison against clinical benchmarks across three external validation cohorts and one TCGA cohort. In a prospective active surveillance cohort (n=182), Kaplan-Meier cumulative incidence curves with log-rank tests and multivariable Cox proportional hazards regression assessed the classifier's prognostic value. Results were reported with 95% confidence intervals for AUCs and hazard ratios.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Differential expression analysis (specific method not named; applied to RNA-seq count data) | Discovery of lncRNAs distinguishing high-grade PCa from BPH, healthy leukocytes, and normal adjacent tissue (Figures 1, S1, Table S1) | 11 high-grade PCa vs 11 BPH (urine EVs); 499 PCa tissues vs 32 normal blood samples; 52 matched PCa vs NAT pairs | not stated |
| LASSO-penalized logistic regression (feature selection and classifier construction) | Selection of 3 lncRNAs from 6 candidates and derivation of C_lnc score in training cohort (Figure S4) | 350 (139 high-grade PCa, 211 controls) | not stated |
| AUC comparison (method for pairwise AUC testing not named; likely DeLong test based on reported p-values) | Comparison of C_lnc vs PCPT-RC 2.0, ERSPC-RC, mpMRI, and PCA3 in training and validation cohorts (Table 2, Figure 3) | 350, 232, 251, and 499 across cohorts | not stated |
| Spearman correlation | Correlation of lncRNA levels between matched urine extracellular vesicles and cancer tissue (Figure 2B, Figure S3) | 30 matched pairs | not stated |
| One-way ANOVA with Dunnett's post-hoc test | Comparison of ENSG00000255007 expression across PCa tissue, matched urine EVs, and NAT (Figure 2D) | 30 per group | not stated |
| Mann-Whitney U test and chi-squared test | Comparison of continuous and categorical baseline characteristics between C_lnc high-risk and low-risk groups in AS cohort (Table 3) | 182 (119 low-risk, 63 high-risk) | stated |
| Log-rank test (on cumulative incidence curves) | Overall AS progression, PSA progression, clinical progression, and histopathologic progression by C_lnc risk group (Figure 4A, Figure S12) | 182 | not stated |
| Multivariable Cox proportional hazards regression | Independent predictors of overall AS progression; C_lnc HR reported after adjustment for age, PSA density, cT stage, percentage positive biopsies, and maximum tumor involvement (Table S3) | 182 | not stated |
-
Multiple pairwise AUC comparisons were made across biomarkers and cohorts without a multiplicity correction↳ Could also: Apply Benjamini-Hochberg FDR correction or Bonferroni adjustment across the family of AUC comparisons — When many simultaneous comparisons are performed, a multiplicity correction controls the expected proportion of false discoveries; this is commonly expected in diagnostic accuracy studies comparing numerous biomarkers
-
The differential expression method used in the discovery phase is described only as 'differential analysis' without naming the algorithm or software↳ Could also: Specify the tool (e.g., DESeq2, edgeR, limma-voom) along with normalization strategy and filtering thresholds — Naming the differential expression method allows readers to assess the distributional assumptions made (e.g., negative binomial vs. empirical Bayes) and facilitates independent reproduction of the discovery step
-
LASSO logistic regression was used for simultaneous feature selection and classifier construction within the training cohort↳ Could also: Use k-fold cross-validation or bootstrap resampling to estimate optimism-corrected AUC internally, or apply elastic-net regression as a complementary approach — Internal validation with cross-validation or bootstrapping yields a bias-corrected performance estimate from the training data, helping distinguish optimistic in-sample AUC from likely out-of-sample performance before external validation
-
A single fixed cut-point (0.3657) derived from the training cohort was applied to all validation cohorts as the primary decision threshold↳ Could also: Report full operating characteristics across a range of thresholds, and/or use decision curve analysis as the primary utility summary with the fixed cut-point as a secondary illustration — A single training-derived cut-point may perform differently in populations with different disease prevalence; the net benefit across thresholds (decision curve analysis, which the authors did perform) gives a richer picture of clinical utility independent of prevalence
-
Kaplan-Meier curves and log-rank tests were used to compare cumulative incidence of AS progression between risk groups↳ Could also: Apply competing-risks regression (e.g., Fine-Gray subdistribution hazard model) if treatment initiation or death could preclude observation of the progression endpoint — In an AS cohort, patients may exit surveillance due to definitive treatment or death before a progression event is observed; competing-risks methods account for these competing events and yield cumulative incidence estimates that do not assume independence of risks
-
Continuous variables in the AS cohort table were summarized with median and IQR only↳ Could also: Also report 95% confidence intervals for between-group differences, or include mean and SD alongside median and IQR — IQR effectively characterizes skewed distributions, and the addition of a 95% CI for the difference or effect size (e.g., rank-biserial correlation for Mann-Whitney) would give readers a direct quantitative sense of the magnitude of between-group differences
Result convergence & founder nodes
Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.
-
C_lnc classifier detects high-grade prostate cancer (≥GG2) with AUC 0.776 and 0.761 in two independent validation cohortsqPCR human urine extracellular vesicle 2023×1papers★ This paper is the founder (earliest)
-
High C_lnc score at diagnosis independently predicts active surveillance progression in prostate cancer (HR 2.10, 95% CI 1.16–3.81, p=0.0146) in multivariable Cox modelqPCR human urine extracellular vesicle up 2023×1papers★ This paper is the founder (earliest)
-
Three lncRNAs (ENSG00000224746, ENSG00000255007, ENSG00000254027) are significantly upregulated in high-grade prostate cancer versus biopsy-negative/GG1 controls (p<0.0001)qPCR human urine extracellular vesicle up 2023×1papers★ This paper is the founder (earliest)
-
C_lnc detects ≥GG2 prostate cancer with AUC 0.733 in TCGA cohort, outperforming PCA3 (AUC 0.655, p=0.0046)RNA-seq human prostate 2023×1papers★ This paper is the founder (earliest)
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
scope.md — pmid-37852185
Paper: Tao W et al. A urine extracellular vesicle lncRNA classifier for high-grade prostate cancer and increased risk of progression: A multi-center study. Cell Rep Med 2023. PMID 37852185 · PMCID PMC10591064 · DOI 10.1016/j.xcrm.2023.101240
Code / data availability as printed
- Code: "This paper does not report original code." (Resource availability). The
text-mined GitHub link
mdbrown/rmdais a generic third-party decision-curve-analysis R package (last push 2018) — used by the paper only to draw DCA plots. P16 third-party case; not the analysis pipeline of the classifier. - Data: "RNA sequencing data (GSE147761) ... deposited in GEO." Also GSE100206 (normal blood), and TCGA via Xena.
Pipeline-derived results & in/out of scope
| Result | Pipeline | Public data? | Scope |
|---|---|---|---|
| 3-lncRNA LASSO+logistic classifier (Clnc); coeffs given explicitly | LASSO selection → logistic regression | training/validation cohorts = RT-qPCR, NOT deposited | OUT (data not public) |
| AUC training 0.756 / val1 0.776 / val2 0.761 (Table 2) | logistic score → ROC | qPCR cohorts not deposited | OUT (data not public) |
| AUC TCGA 0.733 (95%CI 0.699-0.774), ≥GG2 n=453 vs GG1/normal n=98 (Table 2) | published logistic formula applied to TCGA-PRAD FPKM | TCGA-PRAD public (GDC) | IN → C2 |
| 1,681 DE urine-EV lncRNAs (DESeq2, | log2FC | >1, p<0.05) in discovery GSE147761 | DESeq2 on 11 HG-PCa vs 11 BPH |
| Decision curve analysis (net benefit vs PCA3/mpMRI/PCPT/ERSPC) | rmda DCA | qPCR cohorts not deposited | OUT (data not public) |
| Clinical-utility cutpoint 0.3657 (NPV 70.69%, 46.86% biopsies avoided) | thresholding on training | qPCR not deposited | OUT |
KEY DATA-AVAILABILITY FINDING (auditable, possible-fabrication-relevant)
GEO GSE147761, cited as "RNA sequencing data reported in this paper," contains a
circRNA quantification matrix (GSE147761_count.txt.gz: header col circRNA, 2230
hsa_circ_* rows × 22 samples). That accession actually underpins the authors' sister
2021 paper (urine-EV circRNA classifier, PMC8299620). The lncRNA expression on which
this paper's 3-lncRNA classifier is built — including the three classifier features
ENSG00000224746, ENSG00000255007, ENSG00000254027 — is absent from the deposited data.
Consequence: the discovery DE result (1,681 lncRNAs) and the qPCR-cohort AUCs are not
reproducible from any deposited/public data. Only the TCGA validation point (C2) uses
fully public data and the explicitly-published formula.
What we attempt
- C2 (primary): apply the paper's explicit logistic formula
logit = -1.152 + 6.541·ENSG254027 + 9.950·ENSG255007 + 15.268·ENSG224746to TCGA-PRAD FPKM (GDC STAR-Countsfpkm_unstranded), label high-grade = tumor Gleason sum ≥7 (≥GG2) vs normal+Gleason6, compute ROC-AUC + bootstrap CI; compare to 0.733. - C3 (supporting): univariate AUC of each of the 3 lncRNAs in TCGA (scale-invariant signal check).
Not attempted (and why)
- Training/val1/val2 AUCs, DCA, cutpoint metrics — underlying urine qPCR cohorts not deposited.
- Discovery 1,681-lncRNA DE — lncRNA matrix not in GSE147761 (only circRNA deposited).
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.