Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

A urine extracellular vesicle lncRNA classifier for high-grade prostate cancer and increased risk of progression: A multi-center study.

Cell Rep Med · 2023
65/100 PQI 87
⚑ Flagged for review — a reproduced result did not match the reported value

Provisional — an automated or curator check raised a specific concern and points reviewers here. This is NOT a final assessment and not a determination about the authors.

Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
How its reproducibility compares
65/100
Reproducibility score
0.5 SD below mean
vs. all fields · 1173 studies
🎯 Scores higher than 27% of all assessed papers rank 843 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

DESCRIBED WELL ENOUGH? Partially. The classifier is fully specified (explicit LASSO->logistic formula with coefficients and 3 named lncRNAs), but most of its inputs are NOT public. CODE: the paper states verbatim 'This paper does not report original code'; the harvested github.com/mdbrown/rmda is a generic third-party decision-curve-analysis package (P16), not the analysis pipeline. DATA: the central pipeline-derived results (discovery DE of 1,681 lncRNAs; training/val1/val2 AUCs 0.756/0.776/0.761; DCA; cutpoint metrics) rest on a urine RT-qPCR multi-center cohort (n=350/232/251) that was never deposited, and on a lncRNA RNA-seq matrix that was never deposited. KEY AUDITABLE FINDING (possible data-availability flag): GSE147761, the only accession cited for 'RNA sequencing data reported in this paper', actually contains a circRNA matrix (the authors' 2021 sister paper's data) with NONE of this paper's classifier lncRNAs -> C1 is not reproducible because the wrong RNA species was deposited. WHAT WAS REPRODUCED (fully public inputs, end-to-end on «our HPC»): C2 = the TCGA validation AUC. Applying the published fixed-weight formula to TCGA-PRAD STAR-Counts FPKM (GDC API; UCSC Xena bulk files 302-redirect to an S3 bucket that «infra» egress blocks) over all 554 samples with Gleason-grade labels gives AUC 0.695 (log2, CI 0.625-0.760) / 0.672 (raw, CI 0.602-0.737). The reported 0.733 falls INSIDE our log2 bootstrap CI; the negative class n=98 matches the paper exactly and high-grade n=456 vs 453. The point estimate is ~0.04-0.06 lower, consistent with the unstated FPKM-vs-qPCR normalization (coefficients were fit on 2^-dCt). C3 = each marker independently carries high-grade signal (univariate AUC 0.633-0.709), strongest marker = largest coefficient. VERDICT: partial reproduction with one within-tolerance validation point (C2), one confirmed supporting result (C3), and one strong negative data-availability finding (C1). NOT ATTEMPTED: qPCR-cohort training/val AUCs, DCA net-benefit, clinical-utility cutpoint (inputs not deposited); discovery DE (lncRNA matrix not deposited).

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 65
    assessed: 2026-06-16 ⛓ fc59f4c3af6c
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-16
Rubric version
not recorded
Assessed by
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

The study tests whether a urine extracellular vesicle lncRNA classifier can detect high-grade prostate cancer (grade group ≥2) and predict risk of progression during active surveillance, with performance superior to PSA-derived tools like PCA3, mpMRI, and standard risk calculators.

Core claims
  • A 3-lncRNA urine extracellular vesicle classifier (Clnc: AC015987.1, CTD-2589M5.4, RP11-363E6.3) detects high-grade PCa with higher accuracy than PCA3, mpMRI, PCPT-RC 2.0, and ERSPC-RC finding
  • Clnc at diagnosis is an independent predictor of overall active surveillance progression after adjustment for clinicopathological factors finding
  • Six urine extracellular vesicle lncRNAs (ENSG00000224746, ENSG00000233255, ENSG00000254027, ENSG00000225489, ENSG00000255007 upregulated; ENSG00000245025 downregulated) discriminate high-grade PCa from benign/healthy tissue finding
  • Urine extracellular vesicle lncRNA levels correlate with matched cancer tissue lncRNA levels finding
  • LASSO-based logistic regression was used to select 3 of 6 candidate lncRNAs and build the Clnc diagnostic model method
  • Combining Clnc with ERSPC-RC or PCPT-RC improves predictive performance over the risk calculators alone finding
  • Random urine samples show high consistency with first-morning urine for Clnc detection finding
  • ENSG00000255007 expression is exclusively increased in PCa tumor tissue relative to 33 other cancer/tissue types in TCGA finding
Experimental setups
Assay System Perturbation Readout Platform
RNA sequencing urinary extracellular vesicles, human patients none (case-control: high-grade PCa vs benign prostatic hyperplasia) differential lncRNA expression
lncRNA expression analysis (public dataset) TCGA PCa tumor tissue vs normal adjacent tissue; GEO healthy blood leukocytes none differentially expressed lncRNAs
qRT-PCR paired human high-grade PCa tissue, NAT, and matched urine extracellular vesicles (n=30 each) none expression levels of 6 candidate lncRNAs
qRT-PCR + logistic regression/LASSO modeling urine extracellular vesicles from training cohort (n=350: 139 high-grade PCa, 211 controls) none Clnc classifier score, AUC
qRT-PCR classifier validation urine extracellular vesicles, validation cohort 1 (n=232) and validation cohort 2 (n=251) none AUC, NPV, biopsy avoidance rate
lncRNA expression analysis (TCGA) TCGA PCa tissue (n=499) vs healthy/GG1 tissue none AUC of Clnc vs PCA3
Spearman correlation, decision curve analysis, calibration analysis matched urine EV and tumor tissue, multiple cohorts none correlation coefficients, net benefit, calibration agreement
Cox proportional hazards regression, cumulative incidence analysis prospective active surveillance cohort (n=182) risk stratification by Clnc cut point (0.3657) hazard ratio and cumulative incidence of AS progression (overall, PSA, clinical, histopathologic)
Key results
  • Clnc AUC in training cohort 0.756 (95% CI, 0.706-0.806)
  • Clnc AUC in validation cohort 1 0.776 (95% CI, 0.713-0.838)
  • Clnc AUC in validation cohort 2 0.761 (95% CI, 0.699-0.822)
  • Clnc AUC in TCGA cohort exceeds PCA3 AUC Clnc 0.733 vs PCA3 0.655, p=0.0046
  • Clnc at 0.3657 cutoff avoids unnecessary biopsies with negative predictive value NPV 70.69%-80.15%; avoided 41.83%-48.28% of biopsies across cohorts
  • Clnc at diagnosis independently predicts overall AS progression in multivariable Cox model HR 2.10 (95% CI, 1.16-3.81), p=0.0146
  • 2-year cumulative incidence of overall AS progression higher in high-risk vs low-risk Clnc group 38.12% (95% CI 24.67-51.57) vs 19.42% (95% CI 11.27-27.57)
  • Five of six candidate lncRNAs upregulated and one downregulated in high-grade PCa tissue/urine vs NAT
Key statistics
  • other AUC 0.756 (95% CI, 0.706-0.806) (Clnc performance, training cohort)
  • other AUC 0.776 (95% CI, 0.713-0.838) (Clnc performance, validation cohort 1)
  • other AUC 0.761 (95% CI, 0.699-0.822) (Clnc performance, validation cohort 2)
  • other AUC 0.733 (0.699-0.774) (Clnc performance, TCGA cohort)
  • other HR 2.10 (95% CI, 1.16-3.81), p=0.0146 (Clnc as independent predictor of overall AS progression)
  • pvalue p < 0.0001 (Median Clnc differs between negative/GG1 and ≥GG2 PCa in training and validation cohorts)
  • count NPV 70.69%; avoided 46.86% of all biopsies or 77.73% of unnecessary biopsies (Clnc cut point 0.3657 biopsy avoidance, training cohort)
  • other 2-year cumulative incidence: 19.42% (low-risk) vs 38.12% (high-risk) (Overall AS progression by Clnc risk group)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The paper used differential expression analysis of RNA-seq data to discover candidate lncRNAs, then applied LASSO-penalized logistic regression in a training cohort (n=350) to build a 3-lncRNA classifier (C_lnc) for high-grade prostate cancer detection. Classifier performance was evaluated by AUC comparison against clinical benchmarks across three external validation cohorts and one TCGA cohort. In a prospective active surveillance cohort (n=182), Kaplan-Meier cumulative incidence curves with log-rank tests and multivariable Cox proportional hazards regression assessed the classifier's prognostic value. Results were reported with 95% confidence intervals for AUCs and hazard ratios.

Replicationbiological Sample sizeSample sizes stated per cohort; no formal a priori power calculation described GroupsHigh-grade PCa (≥GG2) vs. benign/GG1 PCa or BPH in diagnostic cohorts; C_lnc high-risk vs. low-risk in prospective AS cohort Pairingmixed Randomization/blindingnot stated DispersionIQR Exact p-valuesyes Effect sizesyes Confidence intervalsyes Multiplicity correctionnone stated
Statistical tests used
Test Applied to n Assumptions
Differential expression analysis (specific method not named; applied to RNA-seq count data) Discovery of lncRNAs distinguishing high-grade PCa from BPH, healthy leukocytes, and normal adjacent tissue (Figures 1, S1, Table S1) 11 high-grade PCa vs 11 BPH (urine EVs); 499 PCa tissues vs 32 normal blood samples; 52 matched PCa vs NAT pairs not stated
LASSO-penalized logistic regression (feature selection and classifier construction) Selection of 3 lncRNAs from 6 candidates and derivation of C_lnc score in training cohort (Figure S4) 350 (139 high-grade PCa, 211 controls) not stated
AUC comparison (method for pairwise AUC testing not named; likely DeLong test based on reported p-values) Comparison of C_lnc vs PCPT-RC 2.0, ERSPC-RC, mpMRI, and PCA3 in training and validation cohorts (Table 2, Figure 3) 350, 232, 251, and 499 across cohorts not stated
Spearman correlation Correlation of lncRNA levels between matched urine extracellular vesicles and cancer tissue (Figure 2B, Figure S3) 30 matched pairs not stated
One-way ANOVA with Dunnett's post-hoc test Comparison of ENSG00000255007 expression across PCa tissue, matched urine EVs, and NAT (Figure 2D) 30 per group not stated
Mann-Whitney U test and chi-squared test Comparison of continuous and categorical baseline characteristics between C_lnc high-risk and low-risk groups in AS cohort (Table 3) 182 (119 low-risk, 63 high-risk) stated
Log-rank test (on cumulative incidence curves) Overall AS progression, PSA progression, clinical progression, and histopathologic progression by C_lnc risk group (Figure 4A, Figure S12) 182 not stated
Multivariable Cox proportional hazards regression Independent predictors of overall AS progression; C_lnc HR reported after adjustment for age, PSA density, cT stage, percentage positive biopsies, and maximum tumor involvement (Table S3) 182 not stated
Approaches that could also have been used
  • Multiple pairwise AUC comparisons were made across biomarkers and cohorts without a multiplicity correction
    Could also: Apply Benjamini-Hochberg FDR correction or Bonferroni adjustment across the family of AUC comparisons — When many simultaneous comparisons are performed, a multiplicity correction controls the expected proportion of false discoveries; this is commonly expected in diagnostic accuracy studies comparing numerous biomarkers
  • The differential expression method used in the discovery phase is described only as 'differential analysis' without naming the algorithm or software
    Could also: Specify the tool (e.g., DESeq2, edgeR, limma-voom) along with normalization strategy and filtering thresholds — Naming the differential expression method allows readers to assess the distributional assumptions made (e.g., negative binomial vs. empirical Bayes) and facilitates independent reproduction of the discovery step
  • LASSO logistic regression was used for simultaneous feature selection and classifier construction within the training cohort
    Could also: Use k-fold cross-validation or bootstrap resampling to estimate optimism-corrected AUC internally, or apply elastic-net regression as a complementary approach — Internal validation with cross-validation or bootstrapping yields a bias-corrected performance estimate from the training data, helping distinguish optimistic in-sample AUC from likely out-of-sample performance before external validation
  • A single fixed cut-point (0.3657) derived from the training cohort was applied to all validation cohorts as the primary decision threshold
    Could also: Report full operating characteristics across a range of thresholds, and/or use decision curve analysis as the primary utility summary with the fixed cut-point as a secondary illustration — A single training-derived cut-point may perform differently in populations with different disease prevalence; the net benefit across thresholds (decision curve analysis, which the authors did perform) gives a richer picture of clinical utility independent of prevalence
  • Kaplan-Meier curves and log-rank tests were used to compare cumulative incidence of AS progression between risk groups
    Could also: Apply competing-risks regression (e.g., Fine-Gray subdistribution hazard model) if treatment initiation or death could preclude observation of the progression endpoint — In an AS cohort, patients may exit surveillance due to definitive treatment or death before a progression event is observed; competing-risks methods account for these competing events and yield cumulative incidence estimates that do not assume independence of risks
  • Continuous variables in the AS cohort table were summarized with median and IQR only
    Could also: Also report 95% confidence intervals for between-group differences, or include mean and SD alongside median and IQR — IQR effectively characterizes skewed distributions, and the addition of a 95% CI for the difference or effect size (e.g., rank-biserial correlation for Mann-Whitney) would give readers a direct quantitative sense of the magnitude of between-group differences
Software: not stated

Result convergence & founder nodes

Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
17
Impact: medium
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

GSE100206 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE147761 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

scope.md — pmid-37852185

Paper: Tao W et al. A urine extracellular vesicle lncRNA classifier for high-grade prostate cancer and increased risk of progression: A multi-center study. Cell Rep Med 2023. PMID 37852185 · PMCID PMC10591064 · DOI 10.1016/j.xcrm.2023.101240

Code / data availability as printed

  • Code: "This paper does not report original code." (Resource availability). The text-mined GitHub link mdbrown/rmda is a generic third-party decision-curve-analysis R package (last push 2018) — used by the paper only to draw DCA plots. P16 third-party case; not the analysis pipeline of the classifier.
  • Data: "RNA sequencing data (GSE147761) ... deposited in GEO." Also GSE100206 (normal blood), and TCGA via Xena.

Pipeline-derived results & in/out of scope

Result Pipeline Public data? Scope
3-lncRNA LASSO+logistic classifier (Clnc); coeffs given explicitly LASSO selection → logistic regression training/validation cohorts = RT-qPCR, NOT deposited OUT (data not public)
AUC training 0.756 / val1 0.776 / val2 0.761 (Table 2) logistic score → ROC qPCR cohorts not deposited OUT (data not public)
AUC TCGA 0.733 (95%CI 0.699-0.774), ≥GG2 n=453 vs GG1/normal n=98 (Table 2) published logistic formula applied to TCGA-PRAD FPKM TCGA-PRAD public (GDC) IN → C2
1,681 DE urine-EV lncRNAs (DESeq2, log2FC >1, p<0.05) in discovery GSE147761 DESeq2 on 11 HG-PCa vs 11 BPH
Decision curve analysis (net benefit vs PCA3/mpMRI/PCPT/ERSPC) rmda DCA qPCR cohorts not deposited OUT (data not public)
Clinical-utility cutpoint 0.3657 (NPV 70.69%, 46.86% biopsies avoided) thresholding on training qPCR not deposited OUT

KEY DATA-AVAILABILITY FINDING (auditable, possible-fabrication-relevant)

GEO GSE147761, cited as "RNA sequencing data reported in this paper," contains a circRNA quantification matrix (GSE147761_count.txt.gz: header col circRNA, 2230 hsa_circ_* rows × 22 samples). That accession actually underpins the authors' sister 2021 paper (urine-EV circRNA classifier, PMC8299620). The lncRNA expression on which this paper's 3-lncRNA classifier is built — including the three classifier features ENSG00000224746, ENSG00000255007, ENSG00000254027 — is absent from the deposited data. Consequence: the discovery DE result (1,681 lncRNAs) and the qPCR-cohort AUCs are not reproducible from any deposited/public data. Only the TCGA validation point (C2) uses fully public data and the explicitly-published formula.

What we attempt

  • C2 (primary): apply the paper's explicit logistic formula logit = -1.152 + 6.541·ENSG254027 + 9.950·ENSG255007 + 15.268·ENSG224746 to TCGA-PRAD FPKM (GDC STAR-Counts fpkm_unstranded), label high-grade = tumor Gleason sum ≥7 (≥GG2) vs normal+Gleason6, compute ROC-AUC + bootstrap CI; compare to 0.733.
  • C3 (supporting): univariate AUC of each of the 3 lncRNAs in TCGA (scale-invariant signal check).

Not attempted (and why)

  • Training/val1/val2 AUCs, DCA, cutpoint metrics — underlying urine qPCR cohorts not deposited.
  • Discovery 1,681-lncRNA DE — lncRNA matrix not in GSE147761 (only circRNA deposited).
Figures / tables: Table
C1_discovery_DE_lncRNAs
Reported
1,681 differentially expressed urine-EV lncRNAs (DESeq2 |log2FC|>1 & p<0.05) between 11 high-grade PCa and 11 BPH (GSE147761)
Reproduced
NOT REPRODUCIBLE from public data: GSE147761 deposits only a circRNA matrix (header 'circRNA', 2,230 hsa_circ_* rows x 22 samples); the 3 classifier lncRNAs (ENSG00000224746/255007/254027) and any lncRNA-level matrix are absent. This accession underpins the authors' 2021 circRNA sister paper (PMC8299620).
did not match
C2_TCGA_classifier_AUC
Reported
TCGA AUC 0.733 (95% CI 0.699-0.774), high-grade >=GG2 (n=453) vs GG1/normal (n=98) [Table 2]
Reproduced
AUC 0.695 (95% CI 0.625-0.760) under log2(FPKM+1); AUC 0.672 (95% CI 0.602-0.737) under raw FPKM. Computed on 554 TCGA-PRAD STAR-Counts samples (456 high-grade vs 98 low-grade/normal; 554/554 files parsed, 0 download failures) via the GDC API and the paper's explicit logistic formula. The reported point estimate 0.733 lies INSIDE the log2 bootstrap 95% CI. Cohort negative-class n=98 matches the paper exactly; high-grade n=456 vs reported 453 (delta 3).
within tolerance
C3_TCGA_univariate_marker_AUC
Reported
supporting (no exact reported value)
Reproduced
Univariate ROC-AUC in TCGA-PRAD: ENSG00000224746=0.709, ENSG00000254027=0.650, ENSG00000255007=0.633 (all >0.5). The strongest marker is the one carrying the largest classifier coefficient (15.268).
exact

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

No assessment has been recorded yet.
🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

322.5 k
tokens (I/O) · 18.7 M incl. cache
80 min
runtime · 0.03 CPU-h
0.6 GB
peak RAM
2 (1 failed)
HPC jobs
hummel
machine