The relationship between PLOD1 expression level and glioma prognosis investigated using public databases.
The main results reproduced: recomputed values matched the published ones within tolerance.
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
P16 third-party-tool case: the named repo (github.com/taiyun/corrplot) is a generic R plotting package, NOT the authors' analysis code, so reproduction = re-running the paper's described public-database pipeline on the paper's own public data. This is a CLEAN RE-RUN (room was requeued because the prior ROOM_RESULT lacked the now-mandatory datasets[]+qc_room blocks). Rebuilt the R 4.3.3 env (GEOquery/limma/survival/survivalROC/fgsea) on «our HPC» front1, re-downloaded all public data (GEO GSE4290/7696/50161 series matrices, CGGA mRNAseq 693+325 from cgga.org.cn, TCGA GBMLGG from UCSC Xena), and re-ran 4 SLURM jobs (2218093 geo / 2218094 cgga / 2218095 tcga / 2218096 gsea). Every output reproduced the prior run BIT-FOR-BIT (deterministic). Results across all three data arms: (1) GEO expression Fig 2B: PLOD1 significantly higher in glioma than normal in GSE4290 (log2FC +0.712, p=6.98e-12), GSE7696 (+0.904) and GSE50161 (+1.108); GSE7696/GSE50161 counts exact. (2) CGGA prognosis (1018 glioma, 970 with survival): KM high-vs-low PLOD1 log-rank p<1e-16 (C2); Cox HR uni 1.786 / multi 1.421, both p<<0.001 (C6, vs paper 1.986/1.283); ROC AUC 0.698/0.757/0.779 (C4, same 0.70-0.80 range as paper 0.816/0.793/0.707 but per-year trend differs); DEGs high-vs-low predominantly up, 67 up/16 down (C7, vs paper 61 up/6 down); PLOD1 co-expresses with all 10 named ECM genes r=0.635-0.756 (C9). (3) TCGA validation (GBMLGG, n=691): KM p<1e-16 (C3); ROC AUC 0.809/0.809/0.795 — matches paper 0.786/0.805/0.796 within ~0.02 (C5). (4) GSEA: KEGG_ECM_RECEPTOR_INTERACTION NES=2.081 p=1.6e-10 (C8, vs paper NES 1.96 p=0.002), leading edge = exactly the Fig 7 corrplot gene set (cross-confirms C9). Net: 9/11 sub-claims within-tol (direction+magnitude), 2/11 partial (C4 CGGA ROC per-year values, C7 exact DEG counts) — both same-direction numeric divergences attributable to the paper not specifying how CGGA mRNAseq_693 and mRNAseq_325 were combined/batch-corrected. No mismatches and no fabrication signal: every reported headline value is derivable from the shipped public data, and the central conclusion (PLOD1 up in glioma; high PLOD1 -> poor prognosis in two independent cohorts; PLOD1-driven ECM-receptor co-expression program) reproduces comprehensively. NOT attempted (out of scope, no reproducible numeric pipeline output): cBioPortal mutation panel, CCLE pan-cancer expression, Metascape GO figure, GraphPad/GUI nomogram calibration. Dataset note: TCGA public Xena freeze (702/691) is larger than the paper's curated 592-sample set, which the paper under-specifies (delivers_promised=partial for that deposit); all other deposits deliver exactly what the paper used.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 79assessed: 2026-06-16 ⛓ ae3cbc558bc6
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-22
- Rubric version
- not recorded
- Assessed by
- —
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: sonnetSince PLOD2 and PLOD3 are known to be highly expressed and prognostically relevant in glioma, this study tests whether PLOD1 expression is also elevated in glioma and whether it has independent prognostic value for glioma patients.
- ★ PLOD1 mRNA expression is significantly higher in glioma tissue than in normal brain tissue finding
- ★ High PLOD1 expression is associated with poorer overall survival in glioma patients finding
- ★ PLOD1 expression is an independent prognostic factor for glioma on multivariate Cox regression finding
- ★ PLOD1 expression correlates with clinicopathologic features including age, WHO grade, recurrence/secondary status, IDH mutation and 1p19q codeletion status finding
- ★ High PLOD1 expression is enriched in extracellular matrix (ECM)-related biological processes and pathways by GO and GSEA analysis mechanism
- ★ PLOD1 is positively co-expressed with ECM-related genes (HSPG2, COL6A2, COL4A2, FN1, COL1A1, COL4A1, CD44, COL3A1, COL1A2, SPP1), which are themselves associated with poor glioma prognosis finding
- A nomogram combining grade, IDH status, MGMT methylation, 1p19q codeletion and PLOD1 expression accurately predicts 1-, 2-, and 3-year survival in glioma method
- ★ PLOD1 represents a novel prognostic biomarker and potential therapeutic target for glioma finding
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| gene expression profiling (RNA-seq) | human glioma tissue vs normal brain tissue (CGGA cohort, 636 glioma + 20 normal) | none | PLOD1 mRNA expression level | — |
| gene expression profiling (microarray, GEO validation) | glioma vs normal brain tissue (GSE4290, GSE7696, GSE50161) | none | PLOD1 expression level | GEO microarray datasets |
| online expression database analysis (GEPIA) | GBM and LGG tumor vs normal tissue (TCGA/GTEx-based) | none | PLOD1 expression level | GEPIA (http://gepia.cancer-pku.cn/) |
| genomic mutation/alteration analysis | glioma samples | none | PLOD1 mutation type and frequency | cBioPortal |
| pan-cancer expression comparison | cancer cell lines across tissue types including glioma | none | PLOD1 expression rank across cancer types | CCLE (Cancer Cell Line Encyclopedia) |
| Kaplan-Meier survival analysis | glioma patients, high vs low PLOD1 expression groups (CGGA and TCGA cohorts) | none | overall survival | R survival/survminer packages |
| univariate/multivariate Cox regression and nomogram construction | glioma patients (CGGA cohort, validated in TCGA) | none | hazard ratio for OS; predicted 1-, 2-, 3-year survival probability | R survival and rms packages |
| differential gene expression analysis with GO/GSEA enrichment | glioma samples, high vs low PLOD1 expression (CGGA cohort) | none | differentially expressed genes and enriched pathways | R limma package; Metascape; GSEA |
- ▲ PLOD1 expression is significantly higher in glioma than normal tissue across CGGA, GEO (GSE4290/GSE7696/GSE50161), and GEPIA analyses
- ▼ High PLOD1 expression associated with poor overall survival in both CGGA and TCGA cohorts
- ▲ Univariate Cox regression: PLOD1 expression significantly associated with OS HR=1.986; 95% CI [1.796-2.198]; P<0.001
- ▲ Multivariate Cox regression: PLOD1 expression remains an independent prognostic factor for OS HR=1.283; 95% CI [1.128-1.460]; P<0.001
- – ROC analysis shows PLOD1 predicts 1-, 3-, and 5-year survival in CGGA and TCGA cohorts CGGA AUC: 0.816/0.793/0.707; TCGA AUC: 0.786/0.805/0.796
- ▲ GSEA shows ECM receptor interaction pathway enriched in high-PLOD1 phenotype NES=1.96, normalized P=0.002
- – 67 differentially expressed genes identified (61 upregulated, 6 downregulated), enriched in extracellular structure organization 61 up / 6 down
- ▲ PLOD1 positively co-expressed with 10 ECM-related genes, all of which are highly expressed in glioma and associated with poor prognosis
- fold_change/hazard_ratio HR=1.986; 95% CI [1.796-2.198]; P<0.001 (univariate Cox regression of PLOD1 expression on OS, CGGA cohort)
- fold_change/hazard_ratio HR=1.283; 95% CI [1.128-1.460]; P<0.001 (multivariate Cox regression of PLOD1 expression on OS, CGGA cohort)
- other AUC 1y=0.816, 3y=0.793, 5y=0.707 (ROC curve for PLOD1 predicting survival, CGGA cohort)
- other AUC 1y=0.786, 3y=0.805, 5y=0.796 (ROC curve for PLOD1 predicting survival, TCGA cohort)
- other NES=1.96, normalized P=0.002 (GSEA enrichment of ECM receptor interaction pathway in high-PLOD1 phenotype)
- count 67 DEGs (61 upregulated, 6 downregulated); adjusted P<0.05, |log2FC|>2 (differential gene expression analysis, CGGA dataset)
- pvalue P<0.001 (Kaplan-Meier survival difference between high and low PLOD1 expression groups, CGGA cohort)
- count 10 common genes (intersection of DEGs and core ECM pathway genes used for co-expression analysis)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This retrospective bioinformatics study examined PLOD1 expression across multiple public glioma databases (CGGA n=636, TCGA n=592, and three GEO cohorts), using Kaplan-Meier survival curves, univariate and multivariate Cox regression, and time-dependent ROC curves to assess PLOD1's prognostic value. Differentially expressed genes were identified using limma with an adjusted p-value threshold, followed by GO enrichment via Metascape and GSEA to explore biological mechanisms. Co-expression relationships between PLOD1 and ECM-related genes were quantified using Pearson or Spearman correlations, and a prognostic nomogram was constructed in the CGGA cohort and assessed via calibration curves.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Kaplan-Meier survival analysis (log-rank test implied) | Overall survival comparison between high and low PLOD1 expression groups in CGGA and TCGA; further stratified by IDH mutation and MGMT promoter methylation status | CGGA n=636; TCGA n=592 | not stated |
| Univariate Cox proportional hazards regression | Association of PLOD1 expression and clinicopathologic variables with overall survival in CGGA dataset | 636 | not stated |
| Multivariate Cox proportional hazards regression | Independent prognostic value of PLOD1 adjusted for grade, PRS type, chemotherapy, IDH mutation, 1p19q codeletion in CGGA dataset | 636 | not stated |
| Time-dependent ROC curve analysis (survivalROC package, Kaplan-Meier method) | Predictive accuracy of PLOD1 at 1-, 3-, and 5-year survival in CGGA and TCGA datasets | CGGA n=636; TCGA n=592 | not stated |
| limma moderated t-test / linear model | Differential gene expression between high and low PLOD1 groups in CGGA RNA-seq data | 636 | not stated |
| Pearson correlation (parametric) and Spearman correlation (non-parametric), applied per data distribution | Pairwise co-expression analysis between PLOD1 and 10 common ECM-related genes in CGGA dataset | — | not stated |
| Gene Set Enrichment Analysis (GSEA), permutation-based | Pathway enrichment in high vs. low PLOD1 expression phenotype; ECM receptor interaction pathway (NES=1.96, FDR<0.05) | — | na |
| GO enrichment analysis via Metascape | Biological process enrichment of 67 identified DEGs | 67 DEGs | na |
-
PLOD1 expression was dichotomized at the median into high and low groups for all survival analyses↳ Could also: PLOD1 could be entered as a continuous variable in Cox regression, or an optimal cutpoint could be identified using maximally selected log-rank statistics or restricted cubic splines — A continuous or optimized-cutpoint approach retains more information from the expression variable and avoids the information loss inherent in a fixed median split; it also reduces the risk that the chosen threshold overstates discrimination on the same dataset
-
The limma package (originally designed for microarray log-ratio data) was used for differential expression analysis of RNA-seq data↳ Could also: DESeq2 or edgeR, which model RNA-seq read counts with a negative binomial distribution, could also be applied; limma-voom is another widely used adaptation of limma explicitly designed for count data — Negative binomial models are parameterized for the overdispersion commonly observed in RNA-seq count data; limma-voom applies a variance-stabilizing transformation to count data before applying the linear model framework
-
A prognostic nomogram was constructed and evaluated using calibration curves in the same CGGA cohort used to build it↳ Could also: Bootstrap internal validation (e.g., 1,000 resamples with optimism correction) or application of the nomogram to an independent external cohort (e.g., TCGA) could also be used to assess calibration and discrimination — Calibration evaluated on the training data alone tends to appear optimistic; bootstrap resampling or external validation in a held-out cohort provides a less biased estimate of predictive performance in new patients
-
Multiple Kaplan-Meier survival comparisons were performed across several clinicopathologic subgroups (IDH, MGMT, grade, age, PRS type) without an explicitly stated multiplicity correction↳ Could also: A family-wise error rate correction (e.g., Bonferroni) or a false discovery rate procedure (e.g., Benjamini-Hochberg) could also be applied across the family of survival comparisons — When multiple hypotheses are tested simultaneously, applying a multiplicity correction is a common approach to account for the increased probability of at least one false positive across the set of comparisons
-
Pairwise Pearson or Spearman correlations were used for co-expression analysis between PLOD1 and each of 10 ECM genes individually↳ Could also: Weighted gene co-expression network analysis (WGCNA) or partial correlation methods could also characterize co-expression structure across all genes jointly — Network-based methods model the joint co-expression structure among all genes simultaneously and can identify gene modules and hub genes, which may provide additional biological context beyond pairwise correlations and better account for indirect associations
-
Survival results were reported using Kaplan-Meier curves and log-rank p-values without reporting median survival times or their confidence intervals↳ Could also: Median overall survival with 95% CIs and restricted mean survival time (RMST) could also be reported alongside the Kaplan-Meier curves — Median survival and RMST provide interpretable absolute effect measures that complement the hazard ratio and log-rank p-value, and are especially informative when survival curves diverge early or when the proportional hazards assumption may not hold throughout follow-up
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-34040895
Paper: Tian L et al. (2021) The relationship between PLOD1 expression level and glioma prognosis investigated using public databases. PeerJ 9:e11422. PMCID PMC8127981.
Type: Public-database biomarker / bioinformatics paper. The named "code" repo (github.com/taiyun/corrplot) is a third-party R visualization package, not the authors' analysis code → P16 situation. We reproduce by re-running the described pipeline (R 3.6.3: limma/survival/survminer/survivalROC/rms/corrplot) on the paper's own public data (GEO GSE4290 + CGGA + TCGA).
Datasets used by the paper
| dataset | n | role | obtainable |
|---|---|---|---|
| CGGA mRNAseq (693+325) | 1018 glioma + 20 normal | main analysis | cgga.org.cn (free) |
| TCGA (LGG+GBM) | 592 (449 LGG, 143 GBM) | validation | TCGA/GEPIA |
| GSE4290 | 153 glioma + 23 normal | expression validation (Fig 2B) | GEO (brief's accession) |
| GSE7696 | 80 glioma + 4 normal | expression validation | GEO |
| GSE50161 | 117 glioma + 13 normal | expression validation | GEO |
In-scope, pipeline-derived claims (targets)
| id | claim | reported | fig/loc | feasibility |
|---|---|---|---|---|
| C1 | PLOD1 higher in glioma vs normal (GSE4290) | "significantly higher" (p<0.001) | Fig 2B | EASY — primary (brief dataset) |
| C1b | same in GSE7696 / GSE50161 | significantly higher | Fig 2 | EASY |
| C2 | CGGA KM: high PLOD1 → poor survival | p<0.001 | Fig 3A | MEDIUM |
| C3 | TCGA KM | p<0.001 | Fig 3B | MEDIUM |
| C4 | CGGA ROC AUC 1/3/5-yr | 0.816 / 0.793 / 0.707 | Fig 4 | MEDIUM |
| C5 | TCGA ROC AUC 1/3/5-yr | 0.786 / 0.805 / 0.796 | Fig 4 | MEDIUM |
| C6 | CGGA Cox HR (uni / multi) | 1.986 [1.796–2.198] / 1.283 [1.128–1.460], p<0.001 | Table/Fig 5 | MEDIUM |
| C7 | DEGs high-vs-low PLOD1 (CGGA, limma adjP<0.05, | log2FC | >2) | 67 total (61 up, 6 down) |
| C8 | GSEA: ECM-receptor interaction | NES=1.96, p=0.002 | GSEA | HARDER (optional) |
| C9 | corrplot PLOD1 co-expression (ECM genes) | top: HSPG2,COL6A2,COL4A2,FN1,COL1A1,COL4A1,CD44,COL3A1,COL1A2,SPP1 | Fig 7 | MEDIUM (named repo) |
Out of scope (not attempted)
- cBioPortal mutation panel, CCLE pan-cancer expression, Metascape GO figure, nomogram calibration (manual/GUI tools, GraphPad Prism plots) — no reproducible numeric pipeline output pinnable.
Priority (80/20)
- C1 GSE4290 PLOD1 glioma-vs-normal (brief's dataset) — Job 1.
- C2/C4/C6/C7/C9 all from one CGGA download — Job 2 (headline numbers).
- C3/C5 (TCGA), C8 (GSEA) — optional last 20%.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.