Unveiling prognostics biomarkers of tyrosine metabolism reprogramming in liver cancer by cross-platform gene expression analyses.
Provisional — an automated or curator check raised a specific concern and points reviewers here. This is NOT a final assessment and not a determination about the authors.
The main results reproduced, with only marginal, non-material deviations.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Same input data as the authors
- ✓Reported values were directly comparable
- 🔴A deviation arose in the data or preprocessing
- 🔴A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The deviation was non-trivial in magnitude
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Described well enough for a clean 1:1 on the pipeline output. The cited repo (nguyenquyha/IHC-method, commit 41db5bb) is the authors' own self-contained IHC quantification tool (P16) that ships input images (112 PNGs: FAH/GSTZ1/HGD/HPD), the MATLAB code (brown_calc.m), the intermediate masks, and the expected output (brown_ratio_table.csv). Ported brown_calc.m to Python and re-ran on «our HPC» (SLURM «job»). The shipped 112-value brown-ratio table reproduces essentially EXACTLY: Pearson r=0.99999971, all 112 within 1e-2, 110/112 within 1e-3, and the regenerated tissue masks overlap the shipped masks at Jaccard 0.998 — so both the color step and the morphology port are faithful. Of the three Figure-4 derived fold-changes, GSTZ1 reproduces EXACTLY (2.27-fold, p=0.0007 -> 2.269, p=0.00068) using mean_N/mean_T + unpaired t; but HGD (1.67) and HPD (2.26) are NOT recoverable from the shipped table under that same recipe (got 2.05 and 1.98) and no tested alternative (median, outlier exclusion, per-sample fold) recovers them — flagged as a possible discrepancy for human adjudication (likely a different image subset/version behind those two panels; not asserted as fabrication). NOT attempted: the paper's main cross-platform GSE89377 microarray DEG / tyrosine-metabolism / prognostic-survival analysis — no analysis code is shipped for any of it (the repo is only the IHC tool), so that ~20% is out of scope (no_code sub-result).
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 51assessed: 2026-06-15 ⛓ da501b6985b1
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-15
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusThe study tests whether the five tyrosine catabolic enzymes (TAT, HPD, HGD, GSTZ1, FAH) are dysregulated in hepatocellular carcinoma and whether their expression has clinical and prognostic significance, hypothesizing that tyrosine metabolism is reprogrammed during HCC development.
- ★ Five tyrosine catabolic enzymes (TAT, HPD, HGD, GSTZ1, FAH) are downregulated in HCC versus normal liver at mRNA and protein levels finding
- ★ Low expression of tyrosine catabolic genes (notably TAT, HGD, GSTZ1) correlates with poorer overall and disease-free survival in HCC patients finding
- ★ GSTZ1 overexpression in HCC cells enriches metabolism/oxidative phosphorylation pathways and downregulates glycolysis and cancer pathways, suggesting a tumor-suppressive role mechanism
- ★ miR-539 is upregulated in HCC, negatively correlates with TAT/HPD/GSTZ1/FAH, and predicts worse survival, acting as an upstream regulator mechanism
- Mutation and copy number alterations of tyrosine catabolic genes are infrequent and not strongly associated with their mRNA downregulation in HCC finding
- Cross-platform integrated gene expression analysis (TCGA, GEO, GEPIA, Oncomine, KM plotter) provides a resource framework for identifying HCC biomarkers resource
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Pan-cancer transcriptome mRNA differential analysis (Oncomine) | Multiple cancer types including HCC (LIHC), cancer vs normal tissue | none | Number of datasets with significant gene up/downregulation; gene rank percentile | Oncomine database |
| RNA-seq transcriptome expression analysis (GEPIA) | TCGA-LIHC plus GTEx; 369 HCC tissues vs 160 normal liver | none | log2(TPM+1) gene expression, one-way ANOVA | GEPIA / TCGA / GTEx |
| Stage-wise gene expression (microarray/RNA-seq) | GSE89377: normal (n=13), early HCC (n=5), stage 1 (n=9), stage 2 (n=12), stage 3 (n=14) | none | log2(TPM+1), Student's t-test | GSE89377 (GEO) |
| Kaplan-Meier survival analysis | TCGA-LIHC cohort of 364 liver cancer patients | none | Overall and disease-free survival by high vs low gene expression; HR, log-rank p | GEPIA / KM plotter |
| Immunohistochemistry quantification | HCC tumor vs normal liver tissue | none | Positive staining intensity fold-change, Student's t-test | Human Protein Atlas (HPA) |
| Bulk RNA-seq differential expression and GSEA | Huh7 HCC cell line | GSTZ1 overexpression by adenoviral transfection vs control vector | DEGs (DESeq2), pathway enrichment | GSE117822; R/DESeq2; GSEA; Cytoscape Enrichment Map |
| Mutation and copy number analysis | 353 HCC patients (TCGA) | none | Mutation type/SIFT impact, CNA frequency, GISTIC Q-value | cBioPortal / GISTIC |
| miRNA target prediction and co-expression/survival analysis | TCGA-LIHC (370 HCC, 50 normal); CapitalBio miRNA Array liver dataset | none | Predicted miRNA targets, miRNA-mRNA correlation (r), miRNA expression fold-change, survival HR | TargetScan / starBase / KM plotter |
- ▼ TAT, HPD and GSTZ1 mRNA significantly decreased in HCC vs normal liver (TCGA/GEPIA); HGD and FAH virtually unchanged |Log2FC|=1, p<0.01
- ▼ TAT, HPD, HGD, GSTZ1 and FAH transcripts significantly reduced in HCC stage 2 and 3 vs normal liver, but unchanged in early HCC p<0.05
- ▼ Lower TAT, HGD and GSTZ1 expression associated with worse overall survival in HCC p=0.0067, p=0.0039, p=0.036
- ▼ Lower TAT, HGD and GSTZ1 expression associated with worse disease-free survival p=0.011, p=0.0038, p=0.036
- ▼ HCC tumor IHC staining reduced for HPD, HGD and GSTZ1 vs normal liver HPD 2.26-fold, HGD 1.67-fold, GSTZ1 2.27-fold
- – GSTZ1 overexpression in Huh7 yielded 3163 DEGs and downregulated glycolytic genes HK2 and PDK2 HK2 1.88-fold, PDK2 2.05-fold downregulated
- ▲ miR-539 increased in HCC and negatively correlated with tyrosine catabolic gene expression 2.84-fold (p=0.05); r=-0.221, -0.193, -0.123, -0.166
- ▲ High miR-539 expression led to worse overall survival in HCC patients
- count 390, 428, 431, 457, 448 Oncomine datasets for TAT, HPD, HGD, GSTZ1, FAH (Pan-cancer Oncomine datasets meeting threshold (p≤1e-04, FC=2))
- fold_change 2.26-fold (HPD), 1.67-fold (HGD), 2.27-fold (GSTZ1) decrease (IHC staining HCC vs normal; p=0.0388, 0.0423, 0.0007)
- pvalue p=0.0067 (TAT), 0.0039 (HGD), 0.036 (GSTZ1) (Overall survival log-rank, TCGA-LIHC)
- count 3163 DEGs (1742 up, 1421 down) (GSTZ1 overexpression vs control in Huh7, p<0.01)
- fold_change miR-539 increased 2.84-fold (370 HCC vs 50 normal TCGA-LIHC, p=0.05)
- correlation r=-0.221, -0.193, -0.123, -0.166 (miR-539 vs TAT, HPD, GSTZ1, FAH expression (starBase))
- fold_change HK2 1.88-fold, PDK2 2.05-fold downregulated (Glycolytic genes in GSTZ1-overexpressing Huh7)
- other each gene mutated in <1.1% of patients; HPD CNA Q-value=0.019 (Mutation/CNA profiles in TCGA HCC patients)
Statistical methods review
Model: opusA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This is a cross-platform, in silico bioinformatics study that re-analyzes publicly available transcriptomic, proteomic, mutation, and survival datasets (TCGA-LIHC, GEO, GEPIA, Oncomine, KM plotter, Human Protein Atlas, cBioPortal, starBase) to characterize five tyrosine catabolic genes in hepatocellular carcinoma. Differential expression between tumor and normal tissue was assessed with one-way ANOVA (GEPIA) and Student's t-tests (GEO/IHC quantification), prognosis with Kaplan-Meier curves and log-rank tests plus Cox regression, and pathway changes via DESeq2-derived DEGs followed by GSEA. Group differences are reported with significance thresholds and asterisk tiers, hazard ratios with 95% CIs for survival, and bar/box/violin summaries shown as mean ± SEM.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| one-way ANOVA (GEPIA differential analysis) | tumor vs normal expression of tyrosine catabolic genes, Fig 2A (TCGA + GTEx) | tumor n = 369, normal n = 160 | not stated |
| Student's t-test (two-group, asterisk-tiered p<0.05/0.01/0.001) | expression difference between tumor/stage and normal in GSE89377, Fig 2B | normal n=13, early HCC n=5, stage1 n=9, stage2 n=12, stage3 n=14 | not stated |
| Student's t-test | IHC staining quantification tumor vs adjacent normal (HPD, HGD, GSTZ1, FAH), Fig 4 | — | not stated |
| Kaplan-Meier with log-rank test | overall survival and disease-free survival by high/low gene expression, Fig 3 and S2 (TCGA-LIHC, n=364) | 364 patients | na |
| Cox regression (proportional hazards) | relapse-free survival prediction for miR-539 and miR-661, Fig 6B (KM plotter) | — | not stated |
| DESeq2 Wald test for differential expression | DEGs between GSTZ1-overexpressing and control Huh7 cells, GSE117822 (3163 DEGs, p<0.01) | — | na |
| correlation analysis (Pearson/Spearman r, starBase) | miR-539 vs target gene expression in HCC, S6 Fig (r = -0.221, -0.193, -0.123, -0.166) | 370 HCC, 50 normal | not stated |
-
Tumor-vs-normal and stage-vs-normal expression differences were compared with Student's t-test (and one-way ANOVA in GEPIA), with significance shown as asterisk tiers.↳ Could also: A non-parametric test such as Mann-Whitney U (two groups) or Kruskal-Wallis (multiple stages) could also be used, and exact p-values could be tabulated alongside the asterisks. — Rank-based tests make fewer distributional assumptions, which can be informative for small per-stage groups (n = 5–14), and exact p-values convey the precise strength of evidence.
-
Several genes and stages were each compared to the normal group using separate t-tests.↳ Could also: A single ANOVA across all stages followed by a post-hoc multiple-comparison correction (e.g., Tukey HSD or Dunnett's vs. control) could also frame these comparisons within one model. — A unified model with post-hoc correction also controls the family-wise error rate across the related comparisons and reports them together.
-
Dispersion in bar/box/violin summaries was reported as mean ± SEM.↳ Could also: Standard deviation or a 95% confidence interval could also be reported. — SD or a CI directly conveys the spread of the data (rather than precision of the mean) and is often preferred, especially for small sample sizes.
-
DEG and ANOVA/t-test comparisons were reported against raw p-value thresholds (e.g., p<0.01).↳ Could also: An explicitly stated FDR-adjusted threshold (e.g., Benjamini-Hochberg q-value) could also be applied and reported for the high-dimensional screens. — Reporting adjusted q-values also makes the false-discovery control explicit across the many genes/pathways tested simultaneously.
-
Survival subgroups were defined by dichotomizing expression at the median or quartile cut-points before log-rank testing.↳ Could also: A Cox model treating expression as a continuous variable (optionally with covariate adjustment for stage/age) could also be fit. — Modeling expression continuously avoids the loss of information from dichotomization and can also adjust for potential confounders.
-
Correlations between miRNA and target gene expression were reported as r values.↳ Could also: Reporting whether Pearson or Spearman was used, with 95% CIs on r, could also accompany the coefficients. — Specifying the method and adding CIs clarifies the linear vs. monotonic assumption and conveys the precision of the correlation estimate.
Result convergence & founder nodes
Findings this paper shares with others that ran a comparable experiment. A node’s strength is how many independent papers report it (replication breadth) — not how often it is cited, so a heavily-replicated but under-cited founder still stands out.
-
GSTZ1 protein staining by IHC is reduced 2.27-fold in HCC tumor tissue compared to normal liver.imaging human liver down 2020×1papers★ This paper is the founder (earliest)
-
miR-539 is upregulated 2.84-fold in HCC and negatively correlated with tyrosine catabolic gene expression.microarray human liver up 2020×1papers★ This paper is the founder (earliest)
-
Low GSTZ1 expression is significantly associated with worse overall survival in HCC patients (TCGA-LIHC, p=0.036).other human liver down 2020×1papers★ This paper is the founder (earliest)
-
High miR-539 expression is associated with worse overall survival in HCC patients.other human liver up 2020×1papers★ This paper is the founder (earliest)
-
Low TAT expression is significantly associated with worse disease-free survival in HCC patients (TCGA-LIHC, p=0.011).other human liver down 2020×1papers★ This paper is the founder (earliest)
-
GSTZ1 overexpression in Huh7 HCC cells downregulates glycolytic gene HK2 (1.88-fold) by RNA-seq/DESeq2.RNA-seq huh7 down 2020×1papers★ This paper is the founder (earliest)
-
GSTZ1 (along with TAT and HPD) mRNA is significantly decreased in HCC compared to normal liver across TCGA-LIHC and GEPIA datasets.RNA-seq human liver down 2020×1papers★ This paper is the founder (earliest)
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
- CoINcIDE: A framework for discovery of patient... L1 87/100
- A curated collection of transcriptome datasets... L1 62/100
- An NMF-Based Methodology for Selecting Biomark... L1 84/100
- Meta-analysis of gene expression profiles of l... L1 78/100
- Colorectal Cancer Prediction Based on Weighted...⚑ L1 80/100 ⚑
- Curation of over 10 000 transcriptomic studies... L1 80/100
- Construction and Validation of an Immune Infil...⚑ L1 51/100 ⚑
- Identification of a novel 10 immune-related ge...
- Exploration of the shared diagnostic genes and... L1 76/100
- IRSN-23 gene diagnosis enhances breast cancer... L1 71/100
- Molecular Classification Models for Triple Neg... L1 86/100
- Predicting Bone Metastasis Using Gene Expressi... L1 62/100
- Autoencoder Networks Decipher the Association... L1 74/100
- Comprehensive analysis of a novel RNA modifica... L1 71/100
- Discovery and validation of molecular patterns... L1 83/100
- Comparative profiling of skeletal muscle model... L1 64/100
Downstream reach in the literature
70 downstream papers · 2 datasetsHow widely the datasets deposited by this paper are reused across the whole literature (Europe PMC), beyond our assessed set. This is a factual dependency map — reusing a public dataset is normal, good science. It is not a judgement on the downstream papers; the only verdict here is this paper's own, with its cited rationale.
- Cancer-associated fibroblast-derived secreted phosph... 2023 · 148 cites
- A hypoxia-related signature for clinically predictin... 2020 · 123 cites
- Aryl Hydrocarbon Receptor Signaling Prevents Activat... 2019 · 82 cites
- High EIF2B5 mRNA expression and its prognostic signi... 2018 · 61 cites
- Enhanced DNA Repair Pathway is Associated with Cell... 2021 · 51 cites
- Bioinformatics-based screening of key genes for tran... 2020 · 46 cites
- PCK1 negatively regulates cell cycle progression and... 2019 · 69 cites
- PCK1 Downregulation Promotes TXNRD1 Expression and H... 2018 · 44 cites
- Dysregulated glucuronic acid metabolism exacerbates... 2022 · 19 cites
- GSTZ1-1 downregulates Wnt/β-catenin signalling in he... 2020 · 10 cites
- Transcriptomic changes associated with PCK1 overexpr... 2020 · 9 cites
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-32542016
Paper: Nguyen TN, Nguyen HQ, Le DH. Unveiling prognostics biomarkers of tyrosine metabolism reprogramming in liver cancer by cross-platform gene expression analyses. PLoS One 2020. PMID 32542016 / PMC7295234 / doi:10.1371/journal.pone.0229276.
Code: https://github.com/nguyenquyha/IHC-method (commit 41db5bbfe574d789a86181704c0adad5000055f7, branch master, pushed 2019-04-17).
Data: GSE89377 (microarray, for the cross-platform DEG part) and the IHC images
shipped inside the repo (FAH/GSTZ1/HGD/HPD folders, 112 PNGs + 112 masks).
What the repo actually is (P16: own code, ships data + expected output)
The repo is the authors' own IHC quantification method: a MATLAB script
brown_calc.m that, for each immunohistochemistry image, builds a tissue mask
(grayscale threshold → morphological open/close → largest connected component →
convex hull) and counts the fraction of "brownish" pixels (DAB chromogen) inside
that mask via an HSV color threshold. The repo ships:
- input data: 112 IHC images (FAH 31, GSTZ1 15, HGD 32, HPD 34), each with its
computed
*_mask.png; - code:
brown_calc.m(the full pipeline, self-contained); - expected output:
brown_ratio_table.csv— one brown-ratio per image (112 rows).
This is the ideal reproduction unit: data + code + expected numeric output all shipped together, deterministic, no randomness.
In scope (pipeline-derived, attempted)
- R1 —
brown_ratio_table.csv(primary, direct pipeline output). Re-run the brown-pixel quantification on the shipped images and compare all 112 per-image brown ratios to the shipped table. Two variants for auditability:- R1a (shipped-mask): compute the ratio using the shipped
*_mask.png, isolating the deterministic color step (no morphology uncertainty). - R1b (full-pipeline): regenerate the mask from scratch (port of the MATLAB morphology) and recompute. Also compare regenerated mask vs shipped mask (Jaccard) to validate the morphology port.
- R1a (shipped-mask): compute the ratio using the shipped
- R2 — Figure 4 IHC fold-changes & t-test p-values (derived from R1). The paper reports decreased staining in tumor: GSTZ1 2.27-fold (p=0.0007), HGD 1.67-fold (p=0.0423), HPD 2.26-fold (p=0.0388); FAH not quantified in IHC. Fold = mean(N)/mean(T); unpaired Student t-test. Recompute from the reproduced ratios. (Aggregation of the "± x.xx" term is under-specified in the paper; noted.)
Out of scope (not attempted, and why)
- Cross-platform microarray DEG / tyrosine-metabolism analysis on GSE89377 and the
prognostic-biomarker / survival modeling (the paper's main bioinformatic thread):
no analysis code is shipped for this part — the repo contains only the IHC
image method. Reproducing it would require re-deriving an unspecified pipeline; per
the 80/20 rule this is the hard, under-specified ~20% and is skipped (
no_codefor that sub-result, not the paper as a whole). - Wet-lab IHC staining itself (out of scope by definition — manual/experimental).
Compute
MATLAB is proprietary; the port is run in Python (numpy/scipy/scikit-image) on «our HPC»
SLURM, matching MATLAB's rgb2gray coefficients, mat2gray global rescale, rgb2hsv,
and disk-morphology as closely as feasible. R1a (shipped masks) is mathematically
independent of the morphology port and is the high-confidence anchor.
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
The deterministic pipeline output (112-image brown_ratio_table.csv) reproduces essentially exactly (Pearson 0.99999971) and GSTZ1's Fig-4 fold and p reproduce to three decimals (2.269, p=0.00068), confirming both the data and the recipe are correct. Yet applying that identical recipe to HGD (1.67) and HPD (2.26) yields 2.05 and 1.98 — not recoverable from the shipped table by any tested aggregation, with reported borderline p~0.04 vs computed p<0.01. This sits on the authors' side (a likely different image subset/version or transcription slip behind those panels; the repo also ships FAH data the paper says it didn't quantify), not on our method, and is moderate (direction and significance hold). Overall a solid, well-anchored reproduction with two clearly flagged, non-derivable figure values; the paper's main cross-platform analysis was out of scope (no code).
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.