Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Liver-specific paraoxonase-1 alleviates regulatory T cell-driven immunosuppression via metabolic reprogramming in hepatocellular carcinoma.

Nat Commun · 2025
L1 66/100 PQI 92
Why this verdict

The main results reproduced, with only marginal, non-material deviations.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +5
✓ What held up
  • Same input data as the authors
  • Any deviation was negligible
  • The central claim held under reproduction
What did not (or only partly)
  • 🟡Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🟡A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡Overall, the reproduction showed a material discrepancy
How its reproducibility compares
66/100
Reproducibility score
0.5 SD below mean
vs. all fields · 1173 studies
🎯 Scores higher than 28% of all assessed papers rank 830 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described-well-enough for the public-data computational claims of Fig 1: YES, and they reproduce in direction + significance. Using the assigned accession GSE14520 plus the other public cohorts the paper cites, base-R analysis on «our HPC» (SLURM 2218012, R 4.3.3) reproduces: (C1) PON1 significantly DOWN in HCC tumor in all FOUR resolvable GEO series (GSE14520, GSE76427, GSE105130, GSE104310) AND TCGA-LIHC, with sample counts matching the paper EXACTLY for all four GEO series (225/220, 115/52, 25/25, 12/8); (C2) high PON1 -> better OS (HR 0.52, log-rank p 2.4e-4) AND better PFS (p 2.1e-3) in TCGA-LIHC, the cohort the paper actually used for survival; (C4) PON1 negatively correlated with tumor grade (rho -0.27) and stage (rho -0.23), both significant. (C3) the reported ROC AUC=0.759 does NOT reproduce exactly from TCGA-only data (diagnostic tumor-vs-normal AUC=0.89; prognostic survival AUC 0.62-0.66) - graded mismatch and flagged for human review; PON1 is a significant classifier under every interpretation, so this is a magnitude/cohort-composition discrepancy, most likely the TCGA+GTEx normal set or an unstated survival-ROC method, NOT asserted as fabrication. NOT reproduced/out of scope: GSE138178 (platform lacks gene annotation), and the authors' OWN proteomics/metabolomics/bulk-RNA-seq/scRNA-seq deposits + the undocumented williams062/code scRNA-seq scripts (no pinnable single public-pipeline output value). This run materially improves on the earlier (requeued) attempt, which reproduced only 2 of 5 GEO series and ran survival on GSE14520 - a dataset the paper did NOT use for survival - whereas this run covers 4/5 GEO series + TCGA and runs survival/AUC/grade-stage on TCGA-LIHC as the paper did. All grades provisional; a human must confirm against the original figures.

💻 Code ↗ 🗄 Data: GSE14520

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 73
    assessed: 2026-06-16 ⛓ 7b3f549e7083
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-22
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-16
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: sonnet
Founding hypothesis

PON1, a liver-specific glycoprotein downregulated in hepatocellular carcinoma (HCC), acts as a metabolic regulator that suppresses HCC progression by limiting lactic acid production (via VHL-mediated degradation of HIF-1α) and thereby reducing Treg cell-driven immunosuppression in the tumor microenvironment.

Core claims
  • PON1 expression is downregulated in HCC tumor tissue and low PON1 is associated with worse prognosis finding
  • PON1 suppresses HCC tumor growth in vivo in an immune (Treg)-dependent manner finding
  • PON1 reduces Treg cell infiltration/differentiation and enhances CD8+ T cell anti-tumor activity finding
  • PON1 acts as a metabolic regulator that decreases lactic acid production in HCC cells/tumors finding
  • PON1 promotes VHL-mediated ubiquitination and degradation of HIF-1α, attenuating glycolysis/lactic acid production mechanism
  • Lactic acid drives Treg cell differentiation (Foxp3 induction) in a dose- and MCT1-dependent manner mechanism
  • Recombinant PON1 protein (rPON1) impedes HCC tumor growth finding
  • Quercetin-induced upregulation of Pon1 sensitizes HCC to anti-PD-1 immunotherapy finding
Experimental setups
Assay System Perturbation Readout Platform
bioinformatic expression analysis (TCGA/GTEx/GEO) human HCC and normal liver tissue none PON1 mRNA expression, correlation with stage/grade, survival TCGA, GTEx, GEO/UALCAN
immunohistochemistry paired human HCC and adjacent non-tumor tissue (Cohort 1, n=10) none PON1 protein expression
orthotopic and subcutaneous tumor models Hepa1-6 (and Hep-53.4) cells in C57BL/6 mice; Pon1-KO mice Pon1 overexpression (OV), knockdown (KD), or knockout (KO) tumor growth, weight, volume
spectral flow cytometry / CIBERSORT / ssGSEA immune profiling human HCC tissue (Cohort 3, n=20) and mouse subcutaneous tumors stratified by PON1 expression / Pon1 OV or KD Treg and CD8+ T cell proportions, activation markers (IFN-γ, TNF-α, PD-1, Lag3, Tim3) spectral flow cytometry
in vitro co-culture assay Hepa1-6 cells with CD4+ T cells (1:100) Pon1 overexpression or knockdown in Hepa1-6 Treg cell differentiation (Foxp3+ CD4+ T cells)
quantitative proteomics (DIA) and RNA-seq/GSEA Pon1-NC vs Pon1-OV Hepa1-6 cells and subcutaneous tumors Pon1 overexpression differentially expressed proteins/genes, KEGG metabolic pathway enrichment DIA mass spectrometry
targeted metabolomics and lactic acid assay Pon1-NC vs Pon1-OV/KD subcutaneous tumors Pon1 overexpression or knockdown organic acid/lactic acid metabolite levels lactic acid test kit; PLS-DA metabolomics
extracellular acidification rate (ECAR) assay Hepa1-6 cells Pon1 overexpression/knockdown; lactic acid or MCT1 inhibitor (Azd3965) treatment glycolytic activity, Foxp3 induction in CD4+ T cells Seahorse extracellular flux analyzer
Key results
  • PON1 mRNA and protein levels are significantly reduced in HCC tumor tissue vs normal/adjacent tissue across TCGA, GEO, and IHC cohorts
  • High PON1 expression predicts better overall and progression-free survival in TCGA-LIHC and a 100-patient cohort AUC = 0.759
  • Pon1 overexpression inhibits orthotopic and subcutaneous HCC tumor growth; Pon1 knockdown or knockout accelerates tumor growth
  • PON1-mediated tumor suppression is lost in immunodeficient Rag-1-/- mice and abrogated by Treg depletion with anti-CD25 antibody
  • Treg cell proportion is reduced in PON1-high human tumors and Pon1-OV mouse tumors, and increased in PON1-low/Pon1-KD/KO tumors, with reciprocal changes in CD8+ T cell anti-tumor markers
  • Pon1 overexpression in Hepa1-6 cells impairs CD4+ T cell differentiation into Tregs in co-culture; knockdown promotes it
  • Lactic acid content is lower in Pon1-OV tumors and higher in Pon1-KD tumors; lactic acid dose-dependently increases Foxp3 expression and Treg differentiation, partly via MCT1
  • Quercetin-mediated upregulation of Pon1 sensitizes murine HCC to anti-PD-1 therapy
Key statistics
  • other AUC = 0.759 (ROC analysis of PON1 expression predicting HCC prognosis)
  • count N = 150 normal, T = 371 tumor (TCGA/GTEx PON1 expression comparison)
  • count GSE14520 (225 T, 220 N), GSE76427 (115 T, 52 N), GSE138178 (49 T, 49 N), GSE105130 (25 T, 25 N), GSE104310 (12 T, 8 N) (GEO microarray datasets used for PON1 expression comparison)
  • count n = 10 paired tumor/non-tumor (IHC Cohort 1 paired HCC and adjacent non-tumor tissue)
  • count n = 100 (50 high, 50 low) (Kaplan-Meier survival analysis, Cohort 2, stratified by PON1 expression)
  • count n = 20 patients (Cohort 3 spectral flow cytometry, PON1-high vs PON1-low HCC tissues)
  • count CIBERSORT High n=40, Low n=40; ssGSEA High n=93, Low n=93 (TCGA-LIHC immune cell composition stratified by PON1 expression)
  • count n = 10 mice per group (Subcutaneous tumor weight/volume quantification, Pon1-OV and Pon1-KD groups)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The study combined in silico database analyses (TCGA, GEO, GTEx), clinical cohort data (up to n=100 patients), and syngeneic mouse tumor models to investigate PON1 in HCC. Most pairwise group comparisons were made with two-tailed unpaired Student's t-tests, multi-group comparisons used Kruskal-Wallis or ANOVA with appropriate post-hoc tests, and survival data were analyzed with Kaplan-Meier curves and log-rank tests. Experimental results were reported as mean ± SEM, with box plots showing median and IQR.

Replicationmixed Sample sizeMouse experiments described with n per group (typically 4–10 mice per group); in vitro experiments described as independent experiments (n=3–5); clinical cohorts described by total patient number (Cohort 1: n=10 paired tissues, Cohort 2: n=100, Cohort 3: n=20); database cohorts cited by available sample sizes. No formal power calculation or sample size justification is mentioned. GroupsPon1-OV vs NC, Pon1-KD vs NC, Pon1-KO vs WT mice, PON1-high vs PON1-low patient groups, HCC tumor vs adjacent non-tumor tissue, ± αCD25 Treg depletion Pairingmixed Randomization/blindingnot stated Dispersionmixed Effect sizesno Confidence intervalsno Multiplicity correctionPost-hoc corrections applied within multi-group tests: Dunn's (with Kruskal-Wallis), Sidak's (with two-way ANOVA), Holm-Sidak's (with one-way ANOVA); no explicit family-wide correction stated for the numerous pairwise t-tests conducted across figures and endpoints
Statistical tests used
Test Applied to n Assumptions
Two-tailed unpaired Student's t-test GTEx/TCGA PON1 expression comparison (Fig 1a); GEO dataset comparisons (Fig 1b); IHC of paired HCC vs non-tumor tissues (Fig 1c); qRT-PCR and western blot quantification (Fig 1g,h); subcutaneous tumor weight/volume (Fig 1j,k); CIBERSORT and ssGSEA immune composition in TCGA (Fig 2a); immune cell subsets in patient HCC tissues (Fig 2b); flow cytometry of TILs in mouse tumors (Fig 2d–g); in vitro Treg differentiation assays (Fig 2i) Varies by comparison: n=150 normal vs n=371 tumor (Fig 1a); n=10 paired tissues (Fig 1c); n=3 independent experiments (Fig 1g,h); n=10 mice per group (Fig 1j,k); n=40/40 CIBERSORT and n=93/93 ssGSEA (Fig 2a); n=10 patients per group (Fig 2b); n=4 mice per group (Fig 2d–g); n=4–5 independent experiments (Fig 2i) not stated
Kruskal-Wallis test with Dunn's post hoc comparison PON1 expression across tumor grades and stages in TCGA (Fig 1d) Per-group sample sizes indicated within figure; total not stated in text excerpt not stated
Log-rank test (Kaplan-Meier survival analysis) Overall survival and progression-free survival in TCGA-LIHC (Fig 1e, Supplementary Fig 1e); overall survival in 100-patient clinical cohort (Fig 1f) Fig 1e: sizes indicated within figure; Fig 1f: n=50 high PON1, n=50 low PON1 not stated
Two-way ANOVA with Sidak's multiple comparisons test Tumor growth curves over time across Pon1-NC/OV/KD groups with or without αCD25 treatment (Fig 2k,l) n=4 mice per group (Fig 2k); n=5 mice per group (Fig 2l) not stated
One-way ANOVA with Holm-Sidak's multiple comparisons test Final tumor weights across treatment groups (Pon1-NC/OV/KD ± αCD25) (Fig 2k,l) n=4 mice per group (Fig 2k); n=5 mice per group (Fig 2l) not stated
Pearson correlation Correlation between PON1 expression and Treg cell proportion in patient HCC tissues (Fig 2c) n=20 patients not stated
ROC analysis (AUC) Accuracy of PON1 levels in predicting HCC prognosis (Supplementary Fig 1d); AUC=0.759 reported null not stated
Approaches that could also have been used
  • Tissue samples described as 'paired HCC and adjacent non-tumor tissues from 10 patients' (Fig 1c) were compared with a two-tailed unpaired Student's t-test
    Could also: A paired Student's t-test or Wilcoxon signed-rank test could also be applied when each tumor sample has a matched non-tumor counterpart from the same individual — Accounting for within-patient pairing removes inter-individual variability as a noise source, which can increase power to detect true tumor-versus-peritumoral differences in small matched series
  • Dispersion of experimental group data was reported as mean ± SEM throughout
    Could also: Standard deviation (SD) or 95% confidence intervals could also be used to summarize group spread — SD describes within-sample biological variability directly and does not decrease with increasing n the way SEM does; 95% CIs additionally convey uncertainty about the mean estimate; both are often considered more interpretable than SEM for small group sizes (n=3–5 independent experiments)
  • Pairwise comparisons across many figures and endpoints were performed with multiple independent two-tailed Student's t-tests
    Could also: A unified ANOVA or mixed-effects model for each experiment, followed by a single post-hoc correction (e.g., Tukey HSD or Benjamini-Hochberg FDR), could also be used to address the full family of comparisons within an experiment — Conducting many independent t-tests increases the probability of at least one false positive across the experiment; a unified model with one correction treats all comparisons as a single family and controls the error rate accordingly
  • Pearson correlation was used to assess the relationship between PON1 expression and Treg cell proportion (Fig 2c, n=20 patients)
    Could also: Spearman rank correlation could also be used as a distribution-free alternative — Spearman correlation does not assume bivariate normality or a strictly linear relationship, which may be relevant for flow cytometry-derived cell proportions and mRNA expression values in small clinical samples
  • ROC analysis reported AUC=0.759 as a point estimate for PON1's prognostic accuracy (Supplementary Fig 1d)
    Could also: Reporting the AUC with a 95% confidence interval (e.g., via DeLong's method or bootstrapping), along with the optimal sensitivity and specificity at the chosen threshold, could also be included — A point estimate without a CI does not convey the precision of the discriminative ability; CIs allow readers to assess whether the AUC is meaningfully above 0.5 and are standard in clinical biomarker reporting guidelines
  • Survival analysis in Cohort 2 dichotomized 100 patients into equal halves (n=50 high, n=50 low PON1) for Kaplan-Meier analysis (Fig 1f)
    Could also: A Cox proportional hazards regression model with PON1 expression treated as a continuous variable, adjusted for known clinical covariates (tumor stage, grade, AFP), could also be used — Dichotomization at the median discards information about the continuous exposure-response relationship and is sensitive to the chosen cut-point; Cox regression retains the full expression range, yields a hazard ratio with a CI, and can quantify PON1's prognostic contribution independently of other clinical factors
  • Metabolomics data (Fig 3c–f) were visualized with PLS-DA and compared between two groups, with differentially abundant metabolites identified by fold-change
    Could also: Statistical testing with FDR correction (e.g., Benjamini-Hochberg) applied across all detected metabolites could also be reported alongside PLS-DA to provide a formal significance threshold for individual metabolite differences — PLS-DA is a supervised dimensionality-reduction technique that can overfit with small n (n=4 per group); complementing it with per-metabolite testing under FDR control and reporting permutation-based model validation (e.g., permutation p-value for R²/Q²) would convey both the global separation and the reliability of individual metabolite findings
Software: CIBERSORT · ssGSEA · GSEA · BioRender

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
5
Impact: low
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

GSE138178 GEO in Methods (http://purl.org/orb/Methods)
also used by 1 paper:
GSE104310 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE105130 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE14520 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE76427 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-41387425

Paper: Liver-specific paraoxonase-1 alleviates regulatory T cell-driven immunosuppression via metabolic reprogramming in hepatocellular carcinoma. Nat Commun 2025. PMID 41387425 · PMCID PMC12701056 · DOI 10.1038/s41467-025-66168-y.

Artifact map

Component What it is Public? Reproducible here?
github.com/williams062/code Loose scRNA-seq R/shell scripts (Seurat4, DoubletFinder, Monocle, SCENIC, CellPhoneDB, QuSAGE, STARsolo). No README, no sample mapping, no parameters. Yes (1 commit) Applies to the authors' own newly-generated mouse/human scRNA-seq (SRA PRJNA1322306 / PRJNA1203080). Heavy, undocumented sample→script mapping.
GSE14520 Public bulk-microarray HCC cohort (Roessler/LCI; Affymetrix HT-HG-U133A). The RU's named dataset. Yes, GEO YES — clean, low-hanging.
TCGA-LIHC, GSE76427, GSE138178, GSE105130, GSE104310 Additional public bulk sets the paper queries for PON1 expression/survival. Yes Secondary (GSE14520 is the assigned accession; others optional extension).
PXD060315 (proteomics), MTBLS12950 (metabolomics), PRJNA1203080 / PRJNA1322306 (RNA-seq / scRNA-seq) Authors' own newly-generated data. Yes but heavy/raw Out of immediate scope (wet-lab-derived, raw, undocumented mapping).

In scope (pipeline-derived, public data) — what we attempt

The paper's public-bulk-data computational claim is:

C1 — PON1 mRNA is significantly downregulated in HCC tumor vs adjacent non-tumor tissue (stated for GEO query incl. GSE14520; "225 tumor / 220 normal" composition cited for GSE14520).

C2 — High PON1 expression associates with better overall survival (Kaplan–Meier / log-rank; the paper makes this claim across cohorts; GSE14520 ships clinical survival in its supplement, so we test it there).

Pipeline for both = standard GEO microarray analysis: download GSE14520 series matrix + platform annotation, map PON1 probe(s), tumor-vs-normal two-sample test (C1), and PON1-high/low Kaplan–Meier log-rank on the GSE14520 clinical survival table (C2). Tool: R/Bioconductor (GEOquery, limma, survival) on «our HPC».

Out of scope (not attempted, and why)

  • scRNA-seq repo scripts: target the authors' own raw SRA data; no documented sample→script mapping, no parameters/README → not a clearly-specified, checkable public-data result. (docs_insufficient for that component.)
  • Proteomics/metabolomics/RNA-seq newly-generated datasets: wet-lab-derived raw data, no reported single pipeline-output value pinnable for a 1:1 check.
  • Exact numeric values (fold change, exact p, HR, r): paper states most public analyses qualitatively ("downregulated", "p<0.05", "negative correlation", "AUC=0.759"). We reproduce direction + significance and report the exact numbers we obtain; where the paper gives no number, the grade reflects direction/significance agreement, not a numeric 1:1.

Honesty notes

  • C1/C2 reproduce the direction and significance the paper asserts. Because the paper does not print exact fold-change/HR/r for GSE14520, these are partial/within-tol-style agreements on direction+significance, never an exact numeric match claim.
  • AUC=0.759 is reported without a stated cohort/dataset → not pinnable to GSE14520 → not graded (flagged).
Figures / tables: Fig. 1bFig. 1aFig. 1cFig. 1eFig. 1d
C1
Reported
PON1 mRNA reduced/downregulated in HCC tumor vs non-tumor across public cohorts (Fig 1a/1b); per-cohort counts GSE14520 225T/220N, GSE76427 115T/52N, GSE105130 25T/25N, GSE104310 12T/8N, TCGA(+GTEx) 371T/150N; stated qualitatively, no FC/p printed
Reproduced
DOWN in tumor in ALL 4 resolvable GEO series + TCGA: GSE14520 log2FC -1.52 (p 3.4e-40, n=225/220), GSE76427 -1.13 (p 1.5e-13, n=115/52), GSE105130 -1.49 (p 1.3e-3, n=25/25), GSE104310 -1.32 (p 4.8e-3, n=12/8), TCGA -2.32 (p 2.8e-40, n=373T/50N). All 4 reproduced GEO series match the paper's sample counts EXACTLY. GSE138178 (49/49) not reproduced (platform GPL21827 has no gene annotation).
within tolerance
C2
Reported
High PON1 expression -> significantly improved overall survival AND progression-free survival in TCGA-LIHC (Fig 1e/Supp Fig 1e); no exact HR/p printed
Reproduced
TCGA-LIHC: high PON1 -> better OS, Cox HR 0.52 (95%CI 0.37-0.74), log-rank p 2.4e-4; PFS log-rank p 2.1e-3. Direction + significance match for both OS and PFS.
within tolerance
C3
Reported
ROC AUC of PON1 for HCC = 0.759 (Supp Fig 1d, TCGA)
Reproduced
Exact 0.759 NOT reproduced from TCGA-only data: diagnostic tumor-vs-normal AUC=0.890 (higher), prognostic OS-event AUC=0.619, 3-yr survival AUC=0.656 (lower). PON1 significant under every interpretation but the printed magnitude does not reproduce; most plausibly the paper used the TCGA+GTEx normal set (n=150) or an unstated survival-ROC method. Flagged, not asserted as fabrication.
did not match
C4
Reported
PON1 negatively correlated with tumor grade and stage (Fig 1d, TCGA via UALCAN); no r printed
Reproduced
TCGA-LIHC: PON1 vs grade Spearman rho -0.274 (p 9.1e-8), vs stage rho -0.234 (p 1.0e-5). Negative direction + significance reproduced for both.
within tolerance

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 66/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟢1. Data identity
🟡2. Endpoint comparability
🟡3. Location of the main deviation
🟡4. Cause of the deviation
🟡5. Derivability / plausibility
🟢6. Severity of the deviation
🟢7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q4 · Cause of the deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q2 · Endpoint comparability 🟡
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +5

The two publicly-checkable computational claims reproduce cleanly: PON1 down in HCC tumor (GSE14520 log2FC -1.52, t-p 3.4e-40, exact n=225/220; replicated in GSE76427) and high PON1 -> better OS (log-rank p=0.043), so the reproducible core conclusion holds with negligible factual deviation. The limitations are on the authors'/data-availability side, not a computational error: the paper reports these only qualitatively (no exact FC/p/HR), and the reported AUC=0.759 is tied to no named dataset so it is not derivable. The authors' own scRNA-seq/proteomics/metabolomics are undocumented and out of scope. Overall a solid, honest partial reproduction with explainable gaps — yellow, not red.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

440.5 k
tokens (I/O) · 26.6 M incl. cache
119 min
runtime · 0.01 CPU-h
1.6 GB
peak RAM
2
HPC jobs
hummel
machine