NUSAP1 Could be a Potential Target for Preventing NAFLD Progression to Liver Cancer.
Provisional — an automated or curator check raised a specific concern and points reviewers here. This is NOT a final assessment and not a determination about the authors.
The main result did not reproduce in this reproduction attempt. Where our recomputation produced values that differ from the published ones, those discrepancies are listed below. This is a single automated attempt — not peer review and not a finding of error or misconduct — and differences can also arise from data access, undocumented parameters or the computing environment. The verdict can be contested via “report an error”.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- ✓Reported values were directly comparable
- 🟡Could not use the authors’ exact input data
- 🔴A deviation arose in the data or preprocessing
- 🔴A deviation was attributed to the published material
- 🔴Reported values were not (fully) derivable from the shared data
- 🔴The deviation was non-trivial in magnitude
- 🟡The central claim did not (fully) hold under reproduction
- 🔴Overall, the reproduction showed a material discrepancy
A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.
▸Reproduction agent’s raw note
Described well enough to attempt, NOT to match 1:1. Pipeline = limma DEG on two public GEO microarrays (GSE89632 Illumina, GSE49541 Affymetrix) + cross-platform gene-symbol intersection visualized with the third-party ggvenn package (P16). Reproduced faithfully on «our HPC» with the stated thresholds (|log2FC|>=0.5, adj.P<0.05 BH). RESULT: the qualitative core holds -- NUSAP1, the paper's headline gene, IS recovered as a DEG common to all three NAFLD-progression comparisons and thus a valid hub-gene candidate. GSE49541 grouping is EXACT (40 mild/32 advanced). BUT exact DEG counts do not reproduce: got 3277/3469/384 genes vs reported 5510/3913/739. Only fibrosis-vs-healthy (3913) matches, and only under UNADJUSTED p<0.05 (3932), contradicting the paper's stated BH adjustment. Venn intersection 64-88 vs reported 112. FABRICATION CONCERN (possible): paper claims '11 healthy controls' but GSE89632 has 24 HC; reported non-fibrosis(5510)>fibrosis(3913) is the reverse of our (and the expected) ordering; counts 5510/739 unreachable under any reasonable knob. NOT ATTEMPTED (the ~20%): full STRING-PPI + Cytoscape-MCODE 6-hub-gene table (GUI tool), wet-lab validation (qPCR/WB/cell assays), and external GEPIA/KM-plotter HCC survival on TCGA -- all out of scope per the brief.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessment Score 45assessed: 2026-06-15 ⛓ 6964bdd93f18
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-15
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusCan integrated bioinformatics analysis and experimental validation identify a key gene that links NAFLD progression from non-fibrosis through advanced fibrosis to hepatocellular carcinoma, serving as a potential therapeutic target? The authors hypothesize that NUSAP1 is such a gene.
- ★ NUSAP1 is a hub gene linking NAFLD fibrosis progression and HCC and may be a therapeutic target to prevent NAFLD progression to liver cancer finding
- ★ 112 common DEGs across NAFLD stages were enriched in glucocorticoid receptor pathway, transmembrane transporter regulation, peroxisome, and proteoglycan biosynthetic processes finding
- ★ Six hub genes (KIF22, ZWINT, NUSAP1, KIAA0101, UHRF1, RAD51AP1) were identified via PPI/MCODE network analysis finding
- ★ NUSAP1 is upregulated in vitro and in vivo NAFLD models at mRNA and protein levels finding
- ★ NUSAP1 silencing inhibits cell proliferation, migration, and lipid accumulation under high-fat conditions mechanism
- ★ NUSAP1 is upregulated in NAFLD-associated HCC and associated with poor survival and advanced tumor stage finding
- Integrated bioinformatics pipeline (limma DEGs, Metascape GO/KEGG, STRING/Cytoscape MCODE) to identify NAFLD progression genes method
- Three GEO datasets provide a reusable resource for studying NAFLD-to-HCC progression resource
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Microarray DEG analysis (limma) | Human liver tissue (GSE89632: 21 non-fibrosis, 18 fibrosis, 11 healthy controls) | none/disease state | differentially expressed genes (|log2FC|≥0.5, adj p<0.05) | Illumina HumanHT-12 WG-DASL V4.0 (GPL14951) |
| Microarray DEG analysis (limma) | Human liver tissue (GSE49541: 40 mild fibrosis, 32 advanced fibrosis) | none/fibrosis stage | DEGs associated with fibrosis progression (adj p<0.05) | Affymetrix HG-U133 Plus 2.0 (GPL570) |
| RNA-seq differential expression | Human liver tissue (GSE164441: 10 NAFLD-associated HCC tumor vs 10 paired adjacent non-tumor) | none/tumor vs non-tumor | NUSAP1 and ZWINT mRNA expression | Illumina HiSeq 4000 (GPL20301) |
| RT-PCR (qPCR) | HL-7702 human hepatocyte cell line | 1 mmol/L free fatty acids (oleic:palmitic 2:1) 24 h | hub gene mRNA levels (normalized to GAPDH) | ABI ViiA 7 Real-time PCR system |
| RT-PCR (qPCR) | MHCC-97H hepatoma cell line | 1 mmol/L FFA 24 h | NUSAP1 and ZWINT mRNA levels | ABI ViiA 7 Real-time PCR system |
| RT-PCR (qPCR) | C57BL/6 mouse liver (HFD vs control diet) | 60% high-fat diet 12 weeks vs 10% fat control diet | hub gene mRNA levels | ABI ViiA 7 Real-time PCR system |
| Western blotting | HL-7702, MHCC-97H cell lines and NAFLD mouse liver | FFA / HFD | NUSAP1 protein level (normalized to GAPDH) | NUSAP1 antibody ProteinTech 12024-1-AP |
| CCK-8 proliferation, wound-healing migration, Oil Red O lipid staining | MHCC-97H and HL-7702 cells (si-NUSAP1 knockdown) | NUSAP1 siRNA silencing under high-fat condition | cell proliferation, migration, lipid content | CCK-8 reagent, microplate reader 450 nm; Oil Red O (Servicebio G1015) |
- – 5510 DEGs in non-fibrosis vs HC and 3913 DEGs in fibrosis vs HC identified from GSE89632 5510 and 3913 DEGs
- – 739 DEGs associated with advanced vs mild fibrosis identified from GSE49541; 112 common DEGs across groups 739 DEGs; 112 common
- – PPI network of 112 nodes and 65 edges yielded six hub genes via MCODE 112 nodes, 65 edges
- ▲ ZWINT, NUSAP1, RAD51AP1 showed increasing expression trend across NAFLD progression; ZWINT and Nusap1 significantly upregulated in HFD mice
- ▲ NUSAP1 and ZWINT significantly higher in NAFLD-associated HCC tumor vs paracancer tissue p<0.001
- ▲ NUSAP1 mRNA and protein elevated in FFA-treated HL-7702 and MHCC-97H cells and NAFLD mouse liver
- ▼ NUSAP1 knockdown significantly reduced migration, proliferation, and lipid content in MHCC-97H/HL-7702 under high fat
- ▲ NUSAP1 higher IHC intensity in HCC vs normal liver (HPA) and associated with poor survival and advanced tumor stage
- count 5510 DEGs (non-fibrosis NAFLD vs healthy controls (21 non-fibrosis, 18 fibrosis, 11 HC))
- count 3913 DEGs (fibrosis NAFLD vs healthy controls)
- count 739 DEGs (advanced vs mild fibrosis (GSE49541: 40 mild, 32 advanced))
- count 112 common DEGs (shared across NAFLD progression comparisons)
- count 112 nodes and 65 edges (PPI network from STRING (confidence >0.4))
- pvalue p<0.001 (NUSAP1/ZWINT higher in NAFLD-HCC tumor vs paired paracancer tissue (10 pairs))
- pvalue p=0.09 and p=0.054 (Rad51ap1 and Uhrf1 increase in NAFLD mice, not statistically significant)
- other MCODE score 4.0 (KIF22, ZWINT); 3.73 (KIAA0101, UHRF1, NUSAP1, RAD51AP1) (hub gene connectivity scores in MCODE cluster 1)
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
The paper combined reanalysis of three public GEO microarray/RNA-seq datasets with experimental validation. Differentially expressed genes (DEGs) were identified using the limma moderated t-statistic with Benjamini-Hochberg (BH) FDR correction; hub genes were selected through PPI network analysis and MCODE clustering. Experimental validation in cell lines (n=3) and a mouse model (n=5 per group) used Student's t-tests, and results were summarized as mean ± SEM. Survival associations in HCC patients were assessed via online Kaplan-Meier tools (GEPIA2 and KM Plotter).
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| limma moderated t-statistic (Empirical Bayes) | DEG identification: non-fibrosis vs HC, fibrosis vs HC (GSE89632); advanced vs mild fibrosis (GSE49541) | GSE89632: 50 samples (21 non-fibrosis, 18 fibrosis, 11 HC); GSE49541: 72 samples (40 mild, 32 advanced) | not stated |
| Student's t-test | Relative mRNA expression of hub genes in FFA-treated cell lines (in vitro) and HFD vs CD mice (in vivo); NUSAP1/ZWINT expression in GSE164441 tumor vs adjacent tissue | n=3 per group (in vitro); n=3–5 per group (in vivo); n=10 paired samples (GSE164441) | not stated |
| Log-rank test (implied by Kaplan-Meier online tools) | Overall survival and disease-free survival of HCC patients stratified by NUSAP1/ZWINT expression via GEPIA2 and KM Plotter | — | not stated |
| Hypergeometric/Fisher enrichment test (via Metascape) | GO term and KEGG pathway enrichment of 112 common DEGs | 112 common DEGs against background gene sets | not stated |
-
Dispersion for experimental data (cell line and mouse assays, n=3–5) was reported as mean ± SEM↳ Could also: Mean ± SD or 95% confidence intervals could also be used to summarize spread — With small samples (n=3–5), SD directly describes the variability in the observed data, while SEM reflects precision of the mean estimate; 95% CIs additionally communicate uncertainty in a way that scales explicitly with n, and are often preferred in small-n experimental biology for conveying both spread and inferential context
-
Multiple independent Student's t-tests were applied across several hub genes and cell-line comparisons without a stated correction for multiplicity↳ Could also: A one-way or two-way ANOVA followed by a post-hoc correction (e.g., Tukey HSD or Dunnett's test) could also be applied when comparing more than two groups or testing multiple genes simultaneously — ANOVA-based approaches with post-hoc correction provide a family-wise error rate control across the set of comparisons, complementing the BH correction already applied at the bioinformatics DEG-identification stage
-
The GSE164441 dataset comprised 10 paired tumor and adjacent non-tumor tissue samples, and expression differences were evaluated with a Student's t-test↳ Could also: A paired t-test or Wilcoxon signed-rank test could also be used to explicitly account for the within-subject pairing — Paired designs reduce between-subject variability and generally increase statistical power; a paired test matches the data structure more directly and is a standard alternative when within-subject data are available
-
Hub genes were selected based on MCODE cluster score (≥3.5) within the PPI network↳ Could also: Degree centrality, betweenness centrality, or closeness centrality rankings could also be used to identify hub nodes in a PPI network — Different centrality measures capture different topological properties (local connectivity vs. bridging role vs. proximity to all other nodes); using one or more complementary measures alongside MCODE clustering is a common approach and can surface different biologically relevant candidates
-
Survival associations in HCC were assessed using two external online Kaplan-Meier tools (GEPIA2 and KM Plotter) with a fixed log2FC threshold of 1 for high/low group dichotomization↳ Could also: Continuous Cox proportional-hazards regression could also be used to model NUSAP1 expression as a continuous predictor of survival, or optimal cutpoint methods (e.g., maximally selected rank statistics) could be applied to determine the dichotomization threshold empirically — Median or arbitrary FC-based dichotomization can reduce statistical power and sensitivity to the threshold chosen; continuous or data-driven cutpoint approaches would additionally quantify the hazard ratio and its confidence interval, providing effect-size information alongside the significance test
-
Exact p-values were not reported; statistical results were conveyed as asterisk-based significance thresholds (* p<0.05, ** p<0.01, *** p<0.001)↳ Could also: Reporting exact p-values (e.g., p=0.032) alongside or instead of threshold symbols is also standard practice — Exact p-values allow readers and meta-analysts to assess effect magnitude relative to the threshold, facilitate replication assessment, and are recommended by many journals and statistical reporting guidelines (e.g., APA, Nature guidelines)
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
Downstream reach in the literature
164 downstream papers · 3 datasetsHow widely the datasets deposited by this paper are reused across the whole literature (Europe PMC), beyond our assessed set. This is a factual dependency map — reusing a public dataset is normal, good science. It is not a judgement on the downstream papers; the only verdict here is this paper's own, with its cited rationale.
- Landscape of Intercellular Crosstalk in Healthy and... 2019 · 656 cites
- Prolonged hypernutrition impairs TREM2-dependent eff... 2023 · 190 cites
- Osteopontin Takes Center Stage in Chronic Liver Dise... 2021 · 171 cites
- A maresin 1/RORα/12-lipoxygenase autoregulatory circ... 2019 · 125 cites
- Deficiency of gluconeogenic enzyme PCK1 promotes met... 2023 · 115 cites
- Overexpression of NAG-1/GDF15 prevents hepatic steat... 2022 · 102 cites
- Opposing roles of hepatic stellate cell subpopulatio... 2022 · 282 cites
- Osteopontin Takes Center Stage in Chronic Liver Dise... 2021 · 171 cites
- Autophagy is a gatekeeper of hepatic differentiation... 2018 · 137 cites
- Interspecies NASH disease activity whole-genome prof... 2017 · 105 cites
- Transcriptional Dynamics of Hepatic Sinusoid-Associa... 2020 · 89 cites
- Human hepatic gene expression signature of non-alcoh... 2017 · 82 cites
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-35431924 (NUSAP1 / NAFLD→liver cancer)
Paper: Zeng T et al. 2022, Front Pharmacol. DOI 10.3389/fphar.2022.823140. "Code" repo per registry = github.com/yanlinlin82/ggvenn — a THIRD-PARTY Venn diagram R package (P16: applying a third-party tool to the paper's own data is equally valid). The reproducible computational core is a limma DEG analysis of public GEO microarrays + a Venn intersection (ggvenn) + downstream PPI hub genes.
Pipeline-derived results (IN SCOPE — attempt)
| id | result | pipeline | reported |
|---|---|---|---|
| deg_nonfib | DEGs non-fibrosis vs healthy (GSE89632) | limma, |log2FC|>=0.5, adj.P<0.05 BH | 5,510 |
| deg_fib | DEGs fibrosis vs healthy (GSE89632) | limma, same thresholds | 3,913 |
| deg_adv | DEGs advanced vs mild fibrosis (GSE49541) | limma, same thresholds | 739 |
| venn_common | common DEGs across the 3 comparisons | intersection (ggvenn) | 112 |
| hub_genes | hub genes from PPI of the 112 common DEGs | STRING PPI + MCODE (Cytoscape) | 6: KIF22,ZWINT,KIAA0101,UHRF1,NUSAP1,RAD51AP1 |
OUT OF SCOPE (not attempted, why)
- Wet-lab validation (qPCR/WB/cell proliferation/migration/lipid) — experimental, not a pipeline.
- HCC survival (OS/RFS/PFS/DFS) via external web tools (GEPIA/Kaplan-Meier plotter) on TCGA — external GUI service, not a shipped pipeline; data is TCGA not the paper's.
- GSE164441 RNA-seq HCC validation — secondary; not part of the 3-way Venn core.
- MCODE hub-gene step needs Cytoscape (GUI, manual). The 80%: DEG counts + the 112 intersection. Hub genes = best-effort 20%.
Group mapping (to confirm from GEO pheno)
- GSE89632: healthy=HC controls; non-fibrosis=simple steatosis (SS); fibrosis=NASH. Confirm counts vs paper (21 non-fib / 18 fib / 11 healthy).
- GSE49541: mild fibrosis (stage 0-1) vs advanced fibrosis (stage 3-4). (40 mild / 32 advanced).
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
The reproducible core holds qualitatively: NUSAP1 is recovered as a common DEG/hub candidate and GSE49541 grouping is exact (40/32). However the quantitative backbone fails — DEG counts (5510/3913/739) are 40–90% above reachable values (3277/3469/384 adj.p), the non-fibrosis>fibrosis ordering is biologically reversed, and only one count matches and only under unadjusted p, contradicting the stated BH method. The defect sits on the authors' side: the values are not derivable from the shared GEO data with the described method, and the '11 healthy controls' claim conflicts with the 24 HC actually present, giving a possible-fabrication concern.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.