Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

NUSAP1 Could be a Potential Target for Preventing NAFLD Progression to Liver Cancer.

Front Pharmacol · 2022
L1 45/100 3/4
⚑ Flagged for review — a reproduced result did not match the reported value

Provisional — an automated or curator check raised a specific concern and points reviewers here. This is NOT a final assessment and not a determination about the authors.

Why this verdict

The main result did not reproduce in this reproduction attempt. Where our recomputation produced values that differ from the published ones, those discrepancies are listed below. This is a single automated attempt — not peer review and not a finding of error or misconduct — and differences can also arise from data access, undocumented parameters or the computing environment. The verdict can be contested via “report an error”.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Main result did not reproduce
Decisive
From: Q5 · Derivability / plausibility 🔴
Main result did not reproduce
Decisive
From: Q8 · Severity of the miss (overall human judgment) 🔴
✓ What held up
  • Reported values were directly comparable
What did not (or only partly)
  • 🟡Could not use the authors’ exact input data
  • 🔴A deviation arose in the data or preprocessing
  • 🔴A deviation was attributed to the published material
  • 🔴Reported values were not (fully) derivable from the shared data
  • 🔴The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🔴Overall, the reproduction showed a material discrepancy
How its reproducibility compares
45/100
Reproducibility score
1.7 SD below mean
vs. all fields · 1173 studies
🎯 Scores higher than 6% of all assessed papers rank 1103 of 1173 scored

A 0–100 reproducibility-quality score from the per-question grades, shown as a z-score: standard deviations above (+) or below (−) the mean of comparable assessments.

Reproduction agent’s raw note

Described well enough to attempt, NOT to match 1:1. Pipeline = limma DEG on two public GEO microarrays (GSE89632 Illumina, GSE49541 Affymetrix) + cross-platform gene-symbol intersection visualized with the third-party ggvenn package (P16). Reproduced faithfully on «our HPC» with the stated thresholds (|log2FC|>=0.5, adj.P<0.05 BH). RESULT: the qualitative core holds -- NUSAP1, the paper's headline gene, IS recovered as a DEG common to all three NAFLD-progression comparisons and thus a valid hub-gene candidate. GSE49541 grouping is EXACT (40 mild/32 advanced). BUT exact DEG counts do not reproduce: got 3277/3469/384 genes vs reported 5510/3913/739. Only fibrosis-vs-healthy (3913) matches, and only under UNADJUSTED p<0.05 (3932), contradicting the paper's stated BH adjustment. Venn intersection 64-88 vs reported 112. FABRICATION CONCERN (possible): paper claims '11 healthy controls' but GSE89632 has 24 HC; reported non-fibrosis(5510)>fibrosis(3913) is the reverse of our (and the expected) ordering; counts 5510/739 unreachable under any reasonable knob. NOT ATTEMPTED (the ~20%): full STRING-PPI + Cytoscape-MCODE 6-hub-gene table (GUI tool), wet-lab validation (qPCR/WB/cell assays), and external GEPIA/KM-plotter HCC survival on TCGA -- all out of scope per the brief.

💻 Code ↗ 🗄 Data: GSE89632

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment Score 45
    assessed: 2026-06-15 ⛓ 6964bdd93f18
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-15
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

Can integrated bioinformatics analysis and experimental validation identify a key gene that links NAFLD progression from non-fibrosis through advanced fibrosis to hepatocellular carcinoma, serving as a potential therapeutic target? The authors hypothesize that NUSAP1 is such a gene.

Core claims
  • NUSAP1 is a hub gene linking NAFLD fibrosis progression and HCC and may be a therapeutic target to prevent NAFLD progression to liver cancer finding
  • 112 common DEGs across NAFLD stages were enriched in glucocorticoid receptor pathway, transmembrane transporter regulation, peroxisome, and proteoglycan biosynthetic processes finding
  • Six hub genes (KIF22, ZWINT, NUSAP1, KIAA0101, UHRF1, RAD51AP1) were identified via PPI/MCODE network analysis finding
  • NUSAP1 is upregulated in vitro and in vivo NAFLD models at mRNA and protein levels finding
  • NUSAP1 silencing inhibits cell proliferation, migration, and lipid accumulation under high-fat conditions mechanism
  • NUSAP1 is upregulated in NAFLD-associated HCC and associated with poor survival and advanced tumor stage finding
  • Integrated bioinformatics pipeline (limma DEGs, Metascape GO/KEGG, STRING/Cytoscape MCODE) to identify NAFLD progression genes method
  • Three GEO datasets provide a reusable resource for studying NAFLD-to-HCC progression resource
Experimental setups
Assay System Perturbation Readout Platform
Microarray DEG analysis (limma) Human liver tissue (GSE89632: 21 non-fibrosis, 18 fibrosis, 11 healthy controls) none/disease state differentially expressed genes (|log2FC|≥0.5, adj p<0.05) Illumina HumanHT-12 WG-DASL V4.0 (GPL14951)
Microarray DEG analysis (limma) Human liver tissue (GSE49541: 40 mild fibrosis, 32 advanced fibrosis) none/fibrosis stage DEGs associated with fibrosis progression (adj p<0.05) Affymetrix HG-U133 Plus 2.0 (GPL570)
RNA-seq differential expression Human liver tissue (GSE164441: 10 NAFLD-associated HCC tumor vs 10 paired adjacent non-tumor) none/tumor vs non-tumor NUSAP1 and ZWINT mRNA expression Illumina HiSeq 4000 (GPL20301)
RT-PCR (qPCR) HL-7702 human hepatocyte cell line 1 mmol/L free fatty acids (oleic:palmitic 2:1) 24 h hub gene mRNA levels (normalized to GAPDH) ABI ViiA 7 Real-time PCR system
RT-PCR (qPCR) MHCC-97H hepatoma cell line 1 mmol/L FFA 24 h NUSAP1 and ZWINT mRNA levels ABI ViiA 7 Real-time PCR system
RT-PCR (qPCR) C57BL/6 mouse liver (HFD vs control diet) 60% high-fat diet 12 weeks vs 10% fat control diet hub gene mRNA levels ABI ViiA 7 Real-time PCR system
Western blotting HL-7702, MHCC-97H cell lines and NAFLD mouse liver FFA / HFD NUSAP1 protein level (normalized to GAPDH) NUSAP1 antibody ProteinTech 12024-1-AP
CCK-8 proliferation, wound-healing migration, Oil Red O lipid staining MHCC-97H and HL-7702 cells (si-NUSAP1 knockdown) NUSAP1 siRNA silencing under high-fat condition cell proliferation, migration, lipid content CCK-8 reagent, microplate reader 450 nm; Oil Red O (Servicebio G1015)
Key results
  • 5510 DEGs in non-fibrosis vs HC and 3913 DEGs in fibrosis vs HC identified from GSE89632 5510 and 3913 DEGs
  • 739 DEGs associated with advanced vs mild fibrosis identified from GSE49541; 112 common DEGs across groups 739 DEGs; 112 common
  • PPI network of 112 nodes and 65 edges yielded six hub genes via MCODE 112 nodes, 65 edges
  • ZWINT, NUSAP1, RAD51AP1 showed increasing expression trend across NAFLD progression; ZWINT and Nusap1 significantly upregulated in HFD mice
  • NUSAP1 and ZWINT significantly higher in NAFLD-associated HCC tumor vs paracancer tissue p<0.001
  • NUSAP1 mRNA and protein elevated in FFA-treated HL-7702 and MHCC-97H cells and NAFLD mouse liver
  • NUSAP1 knockdown significantly reduced migration, proliferation, and lipid content in MHCC-97H/HL-7702 under high fat
  • NUSAP1 higher IHC intensity in HCC vs normal liver (HPA) and associated with poor survival and advanced tumor stage
Key statistics
  • count 5510 DEGs (non-fibrosis NAFLD vs healthy controls (21 non-fibrosis, 18 fibrosis, 11 HC))
  • count 3913 DEGs (fibrosis NAFLD vs healthy controls)
  • count 739 DEGs (advanced vs mild fibrosis (GSE49541: 40 mild, 32 advanced))
  • count 112 common DEGs (shared across NAFLD progression comparisons)
  • count 112 nodes and 65 edges (PPI network from STRING (confidence >0.4))
  • pvalue p<0.001 (NUSAP1/ZWINT higher in NAFLD-HCC tumor vs paired paracancer tissue (10 pairs))
  • pvalue p=0.09 and p=0.054 (Rad51ap1 and Uhrf1 increase in NAFLD mice, not statistically significant)
  • other MCODE score 4.0 (KIF22, ZWINT); 3.73 (KIAA0101, UHRF1, NUSAP1, RAD51AP1) (hub gene connectivity scores in MCODE cluster 1)

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

The paper combined reanalysis of three public GEO microarray/RNA-seq datasets with experimental validation. Differentially expressed genes (DEGs) were identified using the limma moderated t-statistic with Benjamini-Hochberg (BH) FDR correction; hub genes were selected through PPI network analysis and MCODE clustering. Experimental validation in cell lines (n=3) and a mouse model (n=5 per group) used Student's t-tests, and results were summarized as mean ± SEM. Survival associations in HCC patients were assessed via online Kaplan-Meier tools (GEPIA2 and KM Plotter).

Replicationmixed Sample sizeGEO datasets: sample sizes per group stated in Table 1. In vivo: n=5 per group (HFD vs CD, 12 weeks). In vitro: n=3 per group, each measurement repeated three times (CCK-8). No formal power calculation reported. GroupsNon-fibrosis vs HC; fibrosis vs HC; advanced vs mild fibrosis; HFD vs CD mice; FFA-treated vs untreated cells; si-NUSAP1 vs negative-control cells; NAFLD-HCC tumor vs adjacent non-tumor Pairingmixed Randomization/blindingnot stated DispersionSEM Exact p-valuesno Effect sizesno Confidence intervalsno Multiplicity correctionBenjamini-Hochberg FDR (referred to as 'B and H method')
Statistical tests used
Test Applied to n Assumptions
limma moderated t-statistic (Empirical Bayes) DEG identification: non-fibrosis vs HC, fibrosis vs HC (GSE89632); advanced vs mild fibrosis (GSE49541) GSE89632: 50 samples (21 non-fibrosis, 18 fibrosis, 11 HC); GSE49541: 72 samples (40 mild, 32 advanced) not stated
Student's t-test Relative mRNA expression of hub genes in FFA-treated cell lines (in vitro) and HFD vs CD mice (in vivo); NUSAP1/ZWINT expression in GSE164441 tumor vs adjacent tissue n=3 per group (in vitro); n=3–5 per group (in vivo); n=10 paired samples (GSE164441) not stated
Log-rank test (implied by Kaplan-Meier online tools) Overall survival and disease-free survival of HCC patients stratified by NUSAP1/ZWINT expression via GEPIA2 and KM Plotter not stated
Hypergeometric/Fisher enrichment test (via Metascape) GO term and KEGG pathway enrichment of 112 common DEGs 112 common DEGs against background gene sets not stated
Approaches that could also have been used
  • Dispersion for experimental data (cell line and mouse assays, n=3–5) was reported as mean ± SEM
    Could also: Mean ± SD or 95% confidence intervals could also be used to summarize spread — With small samples (n=3–5), SD directly describes the variability in the observed data, while SEM reflects precision of the mean estimate; 95% CIs additionally communicate uncertainty in a way that scales explicitly with n, and are often preferred in small-n experimental biology for conveying both spread and inferential context
  • Multiple independent Student's t-tests were applied across several hub genes and cell-line comparisons without a stated correction for multiplicity
    Could also: A one-way or two-way ANOVA followed by a post-hoc correction (e.g., Tukey HSD or Dunnett's test) could also be applied when comparing more than two groups or testing multiple genes simultaneously — ANOVA-based approaches with post-hoc correction provide a family-wise error rate control across the set of comparisons, complementing the BH correction already applied at the bioinformatics DEG-identification stage
  • The GSE164441 dataset comprised 10 paired tumor and adjacent non-tumor tissue samples, and expression differences were evaluated with a Student's t-test
    Could also: A paired t-test or Wilcoxon signed-rank test could also be used to explicitly account for the within-subject pairing — Paired designs reduce between-subject variability and generally increase statistical power; a paired test matches the data structure more directly and is a standard alternative when within-subject data are available
  • Hub genes were selected based on MCODE cluster score (≥3.5) within the PPI network
    Could also: Degree centrality, betweenness centrality, or closeness centrality rankings could also be used to identify hub nodes in a PPI network — Different centrality measures capture different topological properties (local connectivity vs. bridging role vs. proximity to all other nodes); using one or more complementary measures alongside MCODE clustering is a common approach and can surface different biologically relevant candidates
  • Survival associations in HCC were assessed using two external online Kaplan-Meier tools (GEPIA2 and KM Plotter) with a fixed log2FC threshold of 1 for high/low group dichotomization
    Could also: Continuous Cox proportional-hazards regression could also be used to model NUSAP1 expression as a continuous predictor of survival, or optimal cutpoint methods (e.g., maximally selected rank statistics) could be applied to determine the dichotomization threshold empirically — Median or arbitrary FC-based dichotomization can reduce statistical power and sensitivity to the threshold chosen; continuous or data-driven cutpoint approaches would additionally quantify the hazard ratio and its confidence interval, providing effect-size information alongside the significance test
  • Exact p-values were not reported; statistical results were conveyed as asterisk-based significance thresholds (* p<0.05, ** p<0.01, *** p<0.001)
    Could also: Reporting exact p-values (e.g., p=0.032) alongside or instead of threshold symbols is also standard practice — Exact p-values allow readers and meta-analysts to assess effect magnitude relative to the threshold, facilitate replication assessment, and are recommended by many journals and statistical reporting guidelines (e.g., APA, Nature guidelines)
Software: R/limma R 4.0.5; limma version not stated · R/ggplot2 not stated · R/ggvenn not stated · Metascape not stated · STRING not stated · Cytoscape/MCODE Cytoscape v3.8.1; MCODE v2.0.0 · GEPIA2 (online) not stated · KM Plotter (online) not stated

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
22
Impact: medium
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

GPL20301 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE164441 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE49541 GEO in Abstract (http://purl.org/dc/terms/abstract)
no other assessed paper uses this yet
GSE89632 GEO in Abstract (http://purl.org/dc/terms/abstract)
no other assessed paper uses this yet

Downstream reach in the literature

164 downstream papers · 3 datasets

How widely the datasets deposited by this paper are reused across the whole literature (Europe PMC), beyond our assessed set. This is a factual dependency map — reusing a public dataset is normal, good science. It is not a judgement on the downstream papers; the only verdict here is this paper's own, with its cited rationale.

This paper is currently under reproducibility review (see the verdict above). The map below shows where the data in question has propagated — so reuse can be traced, not so the downstream work is presumed affected.

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-35431924 (NUSAP1 / NAFLD→liver cancer)

Paper: Zeng T et al. 2022, Front Pharmacol. DOI 10.3389/fphar.2022.823140. "Code" repo per registry = github.com/yanlinlin82/ggvenn — a THIRD-PARTY Venn diagram R package (P16: applying a third-party tool to the paper's own data is equally valid). The reproducible computational core is a limma DEG analysis of public GEO microarrays + a Venn intersection (ggvenn) + downstream PPI hub genes.

Pipeline-derived results (IN SCOPE — attempt)

id result pipeline reported
deg_nonfib DEGs non-fibrosis vs healthy (GSE89632) limma, |log2FC|>=0.5, adj.P<0.05 BH 5,510
deg_fib DEGs fibrosis vs healthy (GSE89632) limma, same thresholds 3,913
deg_adv DEGs advanced vs mild fibrosis (GSE49541) limma, same thresholds 739
venn_common common DEGs across the 3 comparisons intersection (ggvenn) 112
hub_genes hub genes from PPI of the 112 common DEGs STRING PPI + MCODE (Cytoscape) 6: KIF22,ZWINT,KIAA0101,UHRF1,NUSAP1,RAD51AP1

OUT OF SCOPE (not attempted, why)

  • Wet-lab validation (qPCR/WB/cell proliferation/migration/lipid) — experimental, not a pipeline.
  • HCC survival (OS/RFS/PFS/DFS) via external web tools (GEPIA/Kaplan-Meier plotter) on TCGA — external GUI service, not a shipped pipeline; data is TCGA not the paper's.
  • GSE164441 RNA-seq HCC validation — secondary; not part of the 3-way Venn core.
  • MCODE hub-gene step needs Cytoscape (GUI, manual). The 80%: DEG counts + the 112 intersection. Hub genes = best-effort 20%.

Group mapping (to confirm from GEO pheno)

  • GSE89632: healthy=HC controls; non-fibrosis=simple steatosis (SS); fibrosis=NASH. Confirm counts vs paper (21 non-fib / 18 fib / 11 healthy).
  • GSE49541: mild fibrosis (stage 0-1) vs advanced fibrosis (stage 3-4). (40 mild / 32 advanced).
Figures / tables: Fig.1figureTable
deg_nonfib
Reported
5510 DEGs (non-fibrosis vs healthy, GSE89632)
Reproduced
3277 genes (adj.p) / 3906 (nominal p)
did not match
deg_fib
Reported
3913 DEGs (fibrosis vs healthy, GSE89632)
Reproduced
3469 genes (adj.p) / 3932 (nominal p) -> near-exact under unadjusted p
within tolerance
deg_adv
Reported
739 DEGs (advanced vs mild, GSE49541)
Reproduced
384 genes (adj.p) / 512 (nominal p)
did not match
venn_common
Reported
112 common DEGs (ggvenn 3-way intersection)
Reproduced
64 (adj.p) / 88 (nominal p); NUSAP1 present in both
partial
hub_nusap1
Reported
NUSAP1 hub gene (MCODE 3.73, Table 2)
Reproduced
NUSAP1 present in reproduced common-DEG set (PPI/MCODE step not run)
partial
sample_healthy
Reported
11 healthy controls in GSE89632
Reproduced
GSE89632 contains 24 HC samples
did not match
sample_groups_GSE49541
Reported
40 mild / 32 advanced (GSE49541)
Reproduced
40 mild / 32 advanced
exact

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 45/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🟡1. Data identity
🟢2. Endpoint comparability
🔴3. Location of the main deviation
🔴4. Cause of the deviation
🔴5. Derivability / plausibility
🔴6. Severity of the deviation
🟡7. Core claim
🔴8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Main result did not reproduce
Decisive
From: Q5 · Derivability / plausibility 🔴
Main result did not reproduce
Decisive
From: Q8 · Severity of the miss (overall human judgment) 🔴

The reproducible core holds qualitatively: NUSAP1 is recovered as a common DEG/hub candidate and GSE49541 grouping is exact (40/32). However the quantitative backbone fails — DEG counts (5510/3913/739) are 40–90% above reachable values (3277/3469/384 adj.p), the non-fibrosis>fibrosis ordering is biologically reversed, and only one count matches and only under unadjusted p, contradicting the stated BH method. The defect sits on the authors' side: the values are not derivable from the shared GEO data with the described method, and the '11 healthy controls' claim conflicts with the 24 HC actually present, giving a possible-fabrication concern.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

110.6 k
tokens (I/O) · 4.9 M incl. cache
12 min
runtime · 0.01 CPU-h
2.5 GB
peak RAM
3
HPC jobs
hummel
machine