Developing a thyroid cancer differentiation state classification system using deep residual networks and metabolic signature profiling.
Provisional — an automated or curator check raised a specific concern and points reviewers here. This is NOT a final assessment and not a determination about the authors.
Part of the results reproduced; minor but material deviations remained.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
- Nothing in this column.
- 🔴Could not use the authors’ exact input data
- 🔴Reported values were only indirectly comparable
- 🟡A deviation arose in the data or preprocessing
- 🔴A deviation was attributed to the published material
- 🟡Reported values were not (fully) derivable from the shared data
- 🟡The deviation was non-trivial in magnitude
- 🟡The central claim did not (fully) hold under reproduction
- 🟡Overall, the reproduction showed a material discrepancy
This paper has a computational component, but its primary data is legally or ethically access-restricted — identifiable patient cohorts, rare-disease genomes, or controlled-access biobanks that cannot be openly shared. The reproduction therefore could not be attempted. That is a neutral verdict: it does not mean the result is wrong or that the authors fell short — only that, for legitimate privacy reasons, it cannot be independently checked from public data. We deliberately do NOT assign a 0–100 score here, because a low number would wrongly read as a failed reproduction.
▸Reproduction agent’s raw note
DROP. The paper (npj Digit Med 2025, thyroid-cancer ResNet differentiation classifier) is described in moderate detail but is NOT reproducible from public artifacts. BRIEF links were both text-mining false positives: real code = github.com/Aceracede3000/Aceracede3000 (a single-author GitHub profile repo, not STAR-Fusion); real public data = GEO GSE29265/GSE33630/GSE53157/GSE65144/GSE76039 (not GSE193581). Every headline model (10-gene '10MG' and 10-metabolite '10M' ResNet1D classifiers) is trained/tested on a merged GEO+FUSCC matrix (n=453); the FUSCC raw sequencing + untargeted metabolomics (158 tumors + 57 normal, participant-level) is RESTRICTED / on-request only via the FUSCC Data-Access Committee -> data_restricted. The repo additionally ships NO input data at any commit, NO trained weights, and no pipeline to rebuild the input matrix from the GEO accessions; models.py (ResNet1D) survives only in git history; the final 10 genes are shown only in Fig 7, not text-enumerated -> docs_insufficient. The 5 public GEO cohorts are downloadable but cannot reproduce the reported numbers (model trained on restricted merged data; no weights), so a GEO-only attempt would be a new analysis, not a 1:1 reproduction, and could spuriously 'match'. Per the 80/20 + no-fabrication rules I recorded an honest drop rather than manufacture agreement. NOT ATTEMPTED: downloading/processing GEO, retraining a surrogate model, OCR of the figure gene list, requesting FUSCC data. NEUTRAL FABRICATION FLAG: reported metrics (92.7% acc, AUCs 0.87-0.98) are not independently derivable from any shipped/public artifact and are currently un-auditable end-to-end.
These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.
Assessment versions
Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.
-
v1 current initial assessmentassessed: 2026-06-14 ⛓ 7cb3d9fcaef7
✎ I am an author of this paper
Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.
Provenance — full disclosure
When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.
- Reproduced
- 2026-06-14
- Rubric version
- v1.0
- Assessed by
-
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15no human curator yet
- Last updated
- 2026-08-05
Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.
Deep full-text extraction
Model: opusBecause metabolic status is closely tied to tumor differentiation, the authors test whether integrating multiomic (metabolomic, genomic, transcriptomic) data with deep residual network (ResNet) models can accurately and interpretably classify thyroid cancer differentiation states and reveal the metabolic reprogramming underlying dedifferentiation.
- ★ A ResNet classifier built on a 10-gene metabolic signature distinguishes all thyroid cancer differentiation states with ~92.7% average accuracy. resource
- ★ A ResNet model on a 10-metabolite signature achieves a mean AUC of 0.98 in the FUSCC cohort and effectively separates PDTC from ATC. resource
- ★ Metabolic reprogramming, particularly lipid and glycolytic shifts, underpins thyroid cancer dedifferentiation and progression. mechanism
- ★ Differentiation status (pathology type) is the most significant independent prognostic factor for overall survival among clinicopathologic variables. finding
- ★ TERT C250T mutation is significantly associated with RAI refractoriness, ATC occurrence, and altered lipid metabolism. finding
- SHAP interpretability analysis confirms the biological relevance of the metabolic signatures driving differentiation states. method
- ★ NADPH and NADH abundance increase in WDTC but markedly decrease in DDTC, with a significant rise in the NAD+/NADH ratio in DDTC. finding
- This is the first comprehensive multiomic investigation spanning the full spectrum of follicular epithelial-derived thyroid carcinomas. finding
| Assay | System | Perturbation | Readout | Platform |
|---|---|---|---|---|
| Untargeted metabolomics/lipidomics | FUSCC thyroid tumors (PTC, FTC, PDTC, ATC) and matched normal tissue | none | 512 annotated polar metabolites and lipids abundance | — |
| Whole-exome sequencing (WES) / Sanger sequencing | FUSCC thyroid cancer cohort (158 patients) | none | gene mutations, mutation burden, gene fusions | — |
| Bulk RNA-seq (transcriptomics) | FUSCC thyroid carcinomas plus integrated GEO cohorts (n=453) | none | differential gene expression, metabolic gene signature | — |
| Single-cell RNA sequencing (analysis of public data) | Epithelial/tumor cells from untreated PTC and ATC samples, GEO GSE193581 | none (immunotherapy-treated patients excluded) | metabolic pathway activity (glycolysis, oxidative phosphorylation) | — |
| ResNet deep learning classification (10-gene / 10-metabolite models) | Aggregated GEO + FUSCC datasets | none | classification accuracy/AUC of differentiation states; SHAP feature importance | — |
| Survival analysis (Cox regression, Kaplan-Meier) | FUSCC thyroid cancer cohort | none | overall survival hazard ratios and prognostic factors | — |
| ssGSEA pathway enrichment | FUSCC thyroid carcinoma samples | none | normalized enrichment scores across differentiation states | Reactome database |
- – 10-gene metabolic signature ResNet distinguished all differentiation states 92.7% average accuracy
- – 10-metabolite ResNet model separated PDTC from ATC in FUSCC cohort mean AUC 0.98
- – Differentiation status had the most significant prognostic impact in multivariate Cox analysis; ATC patients had worst survival p < 0.0001 (KM log-rank)
- – NAD+/NADH ratio significantly increased in DDTC while NAD+/NADPH ratio decreased across subtypes
- ▼ Thyroxine biosynthesis pathway NES highest in FTC then PTC, lowest in PDTC/ATC, tracking dedifferentiation
- – TERT C250T mutation associated with RAI refractoriness and ATC occurrence p < 0.001 to p < 0.00001
- ▲ AC026191.1_SRGAP3 fusion positively related to NADH abundance q < 0.05
- ▲ PC(20:1_18:1) and PC(18:1_18:1) levels correlated with age, DM, ETE, and increased LNM number FDR < 0.05
- other AUC = 0.98 (mean, 10-metabolite model) (PDTC vs ATC separation, FUSCC cohort)
- other 92.7% average accuracy (10-gene signature ResNet across differentiation states)
- count 158 patients (PTC=104, FTC=24, PDTC=19, ATC=11); 215 samples (FUSCC cohort recruited)
- count n = 453 (aggregated GEO + FUSCC datasets for ResNet classifier)
- count 512 polar metabolites and lipids annotated (metabolomic profiling)
- mean 45.92 (±16.25) years (cohort mean age)
- mean 3.52 (±1.97) cm (average tumor size)
- pvalue p < 0.0001 (Kaplan-Meier survival difference WDTC vs DDTC (log-rank))
Statistical methods review
Model: sonnetA neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.
This retrospective multiomic study enrolled 158 thyroid cancer patients from a single centre (PTC n=104, FTC n=24, PDTC n=19, ATC n=11) with 57 matched normal tissues, combining untargeted metabolomics, whole-exome sequencing, and transcriptomics. Differential analyses relied on non-parametric tests (Kruskal-Wallis, Mann-Whitney U) and Welch's t-tests, with Benjamini-Hochberg FDR correction applied to metabolite-phenotype comparisons, while survival was assessed by Cox regression and Kaplan-Meier / log-rank methods. The primary analytic outputs were two deep residual network (ResNet) classifiers: a 10-gene transcriptomic model validated by cross-validation in n=453 (FUSCC + GEO) reaching mean accuracy 92.7%, and a 10-metabolite model achieving mean AUC 0.98 in the FUSCC cohort, both interpreted via SHAP.
| Test | Applied to | n | Assumptions |
|---|---|---|---|
| Kruskal-Wallis rank sum test | Comparison of continuous clinical variables (age, tumor size, LNM.No) across four cancer subtypes — Table 1 | 158 | not stated |
| Fisher's exact test | Comparison of categorical clinical variables (gender, ETE, ENE, T/N/M stage, TNM stage) across four subtypes — Table 1 | 158 | not stated |
| ANOVA | Filtering associations of TERT C250T mutation with RAI refractoriness and ATC occurrence — Fig 2b | — | not stated |
| Welch's t-test | Correlations between metabolites and genomic mutations (Fig 2f) and between metabolites and gene fusions (Fig 2g); q < 0.05 | — | not stated |
| Benjamini-Hochberg-corrected Mann-Whitney U test | Associations of metabolites with clinical phenotypes (ETE, LNM.No — Fig 3e) and multigroup differential metabolite analysis across differentiation states (Fig 4b left panel); q < 0.05, |Log2 FC| > 2 | 158 patients; 512 metabolites profiled | not stated |
| Spearman rank correlation | Correlation of specific lipid metabolites with clinical phenotypes (age, DM, ETE, LNM.No) — Fig 3f; FDR < 0.05 | 158 | not stated |
| Univariate Cox proportional hazards regression | Association of clinical and pathological factors with overall survival — Fig 3a | 158 | not stated |
| Multivariate Cox proportional hazards regression | Independent prognostic impact of differentiation status among all clinical factors — Fig 3b | 158 | not stated |
| Log-rank test | Comparison of Kaplan-Meier overall survival curves across differentiation state groups — Fig 3c | 158 | not stated |
| DESeq2 Wald test | Differential gene expression analysis across differentiation states — Fig 4b right panel | 158 | not stated |
| Sparse partial least-squares discriminant analysis (sPLS-DA) and orthogonal partial least-squares discriminant analysis (OPLS-DA) | Metabolic discrimination between WDTC and DDTC, and pairwise group comparisons — Supplementary Figs 6, 7 | 158 | not stated |
| Single-sample gene set enrichment analysis (ssGSEA) | Per-sample pathway enrichment scoring using Reactome database; top 15 pathways by NES visualized — Fig 4a | 158 | na |
-
The 10-metabolite ResNet classifier was evaluated by cross-validation AUC (mean 0.98) within the FUSCC cohort only, with no independent external metabolomic test set↳ Could also: An independent external validation cohort held out entirely from model development could also be used to estimate generalisability — Internal cross-validation tends to yield optimistic performance estimates, particularly when per-class n is small (PDTC n=19, ATC n=11); an external cohort provides an estimate of how the model performs on genuinely unseen data from a different setting
-
Multiple separate testing frameworks were used for metabolite comparisons: BH-corrected Mann-Whitney U for metabolite-phenotype analyses and Welch's t-test (also with q correction) for mutation-metabolite and fusion-metabolite analyses↳ Could also: A single unified non-parametric framework—e.g., Kruskal-Wallis for multi-group comparisons followed by Dunn's post-hoc test with BH correction—could also cover all multi-group metabolite comparisons consistently — Standardising on one approach across all metabolite comparisons simplifies interpretation and ensures a coherent multiplicity-correction strategy; mixing parametric and non-parametric methods across similar data types can complicate cross-result comparison
-
Survival analysis used Kaplan-Meier / log-rank tests and Cox proportional hazards regression for overall survival↳ Could also: A competing-risks analysis using the Fine-Gray subdistribution hazard model could also be applied, treating non-thyroid-cancer deaths as competing events — In a cohort spanning indolent (PTC, high censoring 94%) to rapidly fatal (ATC, low censoring 18%) cancers, non-disease deaths may be non-trivial in the better-prognosis groups; the Fine-Gray model directly estimates the cumulative incidence of thyroid-cancer death in the presence of competing risks, which differs from the standard Kaplan-Meier estimator when competing events are present
-
Pathway enrichment was performed using ssGSEA, producing per-sample enrichment scores that were then averaged per group↳ Could also: Bulk preranked GSEA (using a between-group differential expression statistic as the ranking metric) or over-representation analysis (ORA via Fisher's exact test on a threshold-selected gene list) could also be used — Bulk GSEA uses the full ranked gene list and is well-powered to detect coordinated pathway shifts between defined groups without requiring a significance threshold; ORA is simpler and more widely familiar, making results easier to communicate to a broad audience
-
Continuous clinical variables (age, tumour size) were summarised as mean ± SD, while the table footnote indicates Median (IQR) and the group comparisons used the Kruskal-Wallis test↳ Could also: Median (IQR) could also be reported consistently for all continuous variables, aligned with the non-parametric test used — Median (IQR) is generally more robust than mean (SD) when distributions are skewed or group sizes are small, and pairing the summary statistic with the corresponding test (median with Kruskal-Wallis) aids interpretability; consistency between the table footnote and the reported values also reduces ambiguity for readers
-
The ResNet classifier performance was compared to the traditional thyroid differentiation score (TDS) descriptively, without a formal statistical test of the performance difference↳ Could also: DeLong's method for comparing AUCs, or bootstrap-based confidence intervals around the difference in accuracy or AUC between the ResNet model and TDS, could also be reported — Formal comparison tests or CIs around classifier performance differences allow readers to assess whether observed gaps are consistent with sampling variability or reflect a reliable performance improvement, which is particularly important when per-class n is small
Citation network
Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.
No assessed neighbours yet — the network grows as more papers are assessed.
Data lineage
The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.
Downstream reach in the literature
222 downstream papers · 6 datasetsHow widely the datasets deposited by this paper are reused across the whole literature (Europe PMC), beyond our assessed set. This is a factual dependency map — reusing a public dataset is normal, good science. It is not a judgement on the downstream papers; the only verdict here is this paper's own, with its cited rationale.
- Characterizing dedifferentiation of thyroid cancer b... 2021 · 136 cites
- Large Scale Gene Expression Meta-Analysis Reveals Ti... 2016 · 103 cites
- Immune Cell Confrontation in the Papillary Thyroid C... 2020 · 91 cites
- METTL3-mediated m6A modification of STEAP2 mRNA inhi... 2022 · 71 cites
- Senescent thyrocytes and thyroid tumor cells induce... 2019 · 67 cites
- miR30a inhibits LOX expression and anaplastic thyroi... 2015 · 65 cites
- Drug target prediction and repositioning using an in... 2013 · 138 cites
- Epithelial tumor suppressor ELF3 is a lineage-specif... 2019 · 57 cites
- Cell Cycle M-Phase Genes Are Highly Upregulated in A... 2017 · 54 cites
- Cancer Associated Fibroblasts and Senescent Thyroid... 2020 · 49 cites
- Cancer-Associated Fibroblasts Positively Correlate w... 2021 · 47 cites
- The molecular and gene/miRNA expression profiles of... 2020 · 42 cites
- Aberrant lipid metabolism in anaplastic thyroid carc... 2015 · 137 cites
- Characterizing dedifferentiation of thyroid cancer b... 2021 · 136 cites
- Large Scale Gene Expression Meta-Analysis Reveals Ti... 2016 · 103 cites
- Cell Cycle M-Phase Genes Are Highly Upregulated in A... 2017 · 54 cites
- Cancer Associated Fibroblasts and Senescent Thyroid... 2020 · 49 cites
- Cancer-Associated Fibroblasts Positively Correlate w... 2021 · 47 cites
- Immune Cell Confrontation in the Papillary Thyroid C... 2020 · 91 cites
- Cancer Associated Fibroblasts and Senescent Thyroid... 2020 · 49 cites
- Cancer-Associated Fibroblasts Positively Correlate w... 2021 · 47 cites
- The molecular and gene/miRNA expression profiles of... 2020 · 42 cites
- A distinct tumor microenvironment makes anaplastic t... 2024 · 42 cites
- Identification of Potential Biomarkers for Thyroid C... 2020 · 38 cites
- Genomic and transcriptomic hallmarks of poorly diffe... 2016 · 925 cites
- Senescent thyrocytes and thyroid tumor cells induce... 2019 · 67 cites
- Hgf/Met activation mediates resistance to BRAF inhib... 2018 · 60 cites
- Cancer Associated Fibroblasts and Senescent Thyroid... 2020 · 49 cites
- Oncogenic BRAF disrupts thyroid morphogenesis and fu... 2017 · 48 cites
- Cancer-Associated Fibroblasts Positively Correlate w... 2021 · 47 cites
What was reproduced
The exact results taken into scope, with each reported value next to the value our attempt produced.
Scope — pmid-40993300
Paper: Developing a thyroid cancer differentiation state classification system using deep residual networks and metabolic signature profiling. npj Digital Medicine 2025. PMID 40993300 · PMCID PMC12460824 · DOI 10.1038/s41746-025-01927-1.
Link correction (BRIEF text-mining false positives)
The BRIEF was instantiated with two incorrect auto-mined links:
- BRIEF
Code: github.com/STAR-Fusion/STAR-Fusion→ WRONG. "STAR-Fusion" was text-mined from the repo README's Tools list (Arriba/STAR-Fusion are cited as upstream fusion-calling tools). The paper's actual Code availability statement points to https://github.com/Aceracede3000/Aceracede3000 (a GitHub profile repo, default branchmain, HEAD3c0eace). - BRIEF
Data: geo:GSE193581→ WRONG. GSE193581 ("Anaplastic Transformation Model … Single Cell Lineage") is thyroid-related but is not cited in this paper's Data availability. The paper's actual public accessions are the five external transcriptomic cohorts GSE29265, GSE33630, GSE53157, GSE65144, GSE76039 (note: Methods text once mistypes GSE53157 as "GSE53167").
Pipeline-derived results (candidate in-scope)
| id | result | reported | pipeline | location |
|---|---|---|---|---|
| C1 | 10-gene (10MG) ResNet classifier — mean CV accuracy | 92.7% | PyTorch ResNet1D + StratifiedKFold (10MGModel.py) |
Abstract; Results; Fig 7 |
| C2 | 10MG test-cohort per-class AUC | Normal 0.92, FTC 0.87, PTC 0.95, PDTC 0.92, ATC 0.96 (range 0.87–0.96) | same | Fig 7f; Results |
| C3 | 10MG validation-cohort mean AUROC | ~0.98 | same | Fig 7c |
| C4 | 10MG external-GEO validation accuracy | 92.7% | same model applied to public GEO cohorts | Results/Discussion |
| C5 | 10-metabolite (10M) ResNet classifier — mean AUC (PDTC vs ATC, FUSCC) | 0.98 | PyTorch ResNet1D (512METASHAP.py/10MModel.py) |
Results; Fig 8 |
| C6 | SHAP-derived 10-gene & 10-metabolite signatures | gene/metabolite rankings | SHAP DeepExplainer | Fig 7e/g, Fig 8b |
Out of scope (not a pipeline reproduction here)
- Untargeted LC-MS metabolomics acquisition (wet-lab).
- WES variant/fusion calling (Trimmomatic/BWA/SAMtools/VarScan2/ANNOVAR/Arriba/ STAR-Fusion/annoFuse) — upstream feature generation, on restricted raw data.
- scRNA-seq Seurat/Harmony mapping of metabolic reprogramming — descriptive, on third-party GEO single-cell data, not the headline classifier.
Reproducibility verdict: DROP (not attempted on «our HPC» — nothing runnable)
Every headline classifier result (C1–C5) is blocked by a combination of:
-
data_restricted(root cause). All ResNet models are trained/tested on the merged GEO+FUSCC matrix (n = 453). The FUSCC raw sequencing + untargeted metabolomics (158 tumors + 57 normal, participant-level, re-identification risk) is explicitly on-request only via the FUSCC Data-Access Committee («email»; ≤30 working-day review). It is the training/test substrate for the 10MG model and the entire substrate for the 10M model. No public copy. -
docs_insufficient(code). The repo is a single-author GitHub profile repo. It ships only training scripts — at no commit does it contain:- the input matrices the scripts hard-code (
202503ResNet.csv,20240907PTCPDTCATC1.csv), - any trained model weights (no
.pt/.pth), so predictions cannot be reproduced without retraining on the (restricted) data, - any instructions/pipeline to rebuild the input matrix from the five named
public GEO accessions (probe→gene mapping, batch correction, label coding,
the train=1/test=2 cohort column, the 2:1 split seed).
models.py(theResNet1Ddefinition the scripts import) is absent from HEAD and only recoverable from git history (commitc867ba0b97, 2024-09-26).
- the input matrices the scripts hard-code (
-
Signature not enumerable from public artifacts. The final 10 genes are shown only in Fig 7
Assessments & scoring basis
Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.
An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.
Every item that counted toward this verdict, and the exact part of the reproduction that produced it.
This is a clean data-unavailable + docs-insufficient DROP, not a discrepancy case. Every headline model (10-gene '10MG' and 10-metabolite '10M' ResNet1D classifiers, claims C1–C5: 92.7% accuracy, per-class AUC 0.87–0.96, ~0.98 AUROC, 0.98 metabolite AUC) is trained on a merged GEO+FUSCC matrix (n=453) whose FUSCC arm is restricted/on-request only, and the repo ships no input data, no trained weights, and no input-build pipeline (models.py only in git history; the final 10 genes only in Fig 7). The blocker sits on the authors'/data-availability side — restricted data plus an incomplete deposit — so the values are not independently derivable from public artifacts, but there is no positive evidence of fabrication (the agent's flag is explicitly neutral). Per the fairness principle, restriction drives q1/q2 to red while q5/q7/q8 stay yellow: un-auditable, explainable, no demonstrated discrepancy.
Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.
Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.
🚩 Report an error in this record
Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.
Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.
Reproduction footprint
claude-opus-4-8Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.