Corpus 1,272 assessed · 1,173 scored · 643 reproduced ≥75 · 168 flagged ·∅ 74.1/100
← New search

Developing a thyroid cancer differentiation state classification system using deep residual networks and metabolic signature profiling.

NPJ Digit Med · 2025
L1 No data access 3/4
⚑ Flagged for review — a reproduced result did not match the reported value

Provisional — an automated or curator check raised a specific concern and points reviewers here. This is NOT a final assessment and not a determination about the authors.

Why this verdict

Part of the results reproduced; minor but material deviations remained.

Reproduced on the brainbox compute brainarbeit.com
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Input / endpoint not comparable 1:1
+1 pts
From: Q1 · Data identity 🔴
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +7
✓ What held up
  • Nothing in this column.
What did not (or only partly)
  • 🔴Could not use the authors’ exact input data
  • 🔴Reported values were only indirectly comparable
  • 🟡A deviation arose in the data or preprocessing
  • 🔴A deviation was attributed to the published material
  • 🟡Reported values were not (fully) derivable from the shared data
  • 🟡The deviation was non-trivial in magnitude
  • 🟡The central claim did not (fully) hold under reproduction
  • 🟡Overall, the reproduction showed a material discrepancy
No data access Data access not granted

This paper has a computational component, but its primary data is legally or ethically access-restricted — identifiable patient cohorts, rare-disease genomes, or controlled-access biobanks that cannot be openly shared. The reproduction therefore could not be attempted. That is a neutral verdict: it does not mean the result is wrong or that the authors fell short — only that, for legitimate privacy reasons, it cannot be independently checked from public data. We deliberately do NOT assign a 0–100 score here, because a low number would wrongly read as a failed reproduction.

Reproduction agent’s raw note

DROP. The paper (npj Digit Med 2025, thyroid-cancer ResNet differentiation classifier) is described in moderate detail but is NOT reproducible from public artifacts. BRIEF links were both text-mining false positives: real code = github.com/Aceracede3000/Aceracede3000 (a single-author GitHub profile repo, not STAR-Fusion); real public data = GEO GSE29265/GSE33630/GSE53157/GSE65144/GSE76039 (not GSE193581). Every headline model (10-gene '10MG' and 10-metabolite '10M' ResNet1D classifiers) is trained/tested on a merged GEO+FUSCC matrix (n=453); the FUSCC raw sequencing + untargeted metabolomics (158 tumors + 57 normal, participant-level) is RESTRICTED / on-request only via the FUSCC Data-Access Committee -> data_restricted. The repo additionally ships NO input data at any commit, NO trained weights, and no pipeline to rebuild the input matrix from the GEO accessions; models.py (ResNet1D) survives only in git history; the final 10 genes are shown only in Fig 7, not text-enumerated -> docs_insufficient. The 5 public GEO cohorts are downloadable but cannot reproduce the reported numbers (model trained on restricted merged data; no weights), so a GEO-only attempt would be a new analysis, not a 1:1 reproduction, and could spuriously 'match'. Per the 80/20 + no-fabrication rules I recorded an honest drop rather than manufacture agreement. NOT ATTEMPTED: downloading/processing GEO, retraining a surrogate model, OCR of the figure gene list, requesting FUSCC data. NEUTRAL FABRICATION FLAG: reported metrics (92.7% acc, AUCs 0.87-0.98) are not independently derivable from any shipped/public artifact and are currently un-auditable end-to-end.

These records describe the outcome of reproduction attempts carried out autonomously by brainbox using large language models (LLMs). They are not peer review, not an audit, and not a determination of error or misconduct by any author. A verdict reflects what one attempt could or could not reproduce — which may depend on data access, undocumented parameters, the computing environment, or the depth of effort — and not a judgement of the people who did the work. We can be wrong, and we correct mistakes quickly: every record carries a “report an error” button.

Assessment versions

Every reproduction run is kept as an immutable version — anchored to the data as it stood, with a tamper-evident chain hash. A rerun (e.g. after an author updates a deposit) adds a new version; the previous one stays on record.

  1. v1 current initial assessment
    assessed: 2026-06-14 ⛓ 7cb3d9fcaef7
✎ I am an author of this paper

Updated or fixed a deposit, or is there an erratum? Ask us to re-run the metrics. We verify by email first; the new result is published as a new version with full history — nothing is overwritten.

Reason for the rerun

We email you a confirmation link first. The rerun is an objective re-measurement — it cannot change the verdict in your favour, only ask us to look again.

Provenance — full disclosure

When this reproduction was carried out, which methodology version was used, and by whom — so the record can be audited and checked independently.

Reproduced
2026-06-14
Rubric version
v1.0
Assessed by
🤖 AI curator · claude (ai-curator room) · v1.0 · run #1 2026-06-15
no human curator yet
Last updated
2026-08-05

Provisional, curator- or AI-assessed, and independently checkable. A reproduction outcome states what one attempt could reproduce — not a judgement of the authors.

Deep full-text extraction

Model: opus
Founding hypothesis

Because metabolic status is closely tied to tumor differentiation, the authors test whether integrating multiomic (metabolomic, genomic, transcriptomic) data with deep residual network (ResNet) models can accurately and interpretably classify thyroid cancer differentiation states and reveal the metabolic reprogramming underlying dedifferentiation.

Core claims
  • A ResNet classifier built on a 10-gene metabolic signature distinguishes all thyroid cancer differentiation states with ~92.7% average accuracy. resource
  • A ResNet model on a 10-metabolite signature achieves a mean AUC of 0.98 in the FUSCC cohort and effectively separates PDTC from ATC. resource
  • Metabolic reprogramming, particularly lipid and glycolytic shifts, underpins thyroid cancer dedifferentiation and progression. mechanism
  • Differentiation status (pathology type) is the most significant independent prognostic factor for overall survival among clinicopathologic variables. finding
  • TERT C250T mutation is significantly associated with RAI refractoriness, ATC occurrence, and altered lipid metabolism. finding
  • SHAP interpretability analysis confirms the biological relevance of the metabolic signatures driving differentiation states. method
  • NADPH and NADH abundance increase in WDTC but markedly decrease in DDTC, with a significant rise in the NAD+/NADH ratio in DDTC. finding
  • This is the first comprehensive multiomic investigation spanning the full spectrum of follicular epithelial-derived thyroid carcinomas. finding
Experimental setups
Assay System Perturbation Readout Platform
Untargeted metabolomics/lipidomics FUSCC thyroid tumors (PTC, FTC, PDTC, ATC) and matched normal tissue none 512 annotated polar metabolites and lipids abundance
Whole-exome sequencing (WES) / Sanger sequencing FUSCC thyroid cancer cohort (158 patients) none gene mutations, mutation burden, gene fusions
Bulk RNA-seq (transcriptomics) FUSCC thyroid carcinomas plus integrated GEO cohorts (n=453) none differential gene expression, metabolic gene signature
Single-cell RNA sequencing (analysis of public data) Epithelial/tumor cells from untreated PTC and ATC samples, GEO GSE193581 none (immunotherapy-treated patients excluded) metabolic pathway activity (glycolysis, oxidative phosphorylation)
ResNet deep learning classification (10-gene / 10-metabolite models) Aggregated GEO + FUSCC datasets none classification accuracy/AUC of differentiation states; SHAP feature importance
Survival analysis (Cox regression, Kaplan-Meier) FUSCC thyroid cancer cohort none overall survival hazard ratios and prognostic factors
ssGSEA pathway enrichment FUSCC thyroid carcinoma samples none normalized enrichment scores across differentiation states Reactome database
Key results
  • 10-gene metabolic signature ResNet distinguished all differentiation states 92.7% average accuracy
  • 10-metabolite ResNet model separated PDTC from ATC in FUSCC cohort mean AUC 0.98
  • Differentiation status had the most significant prognostic impact in multivariate Cox analysis; ATC patients had worst survival p < 0.0001 (KM log-rank)
  • NAD+/NADH ratio significantly increased in DDTC while NAD+/NADPH ratio decreased across subtypes
  • Thyroxine biosynthesis pathway NES highest in FTC then PTC, lowest in PDTC/ATC, tracking dedifferentiation
  • TERT C250T mutation associated with RAI refractoriness and ATC occurrence p < 0.001 to p < 0.00001
  • AC026191.1_SRGAP3 fusion positively related to NADH abundance q < 0.05
  • PC(20:1_18:1) and PC(18:1_18:1) levels correlated with age, DM, ETE, and increased LNM number FDR < 0.05
Key statistics
  • other AUC = 0.98 (mean, 10-metabolite model) (PDTC vs ATC separation, FUSCC cohort)
  • other 92.7% average accuracy (10-gene signature ResNet across differentiation states)
  • count 158 patients (PTC=104, FTC=24, PDTC=19, ATC=11); 215 samples (FUSCC cohort recruited)
  • count n = 453 (aggregated GEO + FUSCC datasets for ResNet classifier)
  • count 512 polar metabolites and lipids annotated (metabolomic profiling)
  • mean 45.92 (±16.25) years (cohort mean age)
  • mean 3.52 (±1.97) cm (average tumor size)
  • pvalue p < 0.0001 (Kaplan-Meier survival difference WDTC vs DDTC (log-rank))

Statistical methods review

Model: sonnet

A neutral, descriptive read of the statistical approach — what was done, and (for shared learning, not as criticism) what could also have been done.

This retrospective multiomic study enrolled 158 thyroid cancer patients from a single centre (PTC n=104, FTC n=24, PDTC n=19, ATC n=11) with 57 matched normal tissues, combining untargeted metabolomics, whole-exome sequencing, and transcriptomics. Differential analyses relied on non-parametric tests (Kruskal-Wallis, Mann-Whitney U) and Welch's t-tests, with Benjamini-Hochberg FDR correction applied to metabolite-phenotype comparisons, while survival was assessed by Cox regression and Kaplan-Meier / log-rank methods. The primary analytic outputs were two deep residual network (ResNet) classifiers: a 10-gene transcriptomic model validated by cross-validation in n=453 (FUSCC + GEO) reaching mean accuracy 92.7%, and a 10-metabolite model achieving mean AUC 0.98 in the FUSCC cohort, both interpreted via SHAP.

Replicationbiological Sample size158 tumour samples and 57 matched normal tissues from FUSCC; GEO transcriptomic cohorts aggregated to n=453 for the 10-gene classifier; per-subtype n stated (PTC 104, FTC 24, PDTC 19, ATC 11); no a priori power calculation mentioned GroupsPTC vs FTC vs PDTC vs ATC (four subtypes); WDTC vs DDTC (binary aggregation); tumour vs matched normal tissue Pairingmixed Randomization/blindingnot stated Dispersionmixed Exact p-valuesno Effect sizesyes Confidence intervalsyes Multiplicity correctionBenjamini-Hochberg FDR
Statistical tests used
Test Applied to n Assumptions
Kruskal-Wallis rank sum test Comparison of continuous clinical variables (age, tumor size, LNM.No) across four cancer subtypes — Table 1 158 not stated
Fisher's exact test Comparison of categorical clinical variables (gender, ETE, ENE, T/N/M stage, TNM stage) across four subtypes — Table 1 158 not stated
ANOVA Filtering associations of TERT C250T mutation with RAI refractoriness and ATC occurrence — Fig 2b not stated
Welch's t-test Correlations between metabolites and genomic mutations (Fig 2f) and between metabolites and gene fusions (Fig 2g); q < 0.05 not stated
Benjamini-Hochberg-corrected Mann-Whitney U test Associations of metabolites with clinical phenotypes (ETE, LNM.No — Fig 3e) and multigroup differential metabolite analysis across differentiation states (Fig 4b left panel); q < 0.05, |Log2 FC| > 2 158 patients; 512 metabolites profiled not stated
Spearman rank correlation Correlation of specific lipid metabolites with clinical phenotypes (age, DM, ETE, LNM.No) — Fig 3f; FDR < 0.05 158 not stated
Univariate Cox proportional hazards regression Association of clinical and pathological factors with overall survival — Fig 3a 158 not stated
Multivariate Cox proportional hazards regression Independent prognostic impact of differentiation status among all clinical factors — Fig 3b 158 not stated
Log-rank test Comparison of Kaplan-Meier overall survival curves across differentiation state groups — Fig 3c 158 not stated
DESeq2 Wald test Differential gene expression analysis across differentiation states — Fig 4b right panel 158 not stated
Sparse partial least-squares discriminant analysis (sPLS-DA) and orthogonal partial least-squares discriminant analysis (OPLS-DA) Metabolic discrimination between WDTC and DDTC, and pairwise group comparisons — Supplementary Figs 6, 7 158 not stated
Single-sample gene set enrichment analysis (ssGSEA) Per-sample pathway enrichment scoring using Reactome database; top 15 pathways by NES visualized — Fig 4a 158 na
Approaches that could also have been used
  • The 10-metabolite ResNet classifier was evaluated by cross-validation AUC (mean 0.98) within the FUSCC cohort only, with no independent external metabolomic test set
    Could also: An independent external validation cohort held out entirely from model development could also be used to estimate generalisability — Internal cross-validation tends to yield optimistic performance estimates, particularly when per-class n is small (PDTC n=19, ATC n=11); an external cohort provides an estimate of how the model performs on genuinely unseen data from a different setting
  • Multiple separate testing frameworks were used for metabolite comparisons: BH-corrected Mann-Whitney U for metabolite-phenotype analyses and Welch's t-test (also with q correction) for mutation-metabolite and fusion-metabolite analyses
    Could also: A single unified non-parametric framework—e.g., Kruskal-Wallis for multi-group comparisons followed by Dunn's post-hoc test with BH correction—could also cover all multi-group metabolite comparisons consistently — Standardising on one approach across all metabolite comparisons simplifies interpretation and ensures a coherent multiplicity-correction strategy; mixing parametric and non-parametric methods across similar data types can complicate cross-result comparison
  • Survival analysis used Kaplan-Meier / log-rank tests and Cox proportional hazards regression for overall survival
    Could also: A competing-risks analysis using the Fine-Gray subdistribution hazard model could also be applied, treating non-thyroid-cancer deaths as competing events — In a cohort spanning indolent (PTC, high censoring 94%) to rapidly fatal (ATC, low censoring 18%) cancers, non-disease deaths may be non-trivial in the better-prognosis groups; the Fine-Gray model directly estimates the cumulative incidence of thyroid-cancer death in the presence of competing risks, which differs from the standard Kaplan-Meier estimator when competing events are present
  • Pathway enrichment was performed using ssGSEA, producing per-sample enrichment scores that were then averaged per group
    Could also: Bulk preranked GSEA (using a between-group differential expression statistic as the ranking metric) or over-representation analysis (ORA via Fisher's exact test on a threshold-selected gene list) could also be used — Bulk GSEA uses the full ranked gene list and is well-powered to detect coordinated pathway shifts between defined groups without requiring a significance threshold; ORA is simpler and more widely familiar, making results easier to communicate to a broad audience
  • Continuous clinical variables (age, tumour size) were summarised as mean ± SD, while the table footnote indicates Median (IQR) and the group comparisons used the Kruskal-Wallis test
    Could also: Median (IQR) could also be reported consistently for all continuous variables, aligned with the non-parametric test used — Median (IQR) is generally more robust than mean (SD) when distributions are skewed or group sizes are small, and pairing the summary statistic with the corresponding test (median with Kruskal-Wallis) aids interpretability; consistency between the table footnote and the reported values also reduces ambiguity for readers
  • The ResNet classifier performance was compared to the traditional thyroid differentiation score (TDS) descriptively, without a formal statistical test of the performance difference
    Could also: DeLong's method for comparing AUCs, or bootstrap-based confidence intervals around the difference in accuracy or AUC between the ResNet model and TDS, could also be reported — Formal comparison tests or CIs around classifier performance differences allow readers to assess whether observed gaps are consistent with sampling variability or reflect a reliable performance improvement, which is particularly important when per-class n is small
Software: DESeq2 · SHAP (Shapley Additive Explanations) · ResNet (deep residual network; framework unspecified) · ssGSEA with Reactome database · sPLS-DA / OPLS-DA (R package unspecified)

Citation network

Where this publication sits in the reproducibility-weighted citation graph — what it is built on, and what is built on it. Citation data from OpenAlex.

Citations
1
Impact: low
Foundation confidence
None of its references are in our reproducibility record yet — its foundation cannot be assessed.
Topics

No assessed neighbours yet — the network grows as more papers are assessed.

Data lineage

The datasets this paper uses (text-mined from the full text via Europe PMC), and which other assessed papers stand on the same data. A shared dataset is a factual link — not a judgement.

GSE33630 GEO in Methods (http://purl.org/orb/Methods)
also used by 1 paper:
GSE65144 GEO in Methods (http://purl.org/orb/Methods)
also used by 1 paper:
GSE29265 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE53157 GEO in Results (http://purl.org/orb/Results)
no other assessed paper uses this yet
GSE53167 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet
GSE76039 GEO in Methods (http://purl.org/orb/Methods)
no other assessed paper uses this yet

Downstream reach in the literature

222 downstream papers · 6 datasets

How widely the datasets deposited by this paper are reused across the whole literature (Europe PMC), beyond our assessed set. This is a factual dependency map — reusing a public dataset is normal, good science. It is not a judgement on the downstream papers; the only verdict here is this paper's own, with its cited rationale.

This paper is currently under reproducibility review (see the verdict above). The map below shows where the data in question has propagated — so reuse can be traced, not so the downstream work is presumed affected.
GSE53167 GEO reused by 1 papers in the literature

What was reproduced

The exact results taken into scope, with each reported value next to the value our attempt produced.

Scope — pmid-40993300

Paper: Developing a thyroid cancer differentiation state classification system using deep residual networks and metabolic signature profiling. npj Digital Medicine 2025. PMID 40993300 · PMCID PMC12460824 · DOI 10.1038/s41746-025-01927-1.

Link correction (BRIEF text-mining false positives)

The BRIEF was instantiated with two incorrect auto-mined links:

  • BRIEF Code: github.com/STAR-Fusion/STAR-FusionWRONG. "STAR-Fusion" was text-mined from the repo README's Tools list (Arriba/STAR-Fusion are cited as upstream fusion-calling tools). The paper's actual Code availability statement points to https://github.com/Aceracede3000/Aceracede3000 (a GitHub profile repo, default branch main, HEAD 3c0eace).
  • BRIEF Data: geo:GSE193581WRONG. GSE193581 ("Anaplastic Transformation Model … Single Cell Lineage") is thyroid-related but is not cited in this paper's Data availability. The paper's actual public accessions are the five external transcriptomic cohorts GSE29265, GSE33630, GSE53157, GSE65144, GSE76039 (note: Methods text once mistypes GSE53157 as "GSE53167").

Pipeline-derived results (candidate in-scope)

id result reported pipeline location
C1 10-gene (10MG) ResNet classifier — mean CV accuracy 92.7% PyTorch ResNet1D + StratifiedKFold (10MGModel.py) Abstract; Results; Fig 7
C2 10MG test-cohort per-class AUC Normal 0.92, FTC 0.87, PTC 0.95, PDTC 0.92, ATC 0.96 (range 0.87–0.96) same Fig 7f; Results
C3 10MG validation-cohort mean AUROC ~0.98 same Fig 7c
C4 10MG external-GEO validation accuracy 92.7% same model applied to public GEO cohorts Results/Discussion
C5 10-metabolite (10M) ResNet classifier — mean AUC (PDTC vs ATC, FUSCC) 0.98 PyTorch ResNet1D (512METASHAP.py/10MModel.py) Results; Fig 8
C6 SHAP-derived 10-gene & 10-metabolite signatures gene/metabolite rankings SHAP DeepExplainer Fig 7e/g, Fig 8b

Out of scope (not a pipeline reproduction here)

  • Untargeted LC-MS metabolomics acquisition (wet-lab).
  • WES variant/fusion calling (Trimmomatic/BWA/SAMtools/VarScan2/ANNOVAR/Arriba/ STAR-Fusion/annoFuse) — upstream feature generation, on restricted raw data.
  • scRNA-seq Seurat/Harmony mapping of metabolic reprogramming — descriptive, on third-party GEO single-cell data, not the headline classifier.

Reproducibility verdict: DROP (not attempted on «our HPC» — nothing runnable)

Every headline classifier result (C1–C5) is blocked by a combination of:

  1. data_restricted (root cause). All ResNet models are trained/tested on the merged GEO+FUSCC matrix (n = 453). The FUSCC raw sequencing + untargeted metabolomics (158 tumors + 57 normal, participant-level, re-identification risk) is explicitly on-request only via the FUSCC Data-Access Committee («email»; ≤30 working-day review). It is the training/test substrate for the 10MG model and the entire substrate for the 10M model. No public copy.

  2. docs_insufficient (code). The repo is a single-author GitHub profile repo. It ships only training scripts — at no commit does it contain:

    • the input matrices the scripts hard-code (202503ResNet.csv, 20240907PTCPDTCATC1.csv),
    • any trained model weights (no .pt/.pth), so predictions cannot be reproduced without retraining on the (restricted) data,
    • any instructions/pipeline to rebuild the input matrix from the five named public GEO accessions (probe→gene mapping, batch correction, label coding, the train=1/test=2 cohort column, the 2:1 split seed). models.py (the ResNet1D definition the scripts import) is absent from HEAD and only recoverable from git history (commit c867ba0b97, 2024-09-26).
  3. Signature not enumerable from public artifacts. The final 10 genes are shown only in Fig 7

Figures / tables: Fig 7Fig 7fFig 7cFig 8Fig 8bFig 4b
C1
Reported
10MG mean CV accuracy 92.7%
Reproduced
not reproduced
did not match
C2
Reported
10MG per-class AUC 0.87-0.96 (N .92/FTC .87/PTC .95/PDTC .92/ATC .96)
Reproduced
not reproduced
did not match
C3
Reported
10MG validation mean AUROC ~0.98
Reproduced
not reproduced
did not match
C4
Reported
10MG external-GEO accuracy 92.7%
Reproduced
not reproduced
did not match
C5
Reported
10-metabolite mean AUC 0.98 (FUSCC)
Reproduced
not reproduced
did not match

Assessments & scoring basis

Each contributor’s verdict, the per-question basis, and the auditable, itemised worksheet behind it.

🤖 AI curator · claude (ai-curator room) · v1.0 L1 31/100

An automated assessment. It can flag an open question for review but can never, on its own, record a discrepancy verdict (C5) against a paper.

🔴1. Data identity
🔴2. Endpoint comparability
🟡3. Location of the main deviation
🔴4. Cause of the deviation
🟡5. Derivability / plausibility
🟡6. Severity of the deviation
🟡7. Core claim
🟡8. Severity of the miss (overall human judgment)
Scoring basis — itemised

Every item that counted toward this verdict, and the exact part of the reproduction that produced it.

Supporting (toward a concern)
Content-critical question only partially held
+2 pts
From: Q5 · Derivability / plausibility 🟡
Content-critical question only partially held
+2 pts
From: Q7 · Core claim 🟡
Content-critical question only partially held
+2 pts
From: Q8 · Severity of the miss (overall human judgment) 🟡
Minor / cosmetic deviation
+1 pts
From: Q3 · Location of the main deviation 🟡
Minor / cosmetic deviation
+1 pts
From: Q6 · Severity of the deviation 🟡
Input / endpoint not comparable 1:1
+1 pts
From: Q1 · Data identity 🔴
Concordant (toward reproduced)
Code + data deposited & functional
-2 pts
From: Data & code availability Available & functional
Total score +7

This is a clean data-unavailable + docs-insufficient DROP, not a discrepancy case. Every headline model (10-gene '10MG' and 10-metabolite '10M' ResNet1D classifiers, claims C1–C5: 92.7% accuracy, per-class AUC 0.87–0.96, ~0.98 AUROC, 0.98 metabolite AUC) is trained on a merged GEO+FUSCC matrix (n=453) whose FUSCC arm is restricted/on-request only, and the repo ships no input data, no trained weights, and no input-build pipeline (models.py only in git history; the final 10 genes only in Fig 7). The blocker sits on the authors'/data-availability side — restricted data plus an incomplete deposit — so the values are not independently derivable from public artifacts, but there is no positive evidence of fabrication (the agent's flag is explicitly neutral). Per the fairness principle, restriction drives q1/q2 to red while q5/q7/q8 stay yellow: un-auditable, explainable, no demonstrated discrepancy.

🤝
Reproduced automatically — and fairly

Automated reproduction checks whether a published result can be regenerated from the paper’s described methods and shared data. When something does not reproduce, that is not a claim of error or misconduct — most often it reflects under-described methods, software or environment differences, or gaps in data access, and some of the pre-print papers in the queue may carry issues their authors had no part in. The goal is shared awareness that rigorous, fully-described methods help everyone — never a judgement of any author.

Are you an author? We would genuinely like to hear from you — to clarify the record, add data or code, re-run the pipeline after an accession update, and publish your response right next to the assessment. Everything here is open and auditable.

🚩 Report an error in this record

Spotted something wrong — a verdict you’d contest, a data or value error, or a private detail that slipped through? Tell us, with a short justification. Authors and readers are equally welcome to write in; we review every report.

Prefer email, or the form below not working? Contact us at support@doesitreproduce.com.

Reproduction footprint

claude-opus-4-8

Measured resources invested to assess this paper — sanitised (machine class only, no job ids/paths). Compute = HPC accounting (SLURM); tokens = the AI agent's session.

89 k
tokens (I/O) · 4.8 M incl. cache
8 min
runtime
Per-job HPC accounting not captured for this run — the runtime shown is the reproduction’s measured wall-clock time.